rl.train starts without building an image

rl.train(tasks, reward) no longer builds an image. The SDK uploads the reward, the lines of its file it uses, its tasks and any pip= packages as one small bundle, and the grader runs it as it is, still with no network. The first run starts as quickly as the second.

pip= packages are packed as Python 3.12 Linux wheels. The whole bundle is at most 512 KB compressed; a larger one, or a package with no wheel, fails on your machine before anything runs and names the largest packages. NumPy, SymPy, regex, jsonschema, Math-Verify, Reasoning Gym and OpenEnv are already installed, so they need no pip= entry.

An environment that downloads its data as it loads cannot run with no network. Such a run now fails with a message that names the environment and says why, instead of an internal error.

run.watch() shows the stages while the model is cached and the run is prepared, and starts printing rewards once training begins, instead of stopping with NotFound right after rl.train returns.