← All updates
Update

Reinforcement learning

Explore, train and evaluate with RL

Explore tasks in Environments, or start directly from your own RL command in New run. Research discovery, custom execution and managed recipes have different availability requirements.

Discover and prepare

Search and filter environments by task, project, category and goal. The research catalog covers Reasoning Gym, SWE-smith, TextArena, MiniWoB++, OfficeQA and AppWorld, with examples, access requirements and integration status. OfficeQA is evaluation-only. These entries are research references that need integration, not a collection of ready-to-launch managed environments.

Save environments and drafts in the current account and team's browser, set the model label, task target and planned spending limit, then review the setup. Open custom workload carries those choices into the command form. You can also export a planning draft. Planning does not allocate compute.

Run your code or an available recipe

Custom RL runs accept your code, runtime, model, datasets, trainer, rewards and checkpoints. Your program controls the algorithm and must save and reload its own state. Selecting RLHF, PPO or GRPO as a workload description does not install that algorithm.

For a managed recipe that the console currently marks available, review its fixed model, task, runtime and outputs. Evaluate the base model measures the baseline. Train and compare measures a baseline, trains and evaluates on held-out tasks, with configuration for task count, training steps, seed, traces and spending. Managed recipes are individually qualified. A recipe's availability applies only to that recipe.

Follow outcomes and keep evidence

Emit task-start and task-completion events to see attempts, pass rate, mean reward, outcomes and optional task traces. Baseline, training and evaluation remain separate. Export evidence, inspect logs and costs, stop the run, or retrieve its declared outputs.

Available managed recipes return task results, manifests and provenance, with an adapter only for training. The console checks saved evidence before showing a managed comparison. Missing evidence stays unavailable, and an improvement in model quality is never assumed from completion or a reward curve.

Explore RL environments · RL run guide