nodus.recipes.rl
View MarkdownReinforcement learning and evaluation on catalog Environments or your own.
From a task list and a reward function to a training run in one call. python -m nodus.examples.rl runs a whole
example (nodus/examples/rl.py), a template to copy:
def reward(completion, answer): return 1.0 if answer in completion else 0.0
run = rl.train([("Spell 'cat' backwards.", "tac"), ...], reward, max_cost=2) # model= picks another base modelrun.watch() # each stage, then reward, loss and KL per steprun.outputs.download("./outputs") # the LoRA adapter and the before/after comparisontrain packages the reward with the statements of its file it uses as an Environment in your project, uploaded
as one small bundle with no image build, then runs GRPO on it, learning from groups of replies.
rl.train("nodus/[email protected]") trains on a catalog Environment the same way. To set every GRPO parameter yourself,
on a catalog Environment or one you build (examples/training/custom-reward):
job = rl.grpo_lora(model="Qwen/Qwen3-0.6B", environment="nodus/[email protected]", steps=50, held_out=64, seed=42)plan = job.preview() # estimate, compiled Job, blocking reasons, ETagrun = plan.run(gpu="RTX-4090", max_cost=5, idempotency_key="gc-001")print(run.wait().summary.comparison) # the numbers the console showsThe trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training.
evaluate
Section titled “evaluate”evaluate(*, model: Any, environment: str | None = None, tasks: Iterable[str] | None = None, held_out: int = 64, seed: int = 0, task_filter: Mapping[str, str] | None = None, **params: Any) -> Anymode: Evaluate on evaluate: an Environment’s held-out tasks, or benchmark tasks such as arc_easy.
grpo_lora
Section titled “grpo_lora”grpo_lora(*, model: Any, environment: str, steps: int = 50, train_tasks: int | None = None, held_out: int | None = None, seed: int = 0, lora: LoRA | None = None, task_filter: Mapping[str, str] | None = None, min_comparison_tasks: int = 16, **params: Any) -> AnyGRPO with a LoRA adapter (grpo-lora): baseline, train, then the final eval on the same held-out tasks.
train_tasks and held_out default to 256 and 64, or the Environment’s split sizes when those are smaller.