Skip to content

nodus.recipes.rl

View Markdown

Reinforcement learning and evaluation on catalog Environments or your own.

From a task list and a reward function to a training run in one call. python -m nodus.examples.rl runs a whole example (nodus/examples/rl.py), a template to copy:

def reward(completion, answer):
return 1.0 if answer in completion else 0.0
run = rl.train([("Spell 'cat' backwards.", "tac"), ...], reward, max_cost=2) # model= picks another base model
run.watch() # each stage, then reward, loss and KL per step
run.outputs.download("./outputs") # the LoRA adapter and the before/after comparison

train packages the reward with the statements of its file it uses as an Environment in your project, uploaded as one small bundle with no image build, then runs GRPO on it, learning from groups of replies. rl.train("nodus/[email protected]") trains on a catalog Environment the same way. To set every GRPO parameter yourself, on a catalog Environment or one you build (examples/training/custom-reward):

job = rl.grpo_lora(model="Qwen/Qwen3-0.6B", environment="nodus/[email protected]",
steps=50, held_out=64, seed=42)
plan = job.preview() # estimate, compiled Job, blocking reasons, ETag
run = plan.run(gpu="RTX-4090", max_cost=5, idempotency_key="gc-001")
print(run.wait().summary.comparison) # the numbers the console shows

The trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training.

evaluate(*, model: Any, environment: str | None = None, tasks: Iterable[str] | None = None, held_out: int = 64, seed: int = 0, task_filter: Mapping[str, str] | None = None, **params: Any) -> Any

mode: Evaluate on evaluate: an Environment’s held-out tasks, or benchmark tasks such as arc_easy.

grpo_lora(*, model: Any, environment: str, steps: int = 50, train_tasks: int | None = None, held_out: int | None = None, seed: int = 0, lora: LoRA | None = None, task_filter: Mapping[str, str] | None = None, min_comparison_tasks: int = 16, **params: Any) -> Any

GRPO with a LoRA adapter (grpo-lora): baseline, train, then the final eval on the same held-out tasks.

train_tasks and held_out default to 256 and 64, or the Environment’s split sizes when those are smaller.