# nodus.recipes.rl

> Reinforcement learning and evaluation on catalog Environments or your own.

Source: https://www.nodus-compute.ai/docs/reference/python/nodus-recipes-rl/
Build revision: 4ebfc6023eeed1bd55d9969af1612182a7b0c7ff

<!-- Generated by scripts/gen-reference.mjs from the SDK docstrings (griffe). Do not edit: run make gen. -->

Reinforcement learning and evaluation on catalog Environments or your own.

From a task list and a reward function to a training run in one call. `python -m nodus.examples.rl` runs a whole example (`nodus/examples/rl.py`), a template to copy:

```plaintext
def reward(completion, answer):
    return 1.0 if answer in completion else 0.0

run = rl.train([("Spell 'cat' backwards.", "tac"), ...], reward, max_cost=2)   # model= picks another base model
run.watch()                                           # each stage, then reward, loss and KL per step
run.outputs.download("./outputs")                     # the LoRA adapter and the before/after comparison
```

`train` packages the reward with the statements of its file it uses as an Environment in your project, uploaded as one small bundle with no image build, then runs GRPO on it, learning from groups of replies. `rl.train("nodus/gsm8k@1.0.0")` trains on a catalog Environment the same way. To set every GRPO parameter yourself, on a catalog Environment or one you build (`examples/training/custom-reward`):

```plaintext
job = rl.grpo_lora(model="Qwen/Qwen3-0.6B", environment="nodus/graph-coloring@1.0.0",
                   steps=50, held_out=64, seed=42)
plan = job.preview()                                  # estimate, compiled Job, blocking reasons, ETag
run = plan.run(gpu="RTX-4090", max_cost=5, idempotency_key="gc-001")
print(run.wait().summary.comparison)                  # the numbers the console shows
```

The trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training.

## `evaluate`

```python
evaluate(*, model: Any, environment: str | None = None, tasks: Iterable[str] | None = None, held_out: int = 64, seed: int = 0, task_filter: Mapping[str, str] | None = None, **params: Any) -> Any
```

`mode: Evaluate` on `evaluate`: an Environment’s held-out tasks, or benchmark `tasks` such as arc_easy.

## `grpo_lora`

```python
grpo_lora(*, model: Any, environment: str, steps: int = 50, train_tasks: int | None = None, held_out: int | None = None, seed: int = 0, lora: LoRA | None = None, task_filter: Mapping[str, str] | None = None, min_comparison_tasks: int = 16, **params: Any) -> Any
```

GRPO with a LoRA adapter (`grpo-lora`): baseline, train, then the final eval on the same held-out tasks.

`train_tasks` and `held_out` default to 256 and 64, or the Environment’s split sizes when those are smaller.
