Skip to content

How Nodus works

View Markdown

This page is the map. Each section links to the concept page that goes deeper.

You describe work as a resource: a small declarative object with a kind, a metadata.name and a spec, the same shape Kubernetes uses. You create it with the CLI, the Python SDK, the console or the HTTP API, and Nodus reports progress in its status. The same object reads the same everywhere, so nodus get, kubectl get and the console show one truth.

You want to Kind Everyday command
Run a command to completion Job (and Pipeline, Sweep to chain or fan out Jobs) nodus run, nodus apply -f job.yaml
Keep an isolated container for untrusted code Sandbox nodus create sandbox, nodus exec -it
Call Python remotely and fan out App, Function, FunctionCall nodus deploy app.py
Run a durable agent Agent, AgentRun, AgentGroup nodus create agentrun
Serve or call a model Model, InferenceEndpoint nodus get models -n nodus
Train or evaluate with a recipe TrainingJob, TrainingRuntime, Environment nodus create trainingjob
Develop on a remote machine Workspace nodus ssh workspace/<name>
Store data, images and secrets Volume, Image, Secret, Connection nodus volume put
Bring your own machines Pool, Node, EnrollmentToken, CloudAccount nodus create pool

TrainingJob and TrainingRuntime are Beta. Every other kind above is generally available. Run nodus api-resources for the full list and nodus explain <kind>.spec for any field.

An org is your billing and access boundary: members, API keys, credit and budgets belong to it. Inside an org, projects group resources (every org starts with default). Pass -p <project> to the CLI, or set NODUS_PROJECT for a whole shell. Names are unique within a project, and labels such as team=nlp let you select resources and break down cost across projects.

You state requirements (an accelerator and count, memory, a region class, a deadline or a cost ceiling) and Nodus chooses an offering that satisfies them, such as h100-sxm-80g-x8-us. You see offerings and Nodus ids, never the machines behind them. Each placement of your container on a machine is an Attempt. When capacity is reclaimed, Nodus starts a new Attempt and restores the files your program saved to its checkpoint directory (NODUS_CHECKPOINT_DIR). Your program reloads its own model, optimizer and progress from those files; Nodus restores files, not process memory.

Use the billing guide to add credits, review usage and set budgets. Review the estimate before starting a run and set the cost limits your task needs.