Attempts and recovery
View MarkdownEvery run on Nodus executes as one or more attempts. An attempt is one incarnation of your command on one machine. When that machine is reclaimed, loses its network or fails, Nodus starts a new attempt somewhere else and your run continues from what it saved. This page explains what carries over, how Nodus decides a machine is gone, and what you pay for along the way.
Attempts and epochs
Section titled “Attempts and epochs”A Job index, a Sandbox or a worker slot holds one running attempt at a time. Each new attempt gets the next epoch, a number that only grows. Nodus accepts reports, checkpoints and outputs only from the current epoch, so a machine that comes back after it was replaced can never overwrite the work of its successor.
You see a run’s attempts with nodus get attempts -l nodus.dev/job=<name>. A run’s attempts share its
name with an epoch suffix, and gang members add a rank suffix (-r1, -r2).
| Attempt phase | Meaning |
|---|---|
Pending, Placing, Acquiring |
Choosing and preparing capacity |
Starting |
The machine is ready; the image is pulled and inputs or a checkpoint are restored |
Running |
Your command is running |
Succeeded |
Your command exited 0 |
Failed |
Your command or its machine failed; reason says which (NodeLost, Preempted, OOMKilled, …) |
Cancelled |
The attempt was stopped on purpose: a suspend, a cancel, a budget or lifetime limit |
Continuity modes
Section titled “Continuity modes”recovery.continuity says what a new attempt starts from:
| Mode | A new attempt starts with | Use it for |
|---|---|---|
Checkpointed |
The latest committed checkpoint of your declared state paths (NODUS_CHECKPOINT_DIR by default) |
Training and long jobs that save their own model, optimizer and progress files |
Restartable |
A cold start plus the progress cursor you reported (NODUS_CURSOR_COMPLETED, NODUS_CURSOR_TOTAL) |
Batch work that can skip what it already finished |
Ephemeral |
A cold start | Short or idempotent work |
Snapshotted |
The latest filesystem snapshot | Sandboxes |
Restoring files restores files only, never process memory. Your program loads its own checkpoint when it starts; Nodus never adds resume flags to your command.
An empty checkpoint never counts as saved progress and never replaces an earlier useful one. If every checkpoint a
run commits is empty, the run shows Checkpointed=False, reason=NotCheckpointable, and Nodus plans and prices it
as Ephemeral from then on.
What survives a preemption
Section titled “What survives a preemption”When a provider reclaims interruptible capacity or a machine stops answering, Nodus recovers make-before-break:
- On a reclaim notice, Nodus asks your attempt for an urgent checkpoint and, at the same time, starts preparing a replacement machine.
- The replacement is prepared up to the point where it could start, but it does not start while the old machine might still be writing.
- It starts only once the old machine is provably gone: the old attempt acknowledged its stop, the provider confirmed the machine terminated, or the old attempt’s lease ran out.
- The new attempt restores according to the continuity mode above.
So a Checkpointed run loses at most the work since its last committed checkpoint, and a checkpoint committed
during the reclaim notice still counts. If the old machine comes back before it is replaced, the replacement is
released and your run simply continues; you are not charged for the replacement.
A machine that stops sending heartbeats gets a replacement prepared after 20 seconds. After 60 seconds of silence an attempt on shared capacity is declared lost and replaced. An attempt on a machine dedicated to your organization may keep running up to the edge of its funding while its replacement waits, because such a machine can only lose its own work.
Recovery limits
Section titled “Recovery limits”Recovery stops, and the run fails with a typed reason, when:
| Reason | Rule |
|---|---|
NoProgress |
Two attempts in a row made no progress. Progress means the command started and either ran for recovery.minProgressDuration (default 2 minutes) or committed a non-empty checkpoint. A preemption after progress never counts |
RestoreFailed |
The same checkpoint failed to restore twice |
ImagePullFailed |
The image failed to pull on a second machine (a pull failure is retried once elsewhere) |
MaxAttemptsExceeded |
recovery.maxAttempts recoveries were used (default 8; 3 for distributed Jobs) |
Preempted, NodeLost |
recovery.onInterruption: Fail was set, so the first interruption ends the run |
Failures caused by your command (a non-zero exit, out of memory, an invalid checkpoint) fail fast and are not retried. At most three recoveries per organization prepare capacity at the same time; the others wait their turn and show an Event.
Suspending
Section titled “Suspending”nodus suspend job/<name> stops the run after a final checkpoint and releases its compute; nodus resume
continues at a new epoch from that checkpoint. If the final checkpoint fails or takes longer than
max(10 minutes, 2 × the shutdown reserve), the run keeps running with Suspended=False, reason=SuspendFailed
and a SuspendFailed Event, so a suspend never throws work away. Stops caused by money (credits, budgets, the
maximum cost) or by a lifetime limit do not wait: they stop at the funded edge with whatever checkpoint exists.
A stop always wins over recovery. If the machine is lost while a suspend, cancel or money stop is in progress, the
attempt ends Cancelled with the stop’s reason and nothing is restarted or billed again; nodus resume starts the
next epoch as usual.
Review recovery attempts
Section titled “Review recovery attempts”A machine that never becomes ready is replaced up to six times per attempt; after that the attempt fails with
ReadinessFailed and the run’s retry rule applies. Use the run’s Attempts and Cost tabs to inspect its history.