# Spot vs on-demand GPUs: what breaks, and how do you make spot safe?

> Spot GPUs cost up to 90% less but can vanish with 30 seconds of notice. What breaks, each cloud's notice window, the cost math and a safe setup.

Source: https://www.nodus-compute.ai/blog/spot-vs-on-demand-gpus/
Build revision: b86f52cf237fbb563fd90f8f7b22497a1c608ff1

**Short answer (October 10, 2026):** spot GPUs are the same hardware as on-demand GPUs at a deep discount (AWS and Azure advertise up to 90% off, GCP up to 91%), but the cloud can take the machine back with 30 seconds to 2 minutes of notice, and anything not saved somewhere durable is gone. Spot is safe when your job checkpoints to storage that outlives the machine, saves again when the notice arrives, resumes on start and can fall back when spot capacity runs out; without that, a few reclaims can make spot cost more than on-demand.

Below: what fails on a reclaim, how much warning each platform gives, when spot really wins, and a copy-paste setup.

## Spot vs on-demand at a glance

||On-demand|Spot (interruptible)|
|-|-|-|
|Hourly price|List price|Discounted, often 60% to 90% below list|
|Runs until|You stop it|You stop it, or the cloud needs the capacity back|
|Warning before a reclaim|None needed|30 seconds to 2 minutes|
|Capacity when you ask|Usually there|Can be empty for hours|
|Best for|Real-time inference, jobs that cannot checkpoint|Checkpointed training, batch inference, sweeps, evals|

## How much notice each platform gives

From each platform’s own docs, read October 10, 2026.

|Platform|Listed discount|Notice before reclaim|How your code finds out|What happens to the machine|
|-|-|-|-|-|
|AWS Spot Instances|Up to 90%|2 minutes|`spot/instance-action` in instance metadata returns JSON (404 until then)|Stopped, hibernated or terminated|
|GCP Spot VMs|Up to 91%|Up to 30 s shutdown by default; an optional 120 s notice period before it|`instance/preempted` metadata flips to `TRUE`, then a soft-off triggers your shutdown script|Stopped or deleted; GPU Spot VMs are also preempted for maintenance|
|Azure Spot VMs|Up to 90%|30 s minimum|A `Preempt` event in Scheduled Events|Deallocated (disks kept and billed) or deleted with its disks|
|Modal|n/a|Grace period after an interrupt signal|Your exit handler|GPU Functions are always preemptible; restarted on the same input|
|[Nodus](https://www.nodus-compute.ai/docs/concepts/attempts-and-recovery/)|n/a|Urgent checkpoint request on reclaim; `SIGTERM`, then `SIGKILL` 30 s later on a stop|Events socket, or a `SIGTERM` handler|New attempt on another machine with your state directory restored|

Sources: [AWS interruption notices](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-instance-termination-notices.html), [GCP Spot VMs](https://cloud.google.com/compute/docs/instances/spot), [Azure Spot VMs](https://learn.microsoft.com/en-us/azure/virtual-machines/spot-vms), [Modal preemption](https://modal.com/docs/guide/preemption).

## What actually breaks when a spot GPU is reclaimed

Most spot guides stop at “use checkpoints”. These failures still bite teams who do.

|What breaks|What you see|Fix|
|-|-|-|
|Work since the last save is lost|Run resumes from an old step, or from step 0|Save on a cadence, and save again when the notice arrives|
|Half-written checkpoint|`torch.load` fails or loads a corrupt file on resume|Write to a temp file, then `os.replace` it; keep the last two|
|Checkpoint on the machine’s local disk|Nothing to resume from; local SSDs and deleted VMs take the data with them|Write to object storage or a volume that outlives the machine|
|Only the weights were saved|Loss spikes or repeated samples after resume|Also save optimizer, LR scheduler, AMP scaler, RNG states, epoch and sampler offset|
|One rank of a multi-node job is reclaimed|Every other rank hangs in an NCCL collective until it times out|Restart the whole gang from a sharded `torch.distributed.checkpoint`|
|Script exits 0 after its emergency save|The scheduler marks the run finished, and it never resumes|Exit non-zero after a notice-triggered save|
|Spot pool is empty on relaunch|Job sits pending for hours|Accept several GPU types or regions, or fall back to on-demand|
|A restart loop that never makes progress|You pay for boots that crash before the first save|Cap retries by progress, not by count|

## The cost math: when does spot actually win?

Spot is billed for everything the job does, including the work it loses:

```text
billed hours = useful hours x (1 + save overhead) + reclaims x (lost work + restart time)
spot wins while billed hours / useful hours < 1 / (1 - discount)
```

At 60% off, spot still breaks even when you pay for 2.5x the useful hours. At 30% off, the limit is 1.43x, so restarts eat the discount fast.

A worked example with illustrative rates: 8 GPUs, 20 hours of useful training, $4.00 per GPU-hour on-demand, spot at 60% off ($1.60), 3 reclaims, a 15 minute restart (new machine, image pull, checkpoint load) and a 1 minute save every 30 minutes (3.3% overhead).

|Setup|Billed hours per GPU|Cost for 8 GPUs|
|-|-|-|
|On-demand, no reclaims|20.0|$640|
|Spot, save every 30 min (about 15 min lost per reclaim)|20.67 + 3 x 0.5 = 22.2|$284|
|Spot, save every 30 min and on notice|20.67 + 3 x 0.25 = 21.4|$274|
|Spot, no checkpoints (reclaimed after 7 h, 12 h and 5 h)|24 + 20 + 3 x 0.25 = 44.75|$573|

Checkpointed spot saves about 56%. Spot without checkpoints saves 10% while gambling the whole run, and a fourth reclaim more than about 5 hours into an attempt would push it past on-demand. For real inputs, the [Nodus pricing page](https://www.nodus-compute.ai/pricing/) lists an H100 SXM at $2.60/hr and a B200 at $5.50/hr (as of October 10, 2026), and AWS’s Spot Instance Advisor reports interruption frequency in bands from under 5% to over 20% per month.

## A notice watcher that works on AWS, GCP and Azure

Run this alongside your training loop (replace `train_step` and `save_checkpoint` with your own). A background thread polls each cloud’s metadata endpoint and also catches `SIGTERM`, which Kubernetes and Nodus send before killing a container.

```python
import signal, sys, threading, time, urllib.request

stop = threading.Event()

def _get(url, headers=None, method="GET"):
    try:
        req = urllib.request.Request(url, headers=headers or {}, method=method)
        with urllib.request.urlopen(req, timeout=1) as r:
            return r.read().decode()
    except Exception:
        return None  # no notice yet, or a different cloud

def notice_pending():
    # AWS: IMDSv2 token, then instance-action exists only after a notice
    token = _get("http://169.254.169.254/latest/api/token",
                 {"X-aws-ec2-metadata-token-ttl-seconds": "300"}, "PUT")
    if token and _get("http://169.254.169.254/latest/meta-data/spot/instance-action",
                      {"X-aws-ec2-metadata-token": token}):
        return True
    # GCP: preempted flips to TRUE on reclaim
    if _get("http://metadata.google.internal/computeMetadata/v1/instance/preempted",
            {"Metadata-Flavor": "Google"}) == "TRUE":
        return True
    # Azure: a Preempt entry in Scheduled Events
    events = _get("http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01",
                  {"Metadata": "true"})
    return bool(events) and '"Preempt"' in events

def watch(every=5):
    while not stop.is_set():
        if notice_pending():
            stop.set()
        time.sleep(every)

def on_sigterm(*_):
    stop.set()

signal.signal(signal.SIGTERM, on_sigterm)
threading.Thread(target=watch, daemon=True).start()

for step in range(start_step, total_steps):
    train_step()
    if step % save_every == 0 or stop.is_set():
        save_checkpoint(step)   # atomic write to durable storage
    if stop.is_set():
        sys.exit(143)           # non-zero: interrupted, not finished
```

Two limits. A full save must fit inside the notice (30 seconds on GCP and Azure by default), so keep regular saves frequent enough that a missed emergency save costs little. Inside Docker on AWS, the IMDSv2 token has a hop limit of 1 by default; raise it to 2 (`aws ec2 modify-instance-metadata-options --http-put-response-hop-limit 2`) or the watcher never sees the notice. The [checkpointing walkthrough](https://www.nodus-compute.ai/blog/resume-training-after-spot-preemption/) covers the save and resume code itself.

## Making spot safe on Nodus

Nodus is an AI compute platform that runs each job on the cheapest capacity that will finish it, with checkpoints and recovery built in. Interruptible capacity is opt-in per job:

Terminal window

```sh
pip install nodus-compute
nodus login
nodus get gpus --interruptible -o wide    # interruptible offerings, startup time, interruption rate
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 --dry-run -- python train.py
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 -d -- python train.py
nodus get attempts -l nodus.dev/job=<name>  # one row per machine, with reasons such as Preempted
```

`--dry-run` prints the expected cost and start time without launching. `allow` uses interruptible capacity only when it is cheaper to finish; `prefer` (the bare `--interruptible`) favors it with a 10% allowance; `never` keeps the job off it. On a reclaim, Nodus requests an urgent checkpoint, prepares a replacement in parallel, starts it only once the old machine is provably gone, and restores `/nodus/state`. A run with no progress twice in a row stops with `NoProgress` instead of looping, and `--max-cost` suspends the job with a checkpoint at the cap. Your code still decides what goes into the state directory and loads it on start. Multi-node gang checkpoints are Beta. New accounts get a $30 starter grant once the email is verified, valid for 30 days.

## FAQ

**How much notice do I get before a spot GPU is reclaimed?** Two minutes on AWS, 30 seconds by default on GCP (with an optional 120 second notice period) and at least 30 seconds on Azure, as of October 10, 2026.

**How often do spot GPUs get interrupted?** It varies by GPU, region and week. AWS publishes monthly interruption bands per instance pool; `nodus get gpus -o wide` shows a measured rate per interruptible offering.

**Can I serve inference on spot GPUs?** Batch and offline inference, yes. For real-time traffic, keep an on-demand floor and use spot only for overflow.

**Is spot cheaper than reserved capacity?** For bursty, checkpointed work, usually. If GPUs stay busy most hours for months, a reservation can beat spot with no reclaim risk; Nodus offers [Reserved Compute](https://www.nodus-compute.ai/reserved-compute/) in Beta.

**What is the difference between spot and preemptible VMs on GCP?** Preemptible VMs are the older model with a 24 hour maximum runtime. Spot VMs have no maximum runtime unless you set one.
