Skip to content

Spot vs on-demand GPUs: what breaks, and how do you make spot safe?

Spot GPUs cost up to 90% less but can vanish with 30 seconds of notice. What breaks, each cloud's notice window, the cost math and a safe setup.

Short answer (October 10, 2026): spot GPUs are the same hardware as on-demand GPUs at a deep discount (AWS and Azure advertise up to 90% off, GCP up to 91%), but the cloud can take the machine back with 30 seconds to 2 minutes of notice, and anything not saved somewhere durable is gone. Spot is safe when your job checkpoints to storage that outlives the machine, saves again when the notice arrives, resumes on start and can fall back when spot capacity runs out; without that, a few reclaims can make spot cost more than on-demand.

Below: what fails on a reclaim, how much warning each platform gives, when spot really wins, and a copy-paste setup.

On-demand Spot (interruptible)
Hourly price List price Discounted, often 60% to 90% below list
Runs until You stop it You stop it, or the cloud needs the capacity back
Warning before a reclaim None needed 30 seconds to 2 minutes
Capacity when you ask Usually there Can be empty for hours
Best for Real-time inference, jobs that cannot checkpoint Checkpointed training, batch inference, sweeps, evals

From each platform’s own docs, read October 10, 2026.

Platform Listed discount Notice before reclaim How your code finds out What happens to the machine
AWS Spot Instances Up to 90% 2 minutes spot/instance-action in instance metadata returns JSON (404 until then) Stopped, hibernated or terminated
GCP Spot VMs Up to 91% Up to 30 s shutdown by default; an optional 120 s notice period before it instance/preempted metadata flips to TRUE, then a soft-off triggers your shutdown script Stopped or deleted; GPU Spot VMs are also preempted for maintenance
Azure Spot VMs Up to 90% 30 s minimum A Preempt event in Scheduled Events Deallocated (disks kept and billed) or deleted with its disks
Modal n/a Grace period after an interrupt signal Your exit handler GPU Functions are always preemptible; restarted on the same input
Nodus n/a Urgent checkpoint request on reclaim; SIGTERM, then SIGKILL 30 s later on a stop Events socket, or a SIGTERM handler New attempt on another machine with your state directory restored

Sources: AWS interruption notices, GCP Spot VMs, Azure Spot VMs, Modal preemption.

What actually breaks when a spot GPU is reclaimed

Section titled “What actually breaks when a spot GPU is reclaimed”

Most spot guides stop at “use checkpoints”. These failures still bite teams who do.

What breaks What you see Fix
Work since the last save is lost Run resumes from an old step, or from step 0 Save on a cadence, and save again when the notice arrives
Half-written checkpoint torch.load fails or loads a corrupt file on resume Write to a temp file, then os.replace it; keep the last two
Checkpoint on the machine’s local disk Nothing to resume from; local SSDs and deleted VMs take the data with them Write to object storage or a volume that outlives the machine
Only the weights were saved Loss spikes or repeated samples after resume Also save optimizer, LR scheduler, AMP scaler, RNG states, epoch and sampler offset
One rank of a multi-node job is reclaimed Every other rank hangs in an NCCL collective until it times out Restart the whole gang from a sharded torch.distributed.checkpoint
Script exits 0 after its emergency save The scheduler marks the run finished, and it never resumes Exit non-zero after a notice-triggered save
Spot pool is empty on relaunch Job sits pending for hours Accept several GPU types or regions, or fall back to on-demand
A restart loop that never makes progress You pay for boots that crash before the first save Cap retries by progress, not by count

The cost math: when does spot actually win?

Section titled “The cost math: when does spot actually win?”

Spot is billed for everything the job does, including the work it loses:

billed hours = useful hours x (1 + save overhead) + reclaims x (lost work + restart time)
spot wins while billed hours / useful hours < 1 / (1 - discount)

At 60% off, spot still breaks even when you pay for 2.5x the useful hours. At 30% off, the limit is 1.43x, so restarts eat the discount fast.

A worked example with illustrative rates: 8 GPUs, 20 hours of useful training, $4.00 per GPU-hour on-demand, spot at 60% off ($1.60), 3 reclaims, a 15 minute restart (new machine, image pull, checkpoint load) and a 1 minute save every 30 minutes (3.3% overhead).

Setup Billed hours per GPU Cost for 8 GPUs
On-demand, no reclaims 20.0 $640
Spot, save every 30 min (about 15 min lost per reclaim) 20.67 + 3 x 0.5 = 22.2 $284
Spot, save every 30 min and on notice 20.67 + 3 x 0.25 = 21.4 $274
Spot, no checkpoints (reclaimed after 7 h, 12 h and 5 h) 24 + 20 + 3 x 0.25 = 44.75 $573

Checkpointed spot saves about 56%. Spot without checkpoints saves 10% while gambling the whole run, and a fourth reclaim more than about 5 hours into an attempt would push it past on-demand. For real inputs, the Nodus pricing page lists an H100 SXM at $2.60/hr and a B200 at $5.50/hr (as of October 10, 2026), and AWS’s Spot Instance Advisor reports interruption frequency in bands from under 5% to over 20% per month.

A notice watcher that works on AWS, GCP and Azure

Section titled “A notice watcher that works on AWS, GCP and Azure”

Run this alongside your training loop (replace train_step and save_checkpoint with your own). A background thread polls each cloud’s metadata endpoint and also catches SIGTERM, which Kubernetes and Nodus send before killing a container.

import signal, sys, threading, time, urllib.request
stop = threading.Event()
def _get(url, headers=None, method="GET"):
try:
req = urllib.request.Request(url, headers=headers or {}, method=method)
with urllib.request.urlopen(req, timeout=1) as r:
return r.read().decode()
except Exception:
return None # no notice yet, or a different cloud
def notice_pending():
# AWS: IMDSv2 token, then instance-action exists only after a notice
token = _get("http://169.254.169.254/latest/api/token",
{"X-aws-ec2-metadata-token-ttl-seconds": "300"}, "PUT")
if token and _get("http://169.254.169.254/latest/meta-data/spot/instance-action",
{"X-aws-ec2-metadata-token": token}):
return True
# GCP: preempted flips to TRUE on reclaim
if _get("http://metadata.google.internal/computeMetadata/v1/instance/preempted",
{"Metadata-Flavor": "Google"}) == "TRUE":
return True
# Azure: a Preempt entry in Scheduled Events
events = _get("http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01",
{"Metadata": "true"})
return bool(events) and '"Preempt"' in events
def watch(every=5):
while not stop.is_set():
if notice_pending():
stop.set()
time.sleep(every)
def on_sigterm(*_):
stop.set()
signal.signal(signal.SIGTERM, on_sigterm)
threading.Thread(target=watch, daemon=True).start()
for step in range(start_step, total_steps):
train_step()
if step % save_every == 0 or stop.is_set():
save_checkpoint(step) # atomic write to durable storage
if stop.is_set():
sys.exit(143) # non-zero: interrupted, not finished

Two limits. A full save must fit inside the notice (30 seconds on GCP and Azure by default), so keep regular saves frequent enough that a missed emergency save costs little. Inside Docker on AWS, the IMDSv2 token has a hop limit of 1 by default; raise it to 2 (aws ec2 modify-instance-metadata-options --http-put-response-hop-limit 2) or the watcher never sees the notice. The checkpointing walkthrough covers the save and resume code itself.

Nodus is an AI compute platform that runs each job on the cheapest capacity that will finish it, with checkpoints and recovery built in. Interruptible capacity is opt-in per job:

Terminal window
pip install nodus-compute
nodus login
nodus get gpus --interruptible -o wide # interruptible offerings, startup time, interruption rate
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 --dry-run -- python train.py
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 -d -- python train.py
nodus get attempts -l nodus.dev/job=<name> # one row per machine, with reasons such as Preempted

--dry-run prints the expected cost and start time without launching. allow uses interruptible capacity only when it is cheaper to finish; prefer (the bare --interruptible) favors it with a 10% allowance; never keeps the job off it. On a reclaim, Nodus requests an urgent checkpoint, prepares a replacement in parallel, starts it only once the old machine is provably gone, and restores /nodus/state. A run with no progress twice in a row stops with NoProgress instead of looping, and --max-cost suspends the job with a checkpoint at the cap. Your code still decides what goes into the state directory and loads it on start. Multi-node gang checkpoints are Beta. New accounts get a $30 starter grant once the email is verified, valid for 30 days.

How much notice do I get before a spot GPU is reclaimed? Two minutes on AWS, 30 seconds by default on GCP (with an optional 120 second notice period) and at least 30 seconds on Azure, as of October 10, 2026.

How often do spot GPUs get interrupted? It varies by GPU, region and week. AWS publishes monthly interruption bands per instance pool; nodus get gpus -o wide shows a measured rate per interruptible offering.

Can I serve inference on spot GPUs? Batch and offline inference, yes. For real-time traffic, keep an on-demand floor and use spot only for overflow.

Is spot cheaper than reserved capacity? For bursty, checkpointed work, usually. If GPUs stay busy most hours for months, a reservation can beat spot with no reclaim risk; Nodus offers Reserved Compute in Beta.

What is the difference between spot and preemptible VMs on GCP? Preemptible VMs are the older model with a 24 hour maximum runtime. Spot VMs have no maximum runtime unless you set one.

More from the blog