# What are the best Modal alternatives for GPU training in 2026?

> Modal alternatives for training compared on listed H100 prices, preemption handling, run limits and cost per run, as of October 9, 2026.

Source: https://www.nodus-compute.ai/blog/modal-alternatives-for-training-2026/
Build revision: 9c47048fc5d272b76910fd3d625e29a8c58f3179

**Short answer (October 9, 2026):** the strongest Modal alternatives for GPU training are Nodus if you want the same Python API with lower listed GPU prices and automatic checkpoint recovery, CoreWeave or Together GPU clusters if you need whole 8-GPU nodes, Baseten Training if you already serve on Baseten, and AWS SageMaker if you are spending AWS credits. Listed H100 prices across these options run from $2.60 to $6.50 per GPU-hour, against $3.95 on Modal.

Modal is a great way to run Python on GPUs, and for short jobs it is hard to beat. Training is where its limits show up, so this guide compares the alternatives on what actually decides a training bill: price per GPU-hour, what happens when a machine is reclaimed, how long one run can last, and how many GPUs one job can use. Every price below was read from the vendor’s public pricing page on October 9, 2026.

## Why people look past Modal for training

These come from Modal’s own docs, not from complaints:

* **Every GPU Function can be preempted.** Modal Functions are preemptible by default, and the `nonpreemptible` flag is not supported for GPU Functions ([Modal preemption docs](https://modal.com/docs/guide/preemption)). Modal restarts the Function on the same input, so a long run without its own checkpoints starts over.
* **24 hours per call, at most.** Function timeouts can be set between 1 second and 24 hours ([Modal timeouts](https://modal.com/docs/guide/timeouts)). Longer runs have to be split and resumed.
* **8 GPUs per container.** Multi-GPU training works on one node; multi-node training is in private beta by email request ([Modal GPU docs](https://modal.com/docs/guide/gpu)).
* **Pinning a region costs extra.** Region selection is billed at 1.15x to 1.75x base prices ([Modal pricing](https://modal.com/pricing)).

None of that matters for a 20-minute fine-tune. It matters a lot for a 30-hour run on 8 H100s.

## Modal alternatives for training, compared

|Platform|H100 per GPU-hour (listed)|How you pay|Notes for training|Best fit|
|-|-|-|-|-|
|Modal|$3.95|Per second, serverless Functions|Preemptible by default; you write the checkpoints; 24 h max per call|Short and medium Python jobs|
|[Nodus](https://www.nodus-compute.ai/pricing/)|$2.60|Prepaid credits, metered usage|State directory saved and restored on a new machine automatically|Code written for Modal, long or interruptible runs|
|Northflank|$2.74|Listed hourly; bring your own cloud|Can use your own commitments and spot capacity|Teams running a full app stack plus GPUs|
|Cerebrium|$3.40|Per second, serverless|Guaranteed capacity needs a monthly minimum spend|Serverless, inference-heavy teams|
|Together GPU clusters|$5.49|Per GPU-hour, HGX nodes|Reserved pricing by quote|Large clusters, plus a fine-tuning API|
|CoreWeave|$6.16 on-demand, about $2.45 spot|Whole 8-GPU HGX nodes|Spot nodes can be reclaimed|Big multi-node runs sold by the node|
|Baseten Training|$6.50|Per minute|Volume discounts available|Teams that already serve on Baseten|
|AWS SageMaker|Varies by instance and region|Per second, training jobs|Managed Spot Training copies checkpoints to S3, up to 90% off on-demand|Teams with AWS credits or commitments|

Sources, all read October 9, 2026: [Modal](https://modal.com/pricing), [Northflank](https://northflank.com/pricing), [Cerebrium](https://www.cerebrium.ai/pricing), [Together](https://www.together.ai/pricing), [CoreWeave](https://www.coreweave.com/pricing), [Baseten](https://www.baseten.co/pricing/) and the [SageMaker managed spot training docs](https://docs.aws.amazon.com/sagemaker/latest/dg/model-managed-spot-training.html). Per-second and per-minute rates are converted to hours, and CoreWeave’s 8-GPU node prices are divided by 8. The Nodus figure is the hourly supply price listed on the [Nodus pricing page](https://www.nodus-compute.ai/pricing/).

## What one training run costs

Take one run: 8 H100s for 10 hours, which is 80 GPU-hours. At listed rates on October 9, 2026:

|Platform|Cost of 80 H100-hours|
|-|-|
|CoreWeave spot (one 8-GPU node)|$195 to $197, interruptible|
|Nodus|$208.00|
|Northflank|$219.20|
|Cerebrium|$271.87|
|Modal|$315.94|
|Together GPU clusters|$439.20|
|CoreWeave on-demand|$492.40|
|Baseten Training|$519.98|

The listed rate is only half the bill. On preemptible or spot capacity, a run with no checkpoint that gets reclaimed at hour 7 pays for those 7 hours twice. On Modal that is an extra $221 on top of $316, about 70% more for the same result. With a save every 30 minutes, the same reclaim costs at most half an hour of work. That is why checkpoint handling belongs in the price comparison, not in a footnote.

## Moving a Modal training Function to Nodus

Nodus is an AI compute platform that runs each job on the cheapest capacity that will finish it, with checkpoints and recovery built in. Its Python SDK keeps Modal’s names where the concept is the same (`App`, `@app.function`, `.remote()`, `.map()`, `.spawn()`, `Image`, `Volume`, `Secret`), so most code moves with one import change ([migration guide](https://www.nodus-compute.ai/docs/getting-started/modal-users/)). For training you add two arguments: a checkpoint path and permission to use interruptible capacity.

```python
import nodus  # was: import modal

app = nodus.App("finetune")
image = nodus.Image.debian_slim().pip_install("torch", "transformers", "datasets")

@app.function(
    gpu="H100:8",
    image=image,
    timeout="12h",
    checkpoint="/nodus/state",  # saved, then restored on the next machine if this one is reclaimed
    interruptible=True,         # allow cheaper interruptible capacity; progress survives via the checkpoint
)
def train(lr: float = 2e-5) -> dict:
    # Load model, optimizer and step from /nodus/state if they exist, then keep training.
    ...

@app.local_entrypoint()
def main() -> None:
    print(train.estimate(2e-5))  # expected cost and start time; nothing runs
    print(train.remote(2e-5))
```

Run it, or skip the SDK and launch any training script as a Job with a hard spending cap:

Terminal window

```sh
pip install nodus-compute
nodus login
nodus run app.py

# any script: preview the estimate, then run with a $250 cap
nodus run --dry-run --gpu H100:8 -- torchrun train.py
nodus run --gpu H100:8 --interruptible --max-cost 250 -d -- torchrun train.py
```

When capacity is reclaimed, Nodus asks your process for an urgent checkpoint, prepares a replacement machine at the same time, and starts it only once the old one is provably gone, so two machines never write the same state ([attempts and recovery](https://www.nodus-compute.ai/docs/concepts/attempts-and-recovery/)). The full pattern, including Hugging Face Trainer, is in the [checkpoints guide](https://www.nodus-compute.ai/docs/guides/checkpoints/).

## When Modal is still the better pick

Be honest about the gaps. Nodus does not support Modal web endpoints (`@modal.web_endpoint`, `@modal.asgi_app`; use `nodus.InferenceEndpoint` for serving), `modal.Dict`, `modal.Queue` or parametrized classes, and its multi-node gangs are Beta with at most 8 nodes ([differences list](https://www.nodus-compute.ai/docs/getting-started/modal-users/)). If your training is short, sits next to web endpoints, or depends on those primitives, staying on Modal is reasonable. For clusters well beyond 8 nodes, Together GPU clusters or CoreWeave sell capacity built for that scale.

## How to choose in one minute

* Runs under an hour, pure Python, happy with Modal: stay.
* Same code, lower listed price, runs that must survive reclaims: Nodus.
* Tens of nodes on one fabric: CoreWeave or Together GPU clusters.
* A managed fine-tuning API instead of your own trainer: Together Fine-Tuning or Baseten Training.
* Spending cloud credits: SageMaker with managed spot training and S3 checkpoints.

## FAQ

**What is the cheapest Modal alternative for GPU training?** On listed H100 rates as of October 9, 2026, the Nodus pricing page lists $2.60 per GPU-hour and Northflank lists $2.74, against $3.95 on Modal. CoreWeave spot is lower per GPU, at about $2.45, but it is sold as interruptible 8-GPU nodes.

**Can I move Modal code to Nodus without rewriting it?** Mostly. Change `import modal` to `import nodus` and keep `App`, `@app.function`, `.remote()`, `.map()`, `Image`, `Volume` and `Secret`. Check the gaps above first.

**Does Modal support multi-node training?** Modal supports up to 8 GPUs per container, and multi-node training is in private beta by request. On Nodus, multi-node training is Beta, open to every organization without an access request, and capped at 8 nodes per gang.

**What happens to a training run when a GPU is preempted?** On Modal the Function restarts on the same input, so you need your own checkpoints in a Volume. On Nodus the files under `/nodus/state` are restored on the replacement machine, and you lose at most the work since the last committed checkpoint.

**Can I try these for free?** Modal’s Starter plan includes $30 a month in credits. A new Nodus org gets $30 of starter credit once your email is verified, valid for 30 days; multi-node runs unlock after the first top-up ([credits](https://www.nodus-compute.ai/docs/guides/billing/credits/)).

**Are these on-demand prices?** They are list prices read from each pricing page on October 9, 2026. The Nodus number is the hourly supply price on its pricing page, and `nodus run --dry-run` prints the expected cost of your exact run before anything starts.
