<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:fh="http://purl.org/syndication/history/1.0"><channel><title>Nodus | Blog</title><description>Docs for Nodus: run Jobs, Sandboxes, Functions, Agents and training on the capacity that finishes them.</description><link>https://www.nodus-compute.ai/</link><language>en</language><fh:complete/><atom:link rel="self" href="https://www.nodus-compute.ai/blog/rss.xml"/><item><title>Cloud GPU pricing, October 2026: what an H100, H200, B200 or B300 costs per hour</title><link>https://www.nodus-compute.ai/blog/cloud-gpu-pricing-october-2026/</link><guid isPermaLink="true">https://www.nodus-compute.ai/blog/cloud-gpu-pricing-october-2026/</guid><description>Listed per GPU-hour prices for B300, B200, H200, H100, A100, L40S, L4 and RTX cards as of October 8, 2026, with a training-run cost example.</description><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Short answer (Oct 8, 2026):&lt;/strong&gt; the &lt;a href=&quot;https://www.nodus-compute.ai/pricing/&quot;&gt;Nodus pricing page&lt;/a&gt; lists an H100 SXM 80 GB at $2.60/hr, H200 SXM at $3.59/hr, B200 at $5.50/hr and B300 at $6.94/hr. Consumer and prosumer cards go much lower: an RTX 4090 is $0.29/hr and an RTX 3090 is $0.18/hr. Modal’s listed rates for the same data center GPUs run from 2% higher (B300) to about 2.5x higher (L40S) depending on the card.&lt;/p&gt;
&lt;p&gt;Prices below are per GPU-hour, USD, taken from the public Nodus pricing page on October 8, 2026, which lists Modal’s published rate beside each one.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;data-center-gpus&quot;&gt;Data center GPUs&lt;/h2&gt;&lt;/div&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;GPU&lt;/th&gt;
&lt;th&gt;VRAM&lt;/th&gt;
&lt;th&gt;Nodus&lt;/th&gt;
&lt;th&gt;Modal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;B300&lt;/td&gt;
&lt;td&gt;288 GB&lt;/td&gt;
&lt;td&gt;$6.943&lt;/td&gt;
&lt;td&gt;$7.099&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B200&lt;/td&gt;
&lt;td&gt;180 GB&lt;/td&gt;
&lt;td&gt;$5.500&lt;/td&gt;
&lt;td&gt;$6.250&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;H200 SXM&lt;/td&gt;
&lt;td&gt;141 GB&lt;/td&gt;
&lt;td&gt;$3.593&lt;/td&gt;
&lt;td&gt;$4.540&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;H100 SXM&lt;/td&gt;
&lt;td&gt;80 GB&lt;/td&gt;
&lt;td&gt;$2.600&lt;/td&gt;
&lt;td&gt;$3.949&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A100&lt;/td&gt;
&lt;td&gt;80 GB&lt;/td&gt;
&lt;td&gt;$1.090&lt;/td&gt;
&lt;td&gt;$2.498&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A100&lt;/td&gt;
&lt;td&gt;40 GB&lt;/td&gt;
&lt;td&gt;$0.904&lt;/td&gt;
&lt;td&gt;$2.099&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX PRO 6000&lt;/td&gt;
&lt;td&gt;96 GB&lt;/td&gt;
&lt;td&gt;$1.350&lt;/td&gt;
&lt;td&gt;$3.031&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L40S&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;$0.793&lt;/td&gt;
&lt;td&gt;$1.951&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L4&lt;/td&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;$0.430&lt;/td&gt;
&lt;td&gt;$0.799&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;div&gt;&lt;h2 id=&quot;budget-and-prosumer-gpus&quot;&gt;Budget and prosumer GPUs&lt;/h2&gt;&lt;/div&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;GPU&lt;/th&gt;
&lt;th&gt;VRAM&lt;/th&gt;
&lt;th&gt;Nodus&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RTX 5090&lt;/td&gt;
&lt;td&gt;32 GB&lt;/td&gt;
&lt;td&gt;$0.480&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX 4090&lt;/td&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;$0.290&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX 3090&lt;/td&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;$0.180&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX 6000 Ada&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;$0.743&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX A6000&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;$0.333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX A4000&lt;/td&gt;
&lt;td&gt;16 GB&lt;/td&gt;
&lt;td&gt;$0.173&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;div&gt;&lt;h2 id=&quot;which-one-should-you-rent&quot;&gt;Which one should you rent?&lt;/h2&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Fine-tuning a 7B to 14B model with LoRA:&lt;/strong&gt; a 48 GB card (L40S, A6000) or a 24 GB card with QLoRA. A100 80 GB if you want full fine-tunes of small models.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Full fine-tunes and pretraining at scale:&lt;/strong&gt; H100 or H200. H200’s 141 GB helps with long context and larger batch sizes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Serving frontier-size open models:&lt;/strong&gt; B200 or B300. B300’s 288 GB per GPU means fewer GPUs per replica, which often lowers cost per token even though the hourly rate is higher.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Experiments, evals and small jobs:&lt;/strong&gt; RTX 4090 / 5090 at under $0.50/hr are hard to beat.&lt;/li&gt;
&lt;/ul&gt;
&lt;div&gt;&lt;h2 id=&quot;what-the-hourly-rate-does-not-tell-you&quot;&gt;What the hourly rate does not tell you&lt;/h2&gt;&lt;/div&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Boot and teardown.&lt;/strong&gt; Some platforms bill from machine creation to deletion. On Nodus the bill is split into &lt;code dir=&quot;auto&quot;&gt;Boot&lt;/code&gt;, &lt;code dir=&quot;auto&quot;&gt;Running&lt;/code&gt; and &lt;code dir=&quot;auto&quot;&gt;Teardown&lt;/code&gt; segments so you can see it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Price changes mid-run.&lt;/strong&gt; Nodus freezes the rate when the machine is chosen for the whole run.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interruptions.&lt;/strong&gt; Spot capacity is cheaper, but only if your job checkpoints. Otherwise one preemption can wipe out the savings. See our &lt;a href=&quot;https://www.nodus-compute.ai/docs/guides/checkpoints/&quot;&gt;preemption guide&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reserved terms.&lt;/strong&gt; For multi-month commitments on B200/B300, reserved pricing is quoted separately and depends on term length, location and interconnect. Nodus handles these through &lt;a href=&quot;https://www.nodus-compute.ai/reserved-compute/&quot;&gt;Reserved Compute&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;div&gt;&lt;h2 id=&quot;faq&quot;&gt;FAQ&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;What is the cheapest H100 per hour in October 2026?&lt;/strong&gt; The Nodus pricing page lists an H100 SXM 80 GB at $2.60/hr as of Oct 8, 2026. Run &lt;code dir=&quot;auto&quot;&gt;nodus run --dry-run&lt;/code&gt; for the estimate for your exact job.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How much does a B300 cost per hour?&lt;/strong&gt; The Nodus pricing page lists a B300 at $6.94/hr as of Oct 8, 2026. Reserved terms are quoted per request.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is there a free tier?&lt;/strong&gt; New Nodus accounts get a $30 starter grant valid for 30 days.&lt;/p&gt;
&lt;p&gt;We refresh this table monthly. Live prices: &lt;a href=&quot;https://www.nodus-compute.ai/pricing/&quot;&gt;nodus-compute.ai/pricing&lt;/a&gt;&lt;/p&gt;
</content:encoded><category>pricing</category><category>h100</category><category>b200</category><category>b300</category><category>gpu</category></item><item><title>How to give Claude Code, Codex or Cursor access to GPUs (MCP setup)</title><link>https://www.nodus-compute.ai/blog/give-claude-code-codex-cursor-gpus-mcp/</link><guid isPermaLink="true">https://www.nodus-compute.ai/blog/give-claude-code-codex-cursor-gpus-mcp/</guid><description>Connect one MCP server and your coding agent can estimate, launch and watch GPU jobs on your account, with dry runs and spending limits.</description><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; connect your coding agent to a GPU platform’s MCP server. The agent can then estimate, launch, watch and fetch results from GPU jobs from the same chat where it wrote the code. With &lt;a href=&quot;https://www.nodus-compute.ai/&quot;&gt;Nodus&lt;/a&gt; that is one command per client, and every paid action shows a dry run and cost estimate before it runs.&lt;/p&gt;
&lt;p&gt;Coding agents are great at writing a training script and terrible at the part after: finding a GPU, shipping code to it, watching logs, copying results back. MCP closes that loop.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;claude-code&quot;&gt;Claude Code&lt;/h2&gt;&lt;/div&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;span&gt;&lt;/span&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;claude&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;mcp&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;add&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;--scope&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;user&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;--transport&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;http&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;nodus&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;https://api-next.nodus-compute.ai/mcp&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;Then run &lt;code dir=&quot;auto&quot;&gt;/mcp&lt;/code&gt;, pick &lt;code dir=&quot;auto&quot;&gt;nodus&lt;/code&gt;, and sign in in your browser.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;codex&quot;&gt;Codex&lt;/h2&gt;&lt;/div&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;span&gt;&lt;/span&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;codex&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;mcp&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;add&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;nodus&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;--url&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;https://api-next.nodus-compute.ai/mcp&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;codex&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;mcp&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;login&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;nodus&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;div&gt;&lt;h2 id=&quot;cursor&quot;&gt;Cursor&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Add this to &lt;code dir=&quot;auto&quot;&gt;~/.cursor/mcp.json&lt;/code&gt;:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;{&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;  &lt;/span&gt;&lt;span&gt;&quot;mcpServers&quot;&lt;/span&gt;&lt;span&gt;: {&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;&quot;nodus&quot;&lt;/span&gt;&lt;span&gt;: { &lt;/span&gt;&lt;span&gt;&quot;type&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;http&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;url&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;https://api-next.nodus-compute.ai/mcp&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt; }&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;  &lt;/span&gt;&lt;/span&gt;&lt;span&gt;}&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;}&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;div&gt;&lt;h2 id=&quot;any-other-mcp-client&quot;&gt;Any other MCP client&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Add a remote HTTP MCP server with URL &lt;code dir=&quot;auto&quot;&gt;https://api-next.nodus-compute.ai/mcp&lt;/code&gt; and follow the sign-in prompt. Prefer a local server? Install the CLI and run &lt;code dir=&quot;auto&quot;&gt;nodus mcp install&lt;/code&gt;, which configures the agents on your machine.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;verify-without-spending-money&quot;&gt;Verify without spending money&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Ask the agent: “List my Nodus jobs.” That is read only. A good first prompt for any agent that can read URLs:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;Read https://nodus-compute.ai/connect.md and help me connect Nodus to this agent.&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;Verify setup by listing my jobs. Do not start paid compute.&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;div&gt;&lt;h2 id=&quot;what-the-agent-can-actually-do&quot;&gt;What the agent can actually do&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;The server exposes generic tools over every resource: &lt;code dir=&quot;auto&quot;&gt;get&lt;/code&gt;, &lt;code dir=&quot;auto&quot;&gt;describe&lt;/code&gt;, &lt;code dir=&quot;auto&quot;&gt;logs&lt;/code&gt;, &lt;code dir=&quot;auto&quot;&gt;estimate&lt;/code&gt;, &lt;code dir=&quot;auto&quot;&gt;apply&lt;/code&gt;, &lt;code dir=&quot;auto&quot;&gt;exec&lt;/code&gt; and more. In practice that means prompts like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Run train.py on a 24 GB GPU with a $5 cap and stream the logs.”&lt;/li&gt;
&lt;li&gt;“Why did job/train fail?”&lt;/li&gt;
&lt;li&gt;“Fine-tune Qwen3 0.6B with LoRA on chats.jsonl and download the adapter.”&lt;/li&gt;
&lt;li&gt;“What did my sandboxes cost this week?”&lt;/li&gt;
&lt;/ul&gt;
&lt;div&gt;&lt;h2 id=&quot;guardrails-that-matter-when-an-agent-holds-your-credit-card&quot;&gt;Guardrails that matter when an agent holds your credit card&lt;/h2&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Writes are dry-run first.&lt;/strong&gt; The agent sees the object as it would be created plus the cost estimate, and it only runs once confirmed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Same limits as you.&lt;/strong&gt; The agent uses your account, projects, budgets and spending caps. A project budget stops it from overspending even if it gets creative.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hard caps per job.&lt;/strong&gt; &lt;code dir=&quot;auto&quot;&gt;maxCostUSD&lt;/code&gt; stops a job gracefully, with a checkpoint, before it passes the cap.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Frozen rates.&lt;/strong&gt; The hourly rate is fixed when the machine is chosen and holds for the whole run.&lt;/li&gt;
&lt;/ul&gt;
&lt;div&gt;&lt;h2 id=&quot;faq&quot;&gt;FAQ&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Do I need a separate “agents” product to use Claude Code with GPUs?&lt;/strong&gt; No. MCP is enough. Nodus Agents is a different feature for running agent conversations inside Nodus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Which GPU does the agent get?&lt;/strong&gt; By default the cheapest offering that fits the request and can start now, within your limits. You can ask for a specific type like H100 or B200.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What does it cost to try?&lt;/strong&gt; New accounts get a $30 starter grant, no card needed for the first runs.&lt;/p&gt;
&lt;p&gt;Setup for every client: &lt;a href=&quot;https://www.nodus-compute.ai/connect/&quot;&gt;nodus-compute.ai/connect&lt;/a&gt;&lt;/p&gt;
</content:encoded><category>mcp</category><category>claude code</category><category>codex</category><category>cursor</category><category>agents</category></item><item><title>How to make GPU training survive spot preemption (without babysitting it)</title><link>https://www.nodus-compute.ai/blog/resume-training-after-spot-preemption/</link><guid isPermaLink="true">https://www.nodus-compute.ai/blog/resume-training-after-spot-preemption/</guid><description>Save model, optimizer and step to one directory, write atomically, resume on start, and let the platform carry that directory across a reclaim.</description><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; save your model, optimizer and step counter to one directory, write those files atomically, load them on startup, and run on a platform that copies that directory before the machine disappears and restores it on the next machine. Do that and a preemption costs you minutes of work, not the whole run.&lt;/p&gt;
&lt;p&gt;Spot and interruptible GPUs are often a fraction of on-demand prices. The catch is that the provider can take the machine back. Here is the pattern we use at &lt;a href=&quot;https://www.nodus-compute.ai/&quot;&gt;Nodus&lt;/a&gt;, and it works anywhere.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;1-put-all-recovery-state-in-one-directory&quot;&gt;1. Put all recovery state in one directory&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Everything you need to continue goes in one place: weights, optimizer state, LR scheduler, RNG state, data loader position, current step. Results you want to download (final model, eval reports) go somewhere else. Mixing the two makes checkpoints big and slow.&lt;/p&gt;
&lt;p&gt;On Nodus that directory is &lt;code dir=&quot;auto&quot;&gt;/nodus/state&lt;/code&gt; (also in &lt;code dir=&quot;auto&quot;&gt;$NODUS_STATE_DIR&lt;/code&gt;), and outputs go to &lt;code dir=&quot;auto&quot;&gt;/nodus/outputs&lt;/code&gt;.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;2-write-atomically&quot;&gt;2. Write atomically&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;A checkpoint taken while you are halfway through writing a file restores a broken state. Write to a temp file, then rename it over the old one. Rename is atomic on POSIX filesystems.&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; json, os&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; pathlib &lt;/span&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; Path&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;
&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;STATE_DIR&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;Path&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;os.environ.&lt;/span&gt;&lt;span&gt;get&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;NODUS_STATE_DIR&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;/nodus/state&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;))&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;STATE&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;STATE_DIR&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;progress.json&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;
&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;def&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;save&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;step&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;model&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;opt&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;STATE_DIR&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;mkdir&lt;/span&gt;&lt;span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;parents&lt;/span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;True&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;exist_ok&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;True&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;tmp &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;STATE_DIR&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;ckpt.pt.tmp&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;torch.&lt;/span&gt;&lt;span&gt;save&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;{&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;model&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: model.&lt;/span&gt;&lt;span&gt;state_dict&lt;/span&gt;&lt;span&gt;()&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;opt&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: opt.&lt;/span&gt;&lt;span&gt;state_dict&lt;/span&gt;&lt;span&gt;()&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;step&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: step}&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; tmp&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;os.&lt;/span&gt;&lt;span&gt;replace&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;tmp&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; STATE_DIR &lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;ckpt.pt&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;div&gt;&lt;h2 id=&quot;3-resume-on-startup&quot;&gt;3. Resume on startup&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Checkpoints restore files, not process memory. Your program starts from the top, so it has to check for saved state:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;start &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;ckpt &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;STATE_DIR&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;ckpt.pt&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; ckpt.&lt;/span&gt;&lt;span&gt;exists&lt;/span&gt;&lt;span&gt;():&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;s &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; torch.&lt;/span&gt;&lt;span&gt;load&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;ckpt&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;model.&lt;/span&gt;&lt;span&gt;load_state_dict&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;s&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;model&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;]); opt.&lt;/span&gt;&lt;span&gt;load_state_dict&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;s&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;opt&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;]); start &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; s[&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;step&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;]&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;print&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;f&lt;/span&gt;&lt;span&gt;&quot;resumed from step &lt;/span&gt;&lt;span&gt;{start}&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;for&lt;/span&gt;&lt;span&gt; step &lt;/span&gt;&lt;span&gt;in&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;range&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;start &lt;/span&gt;&lt;span&gt;+&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; total_steps &lt;/span&gt;&lt;span&gt;+&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt;):&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;...&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;div&gt;&lt;h2 id=&quot;4-pick-a-checkpoint-cadence&quot;&gt;4. Pick a checkpoint cadence&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Too often and you waste GPU time writing files. Too rarely and each preemption throws away a lot of work. A practical rule: aim for at least four checkpoints per expected run, and keep checkpointing under about 10% of runtime. Capacity that is rarely interrupted can be checkpointed less often.&lt;/p&gt;
&lt;p&gt;Nodus does this math for you with &lt;code dir=&quot;auto&quot;&gt;interval: auto&lt;/code&gt;, using how often that capacity actually gets interrupted and how long your saves take.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;5-save-when-the-machine-is-about-to-go-away&quot;&gt;5. Save when the machine is about to go away&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Most providers give a short reclaim notice. Use it. On Nodus your program can subscribe to checkpoint requests over a local socket and acknowledge when files are consistent:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; nodus  &lt;/span&gt;&lt;span&gt;# pip install nodus-compute; no-ops outside Nodus&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;
&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;nodus.checkpoint.&lt;/span&gt;&lt;span&gt;on_request&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;lambda&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;save&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;step&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; model&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; opt&lt;/span&gt;&lt;span&gt;))&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;When the capacity gives notice, Nodus sends an urgent request, takes the checkpoint after your ack, and prepares a replacement machine at the same time. The replacement only starts once the old one is provably gone, so two machines never write the same state.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;6-hugging-face-trainer-users&quot;&gt;6. Hugging Face Trainer users&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Trainer already saves and resumes. Point &lt;code dir=&quot;auto&quot;&gt;output_dir&lt;/code&gt; at the state directory, set &lt;code dir=&quot;auto&quot;&gt;save_steps&lt;/code&gt; and &lt;code dir=&quot;auto&quot;&gt;save_total_limit=2&lt;/code&gt;, and call &lt;code dir=&quot;auto&quot;&gt;trainer.train(resume_from_checkpoint=True)&lt;/code&gt; when &lt;code dir=&quot;auto&quot;&gt;NODUS_RESTORED=1&lt;/code&gt;. Or set &lt;code dir=&quot;auto&quot;&gt;integration: HFTrainer&lt;/code&gt; and Nodus registers the save-on-request callback for you.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;running-it&quot;&gt;Running it&lt;/h2&gt;&lt;/div&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;span&gt;&lt;/span&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;pip&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;install&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;nodus-compute&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;nodus&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;login&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;nodus&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;run&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;--gpu&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;H100&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;--interruptible&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;--checkpoint&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;/nodus/state&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;-d&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;--&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;python&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;train.py&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;&lt;code dir=&quot;auto&quot;&gt;nodus describe job/&amp;#x3C;name&gt;&lt;/code&gt; shows each attempt, why it ended (for example &lt;code dir=&quot;auto&quot;&gt;Preempted&lt;/code&gt;) and the latest checkpoint. New accounts get a $30 starter grant, so you can test a preemption-safe run without a card.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;faq&quot;&gt;FAQ&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Does checkpointing restore GPU memory?&lt;/strong&gt; No. It restores files. Your code reloads its own state.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What if a checkpoint is empty?&lt;/strong&gt; An empty state directory never replaces an earlier useful checkpoint, so a crash before the first save cannot erase progress.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How many times will it retry?&lt;/strong&gt; By default up to 8 recoveries (3 for distributed jobs). Two attempts in a row with no progress stops the run with &lt;code dir=&quot;auto&quot;&gt;NoProgress&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Does this work for multi-node?&lt;/strong&gt; Yes, in beta. Each rank writes its shard with &lt;code dir=&quot;auto&quot;&gt;torch.distributed.checkpoint&lt;/code&gt; and rank 0 commits it.&lt;/p&gt;
&lt;p&gt;Full guide: &lt;a href=&quot;https://www.nodus-compute.ai/docs/guides/checkpoints/&quot;&gt;nodus-compute.ai/docs/guides/checkpoints&lt;/a&gt;&lt;/p&gt;
</content:encoded><category>checkpointing</category><category>spot instances</category><category>pytorch</category><category>training</category></item></channel></rss>