How to make GPU training survive spot preemption (without babysitting it)
Save model, optimizer and step to one directory, write atomically, resume on start, and let the platform carry that directory across a reclaim.
Save model, optimizer and step to one directory, write atomically, resume on start, and let the platform carry that directory across a reclaim.