(01)

Solution

Training

Multi-week training runs that keep their place when a node fails, with checkpoints you can find again.

Training line drawing

0

runs lost to a node failure last quarter

Long runs

What it does

Overhead watches every node in a training job. When one drops, the scheduler pauses the run, moves the shard to a healthy node and resumes from the last checkpoint. You get a log entry, not a page at 3 a.m.

What you set

Checkpoint frequency, which pools a run may use, and the most it may spend. Everything else is handled for you.

Next step

Try it on your own cluster.

Read-only for 30 days, connected with one of our engineers.

Create a free website with Framer, the website builder loved by startups, designers and agencies.