Clockwork raises $31M to stop AI training from wasting GPU hours
What's the deal? Clockwork Systems, whose software keeps AI workloads running through infrastructure failures, has raised $31 million in a late-stage round. The company also announced production deployments at LinkedIn and Together AIDealroom has a profile for this one. Try Dealroom →, plus wider adoption by WhiteFiber.
Why now? Large AI training jobs can span thousands of GPUs that must stay synchronised, and one failed GPU, dropped connection, or frozen server can stall everything. Meta reported unexpected interruptions roughly once every three hours over a 54-day period training Llama 3 on 16,384 GPUs.
What's the endgame? Clockwork sells fault-tolerance tools that platform teams deploy as a layer between hardware and workload. LinkPass reroutes traffic around a broken link so the job never sees the failure, while TorchPass shifts work off a failing GPU onto a healthy one to keep training rather than roll back.
The round backs two new TorchPass features, both announced today. Multi-node snapshots — an industry first, the company says — save an entire running job across all nodes without code changes. Fast, asynchronous application checkpoints run in the background to speed up reinforcement learning by passing updated model weights to inference replicas.
What's the problem it solves? The usual fix for a failure is reloading a checkpoint, a saved copy of progress. Recovery can take up to 90 minutes, idles healthy GPUs, and forces the job to repeat work — meaning customers pay for unused chips and repeated compute while models run slower.
The round drew e& capitalDealroom has a profile for this one. Try Dealroom →, NEA, Premji Invest, Seligman VenturesDealroom has a profile for this one. Try Dealroom →, and Wing Venture CapitalDealroom has a profile for this one. Try Dealroom →.
"Failures are inevitable at AI scale. You shouldn't lose hours of productive work to them," said Suresh Vasudevan, chief executive officer of Clockwork. "Fault tolerance is an effective throughput multiplier."
The signal: As AI clusters grow, keeping expensive GPUs doing useful work through failures is becoming a baseline infrastructure requirement rather than a nice-to-have. Clockwork's raise, with production users like LinkedIn already onboard, points to fault tolerance emerging as a distinct layer in the AI compute stack.
Read more: prnewswire.com
Image credit: Clockwork Systems