Berliner Boersenzeitung - Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

EUR -
AED 4.234604
AFN 75.453452
ALL 94.034984
AMD 422.744908
AOA 1058.505761
ARS 1716.64106
AUD 1.640055
AWG 2.076944
AZN 1.961724
BAM 1.965914
BBD 2.322629
BDT 142.402587
BHD 0.434856
BIF 3424.416844
BMD 1.153057
BND 1.485643
BOB 13.57927
BRL 5.865197
BSD 1.153223
BTN 110.282214
BWP 15.758273
BYN 3.37456
BYR 22599.91481
BZD 2.319433
CAD 1.613888
CDF 2634.735473
CHF 0.927801
CLF 0.027223
CLP 1071.397431
CNY 7.801929
CNH 7.779865
COP 3609.437006
CRC 524.292666
CUC 1.153057
CUP 30.556007
CVE 110.835189
CZK 24.203584
DJF 204.92129
DKK 7.475002
DOP 67.115273
DZD 153.032092
EGP 58.886311
ERN 17.295853
ETB 186.139441
FJD 2.553387
FKP 0.867824
GBP 0.856093
GEL 3.020894
GGP 0.867824
GHS 13.475735
GIP 0.867824
GMD 85.32586
GNF 10123.077972
GTQ 8.799268
GYD 241.277525
HKD 9.042791
HNL 30.897954
HRK 7.536355
HTG 150.785714
HUF 362.296198
IDR 20804.662903
ILS 3.534161
IMP 0.867824
INR 110.128058
IQD 1510.67162
IRR 1585741.471806
ISK 142.61021
JEP 0.867824
JMD 182.388293
JOD 0.817455
JPY 183.340082
KES 149.193722
KGS 100.834792
KHR 4637.015895
KMF 495.814258
KRW 1638.217339
KWD 0.35694
KYD 0.961057
KZT 547.088208
LAK 26143.287174
LBP 103272.920228
LKR 387.36278
LRD 208.736556
LSL 19.122676
LTL 3.404677
LVL 0.697472
LYD 7.380971
MAD 10.820566
MDL 20.308742
MGA 4930.514595
MKD 61.838125
MMK 2421.23874
MNT 4146.650533
MOP 9.317857
MRU 46.105187
MUR 54.435819
MVR 17.82629
MWK 1999.739649
MXN 20.011649
MYR 4.715655
MZN 73.691559
NAD 19.122676
NGN 1571.293523
NIO 42.439618
NOK 10.986989
NPR 176.445189
NZD 1.96086
OMR 0.443349
PAB 1.153268
PEN 3.907804
PGK 5.088176
PHP 70.674371
PKR 320.274647
PLN 4.307255
PYG 6892.458113
QAR 4.203927
RON 5.245945
RSD 117.416965
RUB 91.935432
RWF 1697.555122
SAR 4.302823
SBD 9.306957
SCR 15.494452
SDG 692.390762
SEK 10.984539
SGD 1.477959
SLE 27.817454
SOS 659.084933
SRD 43.522085
STD 23865.949363
STN 24.626477
SVC 10.090912
SZL 19.127401
THB 38.468314
TJS 10.649795
TMT 4.04723
TND 3.409741
TRY 54.671156
TTD 7.827506
TWD 37.405738
TZS 3052.721517
UAH 51.449481
UGX 4319.220142
USD 1.153057
UYU 46.388645
UZS 13844.22132
VES 857.062681
VND 30304.06434
VUV 137.826364
WST 3.173816
XAF 659.372
XAG 0.019614
XAU 0.000281
XCD 3.116193
XCG 2.078456
XDR 0.820974
XOF 659.349008
XPF 119.331742
YER 274.802282
ZAR 19.018232
ZMK 10378.896351
ZMW 21.536448
ZWL 371.283844
  • RBGPF

    3.2100

    69.21

    +4.64%

  • CMSD

    0.0200

    22.06

    +0.09%

  • RYCEF

    1.1600

    19.63

    +5.91%

  • CMSC

    0.0800

    21.77

    +0.37%

  • BCE

    -0.6050

    21.725

    -2.78%

  • BCC

    -2.4700

    75.21

    -3.28%

  • RELX

    -1.8050

    36.425

    -4.96%

  • VOD

    -0.0250

    16.105

    -0.16%

  • NGG

    1.3800

    80.26

    +1.72%

  • AZN

    -1.8000

    171.5

    -1.05%

  • GSK

    -1.1250

    52.095

    -2.16%

  • JRI

    0.1220

    12.882

    +0.95%

  • BP

    0.7250

    44.045

    +1.65%

  • BTI

    -1.6050

    61.465

    -2.61%

  • RIO

    2.8900

    96.55

    +2.99%

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training
Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

New TorchPass solution addresses a multi-million dollar challenge with AI infrastructure; uses Live GPU Migration to keep large-scale AI training running through hardware failures instead of forcing costly restarts

Text size:

PALO ALTO, CA / ACCESS Newswire / March 11, 2026 / Clockwork.io, the leader in Software-Driven AI Fabrics- a programmable, vendor-neutral software layer that optimizes large-scale GPU clusters for real-time observability, fault tolerance, and deterministic performance-today announced the general availability of TorchPass Workload Fault Tolerance. This new class of software-driven fault-tolerance eliminates one of the most costly failure modes in large-scale AI training: catastrophic job restarts caused by infrastructure faults.

Delivered as a core capability of the Clockwork.io FleetIQ platform, TorchPass applies the principles of Software-Driven AI Fabrics to distributed training, using Live GPU Migration to allow workloads to continue running through GPU failures, network disruptions, driver bugs, and even full node crashes-without checkpoint restarts or lost progress.

"Companies are investing billions in next-gen chips, yet the costs of running distributed AI jobs remains grossly inflated because the ecosystem has accepted failure as a constant," said Suresh Vasudevan, CEO of Clockwork.io. "We built TorchPass to fundamentally reject that premise. Instead of treating failure as inevitable and restarting after the fact, TorchPass makes infrastructure faults invisible to the workload-training continues through failures transparently, in software. For a typical 2,048-GPU deployment, that translates into over $6 million a year in recovered compute. This is what our Software-Driven AI Fabric approach was designed to deliver: fault-tolerant AI infrastructure."

Dylan Patel, Founder and CEO of SemiAnalysis agreed that large-scale training jobs are limited by interruptions.

"As Blackwell clusters roll out with an NVL72 domain, and we look to the future with Rubin Ultra's NVL576 domain, the idea that a single GPU error or network link flap can take down an entire run is totally unacceptable," said Patel. "TorchPass solves a huge challenge with cluster reliability: it provides transparent failover and live workload migration that keeps MFU high, which in turn drives better GPU economics."

Why AI Training Fails at Scale

Distributed AI training remains one of the most failure-prone workloads in modern infrastructure. As cluster sizes grow, fragility increases sharply. Research from Meta FAIR shows that mean time to failure drops to 7.9 hours in a 1,024-GPU cluster and to just 1.8 hours at 16,384 GPUs. This means that for most large, AI-focused enterprises or AI clouds, failure-driven restarts are completely inevitable - making this a major barrier to scaling AI's impact.

Each failure forces training jobs to roll back to the most recent checkpoint, discarding minutes or hours of completed work and wasting additional time on manual intervention, reprovisioning resources and restarting training. These restarts silently cap GPU utilization, making reliability one of the largest hidden costs in AI infrastructure.

TorchPass addresses this problem by proactively addressing costly AI workload failures, solving them before the job stops or needs to restart. Vital for enterprises running large AI workloads and AI clouds alike, TorchPass dramatically improves the reliability of workloads and cluster utilization. For AI clouds, who can now address impacted GPUs while preserving the training run as planned, this translates into better customer SLAs and overall AI cloud economics, improving their ability to protect margin and deliver new models sooner.

"Managing compute output across large-scale GPU clusters is vital to ensuring we're delivering reliable capacity to our customers. By using TorchPass we have the support of a company that focuses on resilience like it is a core business function: it replaces any specific failing GPU and keeps the rest of the job moving, rather than making one small problem impact our large-scale operations," said David Power, CTO of Nscale. "In our evaluation, Live GPU Migration preserved both run continuity and throughput under real fault conditions, which is exactly what you need to deliver predictable time-to-train and a better customer experience at scale."

How Live GPU Migration Works: Reliability Without Restart

TorchPass performs transparent, in-flight migration of impacted training ranks to spare resources when failures occur. TorchPass typically completes recovery in approximately three minutes while the training process continues uninterrupted.

It supports resilience across three failure scenarios:

  • Unplanned migration, handling sudden events such as kernel crashes, power failures, or GPU faults by reconstructing state from healthy replicas

  • Pre-emptive migration, triggered by early warning signals such as rising temperatures or ECC memory errors, enabling controlled migration before a hard failure

  • Planned migration, enabling maintenance, patching, and workload rebalancing without interrupting training

This approach reduces wasted training progress by 95%, cutting lost time from approximately three hours per day to under ten minutes in a 1,024-GPU cluster.

Jordan Nanos, Member of Technical Staff and lead author of ClusterMAX-SemiAnalysis' independent benchmark for large-scale AI training-stress tested Clockwork.io TorchPass and found it delivered leading performance and efficiency for large-scale distributed training, enabling users to reduce checkpointing overhead in training. He shared the following results:

"In our testing, Clockwork.io TorchPass delivered the fastest and most efficient fault-tolerant performance for a gpt-oss-120B training run. We used TorchTitan on a Kubernetes cluster with 64x H200 GPUs. During our testing we measured job completion time (JCT) and Model FLOPs Utilization (MFU) against a standard approach (checkpoint-restart) and the leading open-source fault-tolerant training framework (TorchFT). We simulated multiple hardware failures on the cluster in order to stress test the fault-tolerant training frameworks.

When compared to checkpoint-restart, TorchPass was significantly faster to recover from failures. This reduced overall JCT and maintained high MFU. And when compared to TorchFT, TorchPass had a significantly higher MFU. This reduced overall JCT while also maintaining an equal time to recover from failures.

Using TorchPass also has a downstream effect where it provides users with an opportunity to reduce or even remove checkpointing from their training code. This means larger effective batch sizes, lower risk of out of memory errors (OOMs), and less time spent thinking about storage. For a research organization, this can ultimately mean a faster time to reach their training objective," concluded Nanos.

Measurable Business Impact from Software-Driven Fault-Tolerance

For customers operating large AI clusters, the impact is immediate and measurable. In a typical 2,048-GPU H200 deployment, TorchPass Workload Fault Tolerance delivers over $6 million in annual savings by preventing wasted compute.

These savings come from eliminating hundreds of thousands of GPU-hours that would otherwise be lost to failure-driven restarts, cascading retries, and idle recovery time. By keeping training jobs running through infrastructure faults instead of restarting them, TorchPass converts lost GPU time into productive training, significantly improving the return on GPU investments that today often operate at just 30-50% of theoretical performance.

Enabling the Next Generation of AI Infrastructure

By making reliability a software-defined capability rather than a hardware constraint, TorchPass provides the operational confidence required to deploy next-generation, tightly coupled systems such as NVIDIA GB200 and GB300 NVL72 and future rack-scale systems, where dense architectures amplify the cost of even small failures.

TorchPass builds on Clockwork.io's prior release of Network Fault Tolerance, which applies the same Software-Driven AI Fabric principles to network resilience by transparently rerouting traffic around link failures.

Together, these capabilities form Clockwork.io's Software-Driven AI Fabric, a vendor-neutral software layer spanning network, compute, and storage. As modern AI workloads run on tightly coupled clusters where hundreds or thousands of processors must operate in coordinated lockstep, infrastructure behaves as a single system, where reliability and performance directly determine overall efficiency. By managing this complexity in software, Clockwork.io enables operators to run heterogeneous AI infrastructure as a unified platform-maintaining high utilization, predictable performance, and resilience while preserving the flexibility to evolve hardware and improve the economics of large-scale AI deployments.

To learn more about the launch of TorchPass, visit the Clockwork.io team in-person at NVIDIA GTC from March 16-19, Booth #205, or visit https://clockwork.io.

About Clockwork.io
Clockwork.io pioneers Software-Driven AI Fabrics™, delivering a programmable software layer that makes large-scale AI clusters observable, deterministic, and resilient by design to drive continuous workload progress and peak cluster utilization. Its FleetIQ platform enables enterprises to train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost. Companies including Uber, Wells Fargo, DCAI, Nebius, Nscale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io.

Media Contact
Dana Trismen
[email protected]
650-269-7478

SOURCE: Clockwork



View the original press release on ACCESS Newswire

(F.Schuster--BBZ)