Half Your Cluster Is Missing


Everyone who has priced an AI cluster has done this arithmetic. Take the peak, apply an availability number, apply an efficiency number, apply a scheduling number: ninety percent, fifty percent, ninety-five percent, call it 43 percent of peak and move on. Fractions do multiply, but only when they are correctly nested: each one measured against the output of the previous stage, on the same machine, over the same window. The figures people usually reach for are not. They come from different machines, definitions, denominators, and windows, and the loss budgets behind them overlap. The product of unnested fractions is not a conservative estimate. It is not an estimate at all.

The missing capacity is not empty. It lives in three coupled budgets. The first is occupancy: whether the GPUs have work at all. The second is step efficiency, measured as Model FLOPs utilization (MFU): how much of the peak arithmetic becomes model FLOPs once bubbles, stragglers, and throttled clocks take their share. The third is accepted progress: whether completed steps survive validation, checkpointing, and rollback. Improving one budget can cost another. Slack is not always waste. Sometimes it stops local variance from becoming global delay.

MFU loss attribution is therefore a causal graph, not a pie chart. A degrading or flapping optical link produces retries; one rank arrives late to a collective; its peers wait; the delay opens pipeline bubbles across healthy GPUs. A synchronized compute phase produces a power transient; thermal dynamics trigger controls that trim clocks; the slow nodes become the next stragglers. Fill a bubble with another job and fleet occupancy rises, but shared-fabric traffic can lower the main job’s MFU. The same lost millisecond appears as link health, congestion, collective time, straggler time, bubble time, or throttling, depending on which dashboard owns it. Add those categories and the same event is counted repeatedly.

The numbers show the scale and limits of the evidence. MegaScale reported 55.2 percent MFU on 12,288 GPUs, leaving nearly half the peak arithmetic unrealized. A current contrast is less complete. The Information reported in May 2026 that an internal xAI memo put MFU near 11 percent; president Michael Nicolls called it “embarrassingly low” and set a 50 percent target. MFU divides model FLOPs by theoretical peak, not idle GPUs. The report omits the workload, number format, cluster scope, and measurement window, so 11 percent cannot be cleanly compared with 55.2. It shows xAI considered the gap serious, not where the compute went.

Other budgets differ. Meta’s Llama 3 run recorded 466 interruptions across 54 days on 16,384 H100s, about nine a day, most of them hardware, and kept effective training time above ninety percent by automating detection and recovery. A follow-on reliability analysis models job mean time to failure falling roughly in inverse proportion to accelerator count: 1.8 hours at 16,384 GPUs, about fourteen minutes at 131,072. Under that projection, an interruption every quarter hour is the design condition at that fleet size.

Frontier is the control case. In 2024 OLCF reported 98.49 percent overall availability and 86.66 percent system utilization, where utilization means node-hours running user jobs. Frontier’s 1.353 exaflop HPL result is 65.8 percent of FP64 peak, while an earlier configuration’s HPCG result was about 0.8 percent of its then-current peak. None conflict: they measure availability, occupancy, and arithmetic efficiency for different workloads. Put Frontier’s 86.66 percent occupancy beside xAI’s 11 percent MFU and the opening error reappears. The clean aggregate is accepted output per unit time divided by peak capacity over the same window, with workload, precision, interruptions, and checkpoint policy attached.

Some missing capacity is insurance, and measurement without actuation is just another dashboard. The engineering answer is a closed loop: a training library emits phase, bubbles, collective exposure, convergence, and checkpoint state; a fleet orchestrator joins them to clocks, thermals, link health, congestion, and co-tenancy, then reshapes a pipeline, changes precision, moves a job, isolates a straggler, drains a link, or adjusts power and checkpoint policy. Pollux demonstrated this library-plus-Kubernetes pattern and cut completion times by 37 to 50 percent.

The objective is accepted progress per accelerator-dollar and megawatt-hour. For AI clouds and API providers with enough demand, more accepted output on a fixed fleet lowers unit cost and flows into gross margin; elsewhere it appears as unit cost and return on capital. At this capital scale, utilization is an income-statement variable.

The machine you have is the one you measure. Half of the other one is missing.

A companion to On Scaling Effective Compute, one of a series of short posts breaking a long map into teachable pieces.

Next in the series: On Naming the Quantity.


Authorship: AI-generated writing, directed by Jason Hoffman. Published as Fullhoffman AI Staff in AI-directed content.

Pangram 4.0: 0% human · 0% AI-assisted · 100% AI-generated
Last audited September 8, 2026, before this attribution update. The percentages describe the detector’s assessment of the text, including quotations and code. About the measurements.

One response to “Half Your Cluster Is Missing”

  1. […] A series of short companion posts breaks this long map into teachable pieces. It begins with Half Your Cluster Is Missing. […]

Discover more from Jason A. Hoffman

Subscribe now to keep reading and get access to the full archive.

Continue reading