-
The Half-Life of AI Engineering
Applied AI has acquired a strange release rhythm. A model fails in a conspicuous way. Builders find a workaround. Someone appends “engineering” to the problem, and the market announces a new discipline that every serious team is expected to adopt. Then a model or product release absorbs the generic workaround or moves the bottleneck, and…
-
Not Every AI Query Becomes a Datacenter Megawatt
AI demand forecasts often pass through one hidden conversion. More adoption produces more queries; more queries require more centralized inference; more centralized inference requires more datacenter power. The first step can be true while the last two weaken. A new Stanford paper, Intelligence per Watt: Measuring Intelligence Efficiency of Local AI, measures how much single-turn…
-
The Workful Life of a GPU
Every few weeks someone asks me a version of the same question: how long does a GPU actually last? Meta’s 2025 annual report gave it new fuel when it extended the estimated useful life of most of its server and network assets to five and a half years, cutting nearly three billion dollars from depreciation…
-
A GPU Hour Is Not a Barrel of Oil
Compute is the new oil again. Before that it was the new electricity. Before that, cloud computing was supposed to become a commodity traded on neutral exchanges. The latest version is more concrete. CME Group and Silicon Data plan to launch two cash-settled futures contracts on October 5, pending regulatory review, tied to benchmark rental…
-
On Naming the Quantity
Someone tells you a number: a hundred petaFLOP/s, a million tokens, ninety percent utilized. The first question is never whether the number is right. It is which quantity the number describes. Most disagreements about computing infrastructure are people quoting different quantities at each other, and the confusion survives because the units all sound alike. The…
-
The Scientific Evidence for Preferring Coke Zero, and What Preferring Diet Coke Says About You
You may have this same division among your friends and family. At my table it runs between Diet Coke and Coke Zero: a friend drinks the first, I drink the second, and both of us believe the science is on our side. So I went and read the science. A disclosure before the findings: I…
-
Half Your Cluster Is Missing
Everyone who has priced an AI cluster has done this arithmetic. Take the peak, apply an availability number, apply an efficiency number, apply a scheduling number: ninety percent, fifty percent, ninety-five percent, call it 43 percent of peak and move on. Fractions do multiply, but only when they are correctly nested: each one measured against…
-
On Scaling Effective Compute
A series of short companion posts breaks this long map into teachable pieces. It begins with Half Your Cluster Is Missing and continues with On Naming the Quantity. The industry quotes compute in units of intention. Gigawatts announced, accelerators shipped, capital committed, peak FLOP/s on the spec sheet. Contracts are signed in those units, market…
-
On Open Weights
The previous post, On Binary Narratives, used the coverage of Moonshot AI’s Kimi K3 release to show how a story about many interacting variables gets compressed into one variable with two values. That post was about the method. This one applies it to the second binary inside the same story: open versus closed. The sharpest…
-
On Binary Narratives
On Friday the Chinese startup Moonshot AI introduced Kimi K3, which it described as the world’s largest open-source AI model. The model was available through Moonshot’s products and API; the full weights were promised for release by July 27, with the technical report still to come. The New York Times covered it under the headline…