The Workful Life of a GPU


Every few weeks someone asks me a version of the same question: how long does a GPU actually last? Meta’s 2025 annual report gave it new fuel when it extended the estimated useful life of most of its server and network assets to five and a half years, cutting nearly three billion dollars from depreciation expense without changing a single machine. One reading of the number is a durability certificate: the servers must be lasting longer now. Another is an obsolescence confession: the schedule must be hiding faster turnover. Both read an accounting estimate as a measurement of the machines, and the coverage makes the same move, treating the schedule as the machine’s expiration date.

CNBC asked how long before a GPU depreciates and answered in shelf life, quoting a financing lawyer for whom three, five, or seven years of schedule decides how financeable the collateral is. Michael Burry called the extensions one of the more common frauds of the modern era and put 176 billion dollars of understated depreciation on 2026 to 2028, on the argument that a two-to-three-year product cycle is the machine’s life. A product cycle was treated as a lifetime, and the coverage mostly relayed the number. By August, lenders were requiring chip-backed loans to be repaid inside three to five years on the assumption that the collateral would retain little value after that, the Financial Times reported, and the conflation moved from commentary into pricing.

In my home office I have work running on a Titan XP (April 2017), an RTX 2080 (September 2018), a 3080 (September 2020), a 5080 (January 2025), and an RTX 6000 Pro (May 2025), and every one of them is doing something for me. The Titan lives in an external enclosure and gives a laptop its offload. Plenty of people still game on a 1080 Ti. None of these devices became slower because NVIDIA shipped another generation.

Fine has a precise meaning here: the machine still clears the threshold of the workload in front of it.

Then there is the Tesla V100. In May 2020, Lambda launched a roughly comparable eight-V100 instance at $4.40 per hour, or $0.55 per GPU-hour. Lambda now lists its eight-GPU V100 tier at $0.79 per GPU-hour, a 44 percent nominal increase. On August 16, 2026, an independent availability tracker marked both Lambda configurations out of stock.

The configurations are not identical, inflation matters, and a cloud instance is not a loose GPU. Still, this is not the price path suggested by the phrase “rapidly depreciating hardware.” How can an older accelerator become more expensive to rent while newer generations move the frontier far beyond it?

The V100 did not get younger

Start with what the posted price is the price of. Lambda’s listed figure is the price of a configured rental service, not the resale value of one card, and an out-of-stock posted price is not necessarily a price that balances demand with available supply. The public evidence cannot say whether Lambda’s V100 fleet is fully occupied, partly unavailable, being retired, or simply not offered where the tracker looked. Someone in a position to know tells me it is the first one: the fleet is occupied. The eight-year-old card is out of stock because it is full. The mechanism is economically ordinary either way: the supply of a legacy generation can stay fixed or shrink while demand persists among the customers whose workloads fit it, and the price rises. Here it is not only an inference. It is confirmed.

The V100 still has concrete capabilities. Lambda’s configuration supplies 16 GB of stacked high-bandwidth memory per GPU. NVIDIA specifies 900 GB/s of memory bandwidth and 7.8 trillion 64-bit floating-point operations per second. Its Tensor Cores accelerate reduced-precision matrix math, and NVLink connects GPUs at up to 300 GB/s bidirectionally. Those characteristics remain useful for scientific computing, established training pipelines, inference, and other work that fits the memory and software envelope. They do not make it the right answer for frontier training.

The software boundary is already moving. NVIDIA’s CUDA 13 release notes say the current version of its GPU software toolkit removed offline compilation and library support for Volta, while CUDA 12.x remains the supported toolchain path. The silicon did not degrade when that boundary moved. Its feasible workload set did.

The answer is that the phrase combines several quantities that do not move together, and two of them deserve separate names. A GPU has a useful life, which is an accounting estimate: chosen, audited, and concerned only with the period over which its cost is spread. And it has a workful life, the time until it stops clearing workloads, which is a property of the machine and the workload together and is measured continuously. Workful is an archaic English word for hard-working, due for a second career. The useful life belongs to the filing. The workful life belongs to the machine.

Two words, one spectrum

At least five different quantities travel under the name depreciation, and they spread along a spectrum from the filing to the machine to the market.

Book value is an accounting quantity. The purchase cost is allocated across an estimated useful life. The schedule can reach zero while the machine continues working and earning revenue.

Resale value is a market quantity. It reflects the buyers, sellers, substitutes, transaction costs, and supply available at that moment. It can fall, flatten, or rise.

Rental earning power belongs to the service around a device. The price includes the surrounding processors, memory, storage, network, software, control plane, support, location, and terms. It also reflects whether capacity is available when the customer asks for it.

Physical condition is a reliability quantity, and it follows the bathtub curve rather than a straight line. At the far end, thermal cycling, cooling wear, memory errors, solder fatigue, and component aging can raise failure or throttling risk. A new product announcement does not. The near end is not aging at all. It is manufacturing quality: dead-on-arrival units, marginal components, and early-life failures that surface in the first weeks of a new training run. An operator watching a new cluster shed GPUs can call that depreciation, but infant mortality arrives before any aging has happened, and screening rather than age determines how much of it reaches the floor. Aging appears as a change in error rate, availability, or achievable clock, not as an automatic annual subtraction from useful throughput.

Workload relevance is an output quantity. It asks what accepted work the complete system can produce, by the required deadline, at what operating cost. Software support belongs here too, because a healthy device can become harder to use when current toolchains stop targeting it.

The value function is therefore not V(GPU, age). It is V(GPU, workload, system, software, market, time). Age is one input, and often not the binding one.

Of the five, the first is the useful life, the accounting quantity. Two are the workful life: the machine’s condition and its fit to work. Two more are the market pricing workful capacity: what a buyer pays to own it, what a renter pays to use it.

The mechanics of workful life

New accelerators do not make old accelerators disappear. They change what the new old ones are for.

Frontier training values the newest memory formats, the largest tightly interconnected GPU domains, the shortest completion time, and the most output per megawatt. A generation that loses badly on those dimensions can be uneconomic in that particular fleet even if someone gives the GPUs away. Rack space, power, cooling, networking, maintenance, and engineering attention are scarce too.

The constraint also runs in reverse. An older GPU that fits an existing datacenter can remain economic there. At that density, power is not the dominant cost, and nothing pushes it up, while the newest systems will not fit the building’s envelope at all. The facility’s power and cooling budget becomes part of the hardware’s workful life.

The same hardware can then clear a different market: smaller training, fine-tuning, batch inference, simulation, visualization, gaming, or local experimentation. There is no universal ordering. A V100’s FP64 and HBM may make it more useful than a newer consumer card for one scientific workload, while the consumer card wins another task on software support, graphics features, or performance per watt.

The personal replacement decision makes the distinction obvious. The purchase price of my RTX 2080 is sunk, although its resale value remains an opportunity cost. If it already clears my workload, a replacement must justify its entire new price with an incremental gain I can use. A frontier lab asks a different question because an hour of delayed training, a megawatt of power, or a rack of scarce capacity can be worth more than the hardware.

Workful life, in the technologist’s sense, is not a date set at purchase. It is the time until the machine loses its last economically fit workload, which is a property of the machine and the workload together, not of the machine alone. It ends on different dates for the physical condition, the workload fit, and the software support. A V100 can be finished for frontier training, fine for fine-tuning, and useful on a desk at the same age.

“Obsolete” is therefore an incomplete sentence. Obsolete for which workload, inside which system, against which constraint?

The accounting of useful life

The construct exists to solve a signal problem. Expensing a fleet on the day it arrives would show a catastrophic loss in one quarter and false pure profit for years after, even though the machines work the whole time, and the income statement would be tracking the delivery schedule instead of the period’s business. Spreading the cost across the periods the asset serves keeps the statement measuring the period’s economics, and a standard rule keeps one company readable against another. Depreciation is a smoothing rule for a scoreboard, not a sensor on the machine. The sensor reading exists: the cash flow statement, where the capital expense lands once, on the day it clears.

Accounting depreciation matters, but it should not be mistaken for any of the other four quantities. Meta made the separation visible in its 2025 annual report, and it was not alone: Microsoft and Alphabet had lengthened their own schedules in earlier years. Meta extended the estimated useful lives of most server and network assets to 5.5 years, a judgment that the equipment will keep serving longer than the old schedule assumed. That judgment reduced 2025 depreciation expense by $2.92 billion and increased net income by $2.59 billion. The installed machines did not become more reliable, faster, or more valuable in a secondary market on the day the estimate changed. Their accounting cost moved across reporting periods.

The industry shows that the estimate is chosen rather than measured: Meta extended to 5.5 years while Amazon shortened a subset of its servers to five years, citing the increased pace of AI technology development, and Amazon itself had extended those assets to six years only one year earlier. Same machines, same words, opposite numbers. A few kept the quantities apart: one accounting analyst wrote that depreciation is a process of allocation, not of valuation, and that Meta and Amazon can both be reasonable. That is the whole argument in one sentence.

The estimate also rests on a physical assumption. A useful-life schedule assumes a failure curve, and the front of that curve is shaped in a factory the analyst cannot see. Burn-in is the one place aging is accelerated on purpose: the manufacturer runs the first stretch of the reliability bathtub curve at high stress so that weak units fail in the factory, as yield loss, rather than at the customer, as downtime. The hours are a policy choice, and their cost is in the price of the GPU. More hours cost the factory time and yield. Fewer ship the infant mortality into the customer’s first training run. The schedule smooths all of it into one number.

For a fleet owner, old hardware can carry gross margin or destroy it. The profit lives in the gap between the useful life and the workful life: the schedule finishes early, the machine finishes late. A substantially depreciated GPU that stays rented at an adequate price produces cash with no book cost attached, even after power, facility, network, and support costs. The same card destroys value if its performance per watt, failure burden, stranded rack space, or software friction costs more than the work it completes. A posted price and an out-of-stock label reveal neither result, and reported gross margin depends on where depreciation and facility costs are classified, so comparisons across operators require the same naming discipline.

The actionable measure is cohort economics by workload. For each hardware generation and workload, take realized price times occupied hours, then subtract power, facility, network, support, failure recovery, orchestration, and the opportunity cost of the capacity consumed. Keep accounting depreciation visible, but keep it separate. Then track where each generation can still produce accepted output economically.

The only public number is an estimate

Almost everything that would measure a fleet is private. Failure and replacement rates, burn-in yields, occupied hours, realized prices, energy per unit of accepted output, and the retirement dates of specific cohorts sit inside operators’ systems in irregular formats, unaudited and unpublished. The one quantity about those machines that is public, standardized, comparable across companies, and attested by an auditor is an accounting estimate: useful life, and the depreciation line that follows from it. That asymmetry, rather than carelessness, produces the misreading. “Useful life: 5.5 years” has units, precision, and a physical noun, and management may explain a change by citing durability or the pace of technology. It is the only normalized public number anywhere near the machines, so it gets asked questions it was not built to answer. The term finishes the job: “useful life” sounds like an expiration date, while the number allocates historical cost across reporting periods. Treating accounting language as technical language is the predictable result, not a lapse in rigor.

The incentives widen that gap. Accounting needs consistent policies, period recognition, materiality, and controlled evidence, while management also knows that the estimate moves reported earnings. Audit asks whether the estimate and the process behind it are reasonable and supportable. Analysts are rewarded for extracting a view of the underlying economics from sparse, comparable disclosures. SRE teams care about preventing and recovering from the next failure. Infrastructure teams care about capacity, accepted output, and cost. Product teams care about latency, reliability, and customer value. No one has to be irrational for the outputs not to meet. They answer different questions on different timescales.

The boundary therefore has to be crossed in both directions, and the crossing is different each way. From outside, an estimate has to be turned back into propositions that can be tested against machines and workloads. From inside, telemetry has to be turned into evidence a controller and an auditor can rely on. Neither conversion happens on its own, and neither side can perform the other’s.

Reading from outside, the work is to keep every inference attached to a test. Start with what the accounting estimate formally says, then write the stronger technical proposition separately. If a longer useful life is taken to mean greater reliability, seek failure and replacement evidence. If it means slower obsolescence, inspect impairments, early retirements, resale values, software support, and workload migration. If it means better economics, look for realized prices, occupied hours, operating cost by cohort, and cash capex. Match the cohorts, denominators, and windows. Where the evidence is private, stop at uncertainty rather than promoting the estimate into a technical fact.

Reporting from inside, the work is to turn telemetry into controlled evidence a decision can rest on, without dropping its qualifications. Define the cohort by hardware generation, configuration, workload, and time window. Report what was installed, serviceable, schedulable, occupied, and retired; the accepted output it produced; the power it consumed; and its failure, recovery, and support costs. Attach the denominator, source, and owner. Reconcile the physical inventory to the fixed-asset register, and preserve the path from event logs to the summary. Then name the judgment the evidence bears on: useful life, residual value, impairment, deployment, or gross margin. An auditor does not need every sensor. An auditor does need a controlled path back to the evidence.

What both conversions need is not one number. It is a record that attaches every number to its decision, quantity, cohort and workload, denominator, window, evidence, and owner. One direction runs from estimate to machine through those links, the other from machine to material financial claim through the same ones. From outside, the question is what would have to be true in the fleet for an inference to hold. From inside, the answer states what is true, under which definition, and which decision it supports. Both sides can then see where judgment enters, which facts support it, and what would prove it wrong. What they end up sharing is a traceable relationship among several quantities, not agreement on one.

When trillions of dollars of infrastructure are being financed on these assumptions, the conversion is material to capital allocation, gross margin, and valuation. The rule underneath is simple: never let one quantity answer a question that belongs to another.

A GPU can be fully depreciated on the books, more expensive to rent, physically healthy, outside the frontier, and exactly right for another job. There is no contradiction.

Useful life on the books. Workful life in the machine.

A watch post in the On Scaling Effective Compute series, extending On Naming the Quantity and A GPU Hour Is Not a Barrel of Oil.


Sources and further reading


Authorship: AI-generated writing, directed by Jason Hoffman. Published as Fullhoffman AI Staff in AI-directed content.

Pangram 4.0: 4.18% human · 0% AI-assisted · 95.82% AI-generated
Last audited September 8, 2026, before this attribution update. The percentages describe the detector’s assessment of the text, including quotations and code. About the measurements.

Discover more from Jason A. Hoffman

Subscribe now to keep reading and get access to the full archive.

Continue reading