A GPU Has More Than One Depreciation Curve

Jason A. Hoffman, PhD | August 16, 2026


On my desk is a desktop with an RTX 2080 that is fine. A Titan in an external enclosure still gives a laptop the offload I ask of it. Plenty of people still game on a 1080 Ti. None of these devices became slower because NVIDIA shipped another generation.

Fine, here, has a precise meaning: the machine still clears the threshold of the workload in front of it.

Then there is the Tesla V100. In May 2020, Lambda launched a roughly comparable eight-V100 instance at $4.40 per hour, or $0.55 per GPU-hour. Lambda now lists its eight-GPU V100 tier at $0.79 per GPU-hour, a 44 percent nominal increase. On August 16, 2026, an independent availability tracker marked both Lambda configurations out of stock.

The configurations are not identical, inflation matters, and a cloud instance is not a loose GPU. Still, this is not the price path suggested by the phrase “rapidly depreciating hardware.” How can an older accelerator become more expensive to rent while newer generations move the frontier far beyond it?

The answer is that the phrase combines several quantities that do not share a curve.

One word, five clocks

At least five different things travel under the name depreciation.

Book value is an accounting quantity. The purchase cost is allocated across an estimated useful life. The schedule can reach zero while the machine continues working and earning revenue.

Resale value is a market quantity. It reflects the buyers, sellers, substitutes, transaction costs, and supply available at that moment. It can fall, flatten, or rise.

Rental earning power belongs to a service, not just a device. The price includes the surrounding processors, memory, storage, network, software, control plane, support, location, and terms. It also reflects whether capacity is available when the customer asks for it.

Physical condition is a reliability quantity. Thermal cycling, cooling wear, memory errors, solder fatigue, and component aging can raise failure or throttling risk. A new product announcement does not. The failures are not all aging, either. New hardware has an early-life failure mode of its own: dead-on-arrival units, marginal components, and burn-in failures that surface in the first weeks of a new training run, which an operator experiences as the machine that depreciated immediately. That is infant mortality, a manufacturing and screening quantity, not aging and not depreciation. Physical aging usually appears as a change in error rate, availability, or achievable clock, not as an automatic annual subtraction from useful throughput.

Workload relevance is an output quantity. It asks what accepted work the complete system can produce, by the required deadline, at what operating cost. Software support belongs here too, because a healthy device can become harder to use when current toolchains stop targeting it.

The value function is therefore not V(GPU, age). It is V(GPU, workload, system, software, market, time). Age is one input, and often not the binding one.

The V100 did not get younger

Lambda’s posted price is the price of a configured rental service. It is not the resale value of one V100, and an out-of-stock posted price is not necessarily a price that balances demand with available supply. The public evidence cannot tell us whether Lambda’s V100 fleet is fully occupied, partly unavailable, being retired, or simply not offered in the regions where the tracker looked.

It does show why assuming that price only falls with age is a mistake. The available supply of a legacy generation can remain fixed or shrink while demand persists among customers whose workloads fit it. That mechanism is an inference, not a Lambda disclosure, but it is economically ordinary.

The V100 still has concrete capabilities. Lambda’s configuration supplies 16 GB of stacked high-bandwidth memory per GPU. NVIDIA specifies 900 GB/s of memory bandwidth and 7.8 trillion 64-bit floating-point operations per second. Its Tensor Cores accelerate reduced-precision matrix math, and NVLink connects GPUs at up to 300 GB/s bidirectionally. Those characteristics remain useful for scientific computing, established training pipelines, inference, and other work that fits the memory and software envelope. They do not make V100 the right answer for frontier training, where time to solution, memory capacity, number formats, energy, topology, and fleet scale can overwhelm a low acquisition price.

Its software clock is already moving. NVIDIA’s CUDA 13 release notes say the current version of its GPU software toolkit removed offline compilation and library support for Volta, while CUDA 12.x remains the supported toolchain path. The silicon did not degrade when that boundary moved. Its feasible workload set did.

Hardware moves through workloads

New accelerators do not make old accelerators disappear. They change the allocation problem.

Frontier training values the newest memory formats, the largest tightly interconnected GPU domains, the shortest completion time, and the most output per megawatt. A generation that loses badly on those dimensions can be uneconomic in that particular fleet even if someone gives the GPUs away. Rack space, power, cooling, networking, maintenance, and engineering attention are scarce too.

The same hardware can then clear a different market: smaller training, fine-tuning, batch inference, simulation, visualization, gaming, or local experimentation. This is not a universal ladder. A V100’s FP64 and HBM may make it more useful than a newer consumer card for one scientific workload, while the consumer card wins another task on software support, graphics features, or performance per watt.

The personal replacement decision makes the distinction obvious. The purchase price of my RTX 2080 is sunk, although its resale value remains an opportunity cost. If it already clears my workload, a replacement must justify its entire new price with an incremental gain I can use. A frontier lab asks a different question because an hour of delayed training, a megawatt of power, or a rack of scarce capacity can be worth more than the hardware.

“Obsolete” is therefore an incomplete sentence. Obsolete for which workload, inside which system, against which constraint?

The income statement has its own curve

Accounting depreciation matters, but it should not be mistaken for any of the other four clocks. Meta made the separation visible in its 2025 annual report. It extended the estimated useful lives of most server and network assets to 5.5 years. That change reduced 2025 depreciation expense by $2.92 billion and increased net income by $2.59 billion. The installed machines did not become more reliable, faster, or more valuable in a secondary market on the day the estimate changed. Their accounting cost moved across reporting periods.

A change like this attracts two technical readings, and both make the same error. The bullish reading hears a durability certificate: servers must genuinely last longer now. The bearish reading hears an obsolescence confession: the schedule must be hiding faster turnover. Each treats a chosen estimate as a measured quantity, and the mistake is structural rather than stupid. Analysts measure what they are given instruments to measure, and the filing is the only normalized instrument on offer; failure logs, burn-in yields, and cohort economics are private. A figure with units and a decimal carries the form of a measurement, and physical machines make it feel more measured still. But the estimate is chosen, and the industry demonstrates as much: Meta extended to 5.5 years while Amazon shortened a subset of its servers to five years, citing the increased pace of AI technology development, and Amazon itself had extended those assets to six years only one year earlier. Same machines, opposite schedules.

The estimate hides a physical hinge as well. A useful-life schedule assumes a failure curve, and the front of that curve is shaped in a factory the analyst cannot see. Burn-in is the one place aging is deliberately accelerated: the manufacturer runs the first stretch of the bathtub at high stress so weak units fail on its floor, in yield loss, rather than on the customer’s floor, in downtime. It is depreciation prepaid by the manufacturer, and it is a policy choice. More hours cost the factory time and yield; fewer ship the infant mortality into the customer’s first training run. The schedule smooths all of it into one number.

For a fleet owner, old hardware can support gross margin or become a margin trap. A substantially depreciated GPU that stays rented at an adequate price may produce attractive cash contribution, even though power, facility, network, and support costs remain. The same GPU can destroy value if its low performance per watt, failure burden, stranded rack space, or software friction costs more than the work it completes. A posted rental price and an out-of-stock label reveal neither result. Reported gross margin also depends on where depreciation and facility costs are classified, so comparisons across operators require the same naming discipline.

The actionable measure is cohort economics by workload. For each hardware generation and workload, take realized price times occupied hours, then subtract power, facility, network, support, failure recovery, orchestration, and the opportunity cost of the capacity consumed. Keep accounting depreciation visible, but keep it separate. Then track where each generation can still produce accepted output economically. Auditing a schedule means reading every independent clock against it: resale prices and rental realizations, interruption logs and failure rates where they leak, and the cash statement, the one instrument every analyst already holds that carries no estimate. Capex is a fact on the day it clears. When the accounting schedule diverges from every observable clock, that is the finding.

This matters when trillions of dollars of infrastructure are being financed against assumptions about useful life and residual value. One straight-line schedule cannot describe the economic life of a heterogeneous fleet serving a changing distribution of demand. Neither can one benchmark rental price.

On Naming the Quantity supplies the discipline: ask which quantity is moving. A GPU Hour Is Not a Barrel of Oil supplies the market consequence: different configurations and workloads do not collapse into one fungible unit. The underlying fleet model comes from On Scaling Effective Compute: generations coexist, demand is a distribution, and value emerges from feasible machine-workload pairings rather than age.

A GPU can be fully depreciated on the books, more expensive to rent, physically healthy, outside the frontier, and exactly right for another job. There is no contradiction.

One birthday. More than one depreciation curve.


Sources and further reading

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from Jason A. Hoffman

Subscribe now to keep reading and get access to the full archive.

Continue reading