Compute Equilibrium

Gustav Svalander ยท

The price of an hour of compute is not set by what it costs to build a data center. It's set by what the least valuable job running in that data center is worth. That job is the one deciding the price, because it's the one that walks away first when the price rises.

Say it as a rule: in equilibrium, an hour of compute costs what the marginal AI task can capture. Not what the best task is worth. The best tasks earn far more than they pay, the same way the electricity keeping a hospital running is worth more than the electricity price. It's the last job in the queue that clears the market. Everything else earns a surplus on top.

That single line explains most of what looks strange about the compute market right now.

Why the marginal job sets the price

The mechanism is an arbitrage, and arbitrages are stubborn.

If an hour of compute can be turned into more captured value than it costs, someone buys more of it. Buyers keep arriving until the last one is barely breaking even. If an hour costs more than the last job can capture, that job stops running, capacity sits idle, and the price falls until something else is worth doing with it. Either way the market walks toward the same place: price equals the value of the marginal job.

Bitcoin already ran this experiment in public. Revenue per hash converges on the cost of hashing, and it does so no matter what anyone believes about the price of the coin. The difference with AI is only in how the convergence is enforced. Bitcoin has a difficulty adjustment. AI has competition between people who own accelerators, which is slower, noisier, and easier to mistake for something else while it's happening.

The convergence shows up in two regimes, and confusing them is where most bad forecasts come from.

When supply is short, price is set by demand. You can't conjure fabs or substations on demand. Silicon takes years, power takes longer. So during a shortage the price of compute rises to meet whatever the marginal task is worth, and the surplus flows to whoever owns the scarce thing. That is what accelerator margins and premium power contracts have been.

When supply catches up, price falls to cost, and the adjustment moves to quantity. With free entry, the price of an hour of compute settles near the cost of producing that hour: hardware amortized, energy, capital. The equilibrium then holds by expanding usage instead of raising price. More work gets done until the last job is worth only the cost floor.

Both regimes have been visible at once, which is why the market reads as contradictory. Rental prices for a given class of accelerator have fallen sharply from their shortage peak while frontier capability still commands a premium. Those are the same rule seen from two sides: supply arriving on the commodity tier, capability rents surviving only as long as the lead does.

What decides the value of the marginal job

Wages, mostly.

If a task is one that people are paid to do, the compute doing that task is worth about what the task is worth, adjusted for quality, for speed, and for how much supervision the output still needs. That gives the theory a floor and a ceiling that aren't guesses. Capability jumps don't move demand smoothly, they unlock blocks of work in steps, and each step drags a chunk of a wage bill onto the demand curve for compute. Supply answers late, because supply always answers late.

So the convergence runs both ways. The price of compute is pulled up toward the wage-equivalent of the newly automatable work. The wage of that work is pulled down toward the cost of compute. Neither side gets to stay where it was.

One correction is worth making explicitly, because it's where most compute forecasts go wrong: value created and value captured are different quantities, and the equilibrium is about the captured one. Competition on the output side pushes prices toward cost, so most of the value created leaks to whoever buys the output. What you can capture is a function of how long your capability lead lasts, how well you're distributed, and how expensive you are to switch away from. Capability leads decay in months. Distribution and switching costs decay in years.

Where margins sit at any given moment tells you which constraint is binding. Silicon margins mean silicon is scarce. Power premiums mean power is scarce. Model margins mean capability is scarce. Application margins mean distribution is scarce. The scarcity migrates, and the profits migrate with it.

Visible progress and economic progress are not the same measurement

This is where the compute question meets a line we've written about before: what can be checked, and what can't.

Benchmark progress is progress against a scorer. Economic progress is progress at work someone pays for, which means work whose result can be checked well enough that a buyer will accept it. Those two move together for a while and then they come apart. A model can get better at every scorable task in existence while the set of work you can hand it, unsupervised, and get paid for, stops growing, because that set is bounded by verification and not by capability.

When the gap between the two opens, the benchmark signal keeps rising and the revenue signal doesn't. That gap is the thing to watch, and unlike a general feeling that valuations are high, it's measurable. Revenue per deployed accelerator-hour, against what that hour costs to keep on the floor. If capability is compounding into capture, the first number rises against the second. If it isn't, no benchmark result will close it.

The end of the utilization arbitrage

There's a quieter consequence that hits everyone building on shared infrastructure.

Cloud economics have rested on statistical multiplexing. Sell the peak at a premium, fill the trough with elastic work priced near zero, because an idle cycle had almost no opportunity cost. Spot markets, preemptible instances, per-invocation serverless, burst credits: all of them are products built on the spread between peak and trough.

An AI workload that is elastic, patient, and effectively unbounded reprices every idle cycle. The trough now has a bidder, and the bidder is the operator itself. Once that's true, the peak-to-trough spread compresses toward a single price, and anything sold out of the spread loses its basis.

The obvious objection is that this is an accelerator story, and cheap elastic products mostly run on ordinary processors. That holds only until power and floor space are the binding constraint rather than silicon. After that, every workload class is bidding for the same watt, and a watt spent on a cheap invocation is a watt not spent on inference. The reservation price transmits through the power budget even where the silicon never touches.

What you should expect to see, and can already look for: thin discounts on interruptible capacity where they used to be deep, elasticity rationed by quota and negotiation rather than offered as a free option, widening gaps between committed and on-demand pricing, and steady workloads moving back onto owned hardware because predictable demand no longer earns a discount. Providers end up competing with their own customers for their own capacity. That's a different business from renting out slack, and it's a worse one.

Building for it

None of this is a reason to slow down. It's a reason to be exact about what you're buying.

Price your agents in captured value per unit of compute, not in tokens. Tokens are an input. The number that decides whether your system survives the cycle is what one useful, accepted action costs to produce, and how much someone pays for it.

Assume elasticity has a price now. Designs that quietly relied on free trough capacity, wide fan-out, speculative work thrown away, retries as a first resort, get repriced whether or not your architecture changes.

Spend compute where a check exists, and spend judgment where none does. Many attempts against a sound check is what agents are good at and what compute is for. Work that nothing can score doesn't get cheaper when compute does.

Be ready to absorb the surplus before it's priced. We're living inside the build-out, and from inside one the shortage always looks like the permanent condition. It isn't. It ends with capacity in the hands of owners who can't exit and have to run it, which means compute below replacement cost for a while. That window is the whole opportunity, and it closes: once the glut is absorbed, the marginal job sets the price again and the discount is gone. So decide now what you'd run at a third of today's price, keep it specified and waiting, and take the capacity the week it's cheap rather than the year after. The last time this happened with fiber, the people who inherited the glut built the decade that followed. The capital structure is what breaks. The equilibrium stays exactly where it was.