The measurable limits
Recursive self-improvement probably has a ceiling. What runs short first is chip packaging, memory and money, before electricity, and the piece ends in a dated bet.
The thought arrives as an objection to my own enthusiasm. Recursive self-improvement (an AI doing the research that produces the next, better AI) might work in principle, but every marginal improvement seems bound to cost more compute and more energy than the last, so the loop would run out of fuel and bend towards an asymptote instead of exploding. I studied industrial engineering, and the shape of that objection is one I was trained to trust. Carnot's efficiency is the bound no heat engine beats however clever its designer, and this feels like the same kind of ceiling.
When I try to put numbers on it, the objection holds up on the shape and fails on the place. Carnot bounds the efficiency of an engine and says nothing about how many engines you can build. When I go looking for the bound on the loop, the thermodynamic one turns out to be absurdly far away, and the limits I can actually find a number for are physical and financial: the packaging lines and memory every accelerator needs, the order books of turbine makers, the queue to connect a plant to the grid, the money that pays for all of it, the length of a training run, and the compute the automated researcher itself consumes. I have written twice about ceilings that kept moving (the wall and the moving ceiling), both times about scaling curves and benchmarks. This piece is about the loop itself, about which of its parts anyone has measured, and it ends in a bet with a date on it.
Diminishing returns are a bill
The best evidence for the objection is old and comes from chips. Bloom, Jones, Van Reenen and Webb counted the researchers behind Moore's law and found that doubling chip density now takes more than 18 times the researchers it took in the early 1970s, which means research productivity in that field falls about 6.8% a year. Ideas get harder to find, measurably.
And Moore's law did not stop. It held for fifty years because the industry kept paying: more researchers, more fabs, more money every cycle. That is the first thing the intrusive thought gets wrong. Rising cost per improvement produces a ceiling only if the cost grows faster than the resources you can throw at it, and a logarithm grows more slowly every year without ever reaching a limit.
Economists write this down as an idea production function. In the semi-endogenous version Charles Jones proposed in 1995, the speed of progress depends on how much research effort goes in and how far the frontier already is:
dAdt=δ Lλ A1−β\htmlData{sym=0}{\frac{dA}{dt}} = \delta \, \htmlData{sym=2}{L}^{\htmlData{sym=3}{\lambda}} \, \htmlData{sym=1}{A}^{1-\htmlData{sym=4}{\beta}}dtdA=δLλA1−β
dAdt\frac{dA}{dt}dtdAhow fast the technology level grows AAAthe technology level; here, how much capability a unit of compute buys LLLresearch effort, in researchers λ\lambdaλwhat duplicated work wastes (at most 1, and 1 in Bloom's baseline) β\betaβhow fast ideas get harder to find
For Moore's law, Bloom and his coauthors estimate β≈0.2\beta \approx 0.2β≈0.2, against 3.4 for the economy as a whole: chips got harder slowly. With a fixed number of researchers and β>0\beta > 0β>0, growth slows below exponential. The twentieth century stayed exponential because the number of researchers kept growing.
Automating research changes one term of that. If an AI does research as well as a person, the number of researchers stops depending on demography and starts depending on compute and on the efficiency of the software itself: a better model runs more copies on the same chips. With compute fixed, the researchers grow with AAA itself, L=cAL = cAL=cA, where ccc is how many copies the available compute runs per unit of efficiency. Put that into the same equation:
dAdt=δ (cA)λ A1−β=δ cλ A1+λ−β\frac{dA}{dt} = \delta \, (cA)^{\lambda} \, A^{1-\beta} = \delta \, c^{\lambda} \, A^{1+\lambda-\beta}dtdA=δ(cA)λA1−β=δcλA1+λ−β
The growth rate, 1AdAdt=δcλAλ−β\frac{1}{A}\frac{dA}{dt} = \delta c^{\lambda} A^{\lambda-\beta}A1dtdA=δcλAλ−β, now rises with every improvement when λ>β\lambda > \betaλ>β, and falls when λ<β\lambda < \betaλ<β. Eth and Davidson compress the same condition into one number, rrr: how many times software efficiency doubles each time the cumulative research effort doubles. In this model it is
r=λβr = \frac{\lambda}{\beta}r=βλ
so r>1r > 1r>1 and λ>β\lambda > \betaλ>β say the same thing. Above 1 the loop accelerates, at 1 it is exponential, below 1 it fizzles. Their figures for hardware are high, about 7 historically and about 5 for GPUs from 2006 to 2022, with their own warning that these might be somewhat overestimated. Their best guess for AI software is lower and wider: somewhere between 1 and 4. Direct measurements of how fast the algorithms themselves improve disagree by more: about threefold a year in Ho and colleagues' fit, and a best guess of tenfold, with an interval from 2 to 50, in Epoch's 2026 review.
Eight orders of magnitude above Landauer
Before going further I want to know whether physics closes the question on its own, since that is the Carnot instinct. There is a real analogue for computation. Landauer showed that erasing one bit dissipates at least kTln2kT \ln 2kTln2, about 2.9×10−212.9 \times 10^{-21}2.9×10−21 J at room temperature, and it follows from the second law, the same law that gives Carnot's bound.
Then the arithmetic. An H100 delivers 989 dense BF16 teraflops at 700 W, according to NVIDIA's datasheet, so about 7×10−137 \times 10^{-13}7×10−13 J per operation; a whole DGX B200 node, fans and CPUs included, comes to roughly 4×10−134 \times 10^{-13}4×10−13. That is about eight orders of magnitude above Landauer per operation, and still five or six if every operation erases a few hundred bits (a modelling choice, nobody measures it). These are peak figures, and real utilisation is lower, but no correction for utilisation closes a gap of 10510^5105.
The bound is weaker than even that suggests, because Landauer limits the erasure of information and computation does not have to erase: reversible computing has no minimum energy per operation in principle. And the brain runs human-level general intelligence on about 20 W, a fifth of the body's resting energy budget. Whatever stops the loop, thermodynamics is not near the front of the queue.
Two numbers measured on both sides of one
The real argument lives in two parameters, and the uncomfortable fact about both is that their published estimates sit on either side of the value that decides everything.
The first is rrr. Davidson and Houlden, after adjusting for compute bottlenecks, put its median at 1.2 with a range from 0.4 to 3.6, about a 60% chance of exceeding 1. Their headline result carries two conditions: assuming an AI that fully automates AI research is deployed, and holding compute fixed from then on, they give roughly 60% to the explosion compressing more than three years of progress into one, and roughly 20% to more than ten years into one.
The second is how far thinking can substitute for experiments. An AI researcher needs to run things, and runs need GPUs, the same GPUs the researchers themselves run on. Economists model the mix with a CES function, whose elasticity σ\sigmaσ says whether labour and capital are substitutes (above 1: cleverer, smaller experiments make up for missing compute) or complements (below 1: with compute fixed, output has a hard cap however much thinking you add). US manufacturing has sat at about 0.7 since 1970. Feeding that into a CES model with equal weights, Davidson reconstructs a cap of about six times today's speed of software progress, and across the range of elasticities economists report, the cap runs from 2 to 100 times. He calls the extrapolation heroic, since factory data covers ratios of labour to capital that vary about twofold and an automated lab would sit orders of magnitude away.
The only direct estimate inside AI labs, by Whitfill and Wu, uses data from OpenAI, DeepMind, Anthropic and DeepSeek between 2014 and 2024. Their baseline gives σ=2.58\sigma = 2.58σ=2.58, strong substitutes. A specification that accounts for the size of frontier experiments gives σ=−0.10\sigma = -0.10σ=−0.10 with a standard error of 0.18, indistinguishable from zero, strong complements. Davidson's answer to the pessimistic reading is empirical: if frontier-scale experiments were the binding input, algorithmic progress should have slowed as training runs grew to swallow a large share of the world's compute, and it did not. That narrows the bottleneck to experiments near the frontier, which is a smaller claim than compute is the limit.
Amdahl's law puts a ceiling on the same thing from the other side. If a fraction sss of the improvement cycle cannot be sped up (a training run takes the time it takes, a fab takes years), no number of artificial researchers gets you more than 1/s1/s1/s. If half of past AI progress came from scaling compute, which faster thinking does not accelerate, faster thinking alone buys less than a factor of two overall.
The loop has a generation time
The piece I am missing comes from Toby Ord's essay of August 2026. Recursive improvement happens in rounds (train, evaluate, deploy), and the time each round takes is its generation time. A singularity in the strict sense, infinite capability in finite time, requires the generation times to shrink to zero fast enough that their sum converges. Large gains per round do not suffice. If generation time has a floor and each round multiplies capability by a fixed factor, growth tends to an exponential; super-exponential growth without a singularity, even doubly exponential, needs gains per round that grow faster than proportionally.
A toy model of my own shows the floor at work. Each round doubles efficiency, the first round takes six months, and each round is faster than the last by the square root of the efficiency gained, down to a floor:
An+1=2An,Tn=max(Tmin, T0An)A_{n+1} = 2A_n, \qquad T_n = \max\left(T_{\min},\; \frac{T_0}{\sqrt{A_n}}\right)An+1=2An,Tn=max(Tmin,AnT0)
A straight line on this log axis is an exponential. Without a floor the rounds shrink fast enough to add up to 1.71 years, and efficiency has no finite value after that.
The curves are arithmetic on invented parameters and only illustrate Ord's point: how long a takeoff lasts is decided by the floor under the cycle, whatever each cycle gains. Ord adds two cautions that cut against both camps. Growth rates that rise from 10% to 20% to 30% a month give e0.05t2e^{0.05t^2}e0.05t2, brutal and still finite at every date. And whether a curve is singular depends on the units: mean time between failures can go to infinity while reliability only reaches 100%, and chess progress that is linear in Elo is exponential in the odds of winning. He proposes requiring labs to report their generation times, for pre-training and for post-training with reinforcement learning.
Part of that floor is already measured. Epoch AI finds frontier training runs lengthening about 1.4 times a year since 2020, and estimates that runs beyond nine months may be inefficient, a threshold it expects around 2027. That is a floor on the pre-training round. Only the faster loops (post-training, scaffolding, tools) can go below it, and their gains on a fixed base model tend to run out. On the observable side, METR's time-horizon measurements put the doubling of the length of tasks models complete at about seven months over 2019 to 2025, 4.3 months since 2023 and about three months since 2024, all of it on software tasks. A shortening cycle and a one-off boost from reinforcement learning would both produce that curve, and Ord's result is that a few years of local data cannot separate a curve that will saturate from one that will not.
What runs short first
If software efficiency hits the ceiling of its paradigm, the long-run rate is set by the other factors of the product: how much hardware gets built, how efficient it is, and the power and money behind it. A power plant has the same structure (useful output equals efficiency times heat supplied), and Carnot only bounds the first factor. So I look for the rate of each of the others, and for the one that runs short first. The answer is a sequence, and electricity is not at the head of it.
Through 2027 the scarce input is the step between the chip and the server. Every AI accelerator is a logic die packaged together with stacks of high-bandwidth memory, and both the packaging and the memory are sold ahead. TSMC's chief executive said it on the July 2026 call: our packaging capacity is so tight that now it's limiting my customers' growth. Micron has the vast majority of its 2027 HBM output under agreement and expects supply to be much tighter in its fiscal 2027 and 2028. The wafers themselves are not what is missing yet: Epoch finds that the four largest AI chip designers used about 90% of that packaging and memory in 2025 and only 12% of advanced logic wafers. Industry trackers put packaging capacity growing by roughly 60 to 70% a year and memory somewhat less; with each chip generation delivering about 40% more per package, that leaves room for the stock of AI compute to grow perhaps two to two and a half times a year (my estimate from those figures), against the 3.3 times a year it grew from 2022 to 2025. New fabs take two to three years to build and TSMC expects them to add real supply only in 2028 and 2029.
Power is the constraint everyone names, and it binds in a narrower way than the name suggests. Epoch counts 31 GW of AI data-centre capacity worldwide at the end of 2025, growing about 2.3 times a year, far faster than the 15% a year the IEA expects for data centres as a whole. Turbine factories cannot follow that pace. In July 2026 GE Vernova reported 116 GW of gas turbines between backlog and slot reservations, and its chief executive described production as mostly sold out through 2030; Mitsubishi Power schedules the orders it booked last quarter for delivery between 2028 and 2030; connecting a new plant to the US grid takes a median of more than five years (Berkeley Lab), and a large new load in Virginia can wait up to seven. The builders route around it one site at a time, with gas turbines on site, reopened nuclear plants and utilities paid to build for them. And the price of the electricity hardly matters: Epoch's cost breakdown puts it near 7% of the yearly cost of a gigawatt campus, against 60% for the servers. Power decides where the largest clusters can go and when, which binds a single training site tightly and the total loosely.
From about 2028 the binding input is money. The five largest spenders have grown capital expenditure by about 70% a year since GPT-4, while their operating cash flow grows by about 23%, and Epoch has the two lines crossing around the third quarter of 2026: from there on every extra dollar is borrowed. The bond market already carries it, with $109 billion of hyperscaler bonds in 2025 and $194 billion in the first half of 2026 alone (J.P. Morgan). What could keep the money growing is revenue, and it is moving faster than anyone planned: Anthropic's annualised revenue went from about $9 billion at the end of 2025 to $65 billion in July 2026, while OpenAI, by the Financial Times' account of its own projections, runs out of cash in 2028 without new funding. Whether the money binds in 2028 or later depends on which of those curves wins, and the silicon binds first either way.
The researcher runs on the experiment budget
Every power plant spends part of what it generates on its own pumps, fans and auxiliaries, and if that self-consumption grew faster than gross output, net output would peak and then fall. The automated researcher has the same problem: it runs on the GPUs its experiments need. Lennart Heim, quoted by Scientific American in May 2026, cited reports that around 60% of lab compute already goes to research and development, competing with serving users. In late March Anthropic made usage limits burn faster at peak hours, and OpenAI announced it was shutting Sora down in its app and API, which Scientific American places in the same squeeze1.
1Anthropic also lowered Claude Code's default reasoning effort from 4 March to 7 April, but its own postmortem attributes that to interface latency, so it is no evidence of rationing.
The intrusive thought is most precise here, because the most capable models are partly more capable by thinking longer. The data points both ways. Holding capability fixed, the price of reaching it falls by 9 to 900 times a year depending on the task. At the frontier the top score is bought with more compute per task: on ARC-AGI, Gemini 3 Pro goes from 31% at $0.81 a task to 54% at $31 with refinement, according to the ARC Prize team. If the second trend dominates, there is an optimal share of compute for running researchers and adding more beyond it slows progress. If each generation solves frontier tasks more cheaply than the last, the self-consumption shrinks by itself.
How much faster AI already makes research is measured badly. In METR's latest study, developers returning from its first trial were about 18% faster with AI, with an interval from 9% slower to 38% faster, and METR calls the estimate unreliable because developers increasingly refuse to work without it. Inside the labs the figures are larger and self-reported: a median fourfold gain in a survey of Anthropic's researchers, which Anthropic itself calls almost certainly an overstatement. The true number for researchers at a frontier lab sits somewhere between the two, and nobody outside has measured it.
The rest of this section is an interpretation, with no study behind it yet. A heat engine needs a cold reservoir to do work; a loop of self-improvement needs a verifier that can tell the better version from the worse one, and without it, generating variations is heat. Karpathy's autoresearch is the cleanest example I know: an agent edits the training code of a small model, trains for a fixed five minutes of wall clock, and reads one number, validation bits per byte, for about a hundred experiments a night. It works because the verifier is fast, cheap and hard to fool. At work I build the benchmarks that decide whether a clinical model ships, and the verifier there is a set of real clinical cases with confirmed diagnoses, which grows only as fast as clinics produce them. Read this way, Whitfill and Wu's two elasticities are a statement about verifiers: where a small experiment predicts the large one, thought substitutes for compute and σ\sigmaσ is high; where behaviour only appears at scale, the only valid check is the large run and σ\sigmaσ falls towards zero. It would also explain why METR's fast curve is a fact about code, the domain with the best verifier, and why a metric that is too easy to optimise produces reward hacking instead of capability.
The bet
Everything above points one way more firmly than any single number in it, and the direction is narrower than the thought I started with. I am not betting that capability stops growing. Every wall in this field has so far turned out to be a bend, the ceilings in my own earlier pieces moved each time, and capability has paths (better algorithms, more thinking at inference, the next architecture) that the physical inputs do not. The bet is on the engine that has driven everything since 2020. By 31 December 2030, the training compute of the largest runs will be growing at less than three times a year, against five times a year since 2020 (medium-high confidence). It counts as a hit if Epoch's fitted trend over the notable models released in 2029 and 2030 is below 3× a year, or, if nobody fits one, if the largest disclosed run of 2030 used less than nine times the compute of the largest of 2028. A trend of 4× a year or more refutes it.
Two narrower bets ride on it. Electricity will not be what ends the fivefold era: global AI data-centre capacity, as Epoch counts it, will pass 100 GW before the end of 2028 while the stock of AI compute grows by less than 2.5 times a year over 2026 to 2028, which is what power being loose in total and silicon being tight looks like from outside (medium). And the test of the verifier: before the end of 2028 an independent measurement will show AI at least doubling the speed of machine-learning research done on a fixed compute budget, while the frontier does not accelerate past its pace of 2025 and 2026, with METR's doubling time staying above about three months (medium). That pairing is what a low substitution elasticity looks like in public data. If both accelerate together, the software loop is deeper than this piece allows, and that is the result I would most like to be told about.
Three things could be fooling me, in order of how much they worry me. Several of the rates above come from one research group, Epoch, so their agreement is weaker evidence than it looks. The shortage of packaging and memory is described by the companies that sell them, who are paid more while buyers believe it. And this kind of claim has a poor record. What keeps the bet on the engine safe is that it barely depends on the question I am least sure of, whether the money keeps coming: if revenue keeps compounding, the silicon still binds first, and if it does not, the money binds sooner.
Three measurements would move the rest from argument to evidence, and none of them is published in a form anyone can use. The substitution elasticity inside AI labs, with error bars that do not cross 1. The trend in compute per solved task at the frontier, generation over generation. And the generation time of the fast loops, which Ord wants labs to report and which nobody outside them knows.
What none of this can see
Everything above rests on a premise I want to state as one, because it carries more weight than any single number: that no discontinuity of the kind machine learning keeps producing arrives in the window being measured. Machine learning has black swans, and they are inherent to the field. The cleanest measurement of that is recent: Gundlach and colleagues ablated the innovations credited for a decade of language-model efficiency and found that one switch, from recurrent networks to the transformer, accounts for most of what they could explain, and that its value grows with the scale at which it is measured. A decade of algorithmic progress, on that reading, is mostly one event plus a slope. Every growth rate in this piece is a slope, so none of them contains the next event.
The first unknown is the instruments themselves. METR's time horizon is the best public measure of what frontier models can do on their own, and its own staff already say it is reaching its edge: an early version of one model was placed at sixteen hours or more, with a 95% interval from 8.5 to 55, on a suite where only five tasks run that long1. Its study of how much AI speeds up developers had to be redesigned in February 2026 because the developers it recruits increasingly refuse to work without AI, which biases the comparison. Two of the quantities this piece leans on may stop being measurable within the window it talks about, and how capability will be evaluated after that is an open question with no candidate answer yet.
1Reported by The Decoder in May 2026, citing METR.
The second is verification. Section by section, the cheapness of checking a result has decided where progress is fast: code before medicine, small experiments before frontier runs. A method that made a model's internals legible, the program mechanistic interpretability works towards, would change that price directly. If a model's reasoning could be read and checked instead of only its outputs, the verifier would stop being the scarce input in domains where it is scarce today, and the substitution elasticity would move with it. Nothing published so far does this at the scale of a frontier model, and nobody can put a date on it.
The third is the substrate. Reversible or optical computing would remove part of the energy cost that sits under the physical side, and both remain laboratory work with no product timeline.
These are a different kind of uncertainty from the economic one, and I keep them apart on purpose. A credit event at a leveraged data-centre builder, lab revenue that stops compounding or a forced change in how GPUs are depreciated are risks with a mechanism and a price; they belong in the estimates, and the bet above is built to hold whichever way revenue goes. A new architecture, an interpretability method that works or a measurement that breaks have no distribution anyone can defend. The conclusions of this piece hold under the premise that none of them happens before 2030, and the honest reading is conditional: if one does, the slopes stop meaning what they meant, and the analysis has to be redone from the event.
The slowest clock
A loop that improves itself runs at the pace of its slowest part, and in this loop the slow parts are made of matter and money. Thermodynamics sits eight orders of magnitude away, since Landauer bounds the erasure of a bit and Carnot the efficiency of an engine, and neither bounds how many get built. The economics of ideas sets the exponent, the doublings of software per doubling of research against the threshold of one. The elasticity of substitution decides whether thinking can stand in for experiments, which turns out to be the question of whether a cheap verifier exists. Amdahl's serial fraction and Ord's generation time put a floor under every round, and below them the packaging lines, the high-bandwidth memory, the turbines and the capital set how fast that floor itself can move. Software can compress every term made of information; the bet of this piece is that, before 2030, the terms made of matter set the pace.
5×yearly growth of the training compute of frontier runs since 2020Epoch AI ~90% / 12%share of 2025 packaging and HBM, against share of advanced logic wafers, used by the four largest AI chip designersEpoch AI 31 GWAI data-centre capacity worldwide at the end of 2025, growing about 2.3× a yearEpoch AI ~7%share of electricity in the yearly cost of a gigawatt AI campusEpoch AI 116 GWGE Vernova's gas turbine backlog and reservations, sold mostly through 2030GE Vernova, Q2 2026 ~70% / ~23%yearly growth of hyperscaler capex, against their operating cash flowEpoch AI 0.4 to 3.6range for r, the doublings of software per doubling of research effort; explosion above 1Davidson and Houlden, 2025 −0.10 to 2.58estimates of the substitution between compute and research labour in AI labs; complements below 1Whitfill and Wu, 2025