The Second Half of the Board
Nothing looks exponential from the inside. Every square is a reasonable step up from the last one, the recent past always looks flat, and you always feel like you are standing right at the elbow. You are — and so was everyone before you.
The old story: a courtier asks a king for one grain of wheat on the first square of a chessboard, two on the second, four on the third, doubling all the way to the sixty-fourth. The king thinks it is a modest request. How much wheat sits on square 64 alone?
—
In 1965 Gordon Moore looked at five data points and predicted the count would keep doubling. Here is every doubling since, from the first commercial microprocessor to a modern AI accelerator — plotted as multiples of the 1971 chip, so two very different things can share one axis. The dashed line is a perfect two-year doubling. Switch the axis and watch the story invert.
View as table
—
Note what the line is made of. It starts on Intel CPUs and ends on NVIDIA GPUs — a different company, a different architecture, a different purpose, bought by different people for different reasons. The substrate changed completely and the slope did not notice.
Transistor density is not the only curve, and it is not even the one that matters most commercially. What a buyer actually cares about is compute per dollar. Measured independently, on different hardware, over different decades, by people with no stake in Moore's reputation:
Transistors per chip
Measured across the chart above: 2,300 → 208 billion is 26.4 doublings in 53 years.
GPU performance per dollar
Epoch AI, across 470 GPU models released 2006–2021.
AI accelerator compute per dollar
Epoch AI, across 20+ accelerators released 2012–2025 — roughly 30–40% better every year.
Those numbers were produced by unrelated methods, on unrelated hardware, and they land within half a year of each other. That is the finding. Whatever is driving this is not a property of silicon — silicon has been replaced several times over inside that span. It is a property of a competitive industry that reinvests its gains, which is why it survives the death of any particular technology.
The most-quoted obituary for Moore's Law came from the person with the most to gain from its replacement. At GTC in September 2022, NVIDIA's chief executive Jensen Huang told reporters:
"Moore's law is dead. And the ability for Moore's Law to deliver twice the performance at the same cost, or at the same performance and half the cost, every year and a half is over. It's completely over." Jensen Huang, GTC press Q&A, September 2022
Twenty-eight months later, at CES in January 2025, the same person said this:
"Our systems are progressing way faster than Moore's Law." Jensen Huang, CES, January 2025
Read as predictions those two statements contradict each other. Read as descriptions they do not, and the gap between them is the whole point of this page. The 2022 claim is about a mechanism: shrinking transistors on a general-purpose CPU stopped delivering cheaper performance on schedule. That genuinely ended. The 2025 claim is about an outcome: performance per dollar still improves, because the industry stopped getting it from lithography alone and started getting it from co-designing chip, system, libraries and algorithms together.
The mechanism died. The exponent did not. It changed substrate — which, looking at the chart above, is what it has done repeatedly and without ceremony.
Two honest caveats, because this is a vendor's account of why you should buy the vendor's product. First, Huang's stronger framing — often repeated as "Huang's law," sometimes as performance doubling every year — runs ahead of the independent measurements. Epoch AI's 2.2-year doubling for compute per dollar is not only slower than that, it is marginally slower than classical Moore's Law. Second, "faster than Moore's Law" is doing quiet work across different metrics: raw FLOPS, FLOPS at reduced precision, and whole-system throughput are not the same measure, and the headline number usually picks the most flattering one. The curve is real. The steepness claimed for it is marketing.
Verdicts on immature technology
"We tried it and it wasn't good enough" is a true statement about a square you have already left. If a capability doubles every two years, a pilot that failed two years ago was judged against half the system you can buy today, and one from four years ago against a quarter. Re-test the conclusion on a schedule, not the tool.
Business cases built on a straight line
A five-year plan that extrapolates linearly from today's unit costs will be wrong by an order of magnitude in either direction — too timid on what becomes affordable, too generous on the margin you keep by doing it the old way. Model the cost curve, not the cost.
Build-versus-wait decisions
On an exponential, waiting is genuinely cheaper — the capability gets better and the price falls, so the patient buyer really does get more for less. What does not wait is the learning: the organisation that started at square 30 has six years of practice by the time the late starter reaches square 33 holding a better tool and no idea how to use it.
Any compounding you own
Same arithmetic, smaller board. One percent a week is about 1.7× a year and roughly 13× over five. Every doubling makes the entire history before it look flat, which is precisely why sustained improvement is so hard to fund — it looks like nothing is happening right up until it does.
The nearest thing to a general lesson: on an exponential, your intuition is not slightly wrong, it is wrong by orders of magnitude — and it is wrong in the same direction every time. The board demonstrates it in sixty-four squares. See also the Compute Scaling Explorer for the same gap measured across real machines, the Context & Cost Curve for what this has done to the price of a token, and From Horses to AI for the last time a general-purpose technology reorganised an economy.
The substrate changes. The exponent doesn't. AI is not a break in the line — it is where the line went when silicon scaling ran out of room.
What this cannot tell you. A trend that has held for 53 years is evidence, not a guarantee — every exponential in nature is a sigmoid that has not yet bent, and this one has real physical and economic ceilings (atoms, power, capital, heat). The chart plots representative chips rather than an industry average, and transistor counts across CPUs and GPUs are not measuring the same kind of work, which is exactly why it is drawn as multiples on one axis rather than as absolute performance. Compute per dollar is the more decision-relevant curve and it is the one measured independently above. Nothing here says capability improves at the rate compute does — that link is contested and is not claimed on this page.
The chessboard. Grain counts are exact powers of two; mass assumes 65 mg a grain, the standard assumption in the wheat-and-chessboard problem. World harvest comparisons use annual world wheat production of roughly 842 million tonnes (2025/26). The whole board comes to about 1.2 trillion tonnes — over 1,400 times one year's harvest.
Sources.
- Moore, G. E., "Cramming more components onto integrated circuits," Electronics, 19 April 1965.
- Kurzweil, R., The Age of Spiritual Machines, 1999 — origin of the "second half of the chessboard" framing.
- Transistor counts: manufacturer specifications for the chips listed in the table above.
- Huang, GTC press Q&A, September 2022 — reported by CNBC.
- Huang, CES, January 2025 — reported by TechCrunch.
- Epoch AI, "Trends in GPU price-performance" (470 GPUs, 2006–2021) and "Performance per dollar improves around 30% each year" (AI accelerators, 2012–2025).
- Wright, T. P., "Factors affecting the cost of airplanes," 1936 — the experience curve underneath all of this; see Technology Trends & Tradeoffs.