Nobody Owns the Whole Path
You ask an assistant one question and get one answer, from one product, with one company's logo on it. Underneath, your question was handed between a dozen separate businesses — each doing one small thing well and passing it on. Guess how many first. Then cut any one of them and watch the blast radius spread through the ones you have never heard of.
You type one question into an AI assistant and get one answer back. How many separate companies' computers touched that question on the way?
What the customer sees
Twelve providers, all up
Pick any hop above and cut it. Watch how far the damage travels — and notice how often the answer to "who do we call" is somebody we have no contract with.
Hop
Pick a hop
Each of these does one small thing extremely well and hands the work to the next. That is why the whole thing works at all — and why no single organisation can tell you whether it is healthy.
Text alternative — all twelve providers, what they do, and what breaks
| Hop | Specialism | Relationship | What the customer sees if it stops |
|---|
Twelve providers, most of them in series. Twelve components that are each up 99.9% of the time multiply out to 98.8% — about 105 hours a year, or four and a half days. Your availability is the product of everyone else's, and you signed roughly a third of those contracts.
Storage and inference are both "computers doing work for you," and they could not be less alike as businesses. Every strange thing about the AI supply chain — the waiting lists, the token pricing, the fact that your model vendor is renting from a cloud that also competes with it — falls out of this one table.
| Storage — remembering | Inference — thinking | |
|---|---|---|
| You buy | A gigabyte, for a month | A million tokens, once |
| Price direction | Pennies, and falling for forty years | Dollars, falling fast per unit — while total spend climbs |
| Scarcity | Effectively unlimited; add another disk | Rationed by chip supply, and the chips have a queue |
| Cost of sitting idle | Almost nothing — that is the normal state | Brutal. An idle accelerator bills exactly like a busy one |
| Copies | Copied to three continents by default | One hot copy, as close to the chips as physics allows |
| Answer time | Milliseconds, and predictable | Seconds, and variable with how busy everyone else is |
| Failure feels like | Gone. Loudly, permanently, and someone is liable | Slow, then wrong. Quietly, and nobody is sure who is liable |
| So the market became | A utility. Boring, metered, near-commodity | An oligopoly. A handful of frontier labs, renting capacity from a handful of clouds |
This is the reason a frontier model is not simply "installed" the way a database is. Remembering is cheap, patient and copied everywhere; thinking is expensive, hot and rented by the second — so the economics pushed them into separate companies, and your one question now has to cross the boundary between them and come back.
Your status page only knows what you own
It is wired to the two or three things your team runs, so it stays green through most of the failures above. Customers become your outage detection — and the gap between "they noticed" and "you noticed" is the number that actually damages trust.
The contract you never signed
You negotiate an SLA with your vendor. Your vendor rents from a second, who rents chips from a third. Nothing you signed reaches past the first hop, and the credit you receive when it fails will be a rounding error against the revenue you lost.
The model vendor is also a tenant
Frontier models are hosted on rented accelerators, often inside a cloud that sells a competing model. That shapes pricing, capacity during a crunch, and who gets served first. "Which model" is a smaller decision than "whose chips, in whose building, under whose law."
Twelve vendors, three buildings
A dozen logos on the invoice can still land in a handful of regions belonging to two or three hyperscalers. Diversity on the invoice is not diversity in the failure domain — and you only discover the overlap on the day they all go down together.
It looks like a mess because it is a mess — and the mess is load-bearing. Each hop got very good at one narrow thing precisely by refusing to care about the others, and that refusal is what let the whole arrangement grow faster than any single company could have built it. No one is holding the whole thing; each hop only has to hold its own piece. The price is that your uptime, your latency, your costs and your legal exposure are all products of decisions made by companies you cannot name, and the useful question is never "is our system up" but "which of the twelve is having a bad day, and would we even know?"
Keep going: Nine Things Have To Work is this same picture stood upright — the nine layers inside a single one of these hops. The AI Value Chain follows the money through the same arrangement.