Building a bridge costs hundreds of millions, once. Maintaining and operating it costs far less, but every day, for fifty years. In the end, no one can say which of the two items weighed more. AI is living exactly this situation, and the confusion between these two costs explains many misunderstandings about its economics.
It's an elementary distinction, and yet it's missing from most discussions. It sheds light on price cuts, data centre investments, and even orbital computing projects. Let's set it out clearly.
The two phases
Training is the making of the model. You feed it colossal amounts of data, adjust its billions of parameters, and this goes on for weeks or months across tens of thousands of chips. It's the process we described in our dedicated article. The cost is gigantic and concentrated: a single bill, before anyone uses the product.
Inference is usage. Every time you ask a question, the model performs calculations to produce its answer. The unit cost is tiny, often a fraction of a cent. But it repeats with every request, from every user, every day.
A simple image: training is writing the encyclopaedia. Inference is consulting it. Writing costs a fortune once. Consulting costs next to nothing, but a billion times a day.
Why inference has taken over
In the early days, training dominated discussions because few people were using these models. With hundreds of millions of daily users, the balance has flipped: for the most established players, the cumulative cost of inference now far exceeds that of training.
This shift explains a good part of the news we cover.
It explains why optimising execution allows price cuts of 80%: when you serve billions of requests, saving a few percent per request amounts to enormous sums.
It explains the interest in mixture-of-experts architectures, which activate only a fraction of the model per request: this is above all an inference optimisation.
It also explains why the giants design their own chips, often dedicated specifically to inference: at that volume, a custom chip becomes cost-effective.
If you build a product on an API, your cost is entirely inference. What matters, then, is not the model's power but your consumption per completed task. That's why we keep stressing levers like caching, routing to lighter models, and controlling reasoning tokens. None of these levers touches training: they all bear on the second bill, the one that comes back every day.
A technical detail with consequences
The two phases don't have the same requirements, and this shapes the entire industry.
Training demands enormous raw compute, but tolerates latency: whether the result arrives in six weeks or six weeks and two days makes no difference. It can therefore run anywhere, including in remote locations, as long as electricity is available.
Inference, by contrast, requires response speed. No one accepts waiting three extra seconds because the request took a detour. It must therefore run close to users.
This difference explains a point we raised in our article on space data centres: orbital training is conceivable, conversational inference much less so, because of signal travel time. This isn't an engineer's detail; it's what decides what can be offshored and what must stay near you.
What to remember
Keep this formula in mind: training is an investment, inference is an operating cost. The former makes headlines, the latter makes the bills.
This distinction gives you an immediate reading grid. When you read that a model cost hundreds of millions to train, it's impressive, but it's not what decides its viability. What decides is how much each answer costs, multiplied by the number of answers. That's why the current race is less about the most powerful models than about the models that are most efficient to serve. And that's rather good news: this race directly benefits those who use them.