DeepSeek raised the price of its V4 Flash model on 14 August, from around $0.14 to $0.27 per million tokens, a 93% increase. Two weeks earlier, we noted that this same model embodied the pricing pressure on Western players. The tide is turning, and that's worth a look.
We ran an article on this 14-cent model that beat its own house's high-end offerings. It has now doubled in price. Let's look at what that means.
The pricing landscape has become unreadable
Let's sum up the moves of the past two weeks, because they point in every direction.
OpenAI cut the price of its fast model by 80%, to the point of making it the default option in its free tier with unlimited exchanges. Claude Sonnet 5's promotional price ends on 31 August with a 50% increase. And DeepSeek is up 93%.
Three players, three directions. There is no single trend, and that in itself is the story.
The possible explanations
None are confirmed, and they are not mutually exclusive.
The initial price was not sustainable. That's the simplest hypothesis. A conquest price, designed to win users, eventually gets adjusted toward its real cost. Fourteen cents for a model reaching that level of performance was remarkably low.
Demand has exploded. When a service becomes very popular, two options exist: invest heavily in capacity, or raise the price to regulate load. The second is faster.
Infrastructure costs are rising. We have documented the pressure on memory affecting the whole sector. A provider whose components cost more eventually passes that on.
This increase validates what we have been saying for weeks: don't build your economics on a current price. A low price is a commercial choice, not a property of the product. It can double in a single announcement, and the only protection is being able to switch providers without rewriting your application, as we detailed in our selection guide.
What this says about the economics of inference
The underlying point is this: serving a model costs something real, and that cost is nowhere near zero.
We explained in our article on the two costs of AI that inference is an operating expense that recurs with every request. Software optimisations genuinely reduce it, as OpenAI's 80% cut showed. But they don't eliminate electricity, hardware, or its scarcity.
A price war can therefore last as long as players are willing to sell at a loss to gain market share. It stops when someone has to present accounts, which is precisely what is happening right now with Anthropic's first operating profit and OpenAI's IPO.
What to take away
It would be premature to declare the end of falling prices: efficiency gains are real and will continue. But this increase is a reminder that the trend is not mechanical, and that a price reflects commercial strategy as much as production cost.
The practical consequence is simple. Measure your real costs on your own tasks, keep the ability to switch providers, and treat any price as temporary. Those who built an economy on a conquest price will sooner or later discover they built it on a promotion.