Ask someone what 7 times 8 is, and they'll answer instantly, on reflex. Ask them to multiply 347 by 89, and they'll grab a piece of paper, set out the calculation, and work through it step by step. No one does that sum in their head in one go. Reasoning models are exactly that shift applied to AI: instead of answering on first reflex, the model takes the time to lay out its reasoning before concluding.
You may have noticed that an AI sometimes displays a message like "thinking in progress" before answering, and that its reply arrives with a delay of several seconds. That's not a technical slowdown; it's a feature, and probably the most important one of the past two years. Here's what's actually happening during that silence.
The problem it solves
Let's recap the basics: a language model predicts the next word, one after another, as we detailed in our article on LLMs. This mechanism works remarkably well for writing, but it has a structural weakness: each word is produced with no possibility of going back.
Imagine having to solve a complex logic problem by writing your answer in a single pass, with no draft, no way to correct yourself, starting straight with the conclusion. That's exactly the constraint classic models faced. The result: they excelled at writing but stumbled on problems requiring multiple logical steps, where an early error condemns everything that follows.
The solution: thinking out loud
The idea that changed everything is called the chain of thought. Rather than forcing the model to answer immediately, you let it first produce a long intermediate reasoning, a kind of written draft in which it explores the problem, tests avenues, sometimes contradicts itself, corrects itself, before formulating its final answer.
This draft works like your scratch paper. It lets the model break a hard problem down into small manageable steps, each simple enough to handle correctly. A reasoning model is therefore simply a model specifically trained to produce quality intermediate reasoning, and to use it to reach a better conclusion.
Each lab has its own commercial name for the same family of ideas. Google talks about Deep Think, OpenAI about max and ultra modes, Moonshot about a reasoning effort parameter set to maximum by default on Kimi K3. Behind these names, the principle stays identical: giving the model more compute and time before it answers. What varies is how much control you're given over that dial.
Parallel reasoning, the next step
The logic has been pushed even further. Rather than a single draft, some systems now launch multiple simultaneous reasoning paths that explore different avenues, before consolidating the most promising results.
It's this mechanism that allowed an AI to produce a proof of a mathematical conjecture left open for fifty years, by launching 64 sub-agents in parallel, each pushed to attack the problem from a different angle to stop the search from closing in too early on one appealing idea. We've moved from an AI that thinks to an AI that organises a team of internal researchers.
The very real price to pay
This power has a cost, and it's not trivial. All that thinking draft is made up of tokens, those units of text billed by providers, the principle of which we explained in our article on tokens. A model that thinks at length therefore consumes a huge amount, even when the final answer is just three lines long.
The Kimi K3 example is telling: independent testers recorded a consumption of over 13,000 reasoning tokens to generate a simple vector image, around 25 euro cents for an otherwise mundane query. On heavy use, this appetite can wipe out the advantage of an attractive per-token price. That's why companies increasingly think in terms of cost per completed task rather than the advertised price per million tokens.
There's also a subtler limit. A long reasoning isn't automatically good reasoning. A model can lay out fifteen impeccably presented steps and reach a false conclusion, because an initial assumption was wrong. The readability of the draft gives a reassuring impression of rigour that doesn't guarantee accuracy.
What to take away
Reasoning models mark a real shift: by granting compute time before answering, they've unlocked mathematics, complex logic, and multi-step tasks that were long out of reach. That's what made possible the rise of autonomous agents capable of carrying out work over several hours.
In practice, this gives you a useful reflex: for a simple question, a fast model is enough and you save both time and money. For a problem that demands rigour, let the model think, even if it means waiting. Knowing when to demand reasoning and when to settle for a reflex has become, for the user, a skill in its own right.