Skip to content

Kimi K3 burns through twelve times more reasoning tokens than Claude Opus: the hidden cost of thinking

An open model that is 40% cheaper per token can cost more per task. An analysis of Kimi K3's reasoning traces explains both why it is excellent and why it is ruinously expensive.

Advertisement
The essentials in 30 seconds ⚡
An analysis published on 24 July 2026 by the evaluator Design Arena reveals that Kimi K3 consumes more than twelve times more reasoning tokens than Claude Opus 4.8, and more than double its own predecessor. This is not a bug: it is precisely what explains its top spot on a frontend code leaderboard. But it is also what makes its advertised price misleading. A model can be cheaper per token and more expensive per task.

We introduced you to Kimi K3 on Monday, the largest open-weights model ever released, coming out of Beijing with 2.8 trillion parameters and results that rival the best closed models. An analysis published since then fills in the missing piece of the puzzle, and it is instructive for anyone paying an AI bill.

The figure that stands out

The evaluator Design Arena studied the model's reasoning traces, i.e. the internal draft it produces before answering. Verdict: Kimi K3 uses an extreme amount of reasoning tokens, more than twelve times more than Claude Opus 4.8, and more than double Kimi K2.6.

Let's recall what that means in practice. As we explained in our article on reasoning models, a modern model first produces a long intermediate reasoning, a kind of written draft in which it explores the problem, before formulating its answer. This draft is made up of tokens, and those tokens are billed. Twelve times more reasoning means twelve times more billable material generated before the answer even appears.

Why this flaw is also its strength

Here is the interesting twist. This appetite is not waste; it is the mechanism behind its performance. Kimi K3 ranks first on Design Arena's single-attempt frontend code leaderboard, with an Elo score of 1392, ten positions above Kimi K2.6 and sixteen above K2.7 Code. It is the biggest jump ever seen in the Moonshot line.

Reading the traces, analysts understood why. Kimi K3 iterates on its own designs within its chain of thought itself, the way an autonomous agent would chain steps, but without leaving its reasoning. Its traces are multi-stage: it moves from planning to decision-making, then to designing each part individually.

The observable effect on the result 🎨
This deep reasoning produces measurable results. The model thinks about the images it embeds, where others fall back on alt text and defaults. According to the analysis, Kimi K3 missed no visual decoration on tests where Claude Fable 5 and K2.7 Code had to improvise. Same for the use of external libraries: it produces better scroll animations and better charts simply because it thinks about how to use them during its reasoning. The quality literally comes from the time spent thinking.

The economic trap

Now the flip side, and it is serious. On paper, Kimi K3 is significantly cheaper than Western competition. Billed at around $3 input and $15 output per million tokens, it is roughly 40% cheaper than Opus 4.8 on each price line. A simple calculation gives, for a coding agent workload of 100 million input tokens and 20 million output tokens per month, around $600 on K3 versus $1,000 on Opus 4.8.

Except that calculation ignores precisely the subject of this article. If the model generates twelve times more reasoning tokens to accomplish the same task, the 40% per-token saving can be entirely absorbed, or even exceeded. The advertised price measures the cost of a token. What matters to a business is the cost of a completed task.

One detail makes the problem worse: at launch, the effort adjustment parameter only accepted a single value, maximum. The model therefore spent reasoning tokens even on trivial queries. The documentation has since mapped three levels, standard, high and maximum, but the default behaviour remains greedy. Conversely, a model like Claude Opus 5 notably highlights an effort dial allowing task-by-task trade-offs, which directly addresses this problem.

The paradox of the free and ruinous open model

There is a lesson here that goes beyond Kimi K3. An open-weights model is, legally, free: you can download it, modify it, host it without paying anyone. That is the sovereignty promise we described in our article on open models.

But free in licence does not mean free to run. A model that thinks a lot consumes a lot of compute, and compute costs money, whether in API bills or in electricity and hardware if you host it yourself. A royalty-free, compute-hungry model can cost more than a closed, frugal one. The two freedoms, legal and economic, do not necessarily go together.

The habit to adopt

The practical conclusion comes down to one sentence: stop comparing prices per million tokens, compare costs per actually completed task on your use cases.

Concretely, before switching a project to a model on the grounds that it shows an attractive rate, run around twenty representative tasks and measure the total bill, reasoning tokens included. You may find that the cheapest model on the price list is the most expensive in the end. It is the same underlying logic we highlighted regarding the real cost of AI in business.

Kimi K3 remains a remarkable achievement, and its top spot on frontend code is not disputed. But it perfectly illustrates a shift of 2026: the question is no longer which model is the smartest, nor even which shows the lowest price, but which accomplishes your work for the least money. These three questions do not have the same answer.

Advertisement