On 16 July 2026, Beijing-based startup Moonshot AI released Kimi K3, a 2.8-trillion-parameter model, the largest open-weights model ever announced. On several independent benchmarks, it sits just behind Claude Fable 5 and GPT-5.6 Sol, and even takes first place on a frontend coding test, ahead of Fable 5. One major caveat, though: at the time of the announcement, the weights were not yet downloadable.
The compression of time is perhaps the most striking phenomenon in AI in 2026. A year ago, the gap between open models and the best closed Western models was measured in years. With Kimi K3, released on 16 July by Beijing-based startup Moonshot AI, that gap is now measured in months. Here is what this model is really worth, and why its arrival matters to everyone.
The numbers, and what they mean
Kimi K3 boasts 2.8 trillion parameters, making it the first open model in the 3-trillion class and the largest ever released in this category. To put it in context, it is roughly 75% larger than DeepSeek V4. But as we explained in our article on parameters, this raw figure only tells part of the story.
The technical detail that matters more: K3 uses a Mixture-of-Experts architecture that activates only 16 of its 896 internal experts per token, or about 1.8% of its total capacity at any given moment. In other words, a gigantic brain, but one that only calls on the few specialists useful for each question. It also comes with a one-million-token context window and native vision, making it a multimodal model capable of handling text and images in the same stream.
Where it really stands
This is where the exercise gets interesting, because the results are nuanced and partly come from independent sources, which is too rare not to be highlighted.
| Benchmark | Kimi K3 score | Ranking |
|---|---|---|
| GDPval-AA v2 (real-world tasks, 44 professions) | 1,687 | 3rd, behind Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,748), ahead of Opus 4.8 (1,600) |
| AA-Briefcase (knowledge work) | 1,527 | 2nd, ahead of GPT-5.6 Sol Max (1,495) |
| Frontend Code Arena (blind test) | 1,679 | 1st, ahead of Fable 5 |
| Artificial Analysis Intelligence Index | 57 | On par with Opus 4.8 and GPT-5.5, behind Fable 5 and GPT-5.6 Sol |
The honest summary comes down to one sentence: Kimi K3 hasn't taken the global crown, but it is the most powerful open model ever delivered, and it now tops a blind benchmark on frontend coding, ahead of Anthropic's best model. The particular nature of this last test deserves emphasis: developers compare outputs without knowing which model produced them, which removes the brand effect.
Moonshot presents K3 as an open model, but at the time of the announcement, the weights were not yet downloadable. The company has scheduled their release for 27 July. Until those files are available, K3 remains technically a model accessible only via API, exactly like a closed model. The promise is dated, but a promise with a date on the calendar is not yet a file on your hard drive. The real test will play out on 27 July.
The price, and a real flaw
On pricing, K3 is billed at $3 per million input tokens, $0.30 if the content is already cached, and $15 for output. That's the highest rate charged by a Chinese lab, but it still works out to roughly twice as cheap per task as Claude Opus 4.8.
However, a flaw noted by independent testers should be flagged. The model runs permanently in reasoning mode, with only one effort level available: maximum. As a result, it consumes a huge number of reasoning tokens, up to more than 13,000 tokens to generate a simple vector image, or about 25 cents for a trivial request. For heavy usage, this appetite can cancel out the advertised price advantage. A model that's cheaper per token isn't necessarily cheaper per task, a distinction companies are learning to make.
What this release says
Beyond the numbers, Kimi K3 confirms a trend we've been tracking for months with GLM-5.2 and then LongCat-2.0: Chinese labs are making openness a deliberate strategy in response to US restrictions. If the weights are indeed released on 27 July and the results hold up, K3 will become the most capable freely available model in the world, usable by any developer or company without depending on any API. That is exactly the sovereignty issue we detailed in our article on open models.
A tasty geopolitical detail: the lab ran part of its optimisation tests on Nvidia H200 cards as well as on a processor from an alternative supplier it declined to name. Nvidia has in fact begun limited H200 shipments to China after the US Commerce Department approved around ten Chinese companies. The question US policymakers are now asking is less whether export controls work, than whether they can still slow anything down.
Kimi K3 hasn't dethroned the summit. But it has done something more lasting: it has narrowed the distance. Where the gap between open and closed was measured in generations, it now counts in weeks of lag on a few specific benchmarks. That compression is the real news this week, far more than the 2.8-trillion figure.