Imagine a hospital with 900 specialists. When you arrive at A&E with a sore ankle, you don't summon all 900 doctors around your bed. A triage agent assesses your case and sends you to the radiologist and the orthopaedic surgeon. You benefit from the expertise of a vast institution, but only two practitioners work on your file. That, in one image, is what a Mixture of Experts is.
If you follow AI model news, a strange figure keeps coming up. Kimi K3: 2.8 trillion parameters, 48 of them active. LongCat-2.0: 1.6 trillion in total, 48 billion active. Apple's AFM 3 Core Advanced: 20 billion, with 1 to 4 activated. Why this systematic double figure? Because almost all recent models rely on an architecture called Mixture of Experts. Here's how it works.
The problem it solves
Let's start from the basics. As we explained in our article on parameters, these are a model's internal settings. The more it has, the more knowledge and nuance it can in principle store.
The issue is that in a classic model, known as dense, all parameters participate in the computation for every word generated. A model with a trillion parameters would therefore have to put all its trillion settings to work for each token produced. That is astronomically expensive in compute, electricity and time. This constraint long capped model size: growing meant becoming unbearably slow and costly.
The solution: splitting the model into specialists
The idea behind a mixture of experts is to split part of the network into many sub-networks, called experts, and only activate a few of them for each token processed.
The mechanism relies on two components. First, the experts: blocks of specialised parameters. They don't map onto neat human domains, that would be too good to be true: they aren't experts in law or cooking, but blocks that differentiated during training by specialising in statistical patterns whose logic isn't always interpretable.
Then the router (gating network), a small layer whose sole job is to decide, for each token, which experts to summon. It's the triage agent in our hospital. It looks at the incoming information, picks the two, four or eight most relevant experts, and hands them the work. The others stay inactive and consume no compute.
For Kimi K3, the figures are telling: the model only activates 16 of its 896 internal experts per token, roughly 1.8% of its total capacity at any given moment.
An MoE model offers the knowledge of a giant model for the compute cost of a mid-sized one. It's this trade-off that unlocked the race towards models with several trillion parameters. Without this architecture, a 2.8-trillion-parameter model would probably be commercially unviable. MoE is the technical reason models have been able to grow so fast over the past two years without prices exploding proportionally.
The trade-offs, because there are some
This elegance comes at a price, and it's worth knowing.
Memory remains a problem. Here's the trap many miss: even if only 2% of parameters work, all of them generally need to be available, since you don't know in advance which ones the router will call. Compute is saved, not storage. That's exactly why an open 2.8-trillion-parameter model remains out of reach for a personal computer, as we noted in our guide to local AI. Apple's innovation of storing the model in flash memory is precisely aimed at working around this limit.
The router can get it wrong. If the triage agent sends your ankle to the cardiologist, the response will be poor. Bad routing produces inconsistent results, and load balancing between experts is a delicate training problem: some experts risk being over-used and others never learning anything.
Training is more complex. You have to teach the experts to specialise and the router to route well simultaneously, which is more unstable than training a classic dense model.
The reflex to adopt when reading a spec sheet
The practical takeaway is simple. When you read that a model has trillions of parameters, immediately look for the second figure: how many are active. That's the one that tells you about compute cost and speed, while the total tells you about the breadth of stored knowledge and memory requirements.
Marketing releases naturally lead with the most flattering figure, the total. Knowing that these two numbers tell two different stories is enough to read a model announcement with a far sharper eye. And next time someone touts a three-trillion-parameter model, you'll know the real question isn't the size of the hospital, but how many doctors are actually looking at your case.