A lossless audio file reproduces every nuance of a recording and weighs a lot. An MP3 discards information the human ear barely perceives, and cuts the size by ten. Quantization does the same thing with an AI model: it reduces the precision of each parameter to save a huge amount of space, betting that the loss will be imperceptible in the result.
This is the technique that explains how a model designed to run on servers can work on a personal computer, and why Kimi K3's weights come to 1.4 terabytes rather than far more. Let's break it down.
The principle
A model is made up of billions of parameters, which are numbers. Each one has to be stored, and the question is: with what precision?
During training, high precision is typically used, with each parameter taking up 16 or 32 bits. That allows very fine adjustments, which are essential while learning.
But once the model is trained, that precision becomes largely unnecessary for using it. Quantization means rewriting each parameter with fewer bits: 8, 4, sometimes fewer. Going from 16 to 4 bits mechanically divides the size by four.
An image: instead of writing a measurement as 3.14159265, you write 3.14. You lose precision, you gain a lot of space, and for most uses it changes nothing.
Why it works so well
The legitimate question is: how can a model stay performant while losing precision on each of its parameters?
The answer lies in redundancy. A model does not depend on the exact value of any single parameter: its knowledge is spread across countless values, as we explained in our piece on interpretability. A small error on one parameter is absorbed by the billions of others. The system is robust to noise, because it was trained in noise.
8 bits: the loss is usually negligible, including on demanding tasks. This has become a deployment standard.
4 bits: degradation remains slight on everyday uses, and this is the format that makes local AI genuinely accessible. It's the best trade-off for most people.
Below that: the effects become noticeable, especially on complex reasoning and long tasks. The model remains usable for simple purposes, but starts to lose coherence.
One important point: degradation is not uniform. An aggressively quantized model often keeps its everyday conversational abilities while losing ground on tasks that require sustaining a chain of reasoning.
What this enables in practice
Three uses follow directly from it.
Local AI. This is the most visible one for individuals. Without quantization, running a decent model on a personal computer would be out of reach. With it, as we detailed in our guide, a 7 to 9 billion parameter model fits into a few gigabytes of video memory.
Service cost. For a provider, a smaller model uses less memory and compute per request. At the scale of billions of requests, that's a major economic lever, and it contributes to the price drops we're seeing.
Embedded devices. Phones, cars, connected objects: quantization is the prerequisite for a model to run on a device with limited resources. It's one of the ingredients in Apple's approach to on-device AI.
What to take away
Quantization is one of those quiet optimizations that does more for the real-world spread of a technology than many flashy announcements. It doesn't make any model smarter: it makes existing models accessible to machines and budgets that were previously shut out.
The useful habit, if you download an open model, is to look at the format on offer. The same model usually exists in several quantized versions, and picking the right one means weighing what your hardware can load against the quality your use demands. For most people, the 4-bit version is the sweet spot, and there's no shame in using it: it's largely what the services you pay for do too.