Skip to content

How Apple fits a 20-billion-parameter model into an iPhone

The model is stored in flash memory, and a small predictor reads your request to decide which parts to wake up. A trick that could change the future of local AI.

Advertisement
The starting problem 📱
Running a powerful AI on a phone hits a physical wall: RAM. A model normally has to fit entirely in active memory to work, which drastically limits its size on a consumer device. Apple claims to have sidestepped this wall with an elegant trick, and the result deserves explaining, because it concerns the future of all on-device AI.

We noted, in our article on smart glasses, that Apple seemed behind on consumer AI. That's true on products. It's far less so on research, as shown by the third generation of its foundation models, unveiled in June 2026 and expected on devices in the autumn. Here's the technical idea that makes it interesting.

The lineup, in brief

Apple has unveiled a family of five models, split between device and server. Two run locally: AFM 3 Core, a dense model of around 3 billion parameters for everyday tasks, and AFM 3 Core Advanced, the more powerful of the two. Three others run on Private Cloud Compute servers, including an image model and an advanced reasoning model.

One point deserves clarification because it has been misreported. Apple signed a multi-year deal with Google in January 2026, and a joint statement at the time described the future AFM family as built in collaboration with Google and based on its Gemini technology. But Apple's technical presentation insists on its own architecture, and its head of AI clarified that the approach was based on distillation, not on adopting Gemini as-is. In other words, Apple's models learned by observing Google's, a technique we explained in our article on distillation, but they are not rebadged Gemini models.

The trick: the model sleeps in flash storage

Now to the genuinely interesting part. AFM 3 Core Advanced has around 20 billion parameters in total, but only activates 1 to 4 billion per request.

The problem this solves is concrete. Normally, a model has to be fully loaded into RAM, the DRAM, which creates a huge footprint and limits model size on consumer hardware. Apple's solution: store the full model in flash storage (where your photos and apps live, far larger and far cheaper), and load only the pieces needed for the current request into RAM.

The question then becomes: how do you know in advance which pieces to wake up? That's where a technique developed by Apple's researchers comes in, instruction-following pruning. Rather than deciding once and for all, at training time, which parts of the model are active, a small predictor reads your request and dynamically chooses, for that specific query, which portions of the network to activate.

The analogy that clarifies 📚
Imagine a vast library and a tiny desk. You can't fit all the books on the desk, it's too small. So a librarian reads your question, rushes off to fetch the three or four relevant volumes, and brings them to you. The library is flash storage. The desk is RAM. The librarian is the predictor. You get access to all the knowledge without needing a gigantic desk.

The result published by Apple's researchers is striking: their model with 3 billion activated parameters beat a dense baseline of the same size by 5 to 8 points on maths and code, while matching the performance of a dense 9-billion model. In other words, you get the capabilities of a model three times larger, for the memory cost of the small one.

Why it matters beyond Apple

This approach directly addresses the limit we described in our guide to local AI: available memory determines the size of the model you can run at home. If you can store a large model in storage and wake only a fraction of it, that constraint loosens considerably.

The consequences are concrete. A more capable model running on your device means AI that works without a connection, that lets no data out, and that costs nothing to use. Apple is indeed making this its central argument, reaffirming that it does not use personal data to train its models, and that requests are handled either locally or via infrastructure where the information remains inaccessible, including to Apple itself.

A detail that raised eyebrows among observers: for its most powerful model, Apple worked with Google and NVIDIA to extend Private Cloud Compute to NVIDIA GPUs hosted in Google Cloud, while claiming to maintain the same privacy guarantees. A company that built its pitch on keeping everything in-house is therefore accepting that part of its AI runs at its main competitor's. It's an acknowledged compromise, backed by technical verification mechanisms, but it's a compromise.

What to take away

Apple hasn't won the race for the most powerful models, and that clearly wasn't its goal. What it's pursuing is different: the best possible AI within the constraints of a device you carry in your pocket, without your data leaving it. That's an engineering problem more than a scaling problem, and instruction-following pruning is an elegant answer to it.

The question that remains open is product execution. The models are due to arrive in autumn 2026 with system updates, on the newest Macs and iPhones. Apple has a complicated history with AI promises, particularly around Siri, as we touched on regarding the Gemini deal for the new assistant. An elegant architecture on paper only counts for what it delivers in people's hands. See you in the autumn.

Advertisement