Skip to content

What is the "context window"? The working memory that decides what your AI can actually understand

People talk about 1 million, 2 million tokens of context. But what exactly is it? And why does an AI sometimes "forget" the start of your conversation?

Advertisement
The starting analogy 🧠
Imagine you're reading a novel, but you can only keep the last 30 pages in memory. By chapter 20, you've forgotten the murder from chapter 1. That's exactly the limit of an AI: its "context window" is the number of pages it can hold in mind at once. Anything beyond that window is, literally, out of its sight.

If you follow AI news, you've inevitably come across these figures: "1 million tokens of context", "2 million, the largest on the market". They come up with every model release, brandished as selling points. But what do they actually mean? And why has this become a battleground between labs in 2026? Let's break it down.

Simple definition

The context window is the total amount of information an AI can take into account at once to produce its response. It includes everything: your question, the documents you've pasted, the conversation history, and even the response being drafted.

It's measured in tokens, those word fragments that are the basic unit of everything an AI reads (if the term is unfamiliar, we've explained it in detail in our article on tokens). In French, 1 token is worth roughly 0.7 words. So a 1 million token window corresponds to about 700,000 words, or roughly 1,500 book pages that the AI can "have in front of it" simultaneously.

The difference between memory and context

Here's a point many people confuse. The context window is not the AI's long-term memory. It's its working memory, its immediate memory.

Let's go back to the brain analogy. Your accumulated knowledge (your language, your memories, what you learned at school) is your long-term memory. For an AI, that's the equivalent of its training, fixed in its parameters. The context window, on the other hand, is what you actively hold in mind at the present moment, like a phone number you just heard. Limited, temporary, but immediately available.

That's why an AI can "forget" the start of a very long conversation: if the exchange exceeds its context window, the first messages fall out of its view, like the first pages of the novel. It hasn't "forgotten" them in the human sense; it simply no longer sees them.

Why it became a race in 2026

Context windows have exploded. We've gone from a few thousand tokens two years ago to 1 million on most cutting-edge models, and up to 2 million claimed by the newest ones. Why this one-upmanship?

Because a large context unlocks entirely new use cases. With 2 million tokens, an AI can analyse an entire codebase in one go, read a full book and answer precise questions about it, or process hundreds of legal documents in a single request. For autonomous agents (those AIs that work on long tasks without supervision), a large context is vital: it lets them stay on track over hours of work without losing information from the start.

The trap of the large context 🎯
Beware of a misconception: a larger context doesn't mean a smarter AI. Studies show that models tend to retain the beginning and end of a long context better, and to "lose" information from the middle, a phenomenon dubbed "lost in the middle". Giving an AI 2 million tokens doesn't guarantee it uses them all correctly. The window size is a maximum capacity, not a promise of performance.

Why a large context is expensive

There's a technical reason you can't simply put infinite context everywhere. An AI's computational cost rises sharply with context size. Each additional token must be "compared" with all the others, which drives up the bill more than proportionally.

Concretely, the more context you give an AI, the more each request costs and the longer it takes. That's why models charge per token and why companies are careful: sending a huge context with every call can send costs through the roof, as several large groups we mentioned in our article on the cost of AI in business discovered. A large context is a powerful tool, but it comes at a price.

What to remember

The context window is one of the most useful concepts for understanding what your AI can and cannot do. It's its working memory: large, but limited and temporary. The bigger it is, the more material the AI can take in at once, which opens up spectacular use cases. But a large context is expensive, doesn't guarantee perfect use of the information, and doesn't replace the model's intelligence.

Next time a lab sells you its "2 million tokens", you'll know how to read between the lines: it's an impressive capability, useful for certain specific use cases, but it's just one dimension among others. A big working memory doesn't make a great mind, for AIs or for us.

Advertisement