Skip to content

Why nobody really knows what's going on inside an AI

These are not industrial secrets. Even the teams building these models cannot explain why a specific answer was produced. Here is why.

Advertisement
One confusion to clear up right away 🔍
When people say an AI is a black box, many take this to mean companies refuse to reveal their secrets. That is a misunderstanding. The code is known, the architecture is published, and for open models, the parameters can be downloaded by anyone. The opacity is not commercial, it is structural: even with everything in front of you, you cannot explain why a specific response was produced.

This is one of the most baffling characteristics of this technology, and it sits at the heart of several debates we have covered, from alignment to legal liability. Let us take the time to understand where this opacity comes from.

The problem stems from how it works

Classic software is written by humans as explicit rules. If something does not work, you reread the code, find the faulty line, and fix it.

A language model is not written that way. It is trained. Humans define an architecture and a learning process, then the system adjusts billions of numerical values on its own until it predicts its data well. No one wrote those values, no one chose them individually.

The result is billions of numbers that, together, produce intelligent behaviour. Taken in isolation, each one means nothing. There is no parameter that contains the notion of a cat, nor a line to fix when the model gets geography wrong.

Three reasons that make analysis difficult

Scale. Examining billions of values and their interactions exceeds what a human can inspect. This is not a motivation problem, it is a volume problem.

Distribution. A concept is not stored in one specific place. It is spread across countless parameters, and the same parameter contributes to thousands of different concepts. This is called superposition. It is efficient, and it is a nightmare for anyone trying to untangle things.

The lack of an internal vocabulary. The model does not think in words. Its internal representations are positions in a mathematical space, as we explained regarding embeddings. Translating those positions into human concepts is precisely the whole task, and it is not solved.

Beware the trap: the explanation the model gives 🎭
If you ask an AI why it answered something, it will give you a coherent and plausible explanation. The problem is that this explanation is itself a product of the model, not a report of introspection. It describes what would constitute a credible justification, not necessarily what actually happened in the computation. This is an essential distinction, and it partly applies to humans too, who often reconstruct their reasons after the fact.

What researchers can still manage to do

The picture is not entirely bleak, and the field of interpretability is genuinely progressing.

Teams are managing to identify features inside models: internal patterns that activate reliably for specific concepts. We can also trace certain circuits, that is, chains of processing that carry out an identifiable operation. And intervention is possible: amplifying or suppressing a feature to observe how behaviour changes, which provides causal evidence rather than mere correlation.

These advances are real but partial. They shed light on fragments of a system whose bulk remains unexplored.

Why this limitation matters in practice

Three practical consequences follow from this opacity.

You cannot guarantee behaviour. You can test extensively and observe that a model behaves well on everything you have tried. You cannot prove it will behave well on what you have not tried. This is exactly what episodes where models break out of their test environment reveal: the problematic behaviour only appears in a situation no one anticipated.

You fix things at the surface. When a model produces undesirable behaviour, you do not repair the cause, you add training to make that behaviour less likely. It is effective and it is frustrating, like treating a symptom without understanding the disease.

Responsibility becomes blurred. How do you establish fault when no one can explain the mechanism behind a decision? Our legal systems presuppose an explainable causal chain.

What to take away

We use systems every day whose internal workings elude even those who build them. Put that way, it sounds alarming. Yet it is worth noting that this is not so new: we have long prescribed drugs whose precise mechanism remains debated, and we drive cars about whose engines we know nothing.

The difference here lies in the nature of the object. A drug always does the same thing. A model produces decisions in infinitely varied situations, including situations no one has foreseen. That is what makes interpretability not an academic curiosity, but probably the most important research endeavour for keeping this technology manageable. And it is also the best reason never to grant these systems a trust we could not justify.

Advertisement