Photocopy a photo. Then photocopy the photocopy. Then repeat fifty times. In the end, only a grey smudge remains. "Model collapse" is exactly that, but with intelligence: an AI that learns from texts produced by other AIs, generation after generation, until it loses touch with reality.
Here is a paradox nobody anticipated five years ago. Artificial intelligences learn by reading gigantic quantities of texts scraped from the internet. That is their food. But since 2023, the internet has been filling up at full speed with texts written by AIs: articles, product descriptions, comments, automatically generated responses. The food of future AIs will therefore be, increasingly, produced by current AIs.
What happens when an intelligence starts feeding on its own outputs, in a loop? Researchers have a name for this scenario, and a genuine concern.
What is "model collapse"?
The term model collapse refers to the progressive degradation of an AI that repeatedly trains on data generated by other AIs, rather than on content produced by humans. With each generation, the model loses a little of the richness, diversity and accuracy of the real.
Why? Because an AI, by nature, produces average, smoothed, probable responses. It avoids rare cases, original formulations, exceptions. If the next generation learns only from these smoothed outputs, it smooths even further. Extremes disappear, diversity shrinks, and after a few cycles, the model produces nothing but a homogeneous, impoverished mush. The photocopy of a photocopy.
No longer merely theoretical
A recent study conducted by the company Graphite has revived the debate. Its simulations suggest that AI-boosted search systems could gradually become less diverse if they rely increasingly on content that is itself AI-generated. In the experiment, repeatedly retrieving texts written by AIs made responses converge towards increasingly similar recommendations, like an echo that both reinforces and impoverishes itself.
Caution is warranted: the study does not prove that commercial search engines already suffer from this ailment. It points to a risk, not a diagnosis. But the signal is being taken seriously, because it touches on a very real mechanism: the more the web fills with synthetic content, the harder it becomes for an AI to find authentically human material to learn from.
Unexpected consequence: data produced by humans before the generative AI era is becoming a valuable resource, almost like oil. Some are already talking about "pre-2023 data" as a rare deposit, the digital equivalent of pre-nuclear-testing steel (sought after because uncontaminated by radioactivity). Your old blog, your archives, your authentic texts may be worth more than you think.
The philosophical question
Here, the subject goes beyond the technical and becomes dizzying. Model collapse raises a question that applies to machines as much as to humans: can you learn something new by feeding only on yourself?
Think of a culture that only ever quotes itself, never encountering foreign ideas. Think of a mind that only ever reads its own thoughts, over and over. Intuition says it would grow poorer, go round in circles, gradually lose touch with what lies outside it. The real, otherness, surprise: that is what stops an intelligence from closing in on itself.
AIs hold this truth up to us like a mirror. They can only remain intelligent by staying connected to something that exceeds them, namely the immense creative disorder of real human experience. An AI cut off from the human real, running only between AIs, would eventually go intellectually dark. Like us, perhaps.
Is there a way out?
Labs are aware of the risk and are developing countermeasures. Carefully filtering training data to weed out low-quality synthetic content. Tagging AI-generated texts so they can be distinguished. Keeping a guaranteed proportion of authentic human data. And above all, using synthetic data that is high-quality, carefully controlled rather than scooping up anything and everything from the web.
Because the nuance matters: not all AI data is poison. A well-designed, verified, targeted synthetic data point can be useful. The danger is uncontrolled accumulation, the blind harvesting of mediocre AI content. The difference between a dietary supplement and junk food.
Model collapse is therefore not a fatality, but a warning. It reminds us that intelligence, artificial or human, is a conversation with the world, not a monologue. The day it stops listening to what lies outside it, it begins to fade. There is something strangely reassuring in that: even the most powerful machines need us, our disorder, our reality, in order to keep thinking.