Skip to content

What is data poisoning? When a few hundred texts are enough to trap a model.

Artists and authors are slipping elements into their work that are invisible to the eye but toxic to an AI. The technique works, and it worries as much as it interests.

Advertisement
The idea in one image 🧪
Imagine giving away a cookbook for free, but a competitor photocopies it without permission to resell it. You can't legally stop them, so you slip into your recipes instructions that a human cook would instinctively correct, but a machine would blindly reproduce. Data poisoning is exactly that, applied to training corpora.

We mentioned this morning the rush by labs on printed books from before 2022, and one of the selling points was that these works escape this technique. It deserves an explanation, because it has become a real issue.

The principle

Data poisoning involves introducing into a dataset elements specifically designed to degrade the behaviour of a model trained on it.

The key point, and what makes the technique formidable, is the asymmetry of perception. These elements are designed to be imperceptible to a human and significant to a machine. A reader notices nothing; a model absorbs misleading information.

For images, tools developed in academic settings allow artists to imperceptibly alter their works before publication, so that a model trained on them learns erroneous associations. For text, the principle is comparable, playing on characters or structures that the eye automatically corrects.

Why such small quantities suffice

This is the most surprising result, and it goes against intuition. You'd imagine you'd need to poison a significant fraction of a corpus to have an effect. Research suggests that a very small number of carefully crafted documents, on the order of a few hundred, can be enough to install unwanted behaviour in a corpus containing trillions of words.

The reason lies in how learning itself works. A model doesn't just average: it looks for patterns. A rare but highly consistent pattern, repeated identically in a small number of documents, can be learned as a rule, precisely because it resembles nothing else. It's a bit like a password: its strength comes not from its length relative to the dictionary, but from its uniqueness.

The backdoor, the most concerning case 🚪
The most studied form is the backdoor. The model behaves normally in all circumstances, except when a specific trigger appears in the query. That trigger can be an innocuous character sequence that no one would type by chance. The danger is obvious: the problematic behaviour remains invisible during all testing phases, since it only manifests in the presence of a key that only the author knows.

Two uses, two very different judgements

This is where the topic gets interesting, because the same technique serves two opposing causes.

As a defensive tool. Artists and authors use it to protect works they necessarily publish online. The argument is self-defence: copyright exists but proves hard to enforce against massive scraping, so you make the goods less appetising. It's a technical response to a legal failure, and many see it as a form of restored leverage.

As an offensive weapon. The same method allows a malicious actor to place a backdoor in a model that will then be deployed in production. This ties into the security concerns we've documented, particularly around agentjacking: in both cases, you exploit the fact that a system doesn't natively distinguish between what it should believe and what it should execute.

Judging the technique in the abstract therefore doesn't make much sense. A knife is neither good nor bad; what matters is who wields it and against what.

The limits, on the defenders' side

We have to be honest with creators tempted by this route: effectiveness isn't guaranteed over time.

Labs are developing detection filters, and a identified poisoning is simply removed from the corpus. Certain image transformations, like aggressive compression, can dampen the effect. And above all, the race is asymmetric: a few well-funded teams work on detection, against isolated creators using public tools whose signatures end up being known.

The most robust strategy therefore remains a combination: technical protection, explicit prohibition of use for training, and collective action through legal channels or licence negotiation.

What to take away

Data poisoning is a symptom, more than a solution. It appears where the balance of power is skewed: creators who can't prevent the use of their works equip themselves with a technical lever, for lack of an effective legal one.

Its existence also produces a side effect no one anticipated, and which we saw in action this morning: it makes data predating its invention more valuable. A book printed before these tools existed is clean by construction. The defence of some creates the value of what others seek to buy, and sometimes to destroy.

That may be the best illustration of the current state of the data debate: everyone is cobbling together their weapons while waiting for the law to rule, and no one knows how long this period will last.

Advertisement