Skip to content

Claude's watermark bypassed in four hours

Anthropic announces watermarking for its text. Code to remove it circulates the same day. That gap tells you everything about the detection problem.

Advertisement
The essentials in 30 seconds ⚡
Anthropic has announced a watermarking system to identify texts produced by Claude. According to observers, code capable of bypassing this marking circulated about four hours after the announcement. The company also acknowledges that its watermark can only indicate a probability that a text comes from its model, not a certainty.

We covered the entry into force of the California law requiring the marking of generated images, videos and sounds. We had noted a detail: text was excluded from it. Here is why.

Why watermarking text is far harder

An image contains a huge amount of redundant information. You can hide a signal in it by imperceptibly altering pixels, without the eye seeing anything and without changing the content. That is the principle we explained in our article on C2PA.

Text does not have that margin. Every word counts, and changing a word changes the meaning. The available techniques therefore proceed differently: they subtly steer the model's choices toward certain words rather than their synonyms, according to a detectable statistical pattern.

Hence a structural consequence: the signal is statistical. On a long text, the pattern is measurable. On a short text, it drowns in the noise. And it can never give more than a probability, which Anthropic explicitly acknowledges.

Why bypassing is easy

Three mundane operations are enough to degrade the signal.

Rewording. Running the text through another model, or simply rewriting it, destroys the statistical pattern.

Substituting. A tool that automatically replaces certain words with synonyms breaks the regularity being sought.

Mixing. A text that is partly human and partly generated dilutes the signal until it becomes undetectable.

None of these operations requires any particular skill. That is what explains the four-hour delay: it was not about breaking encryption, but about automating known manipulations.

The danger is not bypassing, it is the false positive ⚠️
A system that gives a probability will inevitably be used as proof. A teacher, a recruiter or an employer who gets a high score will conclude that the text is generated. Yet false positives hit first those who write in a highly structured way, those for whom it is not the native language, and those who use a restricted vocabulary. That is exactly the problem we flagged regarding cheating detection at school: accusing an innocent person costs more than letting a cheater through.

What this episode reveals

The lesson goes beyond this specific case, and it is consistent with what we have been writing for months.

Detection is a losing race. Every advance in watermarking invites an advance in bypassing, with a structural advantage to the latter: you only need to degrade a signal, not reproduce it.

Provenance is more robust than detection. That is why the C2PA standard tackles the problem the other way around: rather than unmasking the fake, sign the authentic. An unsigned text proves nothing, but a signed text proves something. The asymmetry finally works in the right direction.

Voluntary commitments are fragile. This episode lands on the same day as OpenAI's announcement of slowing its own development, and it illustrates the same difficulty from another angle: a measure announced without a verification mechanism holds up as well as the technology allows, which is sometimes four hours.

What to take away

One should not conclude that the effort was useless. A watermark discourages careless use, helps spot some volume-generated content, and serves as one signal among others.

What it should not be used as is proof. Neither to accuse someone, nor to reassure oneself. An unmarked text can be generated, a marked text can be a false positive, and both errors have consequences.

We always come back to the same point: in a world where producing costs nothing anymore, what remains verifiable is not the content but the chain. Who produced it, with what, and who is willing to answer for it.

Advertisement