News

Anthropic's Semantic Watermarking: Claude Will Pick Worse Words on Purpose

Anthropic's plan to watermark Claude text by biasing token selection toward 'green' words is a deliberate degradation of output quality — and it applies to every response over 200 tokens, even private chats.

August 17, 2026· 3 min read· Source: Daring Fireball
Anthropic's Semantic Watermarking: Claude Will Pick Worse Words on Purpose

Anthropic announced that all Claude models will soon embed a semantic watermark in generated text, complying with an EU transparency regulation. The original support document claimed the watermark would be "imperceptible" and would not change "meaning, quality, or readability." That claim is now revealed as misleading: the actual mechanism biases token selection at inference time, subtly altering word choice in statistically detectable ways.

The technique, as explained in Anthropic's follow-up post and James Padolsey's interactive essay, works like a biased coin. At each token-generation step, the model is slightly more likely to pick a word from a secret "green" list than a "red" list. The lists are computed on the fly with a secret key, so there's no static vocabulary of preferred words. Detection requires the key — only Anthropic can verify its own watermark, and no provider can detect another's.

The critical detail: this applies to all Claude output longer than 200 tokens (~150 words), including private conversations that no one but the user will ever read. The model will sacrifice precision and clarity — even if only slightly — to embed a provenance signal that serves Anthropic's regulatory compliance, not the user's interests.

The Core Problem: No Two Synonyms Are Equal

"He leaped at the chance" and "He jumped at the opportunity" are not the same sentence. The exact word choice matters in writing. When a model is nudged toward a "green" list, it may pick a word that is slightly less precise, slightly less idiomatic, or slightly off in tone — all in service of a watermark that the user never asked for and may never benefit from.

Anthropic's framing treats this as a harmless trade-off, but it's a trade-off made on the user's behalf, without consent, for a purpose that doesn't improve the output. For a tool that is already criticized for producing bland, generic prose, deliberately making it worse — even marginally — is a step in the wrong direction.

The EU Regulation Angle

The regulation motivating this, the EU's Code of Practice on Transparency of AI-Generated Content, is a bureaucratic attempt to address AI provenance. But the implementation shifts the cost onto users: degraded output quality, with no direct benefit to them. The watermark is probabilistic, so it's not even a reliable detector for short texts. For a 150-word email, the signal is too weak to be confident. So the user pays the quality tax, and the watermark still fails its purpose for the very texts where it might matter most.

Anthropic is saying they're going to make their models' output worse, on purpose, for purposes that do not benefit the user in any way.
Manul X Editorial