Key Takeaways
- Anthropic's watermarking is compliance theater — technically real but practically trivial to defeat
- The EU AI Act forced this feature, not user demand or safety research
- Code generation escapes meaningful watermarking because functioning code admits fewer arbitrary choices
- Users canceling subscriptions over this reveal they were buying plausible deniability, not just text generation
Anthropic published its watermarking details Friday. The blog post reads like a company explaining why the seatbelt it was forced to install cannot actually be buckled.
The EU AI Act's Transparency Code requires AI companies to make generated content identifiable. Anthropic chose Google DeepMind's SynthID-Text, a statistical watermark that biases low-stakes word choices — "overcast" versus "grey" — in patterns invisible to readers but detectable with a key. The company will release a detection API. It insists quality remains untouched. To a human, watermarked and unwatermarked text are indistinguishable.
That last claim is the tell.
Statistical watermarking only works where the model has genuine freedom to choose between equally valid tokens. Creative writing offers that freedom. Code does not. A function either compiles or it doesn't. A variable name either follows convention or it confuses maintainers. Anthropic admits as much: code gets "negligible" watermarking, mostly in comments. The detectability of edited text depends on "how heavily Claude has edited it." Light editing leaves "very little for the watermark to attach to."
In other words, the system works best where it matters least — bespoke prose nobody disputes — and fails where it matters most — code, reports, emails, the daily output of knowledge work.
Anthropic acknowledges the escape hatch. Light editing "probably won't remove the watermark completely." A complete rewrite will. Then it adds a philosophical shrug: "In the latter case, of course, it's arguable whether the text can any longer be described as AI-generated."
That sentence should alarm regulators more than users. The EU wrote a law demanding identification of AI content. The leading implementation admits that any determined actor can strip the identification with a rewrite — and then questions whether the result was ever AI content at all. The regulation assumes a binary: human or machine. The technology delivers a spectrum: watermarked, lightly edited, heavily edited, rewritten. The detection API will return a probability score. Courts will hate this.
Reddit and X erupted anyway. One Redditor called it a conspiracy against innocent users. Another insisted only liars would object. Business Insider reported dozens of cancelled subscriptions. The backlash is disproportionate to the technical reality — unless you understand what those subscribers were actually purchasing.
They were buying clean provenance. A consultant pasting Claude output into a client deck. A student submitting an essay. A developer shipping code with a "written by me" commit message. The watermark doesn't prevent any of this. It merely creates a forensic trail that a motivated investigator could follow. Most clients, professors, and managers will never run the detection API. The subscribers know this. They cancelled because the *possibility* of detection breaks the implicit bargain: here is text that passes as yours.
Anthropic knows this too. The blog post distinguishes its approach from companies like Pangram that hunt for linguistic tells — the "not X, it's Y" constructions that betray synthetic prose. Watermarking is "fundamentally different," Anthropic says. Yes. Pangram tries to catch cheaters. Anthropic builds a system cheaters can defeat, then publishes a blog post explaining how.
The SynthID-Text choice is revealing. Google published it in 2024. Anthropic adopted it rather than developing its own. Speed to compliance beat technical ambition. The detection API will arrive later — no date given. Until then, watermarked text enters the world undetectable to anyone without the key. Anthropic holds the key. It has not committed to sharing it with platforms, schools, or publishers.
That gap matters. A watermark only deters misuse if the verifier can check it. If Anthropic gates the detector behind an API it controls, the company becomes the sole arbiter of whether a given text bears its mark. That is not transparency. That is a service tier.
The EU wanted a label on AI content. It got a watermark that vanishes under editing, barely touches code, and requires a proprietary detector. Anthropic complied. The law is satisfied. The problem is not.
Users who cancelled understood the bargain better than the regulators did. They weren't paying for intelligence. They were paying for deniability. Anthropic just made deniability slightly more expensive — light editing now carries risk — but kept the bulk discount intact. Rewrite heavily and you're clean. The company said so itself.
The rest is theater. Sharp, well-written theater. But theater.