taraonsecurity

Anthropic Watermarked Claude, Then Told You How to Wash It Out

Anthropic Watermarked Claude, Then Told You How to Wash It Out

Every Claude model Anthropic has shipped since 2 August 2026 writes a watermark into its own text. You cannot see it. There is no badge and there are no hidden characters, because the thing lives in the statistical pattern of the word choices themselves, with signed C2PA metadata bolted onto any files it generates. It is Anthropic’s build of Google’s SynthID-Text, switched on to satisfy the EU AI Act. Hold the key and you can run a passage and get back a probability that Claude wrote it.

https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb943bbb5-5e37-4ee5-bc5b-c7a1a5357ae5_1024x606.jpeg

I have feelings about this, and most of them are amusement, because I published a piece about a week before the announcement saying that any watermark a detector can read is one an attacker can eventually remove. The logic is not exciting. Detection and removal act on the same object, so the moment something can tell you the mark is still there, you have a feedback loop. Change the text a little, check whether the score dropped, change it again. With a text watermark the easiest change is just rewriting. Paraphrase it, run it through another language and back, feed it to a different model, whatever gets the words to move.

I did not even have to make the argument, because Anthropic made it for me. Their own FAQ says light editing probably will not fully remove the mark, but a full rewrite where every word gets replaced will. One of their engineers put it more bluntly, that it is not perfect, you can edit it, it is a first step. So that is my article’s entire point, admitted by the vendor, on launch day.

None of this means the sky is falling, and it does not mean watermarking is fake or useless either. The real picture is just less dramatic than either take. The mark is genuinely there and it survives ordinary use, so copy the text out or fix a typo and it stays put, and on anything long it takes real work to strip. It is a provenance signal and not much more than that, a way to ask whether a machine had a hand in something. The problems start when people treat it as proof. The more weight you put on it, the worse it goes the day someone rewrites their way straight past it. A removable watermark is not an access control, even though people will keep wiring it in like one.

The detail that actually got me is buried in the coverage. Anthropic went out of its way to say the watermark is a different thing from the AI tells that detectors like Pangram look for, the little stylistic giveaways that scream language model. The example they picked was the “this isn’t X, it’s Y” construction. I have spent a genuinely stupid amount of time this month pulling that exact phrasing out of drafts, and here it is showing up in a watermarking FAQ as the thing the watermark specifically is not. So there are two separate layers running at once. The obvious stylistic tells, which you can scrub, and the statistical mark sitting underneath in the token choices, which rides along whether or not the surface reads human.

Which is the part anyone building on this should actually care about. You can take a Claude draft, strip out every obvious fingerprint, edit it until it sounds completely like you, and the statistical mark can still be sitting there the whole time. It only really goes once you have rewritten it deeply enough. Both layers come off in the end, they just cost different amounts of effort, and if your plan depends on “the text is marked so we know where it came from,” one motivated person with a rewrite habit breaks that plan.

Anyway. None of this changes much on my end, I will keep writing and you will keep reading. It only becomes a problem if you were relying on the mark to be more than a hint, which, going by Anthropic’s own FAQ, it was never really going to be.

How do you rate this article?

3


Tara reads security
Tara reads security

Cybersecurity researcher


taraonsecurity
taraonsecurity

Security researcher writing about web application, cloud, and smart contract security. Notes on vulnerabilities I find interesting and how they get fixed.

Publish0x

Send a $0.01 microtip in crypto to the author, and earn yourself as you read!

20% to author / 80% to me.
We pay the tips from our rewards pool.

Page not displaying correctly?