Anthropic plans to add invisible watermarking to text generated by future Claude models, making it easier to identify AI-written content.
Claude text watermarking will launch worldwide because Anthropic does not yet have a reliable way to limit the feature by region. The move supports the EU AI Act and the EU’s Code of Practice for AI providers.
Watermarking will not add hidden characters
Anthropic says readers will not see a watermark in Claude’s responses. The system does not insert visible labels, hidden characters or extra tokens into generated text.
Instead, it changes how Claude makes some low-stakes word choices while producing a response.
Language models generate text one token at a time. Often, several possible words could fit naturally at the same point in a sentence. Claude text watermarking uses a secret key and the preceding context to influence the randomness behind some of those choices.
The final wording should still read normally. However, across a longer response, the choices create a statistical pattern that an authorised detector can recognise.
Anthropic says internal testing found no practical effect on the quality, creativity or readability of Claude’s output. It also expects a negligible impact on generation speed.
Claude uses a method based on SynthID-Text
Anthropic’s approach builds on Google DeepMind’s SynthID-Text research.
Rather than changing a completed response, the technique operates while the model generates it. A detection system can then analyse the sequence of word choices and assess whether the text matches the pattern expected from a Claude response created with the secret key.
The detector does not need to access Claude’s underlying model to carry out that assessment. However, it does need the relevant watermarking key.
The result is a probability estimate, not definitive proof of authorship. A watermark may indicate that Claude likely contributed to a piece of content, but it cannot determine whether Claude wrote the entire text or heavily edited a human draft.
Code and fixed factual answers receive less watermarking
Claude text watermarking does not apply equally to every type of output.
Anthropic says the system does not interfere where a response requires an exact answer. For example, a model completing “2 + 2 =” has one clearly correct next token.
The same limitation applies to much of code generation. Replacing a term in code with another plausible word could make the output fail, so Claude has fewer opportunities to apply a watermark.
Some code content, such as comments, may still contain watermarked choices. However, Anthropic says the system should not materially affect the code itself.
Watermark detection also works better when text contains more flexible language choices. Longer creative or conversational passages provide more evidence than short, highly factual answers.
Anthropic plans a detection API
Anthropic is developing an API that will estimate whether Claude likely contributed to a submitted text sample.
Detection may become less reliable if a person rewrites the content extensively or provides only a very short sample. Light proofreading may leave some detectable evidence, while a full rewrite that replaces every word can remove the watermark.
Translations generated by Claude can carry a watermark because the model selects each word in the translated output.
Anthropic will use a different method for AI-generated files. For PNG, JPG and SVG files, Claude will add cryptographically signed C2PA provenance metadata showing that it created or processed the file.


0 responses to “How Anthropic Will Watermark Claude AI-Generated Text”