Anthropic Explains How Claude Text Watermarks Will Operate

Anthropic has outlined how it intends to watermark text produced by Claude, saying it will use the SynthID-Text approach described by Google DeepMind in 2024 and plans to offer a watermark detection API. The company is introducing the system to comply with the European Union AI Act’s Transparency Code, which calls for methods that make AI-generated content identifiable.
The clarification follows debate among Claude users after Anthropic announced its watermarking plans. The company said the mechanism does not alter the quality of Claude’s output and that a reader cannot distinguish a watermarked response from an unwatermarked one.
Watermarks rely on small language choices
Anthropic said Claude can encode a pattern when it makes what it calls low-stakes choices in a response. For example, a model might have several equally suitable words available to describe the weather. Those choices can form a pattern that is imperceptible to a reader but detectable by a party holding the relevant encoding key.
The method is different from AI-detection products that infer AI use from stylistic signals in writing. Anthropic contrasted watermark checking with approaches that search for recurring phrasing or other textual “tells”. In its description, a watermark is deliberately embedded during generation, rather than inferred from the apparent style of finished prose.
Editing, human authorship and code
Anthropic said a user may be able to obscure a watermark through rewriting, but light editing will probably not remove it completely. A complete rewrite in which every word is replaced will remove the signal. The company noted that, in that case, it is arguable whether the resulting text should still be considered AI-generated.
For material written by a person and then proofread or edited with Claude, detectability will depend on both the length of the text and the extent of Claude’s changes. If edits are light, nearly all of the words remain human-authored, leaving little or nothing to which a watermark can attach.
Watermarking should be less pronounced in code because functional code gives the model fewer interchangeable choices than ordinary prose. Anthropic said the technique can still be applied where arbitrary wording is possible, including comments, but will have a negligible effect on the actual code produced.
Compliance context for organisations
Anthropic said other major model developers that signed the same Code of Practice will implement their own watermarks. The company’s initial announcement on Anthropic’s Claude watermarking commitment set out the Claude watermarking commitment, while the new explanation provides more detail on its technical limits and planned detection capability.
For businesses, the practical implication is to distinguish embedded watermark verification from style-based AI detection and to account for editing intensity, human contribution and code-specific limits when defining rules for AI-assisted content.

