Anthropic Unveils New Insights on the Functionality of Claude’s Watermark System

On Friday, Anthropic released a blog post to clarify how it plans to implement watermarking for text generated by its chatbot, Claude. Key questions addressed included the operational mechanics of watermarking, its potential for being concealed through editing, and implications for code generation.
The introduction of watermarking has sparked discussions among Claude users following the company’s announcement earlier in the week, which was made in line with the EU AI Act’s Transparency Code. This regulation mandates that AI organizations utilize methods for identifying AI-generated outputs.
In online discussions, some users on platforms like Reddit have voiced concerns, describing the move as a conspiracy against users. Others have suggested that there must be ulterior motives for opposing the watermarking, with reports indicating a notable number of cancellations from subscribers on social media.
Anthropic’s blog provides a general outline of watermarking, detailing how Claude is capable of establishing a pattern in its responses while making seemingly minor choices—like whether to use “cloudy” or “grey” for the weather. This pattern will be invisible to readers but identifiable to those with the appropriate key.
The company emphasized that watermarking will not affect the quality of Claude’s outputs: “A response with a watermark will appear identical to one without,” they stated.
Anthropic explained that it would implement the SynthID-Text method introduced by Google DeepMind in 2024, along with plans to launch a watermark detection API. They clarified that watermarking is different from AI detection methods used by other firms, which identify specific writing patterns to determine AI involvement. “Identifying these patterns is not the same as verifying a watermark,” they noted.
Regarding the possibility of concealing the watermark by rewriting text, Anthropic acknowledged that while light edits might not fully remove it, a complete overhaul would. “In this case, it can be debated whether the text should still be classified as AI-generated,” they remarked.
On the question of whether the watermark would remain in text that underwent minimal proofreading or editing by Claude, the company indicated that the watermark’s presence would depend on the extent of editing. If it has only seen slight adjustments, “almost all the words” would be human-generated, leaving little for a watermark to cling to.
Code generation is expected to exhibit less watermarking compared to regular text, owing to the model’s necessity to produce functional code without the freedom of choosing from multiple options.
Anthropic added that when choices are arbitrary in code, such as comments, watermarking could still be applied. However, it will have a minimal impact on the actual code produced.
Moreover, Anthropic stated that Claude will not be alone in generating watermarked text, as numerous other AI developers have also committed to implementing their own watermarking practices.



