Anthropic has started embedding invisible watermarks into text generated by Claude models launched on or after 2 August, Nature reported. Images produced by Claude will also include metadata containing a digital signature in most cases to show the system processed the file.
The San Francisco company uses an algorithm that adjusts how its models select words, creating a statistically observable pattern in the generated text. Anthropic stated that the watermark does not alter the "meaning, quality, or readability of Claude’s response" and noted that the trace "may persist through some editing." The company plans to apply the system to Claude outputs worldwide.
European regulators prompted the change through the EU AI Act, which was formally adopted in 2024. Under the law, frontier AI model providers must ensure their outputs are detectable or face fines reaching up to €15 million, roughly $17 million, or 3 percent of global annual turnover. Frontier models released after 2 August must comply immediately, while existing market models have until 2 December to meet the mandate.
Detection limits
Watermarks do not indicate how a model was used, Anthropic noted. Detecting a trace signals Claude's involvement but cannot distinguish between AI-authored work and instances where a human used the tool for translation or summarisation. Short passages, paraphrased drafts, and rewritten text may also lose the statistical signal entirely.
Academic researchers question whether the measure will curb low-quality papers or scientific fraud. Reese Richardson, a metascientist at Northwestern University, told Nature that users can strip watermarks easily with other software tools. Nihar Shah, a computer scientist at Carnegie Mellon University, noted that detection tools promised by Anthropic could still catch careless misuse if false-positive rates remain low.
Shah previously implemented a watermarking system for the International Conference on Machine Learning 2026. Organizers embedded watermarks into papers sent for peer review, generating detectable text if reviewers processed the files through AI. That system identified 506 reviewers who breached the conference's ban on artificial intelligence.
