The AI company Anthropic says that its Claude models will label AI-generated content with a watermark.Credit: Samuel Boivin/NurPhoto via Getty
Anthropic, the firm behind the artificial-intelligence model Claude, has announced that text generated by any models launched on or after 2 August will be invisibly embedded with a watermark that indicates the output was written by AI. Meanwhile, images generated by Claude will in most cases come with metadata that contains a digital signature to show that the model processed the file.
Text-based watermarks use an algorithm to tweak how an AI model selects its wording. When applied to a stretch of text, this process leaves a statistically observable trace in the output. Anthropic, based in San Francisco, California, says that its watermark won’t change the “meaning, quality, or readability of Claude’s response” and that the mark “may persist through some editing”.
The move comes in response to the EU AI Act, which was formally adopted in 2024. As of 2 August this year, providers of frontier AI models must ensure that AI-generated outputs are detectable, or be hit with fines of up to €15 million (about US$17 million) or 3% of their global annual turnover. Models released after 2 August will have to meet the requirements immediately, whereas versions already on the market have until 2 December to do so. Anthropic says that watermarks will be applied on Claude’s outputs worldwide.
The presence of the watermark reveals little about how the model was used. Detecting one “provides a signal” that content was made with Claude, says Anthropic, but is not conclusive: the model might have been used just to summarize or translate an original human idea, for example. Equally, a lack of a watermark doesn’t mean that the text was not generated by AI. Because the watermarks are based on patterns of subtle changes in a model’s word choice, passages that are very short, or that have been paraphrased or rewritten, might no longer carry a signal.
‘Humanizer’ tool can erase signs of AI-written text — alarming scientists
The impact that such watermarks will have on academic integrity remains unclear. Given that watermarks can be stripped from text easily — for example, by using another model — they are unlikely to stop motivated people from using AI to produce fake or low-quality papers, known as AI slop, says Reese Richardson, a metascientist at Northwestern University in Evanston, Illinois.
But if AI firms create tools that allow others to check for the watermark — as Anthropic has said it will do — and if these tools have an acceptably low rate of false positives, some illegitimate uses of AI could be detected, says computer scientist Nihar Shah, who studies the evaluation of science at Carnegie Mellon University in Pittsburgh, Pennsylvania.
Watermarks could, for example, help journal editors or conference organizers to enforce strict ‘no AI’ policies in peer reviews, as the International Conference on Machine Learning (ICML) 2026 did in one of its two possible review streams. Organizers of the July event added a watermark to papers distributed for peer review that generated telltale text when AI was used in review reports. They caught 506 reviewers who violated the no-AI policy. “This experience suggests that while some illegitimate AI uses may be done carefully to evade detection, many others may simply copy-paste AI outputs,” says Shah, who was behind the ICML’s watermarking process.
Invisible ink