Can you tell if it’s AI?
Since August 2026, most major AI providers mark their output, and the biggest name of all, Claude, does it worldwide. This walks through what that actually means: why the mark is not a hidden character you can delete, how a model can be nudged toward a secret pattern of word choices, a second and different real system from Google called SynthID, where the mark breaks down, and what happens when a school or a hiring manager treats a detector's result as proof rather than a hint.
The rule that changed everything
Something changed in the second week of August 2026: text from Claude, wherever in the world you use it, started carrying a mark that identifies it as AI-generated. It was not Anthropic's idea alone. It was a deadline.
Two dates, three months apart
Dates from the European Commission's Digital Strategy site and Anthropic's own Help Center article "How Claude marks AI-generated content," both accessed 14 August 2026; corroborated by TechCrunch, The Register and Euronews, all published 11 August 2026.
Why not just switch it on for Europe?
So there is a mark now. The next question is the one almost everyone gets wrong: what kind of mark is it?
Two things called "watermark"
A generated image and a piece of generated text are marked in two completely different ways, and mixing them up is where most of the online confusion starts.
Files and text, side by side
C2PA is the Coalition for Content Provenance and Authenticity, an open standard Anthropic uses for its generated image, audio and video files. Source: Anthropic Help Center, "How Claude marks AI-generated content," accessed 14 August 2026.
The Notepad test
So how does a model's own word choice become a trackable signature? That takes watching one get written, one word at a time.
Green words, red words
A language model does not plan a sentence. It predicts one next word at a time, weighs several good options, and picks one. A watermark intercepts that exact moment. Anthropic keeps its own version of this mechanism undisclosed, but the published research method it is built on, and this widget's simplified illustration of it, works like this.
Watch a sentence get written
Why nobody notices while reading
Method described here follows the published KGW scheme: John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers and Tom Goldstein, "A Watermark for Large Language Models," ICML 2023 (first posted to arXiv 24 January 2023). The demonstration above is illustrative; it does not reproduce Anthropic's actual undisclosed method.
A different real system: Google's SynthID
Google DeepMind's SynthID Text mechanism: Sumanth Dathathri, Abigail See, Sumedh Ghaisas and colleagues, "Scalable watermarking for identifying large language model outputs," Nature vol. 634, pp. 818 to 823, 23 October 2024. Announcement and detector portal: Google, "Google SynthID: AI content detector," blog.google, referencing the SynthID Detector portal launched at Google I/O, 20 May 2025.
Two real, different mechanisms, both aimed at the same idea. Neither is unbreakable. Here is where they actually fail.
Where the mark breaks
Forcing a model toward a secret word list only works where the model actually has spare choices to give up. That is not always true.
How much choice does the model actually have?
Does reformatting code erase the mark?
What a positive detection actually proves
Direct quote from Anthropic Help Center, "How Claude marks AI-generated content," accessed 14 August 2026.
A weak signal is still useful, until someone treats it as a verdict. That is exactly what has started happening.
Treating a smoke alarm as a verdict
A smoke alarm tells you something got hot enough to smoke. It does not tell you whether that was a house fire or burnt toast. Detectors built on watermarks and writing style work the same way, and several institutions have started using them as if they were the fire itself.
What happens when detectors are trusted at scale
Pangram, "Pangram Predicts 21% of ICLR Reviews are AI-Generated," 18 November 2025. Jabarian and Imas, "Artificial Writing and Automated Detection," NBER Working Paper No. 34223, 26 August 2025. Narayanan's calculation, posted publicly and widely discussed, cites 500 to 1,000 submissions per student over a four-year degree. University detector policies collected from published tracker sites, accessed 14 August 2026.
The quieter problem: it changes minds, invisibly
Sterling Williams-Ceci, Maurice Jakesch, Advait Bhat, Kowe Kadoma, Lior Zalmanson and Mor Naaman, "Biased AI writing assistants shift users' attitudes on societal issues," Science Advances, 11 March 2026.
Why this matters
Sources used. Regulation and rollout: European Commission Digital Strategy, "Transparency obligations under Article 50 AI Act" and "Strong backing for the Code of Practice on Transparency of AI-generated Content," accessed 14 August 2026; Anthropic Help Center, "How Claude marks AI-generated content," accessed 14 August 2026; TechCrunch, The Register and Euronews, all 11 August 2026. Mechanism: Kirchenbauer, Geiping, Wen, Katz, Miers and Goldstein, "A Watermark for Large Language Models," ICML 2023, arXiv 24 January 2023; Dathathri, See, Ghaisas et al., "Scalable watermarking for identifying large language model outputs," Nature 634, 23 October 2024; Google, "Google SynthID: AI content detector," blog.google. Fragility and low-entropy watermarking: arXiv 2405.14604 and arXiv 2512.14753, both on watermarking low-entropy generation. Detection and its limits: Pangram, "Pangram Predicts 21% of ICLR Reviews are AI-Generated," 18 November 2025; Jabarian and Imas, "Artificial Writing and Automated Detection," NBER Working Paper No. 34223, 26 August 2025; Arvind Narayanan's public false-positive-rate calculation, referenced via Hacker News discussion, accessed 14 August 2026; university AI-detector policy trackers, accessed 14 August 2026. Psychological influence: Williams-Ceci, Jakesch, Bhat, Kadoma, Zalmanson and Naaman, "Biased AI writing assistants shift users' attitudes on societal issues," Science Advances, 11 March 2026. Deliberately not used: a claimed "style loophole" where a restrictive style prompt suppresses the watermark, and a cited British Journal of Educational Technology study on "strategy-graded reliance," both of which appeared in early drafting material but could not be traced to any real, findable source.