ntworld.ink · Presentations Press M for menu
An interactive walkthrough

Can you tell if it’s AI?

Since August 2026, most major AI providers mark their output, and the biggest name of all, Claude, does it worldwide. This walks through what that actually means: why the mark is not a hidden character you can delete, how a model can be nudged toward a secret pattern of word choices, a second and different real system from Google called SynthID, where the mark breaks down, and what happens when a school or a hiring manager treats a detector's result as proof rather than a hint. It ends with a playground where you run the mark yourself and read the detector's numbers.

Module 01

The rule that changed everything

Myth we'll unpick: an AI watermark is a hidden character you can find-and-replace away

Something changed in the second week of August 2026: text from Claude, wherever in the world you use it, started carrying a mark that identifies it as AI-generated. It was not Anthropic's idea alone. It was a deadline.

Two dates, three months apart

2 Aug 2026Article 50 takes effectThe EU AI Act's transparency obligations for AI-generated content become enforceable, with a grace period to 2 December 2026 for content already in circulation.
11 to 12 Aug 2026Anthropic switches it on, everywhereAnthropic confirms Claude output is marked "across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered, worldwide," alongside roughly 190 other signatories to the European Commission's Code of Practice on Transparency of AI-generated Content, including Google.

Dates from the European Commission's Digital Strategy site and Anthropic's own Help Center article "How Claude marks AI-generated content," both accessed 14 August 2026; corroborated by TechCrunch, The Register and Euronews, all published 11 August 2026.

What Article 50 actually requires

EU AI Act · Article 50(2) · full text
Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are . Providers shall ensure their technical solutions are , taking into account the specificities and limitations of various types of content, the costs of implementation and the generally acknowledged state of the art, as may be reflected in relevant technical standards. , or .
The whole obligation is three sentences. Click any highlighted phrase to unpack it.

Article 50(2) of Regulation (EU) 2024/1689 (the EU AI Act), applying from 2 August 2026; text as published in the AI Act Explorer at artificialintelligenceact.eu, accessed 17 August 2026. The related recital, 133, discusses watermarks, metadata and cryptographic provenance as acceptable marking techniques.

The carve-out everyone will argue about

Read that third sentence again. If the AI performs "an assistive function for standard editing", or does not substantially alter what you gave it or what it means, the marking duty does not apply to that extent. The law itself distinguishes AI that writes from AI that tidies. Where tidying ends and writing begins, the Article does not say; that boundary is exactly where the arguments will happen.

Keep that carve-out in mind for Module 04, where a spell-checked essay comes back flagged anyway.

The GDPR playbook, again

Running two versions of the same model, one marked and one not, roughly doubles the engineering cost and complicates every API integration a business has built. It is cheaper to build one pipeline than two, so the EU's rule became the practical floor for the entire internet. Europe has run this play before: the GDPR was written for Europeans, but complying once was cheaper than building everything twice, so privacy banners and data-export buttons went worldwide. Article 50 is following the same path: Anthropic says it marks everywhere because it does not yet have a durable way to scope the mark by region, so for a supported model there is no region where an unmarked Claude response is the default.

So there is a mark now. The next question is the one almost everyone gets wrong: what kind of mark is it?

Module 02

Two things called "watermark"

A generated image and a piece of generated text are marked in two completely different ways, and mixing them up is where most of the online confusion starts.

Files and text, side by side

Files: SVG, PNG, JPG
Method C2PA signed metadata
What it is A sidecar record, naming the tool and creation details
Survives Being opened normally
Does not survive a screenshot, which creates a brand-new file with no sidecar, or upload to most social platforms, whose compression routinely strips metadata to save space.
Text
Method A statistical pattern in the word choices themselves
What it is Not a hidden layer; the words actually chosen
Survives Copy, paste, retype into Notepad
You cannot strip it without rewriting the sentence, because there is no separate layer to strip.

C2PA is the Coalition for Content Provenance and Authenticity, an open standard Anthropic uses for its generated image, audio and video files. Source: Anthropic Help Center, "How Claude marks AI-generated content," accessed 14 August 2026.

A third kind of mark: written into the pixels

A photorealistic AI-generated image of a saltwater crocodile sitting at a Darwin cafe table, overlaid with a mock SynthID verification panel reading SynthID verified: Gemini, origin Google Gemini AI
An illustrative mock-up, itself AI-generated, of SynthID verification on a Gemini image. The overlay is not a real detector interface.

For images, Google's SynthID takes the other road: instead of a metadata manifest riding alongside the file, it alters the pixels themselves, invisibly, in a pattern its detector can find later. Google says the mark does not change visible image quality and is designed to survive cropping, filters and lossy compression, the exact operations that strip a C2PA manifest. That is a vendor claim about robustness, but the design idea matters: a mark woven into the content survives what a mark attached to the content does not. Source: Google DeepMind, SynthID, deepmind.google, accessed 17 August 2026.

Hold that idea, because text watermarking applies the same philosophy to words: the mark is the content.

The Notepad test

If a text watermark were a hidden Unicode character or a swapped curly quote, copying the words into a plain-text editor would strip formatting and destroy it in seconds. That is exactly the test a text watermark has to survive, and it does, because the mark is not sitting on top of the words. It is which words got chosen in the first place.

So how does a model's own word choice become a trackable signature? That takes watching one get written, one word at a time.

Module 03

Green words, red words

A language model does not plan a sentence. It predicts one next word at a time, weighs several good options, and picks one. A watermark intercepts that exact moment. Anthropic said on 14 August 2026 that Claude uses Google DeepMind's SynthID-Text method; the simpler published scheme that method descends from, and this widget's simplified illustration of it, works like this.

Watch a sentence get written

The model writes one word at a time. Before each word, a secret key sorts every word in the language into a green list and a red list. The model still has its own favourite word, but the watermark pushes it to pick a green one. Do it in two steps: first see the options and the model's favourite, then reveal what the watermark actually made it choose. Watch the green picks build up while the sentence still reads normally.

Press "Show the options" to begin.

Why nobody notices while reading

Think of the party game Taboo: your card says "brain", but the five most obvious words are banned, so you end up saying "the squishy grey computer inside your bone helmet". The sentence still makes sense; it is just steered away from the writer's absolute first choice, over and over, in a pattern only someone holding the same secret key can measure.

Method described here follows the published KGW scheme: John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers and Tom Goldstein, "A Watermark for Large Language Models," ICML 2023 (first posted to arXiv 24 January 2023). The demonstration above is illustrative; it does not reproduce SynthID-Text's tournament sampling, the method Anthropic adopted for Claude (announced 14 August 2026).

A different real system: Google's SynthID

Green/red list (KGW-style)
Splits candidates into two lists, biases toward one
Detected via a statistical score on green-word rate
SynthID
Runs an elimination-style "tournament" among candidates
Covers text, images, audio and video from Google's models
Over 10 billion pieces of content watermarked so far, per Google. Adopted by Anthropic for Claude's text mark, announced 14 August 2026.

Google DeepMind's SynthID Text mechanism: Sumanth Dathathri, Abigail See, Sumedh Ghaisas and colleagues, "Scalable watermarking for identifying large language model outputs," Nature vol. 634, pp. 818 to 823, 23 October 2024. Announcement and detector portal: Google, "Google SynthID: AI content detector," blog.google, referencing the SynthID Detector portal launched at Google I/O, 20 May 2025.

Two published mechanisms, one teaching model and one in production, both aimed at the same idea. Neither is unbreakable. Here is where they actually fail.

Module 04

Where the mark breaks

Forcing a model toward a secret word list only works where the model actually has spare choices to give up. That is not always true.

How much choice does the model actually have?

Does reformatting code erase the mark?

Yes, usually. The mark is a pattern in which exact words the model picked, in the exact order it picked them. Most coding software runs a code formatter: a built-in tidy-up tool, such as Prettier or Black, that automatically re-spaces the code and standardises its layout the moment you paste it in. In doing that tidy-up it reshuffles the very word-and-spacing pattern the mark was hiding in, so the mark breaks as a side effect, without anyone trying to remove it. Why this is flagged rather than stated flatly: the fragility of watermarking code is well studied, with several 2024 to 2025 research papers on why marking code and other low-choice text is hard. The specific point that a routine formatter pass wipes the mark is a sound conclusion drawn from that research, not a separately tested finding.

What a positive detection actually proves

Anthropic's own documentation states it plainly: "a detected mark provides a signal that content was processed by Claude, but is not fully conclusive." It cannot tell "written entirely from scratch" apart from "pasted in and asked to fix a typo." Write a deeply personal essay yourself, then ask Claude to check your commas, and the whole thing can come back flagged.

Direct quote from Anthropic Help Center, "How Claude marks AI-generated content," accessed 14 August 2026. Note the irony against Module 01: assistive standard editing is the very case Article 50's carve-out excuses from marking, yet a provider marking everything marks it anyway.

A weak signal is still useful, until someone treats it as a verdict. That is exactly what has started happening.

Module 05

When a detector result is treated as proof

A detector can tell you a piece of writing looks AI-generated. It cannot tell you who wrote it, how much of it was AI, or whether any rule was broken. The problem in this module is simple: some institutions have started treating that "looks AI" result as proof of cheating, and the numbers below show why that goes wrong. One way to hold the distinction: a detector is like a smoke alarm, it tells you something is worth a look, not that the house is on fire.

What happens when detectors are trusted at scale

Peer review21 per cent of ICLR reviews, AI-generatedDetection company Pangram analysed roughly 70,000 peer reviews submitted to the ICLR machine-learning conference and found about 21 per cent were fully AI-written, with longer, more sycophantic text and formulaic bolded headers as the tell.
BenchmarkingMost detectors fail under real scrutinyEconomists Brian Jabarian and Alex Imas benchmarked the leading detectors and introduced a "policy cap": fix your maximum tolerable false-positive rate first, then see what survives. At a strict 0.5 per cent cap, most tools collapsed; only one maintained accuracy, and only with backend access an ordinary teacher does not have.
False positives at scaleEven a good detector wrongly flags real peopleComputer scientist Arvind Narayanan has pointed out that even a 1-in-10,000 false-positive rate, applied across the hundreds of pieces of writing a student submits over a degree, works out to roughly 5 to 10 per cent of an entire student body being falsely flagged for cheating at some point.
BiasThe errors are not evenly spreadMany detectors score on how predictable text is; writers using simpler, more formal sentence structures, often non-native English speakers, read as more "AI-like" regardless of who wrote it. NYU, MIT, Vanderbilt and Northwestern have all disabled Turnitin's AI-detection feature for exactly this reason.

Pangram, "Pangram Predicts 21% of ICLR Reviews are AI-Generated," 18 November 2025. Jabarian and Imas, "Artificial Writing and Automated Detection," NBER Working Paper No. 34223, 26 August 2025. Narayanan's calculation, posted publicly and widely discussed, cites 500 to 1,000 submissions per student over a four-year degree. University detector policies collected from published tracker sites, accessed 14 August 2026.

What a detector report actually looks like

Pangram dashboard for one ICLR submission: 18 of 23 segments are AI across 5,344 words, with a segment strip coloured mostly orange for AI and a list of individual segments labelled AI or Human with High, Medium or Low confidence
Pangram's report for one real ICLR 2026 submission, from the public database. Scroll the panel.

This paper came back 18 of 23 segments AI across 5,344 words, with an overall verdict of "AI content detected, but not fully AI-generated": 78 per cent AI, 22 per cent human. Notice what a serious report does: it segments the document, grades its own confidence High / Medium / Low, hedges its verdict as belief rather than fact, and separates AI-generated from AI-assisted.

Notice also what no report can say: who typed, why, or whether the use was permitted. The strip shows a pattern in the text; everything after that is a human judgement.

Pangram ran this over roughly 19,000 ICLR papers as well as the 70,000 reviews. On the paper side, 61 per cent stayed mostly human-written, while 9 per cent were more than half AI. The whole corpus is browsable.

Report exhibit captured from the public database at iclr.pangram.com, 17 August 2026. Paper-side and review-side figures: Pangram, "Pangram Predicts 21% of ICLR Reviews are AI-Generated," 18 November 2025.

The vicious cycle: detect, humanise, repeat

Google search results for AI humaniser, listing Humanize AI, SuperHumanizer, Grammarly, AI Humanize, Clever AI Humanizer, QuillBot and ZeroGPT, several promising to bypass AI detection
Page one of a Google search for "AI humaniser". Scroll the panel.

Search "AI humaniser" and page one is an entire industry, promising to "bypass all AI detection" by rewriting AI text until the statistical fingerprint, watermark included, is gone. Rewriting the words is the one attack every text watermark and every stylistic detector is known to be weakest against, because the mark is the word choices; change the words and the mark goes with them.

Look closer at the names. Grammarly, QuillBot and ZeroGPT all appear, and each also operates an AI detector. The same market sells the smoke alarm and the silencer, and every round of the loop pushes machine text closer to human text, eroding the ground detection stands on.

Screenshot: Google results for "AI humaniser," captured 17 August 2026. Paraphrasing as the standard attack on watermark and detector schemes: analysed in the KGW watermark paper itself (Kirchenbauer et al., ICML 2023) and in Sadasivan, Kumar, Balasubramanian, Wang and Feizi, "Can AI-Generated Text be Reliably Detected?", arXiv 2303.11156, first posted March 2023.

Detection has an arms-race problem on one side. On the other, it has a subtler problem: the tool can change the writer.

The quieter problem: it changes minds, invisibly

A 2026 Science Advances study gave people an AI writing assistant secretly biased toward one side of a real debate, on topics like the death penalty and felon voting rights. Participants' actual private opinions shifted toward the assistant's bias, every word still typed by their own hand, and warning them beforehand did nothing to stop it. Most did not believe they had been influenced at all.

Sterling Williams-Ceci, Maurice Jakesch, Advait Bhat, Kowe Kadoma, Lior Zalmanson and Mor Naaman, "Biased AI writing assistants shift users' attitudes on societal issues," Science Advances, 11 March 2026.

Why this matters

The mark on AI text is real, it is worldwide since August 2026, and it is genuinely hard to strip. None of that makes it a lie detector: it fades when the task gives the model little room to choose its words, such as code; it says nothing about how much of a piece was AI-written; and the detectors built on top of it are already producing real false accusations at scale. Treat a positive flag as one signal to weigh, never as proof, and remember the mark only ever catches what you can see; it says nothing about whether a suggestion quietly nudged what you decided to write in the first place.

Sources used. Regulation and rollout: Article 50(2), Regulation (EU) 2024/1689, full text via the AI Act Explorer at artificialintelligenceact.eu, accessed 17 August 2026; European Commission Digital Strategy, "Transparency obligations under Article 50 AI Act" and "Strong backing for the Code of Practice on Transparency of AI-generated Content," accessed 14 August 2026; Anthropic Help Center, "How Claude marks AI-generated content," accessed 14 August 2026; Anthropic, "How Claude's text watermarking works," 14 August 2026; TechCrunch, The Register and Euronews, all 11 August 2026. Image watermarking: Google DeepMind, SynthID, deepmind.google, accessed 17 August 2026 (robustness description is Google's own claim); the crocodile verification image is an AI-generated illustrative mock-up made for this deck, not a real detector interface. Mechanism: Kirchenbauer, Geiping, Wen, Katz, Miers and Goldstein, "A Watermark for Large Language Models," ICML 2023, arXiv 24 January 2023; Dathathri, See, Ghaisas et al., "Scalable watermarking for identifying large language model outputs," Nature 634, 23 October 2024; Google, "Google SynthID: AI content detector," blog.google. Fragility and low-entropy watermarking: arXiv 2405.14604 and arXiv 2512.14753, both on watermarking low-entropy generation. Detection and its limits: Pangram, "Pangram Predicts 21% of ICLR Reviews are AI-Generated," 18 November 2025, and the public browsable corpus at iclr.pangram.com, report exhibit captured 17 August 2026; Sadasivan, Kumar, Balasubramanian, Wang and Feizi, "Can AI-Generated Text be Reliably Detected?", arXiv 2303.11156, first posted March 2023; Google search results for "AI humaniser," screenshot captured 17 August 2026; Jabarian and Imas, "Artificial Writing and Automated Detection," NBER Working Paper No. 34223, 26 August 2025; Arvind Narayanan's public false-positive-rate calculation, referenced via Hacker News discussion, accessed 14 August 2026; university AI-detector policy trackers, accessed 14 August 2026. Psychological influence: Williams-Ceci, Jakesch, Bhat, Kadoma, Zalmanson and Naaman, "Biased AI writing assistants shift users' attitudes on societal issues," Science Advances, 11 March 2026.

Module 06

Try it yourself: the watermark playground

In this module: run the published watermark on real text and read the detector's numbers

Module 03 showed one sentence being steered. This module hands you the controls. A small stand-in model writes a passage, with or without the mark, and a detector that has never seen the model scores it. Then you paste in human writing and watch the score fall back to chance.

What the detector counts

The detector does not read for style. It re-runs the secret coin toss for every word and counts how many landed on the green list. A writer with no key lands there about γ of the time by chance, where γ is the green share of the vocabulary. The score is a z-score: how many standard deviations the green count sits above that chance level.
z  =  ( g − γT )  /  √( T γ (1 − γ) )
gThe number of words the detector finds on the green list.
TThe number of words it scored.
γThe green share of the vocabulary, 0.5 in the paper's main setting, so chance gives γT green words.

A worked number. Score 300 words with γ = 0.5 and chance predicts 150 green. Find 190 and the count sits 4.6 standard deviations above chance. The paper calls anything above z = 4 detected; an unmarked writer reaches that by chance about 3 times in 100,000.

Equation 3 and the z = 4 threshold: Kirchenbauer, Geiping, Wen, Katz, Miers and Goldstein, "A Watermark for Large Language Models," ICML 2023, sections 2 and 3.1.

The playground

Write
Watermark
Settings
Text under test green listred lista choice point
Press Detect to score this text, or Write it to have the stand-in model write a passage.
Detect
Not scored yet
Green words–
Chance would give–
z-score–
p-value–

Six things to try

BaselineDetect the starting textIt is a human-written passage from Wikipedia. About half its words land on the green list, which is what chance predicts, so the score sits near zero.
HardWrite the wet-season passage with the Hard markNearly every choice point goes green and the z-score climbs past 4. Read the passage: it still makes sense, because the model only ever swapped one good word for another.
Low choiceNow write the CDU history the same wayNames, dates and titles give the model nothing to choose between. Far fewer choice points means a weaker mark, and it stays below the threshold. This is Module 04's low-entropy problem, too little choice, measured.
Wrong keyChange one letter of the detector keyThe detector re-runs the coin toss with the wrong seed, so the green words it finds are just chance again. The mark is still in the text; without the key, nobody can see it.
SoftSwitch to Soft and try δ of 1, 2 and 4A small nudge leaves the model's favourite in charge and the mark faint. The paper's default, 2, is the compromise; 4 behaves almost like Hard.
AttackPress Paraphrase, then press it againEach press rewrites a quarter of the choice points without the mark. The score falls, and the paper's point shows: removing the mark means changing a large share of the words, not one or two.

Paste anything of your own into the text box as well. The detector scores any text; it never needs the model.

What this toy leaves out

A real model chooses from a vocabulary of 50,000 or more word fragments at every step. The stand-in here chooses between a few hand-written wordings at some steps and has no choice at all at the others, so its mark is much weaker than a real one. The green list here is a per-word coin toss seeded by the previous words and the key; the paper draws a fixed-size list from the whole vocabulary. The detector maths is the paper's, unchanged. Anthropic's production scheme is SynthID-Text, disclosed 14 August 2026; this is the earlier published method it descends from.

Sources used in this module. Mechanism, Algorithms 1 and 2, Equation 3, the z = 4 threshold, the low-entropy caveat, the repeated n-gram remedy and the private-key discussion: John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers and Tom Goldstein, "A Watermark for Large Language Models," ICML 2023, arXiv 2301.10226, version 4, 1 May 2024. Human-written test passage: the history section of the Wikipedia article "Charles Darwin University," licensed CC BY-SA 4.0, as supplied on 30 August 2026. The two passages the stand-in model writes were written for this deck; the CDU history one restates the same Wikipedia facts. The panel layout follows the public "LLM Watermarking Playground" demonstration that runs a small Qwen model in the browser; this version replaces the model with the stand-in described above.