Does Claude watermark its text? Models, detection and removal
Which Claude models watermark text, who can detect it, and why Unicode cleaners cannot remove it. My guide to Claude, Gemini and AI text watermarks.
P.S. A working AI text watermark remover from this research is at painintheagent.com/tools/ai-text-watermark-remover.
Kirill Balakhonov · 2026-08-16 · updated 2026-09-08
Anthropic recently announced that its new Claude models will embed an invisible watermark in all generated text, and Reddit exploded. Half of the threads discussed the end of privacy, while the other half featured individuals selling “watermark removers” for a problem they clearly did not comprehend. The practical question is which model produced your text. Anthropic now lists Fable 5.1 and Mythos 5.1 as carrying the watermark; marking for older models is being added later.
I run LLM pipelines every day, both at work and in my own AI writing tool, so this topic hit my own roadmap directly. I spent the last days reading the vendor docs, the actual papers and a depressing amount of reddit. This is the guide I wish someone had written for me. Simple words, no panic, sources mentioned along the way so you can check me.
Which Claude models watermark text?
Claude models launched on or after August 2, 2026 weave a hidden mark into the text they generate. Worldwide, at the model level, so it does not matter if you use the chat app, the API or Claude through some cloud provider. Anthropic also attaches signed provenance metadata (the C2PA standard, the same thing camera makers use) to files Claude produces, like PNG and SVG images.
Anthropic’s current model list includes Fable 5.1 and Mythos 5.1. Watermarking begins at launch for models released on or after August 2, 2026. Marking is still being added to earlier models. My initial version of this guide stated that no selectable model marked its text yet. That statement is now outdated: consult the model list before assuming your output is unmarked.
The reason behind all of it is a new European AI law (the EU AI Act) and a Code of Practice that goes with it. Anthropic signed the code, and so did roughly 190 other organizations, including Google, OpenAI, Microsoft, Meta and Mistral. The loud holdout is xAI, which decided to skip it. My guess is that marking only European traffic would be more expensive than marking everything, so the whole world gets the watermark.
No detector is available to the general public. Anthropic’s detection API is in private preview for eligible organizations, including researchers and educational institutions. Access is expanding, but ordinary readers cannot yet verify arbitrary Claude text through a public tool.
What a text watermark actually is
Forget invisible ink. A text watermark is a statistical trick, and once you see it, it is almost disappointingly simple.
When a model writes, it picks the next word from several good candidates. “The results were great” and “the results were strong” are both fine. A watermarking model carries a secret rule that sorts candidate words into two invisible buckets and then leans, very slightly, toward one bucket. Think of a coin that is bent by one percent. One flip tells you nothing. A thousand flips tell you clearly which coin was used.
A detector with the secret key counts how often the text landed in the “preferred” bucket. Human text lands about half the time, because humans do not know the buckets exist. Marked text lands there noticeably more often. That statistical excess is the whole watermark.
Anthropic says its mechanism is a version of SynthID-Text, a method the DeepMind team published openly in Nature back in 2024 and has been running in Gemini since then. So the science is public, only the secret key is not.
Two things follow directly from this design, and they explain most of the confusion online.
First, the watermark needs length. A short answer is a handful of coin flips, and no statistician on earth reads a signal from five flips. Anthropic itself says short text may be impossible to check reliably.
Second, the watermark lives in word choice, and only in word choice. There is nothing attached to the file, nothing in the clipboard, nothing your text editor could show you. Which brings me to the next chapter.
Can Unicode cleaners remove Claude’s watermark?
No. A Unicode cleaner can remove invisible characters without changing the wording. Claude’s statistical watermark sits in word choices, so character cleanup leaves that mechanism untouched. A deep rewrite changes the wording; my SynthID experiment measures what survived five different ways of doing it under a research key.
The most popular mistake in every thread. Someone pastes ChatGPT output into a character inspector, finds weird invisible Unicode characters (narrow no-break spaces, zero-width spaces), and declares they found the watermark. There was a wave of this in 2025, and it comes back every few months. Over the last month I found exactly one thread where a guy was running a “watermark remover” that proudly reported finding line breaks. Literal line breaks, the things you type with Enter.
Those hidden characters are real, but they are typography quirks and training artifacts. OpenAI directly said they are a side effect of training, and honestly, they would be a comically bad watermark, since find-and-replace kills them in one second. A real statistical watermark survives find-and-replace just fine, because there is no character to replace. The mark sits in the choice between normal words.
The em dash panic belongs here too. Yes, models love em dashes. So do many human writers who learned punctuation before 2022. A writing habit is a tell, maybe, but it is not a watermark. A watermark needs a secret key and a detector. A habit just needs a reader with prejudices.
Who marks text today
I keep this table maintained, with a source for every row.
| Provider | Text watermark | What we know | EU CoP |
|---|---|---|---|
| Anthropic (Claude) | Live on Fable 5.1 and Mythos 5.1, older models pendingsince 2026-08-02 | According to an Anthropic support article, Claude models released on or after August 2, 2026 embed text watermarks at launch. The article lists Fable 5.1 and Mythos 5.1 as the currently supported models. This watermarking is applied at the model level globally, across the Claude Platform API, Claude, Claude Code, Claude Cowork, Claude Tag, and deployments on AWS, Google Cloud, and Microsoft Foundry. Watermarking for models released prior to that date is ongoing and is expected to be implemented over the next few months. The AI Act gives systems already on the market until December 2, 2026 to add marking. The technique is a variant of SynthID-Text, a method published by DeepMind in Nature in 2024. Anthropic states that minor edits likely will not fully remove the mark, whereas a complete rewrite will, and translations produced by Claude retain a watermark. Generated files are signed with C2PA metadata, and the free Claude Content Checker reads this credential from files but does not analyze text content. The watermark detection API is in private preview (announcement updated on September 1, 2026): it is accessible to eligible organizations defined under EU law, such as regulators, law enforcement, media outlets, fact-checkers, researchers, educational institutions, and EU civil society groups, as well as enterprises with their own compliance obligations, via a request form. Anthropic intends to broaden access over time. The Anthropic checker page indicates that text verification occurs through this API. source | signed |
| Google (Gemini) | Livesince 2024-05 | Google states that SynthID-Text marks text generated by the Gemini application and web interface. The company announced the text modality at I/O in May 2024, published the methodology in Nature in October 2024, and open-sourced the watermarking code. I found no public method to verify text for Gemini’s production watermark. The verifier within the Gemini application processes images, video, and audio, and Google is extending this verification to Search and Chrome. The SynthID Detector portal remains an early-tester waitlist for journalists, media professionals, and researchers; Google’s May 2025 launch post stated it accepts text as well as media, whereas the current DeepMind page describes uploading an image, video, or audio file. The Google Cloud AI Content Detection API is in private preview, and its documentation covers images. The open-source detector functions only with the corresponding key, configuration, and tokenizer, so it cannot check Gemini’s production watermark. source | signed |
| OpenAI (ChatGPT) | Built, not deployed | OpenAI has had a working text watermark internally since approximately 2024 (reported as 99.9% detection on long texts) but opted not to release it. Since May 19, 2026 images from ChatGPT, Codex, and the API carry C2PA metadata alongside Google’s SynthID, with audio support added on July 31, 2026. I found no announced text watermark in OpenAI’s API provenance guide. It lists Content Credentials for images and SynthID for images and audio, with no mention of text. The public checker at openai.com/verify and the Content Provenance API accept only image and audio files. The invisible Unicode characters frequently discovered in ChatGPT output are, according to OpenAI’s response to the startup Rumi in April 2025, a side effect of large-scale reinforcement learning, not a watermark. OpenAI has signed the EU transparency code and states its objective is to expand provenance signals to all modalities, including text. source | signed |
| Microsoft (Copilot) | Committed, none known | Signed the EU transparency code. Microsoft 365 Copilot can add visible or spoken watermarks to AI-generated video, audio and images, and writes C2PA-style metadata into generated images regardless of that setting. None of it applies to text. source | signed |
| Mistral | Committed, none known | Signed the EU Code of Practice on transparency. No public evidence of a deployed text watermark yet. source | signed |
| xAI (Grok) | None known | Did not sign the EU transparency code and made a point of it. No public evidence of a deployed text watermark. Article 50 still binds Grok in the EU, so xAI has to solve marking on its own terms or face enforcement without the code's safe harbour. source | no |
| Cohere | Committed, none known | Named by the Commission among providers that signed the 2026 transparency code. No public evidence of a deployed text watermark. (Amazon, IBM and Writer signed the separate GPAI Code of Practice, a different document that does not cover output marking, so they are not in this table's signatory column.) source | signed |
| Meta (Llama, Muse) | None known | Signed the EU transparency code on Jul 28, 2026. The announcement speaks of identifying and labelling AI-generated content on Meta's platforms and points to a research demo that detects images made with Meta AI; it says nothing about marking text. No public evidence of a text watermark in the Llama line or the newer proprietary Muse models. source | signed |
| DeepSeek, Alibaba (Qwen) | None known | Not EU signatories. Open-weight releases ship without output watermarks, which is exactly why these models keep coming up in every watermark thread. Their Chinese consumer services do fall under the CAC labelling measures in force since Sep 1, 2025, but that regime asks for a visible label plus file metadata, not a mark embedded in the words. source | no |
Verified as of 2026-09-05. Manual check of vendor help-center docs, official FAQ pages, the EU Code of Practice signatory list, peer-reviewed papers and press coverage. Each row cites the strongest public source we could find. 'None known' means we found no public evidence of a deployed text watermark, not proof of absence.
Gemini and supported Claude models now include text watermarks. The table distinguishes between deployed features and commitments or research. A signature on a code of practice does not show that a specific model marks text; verify that model’s documentation.
By the way, the top thread on r/LocalLLaMA claimed that Meta and Microsoft signed the code too and concluded that even local models will be forced to watermark by law. When I first checked this, I read a law firm’s July summary that listed Meta as a non-signatory, and I almost repeated it here as a gotcha. Then I found Meta’s own press release from July 28. They did sign. The signatures in that thread were right, the conclusion still is not: Meta frames its commitment around labelling images, video and audio on its platforms, and nobody has explained how you would force a model whose weights sit on your own disk to watermark anything. Which is a good reminder about this whole topic, half the confident claims out there are one press release out of date, including, briefly, mine.
Why is Anthropic the one doing this first
This is a fair question, given that nearly every big lab signed the same code and Anthropic made the change the subject of a public announcement.
Google is the quiet exception. Gemini has carried SynthID marks since 2024, they published the method in Nature and moved on. No drama, and most Gemini users have no idea.
OpenAI is the interesting case. Reporters found out in 2024 that OpenAI had a working watermark and a detector that recognized long ChatGPT text with 99.9 percent accuracy in internal tests. They never shipped it. The reported reasons were fear of false accusations, worry about non-native English speakers getting flagged (and I understand them perfectly, being one!), and a survey where about a third of users said they would use ChatGPT less if it watermarked. I find that last number the most honest fact in this whole story. The watermark did not threaten their safety record. It threatened their revenue. Since then OpenAI started stamping provenance metadata and SynthID marks into ChatGPT images and audio, so the machinery clearly exists over there. I found no public confirmation that ChatGPT has deployed a text watermark. OpenAI’s current verification documentation covers images and audio; that is the boundary of what I can confirm from it.
So why did Anthropic jump first? They have not explained beyond “the EU code asks for it”, so what follows is my speculation. Anthropic sells itself as the safety company, and being the first mover on transparency fits that brand perfectly. The August 2 date also lines up with when the new European obligations start to bite for freshly released models, so someone had to be first, and the company whose marketing is built on doing the responsible thing had the least room to stall. The others promised the same thing on paper. Watch what they ship, and when.
Why all of this started now
The European AI law has a transparency article that says machine-generated content must be marked as machine-generated, in some machine-readable way. The law itself is vague on how, so in June 2026 the EU published a Code of Practice that translates the vague words into concrete practices, and the big labs signed it.
Behind the legal story there is a quieter business story. The internet is filling with AI text, and the labs themselves suffer from it, because training new models on the output of old models (surprise!) is how you get worse models. A reliable way to recognize your own output is useful for the vendor even if no regulator asked. Some folks on reddit present this as a hidden conspiracy. I would call it an obvious aligned interest, and Fortune quoted people saying the same thing out loud.
What a watermark proves, and what it does not
This part matters more than the mechanism, and it is where I expect the most real-world damage.
The watermark lives in words the model chose, and it does not know why it was choosing them. Anthropic’s FAQ is honest about how this scales, and a reader on r/ClaudeAI made me state it more precisely than my first version did. Ask Claude to fix only the grammar in your text, and the mark can live only in the handful of corrections, which their FAQ says might be too few for a detector to even register. Ask it to smooth the style, and the share of machine-chosen words grows, and the mark grows with it. Ask it to translate, and every word of the output is chosen by the model, so the translation is fully marked, their FAQ says exactly that too. And one more line from the same page: a watermark “cannot distinguish ‘Claude wrote this’ from ‘Claude heavily edited this’”. Your ideas, your voice, and the mark cannot see any of it.
So a watermark detector is a provenance tool, and people will inevitably use it as an authorship judge. Those are different jobs. “This text touched Claude” is a fact. “You did not write this” is an accusation, and the mark alone cannot carry it. Nobody running a detector at scale will stop to ask which one they are looking at. Keep this asymmetry in mind, because the last chapters of this guide are about what to do with it.
How can I check a Claude watermark?
Anthropic’s detector is in private preview for eligible organizations, including researchers and educational organizations. There is no general public tool where anyone can paste text and obtain an official Claude watermark reading. The access conditions are in Anthropic’s help article.
A Unicode finding count or a human-writing score does not answer that question. For a remover, ask which watermark key was tested, which texts were used, and whether the result kept the facts. My downloadable SynthID test includes all fifty document-method pairs, with a failed translation kept in the data. It uses a research key, not either vendor’s production key.
What survives, and what kills it
This is the chapter everyone actually wants, so let me be precise about what published research says.
Copy and paste changes nothing. The mark is in the words, and the words came along.
Light editing usually leaves the signal alive. Anthropic’s careful wording is that the mark “may persist through some editing”. You fixed a typo and swapped two sentences, most of the biased word choices are still there.
Substantial rewriting can weaken a statistical watermark. DIPPER and later approaches such as SIRA are part of the published attack literature. In my reference-key study, full LLM rewriting crossed below the fixed threshold in 8/10 or 10/10 cases across the tested batches. A corrected generator also changed Chinese translation from 1/10 to 5/10 below-threshold results on new corpora. Those configured-workflow results do not measure a private Claude or Gemini key. The free tool uses a separate multistage rewrite-and-review pipeline.
The rewriting model matters. Since the watermark originates from the generating model, the specific system used for your “cleaning” pass is critical. If you rephrase text that already contains a watermark using another model that also applies watermarks, you remove one watermark and apply a new one. Each interaction with a marking model marks again. To successfully strip the mark without adding a new one, you must use a model that does not embed watermarks. Currently, this category includes most of the market: DeepSeek, Qwen, Mistral, Meta’s Llama and Muse families, and any open-weight models running locally. Always verify the specific model and endpoint, as supported Claude versions now include markings, and providers may alter their services.
So the marks can be weakened or removed, the research is public about it, and I see no point pretending otherwise. The real question is a different one, and it is the question I think the law got wrong.
Text generated by AI and your own text that AI touched are two fundamentally different things. A machine wrote the first one. You wrote the second one, every idea in it is yours, and the model fixed commas or smoothed the style. The watermark cannot reliably tell these apart. To be fair to the law, it does contain an exemption for standard editing and changes that do not substantially alter your text. The mark just has no idea about any of that: the implementation marks whatever share of the words the model ended up choosing, exemption or not. And the person pointing a detector at your writing will not spend one minute on the difference, you will just be “flagged as AI”.
And another reader on r/ClaudeAI pushed me to state one more thing I had blurred: this law does not put a marking duty on you. Article 50 tells the model providers to mark what their models generate. The only duty it puts on a person is to disclose AI-generated text you publish on matters of public interest, and even that reaches you only if you are in the EU or selling into it. Stripping a mark from your own writing is not the thing the Act regulates. I am not a lawyer, and where the US state copies of this idea end up, with California’s version already being fought in court, is a separate and unsettled story.
So if your texts are genuinely yours and a model only polishes them, I think you are fully in your right to make sure they do not carry a mark, and the clean way to do that exists: run the polishing pass on a model that does not mark. That is a tool choice, it is allowed, and it is exactly what I did with my own pipeline. Hiding machine-written work where machines are banned is a different act, and I am not advising it. The difference between those two situations is the entire point of this guide, and the law has not caught up to it yet.
What this changes for you, by role
If you use Claude at work for ordinary tasks, calm down first. Detection access is limited today, and the mark still cannot establish who wrote your ideas. For most jobs in 2026, “AI touched this” is about as scandalous as “spellcheck touched this”. The people who should actually think are those working under rules that ban AI use. The mark did not create their risk, it only made the existing risk more real.
Now one prediction, and one trap inside it. Once detectors exist, someone will probably build a dashboard that claims to show the “share of AI” in your document, even though a watermark score cannot honestly measure that: a detector reports how confident it is that a mark is present, and that is all. The trap does not need the dashboard, though. Take a text you wrote entirely yourself, in your own language, and translate it with a marking model. Every word of the output was chosen by the model, and Anthropic’s FAQ says plainly that a translation carries the mark for exactly that reason. Fully your ideas, fully your text, and to any detector it is simply marked AI output. That is not a corner case, that is every non-native speaker who translates their own writing, and it is flatly unfair to them. If that is you, translate with a model that does not mark. Today that choice is still yours to make.
If you are a student, your situation is the ugliest one, because the incentive to misuse detectors as authorship judges is strongest in education, and the false-accusation problem (which OpenAI itself cited as a reason not to ship) lands on you. The standard advice says keep your drafts as evidence. I will be honest with you: nobody is going to read your drafts. If your text is genuinely yours and you only used AI to clean it up, be smart instead, and do the cleanup with tools that leave no mark. The list is one chapter up.
If you write and care about your voice, like I do, the meaningful question is quieter. When a model polishes your dictated thoughts, the mark blends your authorship with the machine’s statistics, and you may simply not want that blend in your text. This is a preference, and you are allowed to have it. It is the main reason my own pipeline moved its final writing pass to open-weight models.
If you build AI products, two practical notes. The transparency duties from the European law can land on you as the deployer, so read what your model vendor marks and what it expects you to disclose. And your model choice is now also a provenance choice for every one of your users, which is a strange new kind of responsibility to inherit from an API.
What I expect next
Anthropic says older Claude models will get the watermark later, so the “models launched after August 2” line will quietly expand. The detection tools will show up, first for platforms, then for everyone, and the first public false-accusation scandal will follow shortly after. Other signatories will ship their own marks, because they promised the EU they would. And the paraphrase arms race will continue, since the strongest removal tool is just another language model, and those are not getting worse.
The open-weight escape hatch stays open, as far as anyone can tell. Nobody has shown a mechanism that forces a model running on your own hardware to mark its output, and the companies that did not sign keep releasing weights.
I will keep the table above updated as statuses change. If I got a fact wrong or you know a status changed, send it through the form below, I would honestly rather fix it than be right on the internet.