Can rewriting remove AI text watermarks? My SynthID test
A corrected SynthID generator changed Chinese translation from 1/10 to 5/10 removals. The LLM workflow reached 10/10 in that control. Read the methods and data.
P.S. A working AI text watermark remover from this research is at painintheagent.com/tools/ai-text-watermark-remover.
Kirill Balakhonov · 2026-08-22 · updated 2026-10-07
What the test found
My choice for this tool remains a full rewrite with source checks. The follow-up tests support that choice, with an important correction to my original experiment: I had applied SynthID before temperature, top-k and top-p, which produced a stronger mark than the standard integration did in the control.
With the corrected order, the single-pass Qwen3.7 Plus workflow crossed below the fixed detection threshold in 10/10 reports, DIPPER in 9/10 and Chinese translation in 5/10. A fresh LLM run on the original sources reached 8/10. The historical 10/10 remains a result of its original configuration. The follow-up tables also separate claim preservation, input protection and the current web tool from that research benchmark.
Imagine you wrote a text yourself, from the first word to the last, and handed it to a model (Claude/ChatGPT/Gemini) only to translate it, say, into another language. Or you took a post you had written by hand for one social network, or a transcript of a talk you gave, and asked an AI to turn it into a post for LinkedIn or Reddit. A watermarking model can add an invisible statistical signal during that work. A matching detector may later identify the signal in the output. And that mark will not tell anyone whether you wrote the substance yourself, did the research, put your own experience into words (and assembled it into a readable text for a specific purpose), or simply typed the prompt “write an interesting LinkedIn post”: to an outside observer the mark looks the same.
Judging by the comments under my previous Reddit post, which collected 150k+ views (the post accompanied an article where I gathered everything known about text watermarks at the moment), some people think watermarks are a good thing, and some consider them a serious violation of their digital freedom and do not want any trackers or marks in their texts. I belong to the second group and decided to work the technical side out myself. These measurements concern detection and text fidelity; they do not establish what is permitted in every jurisdiction or writing context.
- Is it possible today to reliably check whether a generated text carries a watermark?
- Is it possible to remove a watermark in some relatively cheap way (without relying on loud claims by people on Twitter who vibe-coded some watermark remover with no evidence that it works at all)?
- Is it possible to verify that the watermark was removed?
- How badly does the text degrade after removal, and what has to be checked in the text afterwards?
- Build a working watermark remover for myself and share it with the community (links to the open-source repository and the live demo are at the end).
1. How a watermark works, in plain words
There are no invisible characters or hidden spaces in the text. The watermark lives in the choice of words itself. When a model writes, at every step it has several suitable next words, and normally it picks between them almost at random. Take a die that has been weighted a little: it can still land on any side, but one side comes up slightly more often. A watermark does roughly that to word choice. Among equally good options the model picks a little more often the ones that sit on a secret list, and the list itself is recomputed each time from a secret key and a few previous words.
On one word the bias is invisible. On five hundred words, anyone holding the same key sees it clearly. The detector walks through the text, recomputes the list for every word from the key and the preceding words, and counts how often the choice matched it. If the matches are noticeably above half, the text is marked. The model itself does not need to run for this; all you need is the text, the key and the same way of splitting text into tokens.
This explanation is simplified but keeps the essential mechanism. The DeepMind team outlined their SynthID Text method in a 2024 Nature article and released the associated code, which I replicated for this study. Google integrates SynthID into Gemini, and Anthropic characterizes Claude’s watermark as a variant of this same method. Anthropic currently identifies Fable 5.1 and Mythos 5.1 as models that apply markings. Their detection API is available only in private preview for qualified entities, not the general public. My study relies on my own watermark key, not either vendor’s key. OpenAI’s textGrain, announced for ChatGPT in October 2026, is a separate scheme that this study did not touch. I maintain the current statuses in a table within my earlier post.
The same mechanism leads to a conclusion I will keep repeating. The detector’s score means agreement with one specific key. It does not say “what percentage of the text was written by AI”, and it cannot tell a translation of your own text from a text the model invented from scratch.
2. How I simulated applying the watermark
At first I took the easy road and built a toy mark: a list of synonym pairs
(big / large), with the key deciding which member of the pair to use. It ran
on a CPU, it was transparent, and experiments on it were fast. The result was
odd: translation and paraphrase removed such a mark, light synonym edits did
not, and at the same time it was clear that this proved nothing about the real
mark. The real mark is smeared as a weak bias over thousands of consecutive
choices and depends on context every time. The toy scheme was answering the
wrong question.
So I reproduced the scheme from the DeepMind paper. The transformers library
contains the reference implementation of SynthID Text: it plugs into generation
and at each step slightly shifts the probabilities of the next token based on
the key and a few previous tokens. I set my own key, generated a corpus with
that shift, and for every text produced a paired clean control from the same
prompt and the same seed but without the mark. Each clean twin was a separate unmarked completion from the same prompt and
seed, with potentially different wording and facts. I used its score as a
comparison baseline for the marked completion.
The later review found that my custom processor was placed before temperature/top-k/top-p. Standard Transformers 5.15.1 puts the SynthID processor after them. The original configuration and results are retained; the new control below uses the standard order.
The first corpus I made with the tiny Qwen2.5-0.5B-Instruct on an RTX 4090.
That was a mistake, and it cost me the most time. A small model with free word
choice could not hold a long English report together: the texts jumped into
Cyrillic and Chinese characters and contained fragments of nonexistent words.
When such a text is handed to a rewriter with the instruction to keep every
number and claim, the rewriter has one safe move left, and that is literal
copying. I got the handsome result “the mark survived 177 of 181 rewrites”,
nearly published the conclusion that no remover was possible, and only then saw
that all the residual signal sat in the copied chunks.
pavement, cosmic number~RúB.B5.N иметь any opportunity that departs, web browser files must synthesize within day.
The final corpus was written by Qwen2.5-14B-Instruct in float16 on a single
NVIDIA A40 rented on RunPod for $0.44 per hour. Two generation runs took about
half an hour and cost under 25 cents. The result was ten coherent fictional
reports of 480 to 639 words by the repository’s word counter: a library switching to a digital visitor log, a
repair cafe tracking appliances that come back twice, and so on. I manually edited this fictional corpus using 24 recorded replacements across
eight documents and their clean twins. The edits changed dates, quantities,
expense categories and operational rules; they were not all deductions from
obvious contradictions. The detector was recomputed after the declared edits
to each text.
Qwen only wrote the sources here; the key and the mark do not depend on it.
3. Which removal methods exist in the literature
Attacks on text watermarks have been studied since 2023, when the first scheme of this kind appeared (the “green list” paper by Kirchenbauer and colleagues). One well-studied approach is to make another model rewrite the text. A dedicated model was even built for it, DIPPER, with 11 billion parameters (Krishna and colleagues, 2023), which can paraphrase with an adjustable strength of change and was built specifically to evade detectors.
The most relevant work for me came out in December 2024 from Martin Vechev’s group at ETH Zurich (Jovanović, Gloaguen, Vechev). They tested SynthID Text specifically, used that same DIPPER as the baseline rewriter, and reported that the success rate of “scrubbing” the mark exceeds 90%. They counted a case as a success when the detector missed the mark at a very strict false-alarm threshold and the automatic meaning-similarity metric P-SP rated the paraphrase as good. Remember that metric, I will come back to it.
I took four simple methods and one hard one. The simple ones: light synonym
replacement, a full paraphrase, translation into German and back, translation
into Chinese and back. All four were run by the same model, Qwen3.7 Plus from
Alibaba, through an API at temperature zero; all forty transformations cost
about 7 cents. The hard method, DIPPER, I had to deploy myself: I brought the
published model up on RunPod on an NVIDIA A100 80GB at $1.39 per hour; the run
over ten texts took about six minutes and cost under 14 cents. The settings came from the ETH work.
The input handling was different. The four API methods masked money amounts, numeric percentages, quotes, URLs and selected other spans, then restored them exactly. DIPPER received the ordinary source text with whitespace normalized, without the same protection or paragraph-preservation wrapper. The fidelity results compare those configured workflows; they do not isolate model quality.
The literature is broader than these five choices. The authors of WASH combine probability distributions during generation; their tested schemes do not include SynthID. They also show that coordinated watermarks can survive the averaging. The ToBlend authors blend token continuations to attack AI-content detectors. SIRA chooses parts to rewrite using self-information. These are distinct procedures, and none establishes guaranteed SynthID removal by merging seven finished texts.
Natzir’s independent implementation also reports effective Chinese translation under different models, keys and calibration. I have not independently rerun those measurements. They reinforce the need to qualify my original translation result.
One more trap that is often forgotten. The rewriting has to be done by a model that does not apply the same kind of mark itself. Otherwise you wash one mark off and immediately get another. Qwen through the API did not add the mark I was studying in my experiment, and the safest option is open weights on your own hardware, where you control the generation procedure.
4. How watermarks are detected and how quality was checked
My detector was exactly the one described in section 1: it walks through the tokens, computes the “list” from the key and the preceding tokens, and counts the share of matches. The main question is where to put the threshold. Clean text matches about half the time, marked text noticeably more, and between them you need a boundary that clean text almost never crosses. I calibrated it with a hundred thousand random-table null samples for each text and chose a threshold targeting one-percent false positives under that conditional simulation. Ten clean controls do not establish a population false-positive rate of one percent. On the ten pairs of the corpus the detector found the mark in all ten marked texts and in none of the clean ones.
Here was the second trap. My first threshold, taken from the old corpus, was one I believed to be a one-percent threshold, but in fact it gave about five percent false alarms. If you see someone’s “remover” or “detector” quote a threshold without describing the calibration, that is a reason not to trust the numbers.
Removing the mark is not enough; the text has to remain your text with the same facts. How do you check that, when an automatic similarity metric can miss a swapped date? Here is what I did. Before any transformation, for each source I fixed ten verifiable claims: names, dates, quantities, causal links, problems, outcomes, recommendations. Then four models from four different companies read every result blind. A judge received one source and one candidate, did not know how the text had been produced, and for each of the ten claims answered: preserved, changed or missing. A claim counted as preserved if at least three judges out of four said so. Separately, the judges rated readability and usability on a 0 to 100 scale in steps of 10.
Here I stepped on the third trap, and it deserves a detailed account, because it concerns anyone who checks texts with models. In the first version of the panel, three judges out of four saw all five candidates of one document in a single prompt. An independent reviewer showed that this way judges carry errors from one neighbour to another: in one document GPT-5.6 attributed to the plain paraphrase two changes (“nearly 24%” became “exactly a quarter”, “from $34 to $30” became “from $36 to $32”) that DIPPER had actually made. That was the only fact out of a hundred the paraphrase had lost. I threw that panel away entirely and collected 200 new judgments over 50 source-candidate pairs, one candidate per prompt.
Before the rerun, every judge went through a canary of three prompts: a byte-for-byte identical text, a text with five deliberately corrupted facts, and an empty text. Claude Haiku 4.5 gave the empty text 10 out of 10 and 100% usability and found only 2 of the 5 corruptions; Grok 4.20 also credited the empty text with all ten claims. I replaced both: the final panel is GPT-5.6 Luna, Claude Sonnet 5, Gemma 4 31B and Grok 4.6, and all four passed the canary completely. The rerun of the panel cost 92 cents.
The final score of each text equals the worst of four required scores: signal removed, facts preserved, readability, usability. If the mark is removed but facts are lost, that result is useless, and the other way round.
5. Follow-up results and the original comparison
Generator correction and follow-up tests
My original generator applied the watermark before temperature, top-k and top-p. The standard Transformers integration applies it after those steps, before sampling a token. A new control used that standard order. The historical results remain available under their original configuration.
| Method | Original A | Repeat B | New corpus C | Standard-order control |
|---|---|---|---|---|
| LLM paraphrase | 10/10 | 8/10 | 10/10 | 10/10 |
| DIPPER | 10/10 | 10/10 | 10/10 | 9/10 |
| Light synonyms | 2/10 | 0/10 | 1/10 | 1/10 |
| Translation via German | 0/10 | 0/10 | 0/10 | 0/10 |
| Translation via Chinese | 0/10 | 0/10 | 1/10 | 5/10 |

B uses the original sources and key with fresh API responses. C and the control use a new key with the same briefs and seed schedule, but different generated source texts. All marked sources passed the common threshold; all clean controls stayed below it. These batches are not forty independent documents.
Selected claims in the standard-order control
| Method | Preserved | Changed | Missing | Uncertain |
|---|---|---|---|---|
| LLM paraphrase | 97 | 2 | 0 | 1 |
| DIPPER | 74 | 9 | 14 | 3 |
| Light synonyms | 100 | 0 | 0 | 0 |
| Translation via German | 99 | 0 | 0 | 1 |
| Translation via Chinese | 94 | 4 | 0 | 2 |
The input handling differed: LLM methods received protected placeholders for amounts, percentages and selected other spans, which were restored exactly. DIPPER received whitespace-normalized source text without the same wrapper. These scores compare the configured workflows; they do not isolate model quality or evaluate a similarly protected DIPPER.
The author reclassified two paraphrase cases as changed qualifications: a minimum became most shifts, and possibly became likely. One uncertain DIPPER case was preserved because the apparent inconsistency was already in its source. The original annotations and the reasons for each change are retained.
The accepted original threshold stays 0.5095383054287164. C and its control use 0.509656862745098, fixed before transformations. The simulated null targets 1% false positives; ten clean controls do not establish a population false-positive rate. The tested rewrite is single-pass Qwen3.7 Plus with protected spans. The live tool runs the same single-pass Qwen3.7 Plus method with a stricter prompt and temperature 0.7 instead of 0; this configuration was not part of these tests. Vendor production keys were not tested.
Summary JSONPair metrics CSVEvery source and resultAuthor adjudications
Evidence reviewed 2026-09-08.
Historical A: the original configuration
| Method | Mark removed | Facts | Usability | Total |
|---|---|---|---|---|
| Light synonym edits | 20% | 100% | 98% | 49.6% |
| Full paraphrase | 100% | 100% | 98% | 93.6% |
| Round trip via German | 0% | 98% | 97% | 63.2% |
| Round trip via Chinese | 0% | 87% | 83% | 61.6% |
| DIPPER from the paper | 100% | 77% | 73% | 71.0% |
This is a test of the published SynthID Text algorithm with a reference key, not of Gemini's private production detector, a Claude detector or OpenAI's textGrain detector for ChatGPT. The corpus is ten fictional reports in one length range, so 10 out of 10 describes only this sample. The share of signal removed is not the share of text written by a human or by a model.
Verified 2026-08-22. Judges: OpenAI, Anthropic, Google, xAI.
In the original configuration, the full LLM workflow crossed below the threshold in all ten texts and removed an average 94.5% of the signal. The four-model panel credited it with all 100 selected claims. DIPPER also removed the mark in all ten and took away 94.7% of the signal, but the original panel did not credit 23 of its 100 selected claims, and rated the usability of its texts at 73%. Neither round-trip translation removed the mark in a single text: the German route took away 63% of the signal, the Chinese route 62%, and that was not enough. Light synonyms removed the mark in two texts out of ten.
Does translation remove a SynthID text watermark?
Translation can weaken or remove a SynthID signal. In the original configuration, neither of my two translation routes produced a result below the threshold. All ten German round trips remained detectable. Nine Chinese round trips completed and remained detectable; the tenth failed and has no detector score. I count that failure in the ten attempted texts. Chinese translation later crossed below the threshold in 1/10 new-corpus cases with my old processor placement and 5/10 with standard placement. My original zero does not justify the claim that translation generally fails.
Translation surprised me the most, and the explanation is visible in the data. Back-translation restores the same familiar English phrasing that was in the source: after the German route 45% of five-token sequences match the original, after the Chinese route 36%. Together with the phrasing, the other die’s rolls come back. Paraphrase leaves only 9% of such sequences.
Does light editing remove a SynthID text watermark?
Light synonym edits removed the reference mark in two of ten texts. The other eight stayed detectable. This result applies to these English reports and this key; it does not predict the outcome for a particular ChatGPT, Claude or Gemini text.
Now about P-SP, the metric used at ETH. For DIPPER it is 91%, meaning that by that metric its texts are close to the source. By facts, 77% was preserved. A shared topic and similar vocabulary do not guarantee that a number, a date or a cause stayed the same. The ETH result “DIPPER removes SynthID in more than 90% of cases” held up for me, 10 out of 10. Where we part is the word “success”: I need a text that can be handed to a person without checking it against the source.
The best combined score in the original sample belonged to the protected single-pass LLM workflow. It was the only one that both removed the mark in every text and kept every verifiable claim; its final score is 93.6% against 71% for DIPPER and 62 to 63% for the translations. “100 out of 100” here means one hundred pre-selected composite claims as decided by a majority of judges, not a full inventory of every fact in the text; all 34 non-unanimous votes were later checked against the texts (more on that in section 6), and for the paraphrase not one of them was confirmed as an error.
Regarding limitations, this assessment focuses on the published scheme using my key on ten fictional reports of one length. Because Google’s key remains confidential and no public detector exists for Gemini, this experiment cannot confirm the removal of the production mark. This tool also has no access to Claude’s private detection API. Shorter texts present different challenges: fewer words result in a weaker cumulative signal and noisier conclusions. The transformations, DIPPER and final-panel subtotal was $1.1239329054; corpus generation, canaries and the later validation are separate costs. All code, data, prompts, outputs, and judgments are available in the repository.
6. How to review the result without collecting errors
For every “changed” status the judges wrote down what exactly changed, and those notes show what breaks most often. All the examples below are real, from my ten texts and the final panel. Here is what I now check in a rewritten text, and in what order.
Numbers and dates. This is the first thing to drift. DIPPER turned the tablet cost “about $7,500” into “$7,000”, the satisfaction rise “from 78% to 83%” into “from 76% to 83%”, “nearly 85%” into a precise “84.9%”, and split the total “nearly $3,500” into “$3,000 plus $500 for training”, although no such breakdown exists in the source. Every number and every date from the source must be found in the result with the same value, with no added precision and no added hedges.
Names, titles and roles. After the Chinese round trip “Florence Heights Middle School” became “High School”, and “president Johnathon Adams” turned into “chair John Adams”. After the German route “head librarian” became “library director”, and “math day” turned into “Pi Day”. DIPPER renamed the company “GreenCycle Delivery Services” to “GreenCycle Bicycle Courier Services”. Check names, job titles and proper nouns letter by letter.
Subject and cause. DIPPER attributed the 40% savings to renting radio equipment, while the source was about maintaining physical audio guides, and replaced “auto-fill for common items” with “automatic reminders”. After the German route, “gear system” became “transmission system”; whether these name the same component depends on context and remains uncertain in the follow-up review.
Strength of wording. The recommendation to consider a science club became, after the original paraphrase, “the document urges school administrators to establish a permanent science club”. Check whether “perhaps” and “consider” turned into “must”.
Losses and inventions. In one text DIPPER kept the old waiting time “more than ten minutes” and lost the new one, “under two minutes”; in another it retained the additional bus but changed the responsible unit from Fleet Services to the maintenance department and made deployment immediate; the Chinese route once returned an empty result, and in another text glued a sentence together so that the figure 85% was left hanging without a subject. Remove everything that was not in the source; look for everything that was.
Verbatim chunks. If large fragments of the source were carried over unchanged, the mark’s signal stayed in them, however new the rest of the text looks. In my data the paraphrase kept under 10% of verbatim five-token sequences, the translations 36 to 45%, and that was enough for the mark to stay.
The fastest way is to put the source and the result side by side and go through exactly this list instead of reading the result “for general impression”. You can use a second model as a judge, I did, but judges make mistakes even without neighbours. In the final panel, 34 of 500 claim-candidate cases had non-unanimous decisions, across 2,000 individual votes; Claude Code (Fable 5) checked each of them against the full texts, and the two overturned decisions I verified by hand: “head librarian → library director” is a real change of role that three judges forgave, while “Environmental Studies → the Environment” in DIPPER’s output does not change the fact. So in my setup the majority decided, a second model re-read the disputed places, and the final word on them belonged to a human, and for your own texts I recommend the same order: the judge marks, you decide.
7. What demo I decided to build
A simple page. The current tool runs the same single-pass Qwen3.7 Plus method as the benchmark above, with a stricter prompt that asks the model to keep terms and formatting and a temperature of 0.7 instead of 0; this configuration was not part of these tests. You paste your text and press one button; the tool paraphrases it and shows the rewrite with every changed word in bold and the share of 5-word sequences that were replaced. Full paid runs also include one meaning check by a second model, which flags possible changes without editing the rewrite. Free runs skip that check. Compare the result with your source before copying it; a model check can still miss errors.
The whole experiment, including the rewriter and judge prompts, is open in the repository on GitHub. Ten fictional reports show the mechanism, they do not prove “it works on everything”, so check the changed places in your own texts. You can try it at painintheagent.com/tools/ai-text-watermark-remover: three free runs on the first 1,000 characters of a text, then $10 a month for 100,000 characters.
For those who want to check the numbers
| What was measured | Result |
|---|---|
| Mean signal, marked texts | 57.03% |
| Mean signal, clean texts | 49.98% |
| Threshold at a 1% false-alarm target | 50.95% |
| Marked texts detected | 100% |
| Clean texts falsely detected | 0% |
| Method | Signal removed | Exact 5-grams | P-SP | Readability | Pipeline |
|---|---|---|---|---|---|
| Light synonym edits | 50.5% | 54.6% | 98.4% | 98% | 100% |
| Full paraphrase | 94.5% | 9.4% | 94.8% | 98% | 100% |
| Round trip via German | 63.2% | 45.0% | 98.5% | 99% | 100% |
| Round trip via Chinese | 61.6% | 36.0% | 88.3% | 86% | 90% |
| DIPPER from the paper | 94.7% | 9.4% | 91.3% | 86% | 100% |
Judges saw one candidate per prompt; the first version of the panel showed three judges five candidates at once and carried errors between them, so it was replaced entirely. Non-unanimous votes: 34; after a manual check against the texts 2 were overturned (claims after adjudication: Light synonym edits 100%, Full paraphrase 100%, Round trip via German 97%, Round trip via Chinese 87%, DIPPER from the paper 78%). Sources of the new corpus: Qwen/Qwen2.5-14B-Instruct. Rewriter for synonyms, paraphrase and translations: qwen/qwen3.7-plus (Alibaba), temperature 0. DIPPER: kalpeshk2011/dipper-paraphraser-xxl, revision c1fbf7a958a2, NVIDIA A100 80GB PCIe. Judges: openai/gpt-5.6-luna, anthropic/claude-sonnet-5, google/gemma-4-31b-it, x-ai/grok-4.6. The old threshold 50.67% gave 5.1% false alarms on the exact null distribution instead of the claimed 1%. Correlation between literal copying and detector score in the old corpus: 0.988. The transformations, DIPPER and final-panel subtotal stayed under $1.12; corpus generation, canaries and follow-up tests are separate.
Data and citation
SynthID Text removal and fact preservation: ten-document retest
Five rewriting methods tested on ten English reports with a reference SynthID Text key. The dataset contains detector scores, four-judge fact checks, manual adjudication and the source and rewritten texts. One failed translation is retained. These are research-key measurements, not tests of Claude's or Gemini's production watermark.
- Method summary (5 rows)
- Per-document scores (50 rows)
- Scores, texts and claim checks (50 rows)
- Methodology and citation
- MIT license
The main table uses the original blinded panel. The downloads also show the two fact-check decisions corrected during manual adjudication. A failed translation stays in the data with no detector score. These texts come from my fictional research corpus.
To cite this version: Kirill Balakhonov (2026), SynthID Text removal and fact preservation: ten-document retest, version 1.0.0. Link to this section. The files use the MIT license.
Measurements: 2026-08-22. Export: 2026-09-06. Pinned source table.