# SynthID Text removal and fact preservation: ten-document retest

Version 1.0.0. Research measured on 2026-08-22; packaged on 2026-09-06.

Five rewriting methods tested on ten English reports with a reference SynthID Text key. The dataset contains detector scores, four-judge fact checks, manual adjudication and the source and rewritten texts. One failed translation is retained. These are research-key measurements, not tests of Claude's or Gemini's production watermark.

## Files

- summary.csv: five method rows, matching the article table. Original panel
  and manually adjudicated fact-preservation scores are separate columns.
- pairs.csv: 50 document-method rows. An incomplete Chinese translation is
  retained with detector_status=not_scored and an empty candidate_mean_g.
- results.json: the same scores with the source and candidate texts, panel
  claim decisions, two adjudication changes and pinned source URLs/hashes.
- LICENSE: the source repository's MIT license.

Signal removed is relative to each source's clean twin; it is not the share
of human-written text. Mark removal is a detector decision under the research
key, not a claim of universal undetectability. The threshold is 0.5095383054287164.
Facts are ten claims fixed per source, evaluated separately for each candidate
by four blinded judges. A claim needs three votes to pass. The final score
uses the original panel and is the worst required score for each document,
averaged across the ten documents; it was not recomputed after adjudication.

The corpus contains fictional reports in a narrow English length range.
The downloadable texts are the controlled research corpus, not texts entered
by users of the web tool. The repository contains the clean twins, watermark
configuration, calibration, prompts and code needed to inspect the experiment.

## Citation

Kirill Balakhonov (2026). SynthID Text removal and fact preservation:
ten-document retest. Version 1.0.0. Pain in the Agent.
https://painintheagent.com/blog/text-watermark-removal-retest/#data-and-citation

Research repository: https://github.com/krllagent/text-watermark-roundtrip
Source table: https://github.com/krllagent/text-watermark-roundtrip/blob/f323239085a7b6e0cf800e153df70dc0f7c93960/results/curated-percent-table-v2.json
Adjudication: https://github.com/krllagent/text-watermark-roundtrip/blob/f323239085a7b6e0cf800e153df70dc0f7c93960/results/curated-panel-adjudication-v1.json
Calibration: https://github.com/krllagent/text-watermark-roundtrip/blob/f323239085a7b6e0cf800e153df70dc0f7c93960/results/quality-synthid-curated-calibration-v1.json

To reproduce this export locally with both repositories checked out:
python3 scripts/export-watermark-research.py --research-repo ../text-watermark-roundtrip
