How the AI detector works
I trained this detector on labelled human and machine writing. It gives a score for the whole document, with the fullest coverage for English. Writing highlights come from a separate review and do not determine that score.
Since 16 September 2026 English and Russian texts are scored by the trained classifier described in the sections below, in its fourth version (local-origin-tfidf-v5): the training split was extended with 120 additional English examples labelled human, and the calibration and thresholds were refit by the same rule. The development, calibration, final-test and transfer partitions are unchanged, so the results below remain a first evaluation of this version.
Tested alternative: agreement with Undetectable and Clever
From 14 to 15 September 2026 the English score came from a different detector, which answered a practical question: would Undetectable or Clever flag this text? A language model (GLM-5.3-Flash) rated eleven explicit signs of AI writing, such as stock vocabulary, uniform sentence rhythm, generic statements and assistant-like tone, and gave its own estimate. A calibration trained on those two detectors' verdicts turned the ratings into the percentage. It is paused while the two are compared on rewritten text.
In grouped cross-validation on 382 texts (human writing from before 2020, AI texts from public datasets and their humanized versions), the verdict at 50% matched "Undetectable or Clever flags it" in 79% of texts, Undetectable alone in 73% and Clever alone in 69%. It separated human from AI source texts with ROC AUC 0.99 and flagged 5 of 40 human texts. The previous classifier, described below, matched in 71% with ROC AUC 0.91.
This is a research checkpoint trained on a few hundred texts. The sections below describe the trained classifier that scores texts now.
The score and its thresholds
The classifiers learn word and character patterns from 8,348 training texts. English combines a character model with a word model. Russian uses a character model. Formatting is normalized before scoring, including HTML tags, typographic quotes and split contractions found in exported corpora.
I used 1,158 development texts to choose the approach. Another 1,452 examples set the calibration and decision thresholds. Neither stage used the final test inputs. The thresholds target a low false-positive rate; texts between them receive an uncertain result.
Calibration used equal numbers of human and machine examples. A score of 80 is not a universal 80% probability for any text someone submits, and it does not mean 80% of the words were written by AI. Different writing settings and different proportions of AI text can change how the score behaves.
Final test: 2,000 new inputs
Each language had 500 human-labelled and 500 machine-labelled inputs. An uncertain or unscored machine input counts as missed in the detection rate below. Seven English inputs fell below the 80-word requirement after formatting and word counting; they received no score and remain in the totals.
| Language | AI detected | Human false positives | AUROC |
|---|---|---|---|
| English | 459/500 (91.8%) | 15/500 (3.0%) | 0.992 |
| Russian | 391/500 (78.2%) | 18/500 (3.6%) | 0.954 |
The 95% sampling interval for the English false-positive rate is 1.8% to 4.9%. For Russian it is 2.3% to 5.6%. The observed rates are below 5%; the intervals do not establish that the rate will stay below 5% in use.
AUROC measures how well the scores rank the labelled groups. It is not an accuracy percentage. AUROC and calibration error use only inputs that received a score.
Differences between writing settings
| Sample | AI detected | Human false positives |
|---|---|---|
| Stories | 57/60 | 0/60 |
| English abstracts | 47/60 | 3/60 |
| Financial Q&A | 40/40 | 3/40 |
| English news | 59/60 | 0/60 |
| Everyday Q&A | 36/40 | 1/40 |
| Reviews | 55/60 | 3/60 |
| Russian news | 343/450 | 16/450 |
| Russian abstracts | 48/50 | 2/50 |
| Essays | 60/60 | 1/60 |
| Technical Q&A | 20/20 | 1/20 |
| How-to guides | 85/100 | 3/100 |
The financial Q&A sample had three false positives among 40 human texts. Small genre samples have wide uncertainty. Russian coverage is mainly news and scientific abstracts, so these results do not establish the same performance on Russian fiction or personal messages.
A generator excluded from training
I removed all DeepSeek-V3 generations from training, development and calibration. A separate set of 100 new DeepSeek texts and 100 human texts tested transfer to that generator. The detector flagged 100/100 machine texts and 2/100 human texts. These are separate from the 2,000 inputs above.
What changed from the earlier detector
The earlier version counted writing-style observations. On its 64 observed examples, treating the stronger style bands as a prediction flagged 25 of 32 machine texts and 10 of 32 human texts. The trained classifier flags 26 of those machine texts and none of the human texts. That is a regression comparison on known examples, not a fresh test.
I also kept the failed trials. The first local classifier flagged only 11 of 100 machine answers in a separate Q&A test. The next trial improved Q&A coverage but flagged only 28 of 100 machine texts in an unseen procedural genre. Those failures prompted changes to training coverage. Each later final test used new examples; the failed results were not reused as an unseen test.
Sources and repeatability
The corpora are DACTYL, AINL-Eval 2025, HC3 and M4. DACTYL includes GPT-4o, Claude 3.5, Gemini 1.5 and other generators. HC3 and parts of M4 use older generators. This is not a validation of every current model.
The manifests fix source revisions, text digests and split membership. Whole source groups stay together where identifiers are available. Exact text matches and shared 16-word spans are excluded across partitions, including recoverable human examples inside generation prompts. DACTYL and AINL omit some source-family identifiers, so semantic overlap cannot be ruled out completely. I rely on the dataset labels rather than independent verification of every author's identity.
The site's JavaScript implementation was checked on all 2,200 final and transfer inputs. Its raw scores matched the Python calculation within 0.000001. No new paid model calls were needed for this classifier evaluation.
Download the benchmark report. It includes confusion counts, calibration error, sampling intervals and model digests. Original research texts are not republished.
The highlights and their percentage
The review selects up to eight passages with writing patterns such as stock phrasing or repeated structure, and a fixed list adds literal stock phrases. Their percentage is highlighted words divided by all source words, overlaps counted once. That editing measure is separate from the trained classifier: it explains what stands out. The Watermark-free AI Humanizer can create a new version and report the classifier score again; its internal rewriting method can change independently.
Each explanation must name exact words from its selected passage. Observations that fail this check are left out. The report groups identical explanations with their matching passages and hides the source preview when there are no highlights.
Use the result alongside other evidence. Mixed authorship, substantial editing and writing outside the tested settings can be harder to classify. A low score is not a certificate of human authorship.
Checked 2026-09-16 UTC. Classifier: local-origin-tfidf-v5. Style coverage: marked-word-coverage-v1.