Benchmark report · Detector v5 · September 2026
How accurate is ParaTrace?
Every figure on this page was measured on text the detector never saw in training, scored by the production service at one fixed setting, the Strict operating point.
99.6 %
of AI-written answers detected
26,086 answers · 89 models · 26 developers
0.8 %
of human texts wrongly flagged
19,989 held-out texts, same setting
97.0 %
still detected under evasion attacks
11 attack types · 16,500 texts
20 ms
to score a 300-word text
on our servers
I · Detection
It catches 99.6 % of AI-written answers
26,086 answers to real user prompts, written by 89 models from 26 developers, at a setting that wrongly flags fewer than 1 in 100 human texts.
| Developer | Detected | Answers |
|---|---|---|
| OpenAI | 99.6 % | 4,597 |
| 99.3 % | 3,783 | |
| Anthropic | 99.5 % | 3,485 |
| Meta | 99.8 % | 2,138 |
| Mistral | 99.7 % | 1,958 |
| Alibaba | 99.8 % | 1,835 |
| DeepSeek | 99.9 % | 1,393 |
| Reka | 99.8 % | 1,200 |
| Microsoft | 99.7 % | 1,072 |
| 01.AI | 99.9 % | 900 |
| Cohere | 99.8 % | 822 |
| xAI | 99.8 % | 579 |
| Zhipu | 99.7 % | 373 |
| Amazon | 100 % | 312 |
| NVIDIA | 99.7 % | 308 |
| Nexusflow | 100 % | 300 |
| MiniMax (small sample) | 100 % | 183 |
| Databricks (small sample) | 99.4 % | 168 |
| Moonshot (small sample) | 100 % | 132 |
| Tencent (small sample) | 100 % | 54 |
| Other developers | 98.8 % | 494 |
† Fewer than 200 answers, so indicative only. “Other developers” groups six smaller developers, a research fine-tune and one anonymous test model. Snapshots and reasoning modes of one model count once.
Late-2025 flagship models
- Anthropic99.4 % of 318
- Google99.3 % of 668
- xAI99.0 % of 102
- OpenAI98.9 % of 357
Each developer’s leading model as of late 2025. Answers were collected up to October 2025. Newer models are not yet benchmarked.
New models it never saw
98.6 %
of 1,172 answers from late-2025 model versions that were not in its training data.
Developers it never saw
99.6 %
of 3,753 answers from developers with no model at all in its training data.
Clear-cut verdicts
98 %
of AI-written answers score 99 % or higher, far from the cut-off, so most verdicts don’t hinge on a close call.
II · False alarms
Fewer than 1 in 100 human texts wrongly flagged
0.84 % of 19,989 held-out human texts were flagged as AI, and the longer the text, the rarer a false alarm.
- Under 50 words in the text, 3.5 % (30 of 859 texts)
- 50–99 words in the text, 1.6 % (26 of 1,599 texts)
- 100–199 words in the text, 0.9 % (55 of 6,157 texts)
- 200–399 words in the text, 0.6 % (50 of 8,419 texts)
- 400+ words in the text, 0.2 % (7 of 2,955 texts)
Under 50 words, 3.5 % are wrongly flagged, which is why the service warns on short texts. From 400 words it is 0.24 %.
| Writing | Flagged | Texts |
|---|---|---|
| Essays | 0.2 % | 5,561 |
| Questions & answers | 0.7 % | 3,369 |
| Reviews | 0.8 % | 1,639 |
| Other web writing | 0.8 % | 1,590 |
| Forum posts | 0.9 % | 1,851 |
| Encyclopedia articles | 1.1 % | 1,682 |
| Scientific abstracts | 1.3 % | 1,595 |
| Recipes (small sample) | 1.6 % | 185 |
| News articles | 1.6 % | 2,319 |
| Book excerpts (small sample) | 5.5 % | 164 |
| Emails (small sample) | 5.9 % | 34 |
† Small sample, so indicative only.
These are held-out documents, never used in training, from the same kinds of sources the detector learned from. Writing unlike those sources can be flagged more often. See Where to be careful below.
III · Evasion
Evasion tricks cost it little
Across 11 common attacks, 97.0 % of AI texts are still detected, against 98.1 % before any attack. For text from chat assistants it is 99.1 %.
- AI paraphrasing, 93.8 % of 1,500 texts, 95 % confidence interval 92.5–94.9 %
- Deleted articles (a, an, the), 95.3 % of 1,500 texts, 95 % confidence interval 94.1–96.3 %
- Random upper/lower case, 96.1 % of 1,500 texts, 95 % confidence interval 95.0–96.9 %
- Deliberate misspellings, 96.7 % of 1,500 texts, 95 % confidence interval 95.7–97.5 %
- Synonym swaps, 96.9 % of 1,500 texts, 95 % confidence interval 95.9–97.6 %
- British/American spellings, 97.7 % of 1,500 texts, 95 % confidence interval 96.8–98.3 %
- Changed numbers, 98.1 % of 1,500 texts, 95 % confidence interval 97.3–98.7 %
- Look-alike letters, 98.2 % of 1,500 texts, 95 % confidence interval 97.4–98.8 %
- Invisible characters, 98.2 % of 1,500 texts, 95 % confidence interval 97.4–98.8 %
- Extra spaces, 98.2 % of 1,500 texts, 95 % confidence interval 97.4–98.8 %
- Extra paragraph breaks, 98.3 % of 1,500 texts, 95 % confidence interval 97.5–98.8 %
1,500 texts per attack. Lines show 95 % confidence intervals, and the axis starts at 90 %.
The hardest attack
93.8 %
of AI texts rewritten by an AI paraphrasing model are still caught.
Character tricks
Look-alike letters, invisible characters and extra spaces are undone before scoring, so they have no effect on the verdict.
Independent documents
The attacked texts come from a public robustness benchmark. None of these documents were used in training.
IV · Beyond training
It holds up beyond its training data
Detection that only works on familiar material is brittle. These tests use genres, models and settings the detector was never shown.
A genre it never studied
92.8 %
of 5,000 AI-written poems detected, although poetry was absent from training. For poems from chat assistants it is 98.6 %.
Models held back on purpose
99.4 %
of 8,118 texts from two widely used model versions that were deliberately kept out of training.
Any sampling setting
95.2 %
for text generated with random sampling, the hardest setting. With a repetition penalty it is > 99.9 %.
V · Formats
Short or long, prose or code
From 50 words up, detection stays high whatever the length or format of the answer.
| Length | Detected | Answers |
|---|---|---|
| 50–99 words | 98.9 % | 3,091 |
| 100–199 words | 99.6 % | 5,251 |
| 200–399 words | 99.9 % | 8,624 |
| 400–799 words | 99.9 % | 6,707 |
| 800 words or more | 99.3 % | 2,378 |
| Content | Detected | Answers |
|---|---|---|
| Prose | 99.7 % | 20,999 |
| With code | 99.8 % | 2,211 |
| With math notation | 98.9 % | 1,566 |
| With tables | 100 % | 807 |
| With no formatting at all | 99.1 % | 6,252 |
It isn’t just spotting formatting. Answers with no headings, lists or bold text at all are detected 99.1 % of the time.
VI · Speed
Fast enough to check everything
Scoring takes milliseconds, even for long documents.
20 ms
to score a 300-word text
70 ms
to score a typical 1,500-word text
Time on our servers from text in to verdict out, text preparation included. Network and queue time are not included.
VII · Limits
Where to be careful
What we measured that falls short, and what we haven’t measured yet.
- Persuasive essays by adults, a kind of writing kept out of training entirely, were wrongly flagged 23.6 % of the time (41 of 174). Human writing unlike the training sources can be flagged far more often than the averages above.
- Human poetry was wrongly flagged 5.7 % of the time, across 1,733 poems.
- Short texts under 50 words were wrongly flagged 3.5 % of the time. The service warns on anything shorter.
- Synonym-swapping tools can mislead it. Human text run through one was flagged 16.1 % of the time.
- AI-paraphrased human writing counts as AI. By design, 48.0 % of it was flagged, because the final wording came from a model.
- English only. We have not yet evaluated documents that mix human and AI writing, human text lightly edited with AI, or paraphrasing and “humanizer” tools beyond those tested here.
- Newer models are not yet benchmarked. Answers were collected up to October 2025.
This is a statistical estimate, not proof. Don't use it as the sole basis for decisions about a person.
VIII · Method
How we measured
- One setting throughout.
- Every figure uses the Strict operating point, calibrated so that 1 % of 20,000 human validation texts are flagged.
- AI-written answers.
- 26,086 responses to real user prompts from public chat logs, collected between 2024 and October 2025, plus a small set generated in-house in September 2026. Prompts used for testing were excluded from training.
- Human writing.
- 19,989 held-out documents, never used in training, from 19 sources that include essays, questions and answers, news, forum posts, reviews, encyclopedia articles and scientific abstracts.
- Evasion and new genres.
- Documents from a public robustness benchmark, none used in training. Each attack was applied to the same 1,500 AI texts.
- Scoring.
- Every text was scored by the production service, text preparation included, exactly as when you submit it.
- Rates.
- The detection rate is the share of AI-written texts flagged as AI. The false-alarm rate is the share of human-written texts flagged as AI. Ranges are 95 % confidence intervals (Wilson).