Skip to content
ParaTrace

Benchmark report · Detector v5 · September 2026

How accurate is ParaTrace?

Every figure on this page was measured on text the detector never saw in training, scored by the production service at one fixed setting, the Strict operating point.

  • 99.6 %

    of AI-written answers detected

    26,086 answers · 89 models · 26 developers

  • 0.8 %

    of human texts wrongly flagged

    19,989 held-out texts, same setting

  • 97.0 %

    still detected under evasion attacks

    11 attack types · 16,500 texts

  • 20 ms

    to score a 300-word text

    on our servers

I · Detection

It catches 99.6 % of AI-written answers

26,086 answers to real user prompts, written by 89 models from 26 developers, at a setting that wrongly flags fewer than 1 in 100 human texts.

Detection rate by developer
DeveloperDetectedAnswers
OpenAI99.6 %4,597
Google99.3 %3,783
Anthropic99.5 %3,485
Meta99.8 %2,138
Mistral99.7 %1,958
Alibaba99.8 %1,835
DeepSeek99.9 %1,393
Reka99.8 %1,200
Microsoft99.7 %1,072
01.AI99.9 %900
Cohere99.8 %822
xAI99.8 %579
Zhipu99.7 %373
Amazon100 %312
NVIDIA99.7 %308
Nexusflow100 %300
MiniMax (small sample)100 %183
Databricks (small sample)99.4 %168
Moonshot (small sample)100 %132
Tencent (small sample)100 %54
Other developers98.8 %494

† Fewer than 200 answers, so indicative only. “Other developers” groups six smaller developers, a research fine-tune and one anonymous test model. Snapshots and reasoning modes of one model count once.

Late-2025 flagship models

  • Anthropic99.4 % of 318
  • Google99.3 % of 668
  • xAI99.0 % of 102
  • OpenAI98.9 % of 357

Each developer’s leading model as of late 2025. Answers were collected up to October 2025. Newer models are not yet benchmarked.

New models it never saw

98.6 %

of 1,172 answers from late-2025 model versions that were not in its training data.

Developers it never saw

99.6 %

of 3,753 answers from developers with no model at all in its training data.

Clear-cut verdicts

98 %

of AI-written answers score 99 % or higher, far from the cut-off, so most verdicts don’t hinge on a close call.

II · False alarms

Fewer than 1 in 100 human texts wrongly flagged

0.84 % of 19,989 held-out human texts were flagged as AI, and the longer the text, the rarer a false alarm.

False alarms by length
  • Under 50 words in the text, 3.5 % (30 of 859 texts)
  • 50–99 words in the text, 1.6 % (26 of 1,599 texts)
  • 100–199 words in the text, 0.9 % (55 of 6,157 texts)
  • 200–399 words in the text, 0.6 % (50 of 8,419 texts)
  • 400+ words in the text, 0.2 % (7 of 2,955 texts)

Under 50 words, 3.5 % are wrongly flagged, which is why the service warns on short texts. From 400 words it is 0.24 %.

False alarms by kind of writing
WritingFlaggedTexts
Essays0.2 %5,561
Questions & answers0.7 %3,369
Reviews0.8 %1,639
Other web writing0.8 %1,590
Forum posts0.9 %1,851
Encyclopedia articles1.1 %1,682
Scientific abstracts1.3 %1,595
Recipes (small sample)1.6 %185
News articles1.6 %2,319
Book excerpts (small sample)5.5 %164
Emails (small sample)5.9 %34

† Small sample, so indicative only.

These are held-out documents, never used in training, from the same kinds of sources the detector learned from. Writing unlike those sources can be flagged more often. See Where to be careful below.

III · Evasion

Evasion tricks cost it little

Across 11 common attacks, 97.0 % of AI texts are still detected, against 98.1 % before any attack. For text from chat assistants it is 99.1 %.

AI texts still detected, by attack
  • AI paraphrasing, 93.8 % of 1,500 texts, 95 % confidence interval 92.5–94.9 %
  • Deleted articles (a, an, the), 95.3 % of 1,500 texts, 95 % confidence interval 94.1–96.3 %
  • Random upper/lower case, 96.1 % of 1,500 texts, 95 % confidence interval 95.0–96.9 %
  • Deliberate misspellings, 96.7 % of 1,500 texts, 95 % confidence interval 95.7–97.5 %
  • Synonym swaps, 96.9 % of 1,500 texts, 95 % confidence interval 95.9–97.6 %
  • British/American spellings, 97.7 % of 1,500 texts, 95 % confidence interval 96.8–98.3 %
  • Changed numbers, 98.1 % of 1,500 texts, 95 % confidence interval 97.3–98.7 %
  • Look-alike letters, 98.2 % of 1,500 texts, 95 % confidence interval 97.4–98.8 %
  • Invisible characters, 98.2 % of 1,500 texts, 95 % confidence interval 97.4–98.8 %
  • Extra spaces, 98.2 % of 1,500 texts, 95 % confidence interval 97.4–98.8 %
  • Extra paragraph breaks, 98.3 % of 1,500 texts, 95 % confidence interval 97.5–98.8 %

1,500 texts per attack. Lines show 95 % confidence intervals, and the axis starts at 90 %.

The hardest attack

93.8 %

of AI texts rewritten by an AI paraphrasing model are still caught.

Character tricks

Look-alike letters, invisible characters and extra spaces are undone before scoring, so they have no effect on the verdict.

Independent documents

The attacked texts come from a public robustness benchmark. None of these documents were used in training.

IV · Beyond training

It holds up beyond its training data

Detection that only works on familiar material is brittle. These tests use genres, models and settings the detector was never shown.

A genre it never studied

92.8 %

of 5,000 AI-written poems detected, although poetry was absent from training. For poems from chat assistants it is 98.6 %.

Models held back on purpose

99.4 %

of 8,118 texts from two widely used model versions that were deliberately kept out of training.

Any sampling setting

95.2 %

for text generated with random sampling, the hardest setting. With a repetition penalty it is > 99.9 %.

V · Formats

Short or long, prose or code

From 50 words up, detection stays high whatever the length or format of the answer.

AI answers detected, by length
LengthDetectedAnswers
50–99 words98.9 %3,091
100–199 words99.6 %5,251
200–399 words99.9 %8,624
400–799 words99.9 %6,707
800 words or more99.3 %2,378
AI answers detected, by content
ContentDetectedAnswers
Prose99.7 %20,999
With code99.8 %2,211
With math notation98.9 %1,566
With tables100 %807
With no formatting at all99.1 %6,252

It isn’t just spotting formatting. Answers with no headings, lists or bold text at all are detected 99.1 % of the time.

VI · Speed

Fast enough to check everything

Scoring takes milliseconds, even for long documents.

  • 20 ms

    to score a 300-word text

  • 70 ms

    to score a typical 1,500-word text

Time on our servers from text in to verdict out, text preparation included. Network and queue time are not included.

VII · Limits

Where to be careful

What we measured that falls short, and what we haven’t measured yet.

  • Persuasive essays by adults, a kind of writing kept out of training entirely, were wrongly flagged 23.6 % of the time (41 of 174). Human writing unlike the training sources can be flagged far more often than the averages above.
  • Human poetry was wrongly flagged 5.7 % of the time, across 1,733 poems.
  • Short texts under 50 words were wrongly flagged 3.5 % of the time. The service warns on anything shorter.
  • Synonym-swapping tools can mislead it. Human text run through one was flagged 16.1 % of the time.
  • AI-paraphrased human writing counts as AI. By design, 48.0 % of it was flagged, because the final wording came from a model.
  • English only. We have not yet evaluated documents that mix human and AI writing, human text lightly edited with AI, or paraphrasing and “humanizer” tools beyond those tested here.
  • Newer models are not yet benchmarked. Answers were collected up to October 2025.

This is a statistical estimate, not proof. Don't use it as the sole basis for decisions about a person.

VIII · Method

How we measured

One setting throughout.
Every figure uses the Strict operating point, calibrated so that 1 % of 20,000 human validation texts are flagged.
AI-written answers.
26,086 responses to real user prompts from public chat logs, collected between 2024 and October 2025, plus a small set generated in-house in September 2026. Prompts used for testing were excluded from training.
Human writing.
19,989 held-out documents, never used in training, from 19 sources that include essays, questions and answers, news, forum posts, reviews, encyclopedia articles and scientific abstracts.
Evasion and new genres.
Documents from a public robustness benchmark, none used in training. Each attack was applied to the same 1,500 AI texts.
Scoring.
Every text was scored by the production service, text preparation included, exactly as when you submit it.
Rates.
The detection rate is the share of AI-written texts flagged as AI. The false-alarm rate is the share of human-written texts flagged as AI. Ranges are 95 % confidence intervals (Wilson).