Head to head · Detector v7 · October 2026
ParaTrace against Pangram
21 results on the same public tests, with Pangram’s figures as Pangram published them and ours measured on the same texts.
ParaTrace matches Pangram on fully AI-written and AI-paraphrased student essays, catching 100 % of them, and misses fewer AI texts than Pangram 3 in a test of writing by well-known authors. Across the 21 results measured on the same tests, it is within 5 points of Pangram 4 on 15 and of Pangram 3 on 16. The remaining gaps are on humanized and very short texts.
Within 5 points
- Against Pangram 415 of 21 results
- Against Pangram 316 of 21 results
Results where ParaTrace is level, ahead, or no more than 5 points behind.
Strengths
Where ParaTrace matches or beats Pangram
Level on AI essays
100 %
of fully AI-written and AI-paraphrased student essays caught, the same as Pangram 4 and Pangram 3.
A perfect AUROC
1.0000
on eight kinds of writing, from news to poetry, level with Pangram 4 and ahead of Pangram 3 at 0.9999.
Ahead of Pangram 3
4.38 %
of AI texts missed in a test of writing by well-known authors, against 5.05 % for Pangram 3.
The breakdown
All 21 results, test by test
| Result | Pangram 4 | ParaTrace | Gap |
|---|---|---|---|
| Fully AI-writtenAUROC | 1.000 | 1.000 | Level |
| Fully AI-writtenAI caught at 1 % false alarms | 100 % | 100 % | Level |
| AI-paraphrasedAUROC | 1.000 | 1.000 | Level |
| AI-paraphrasedAI caught at 1 % false alarms | 100 % | 100 % | Level |
| Polished by AIAUROC | 1.000 | 0.998 | −0.2 |
| Polished by AIAI caught at 1 % false alarms | 100 % | 94.2 % | More than 5 points behind, −5.8 |
| OverallAUROC | 1.000 | 0.999 | −0.1 |
| OverallAI caught at 1 % false alarms | 100 % | 97.7 % | −2.3 |
| Result | Pangram 4 | ParaTrace | Gap |
|---|---|---|---|
| All textsAUROC | 1.0000 | 1.0000 | Level |
| All textsAI caught at 1 % false alarms | 100 % | 99.96 % | −0.04 |
| Result | Pangram 4 | ParaTrace | Gap |
|---|---|---|---|
| Full lengthAUROC | 1.0000 | 0.9995 | −0.05 |
| Full lengthAI caught at 1 % false alarms | 100 % | 98.74 % | −1.26 |
| Under 50 wordsAUROC | 0.9999 | 0.9527 | −4.72 |
| Under 50 wordsAI caught at 1 % false alarms | 99.70 % | 73.02 % | More than 5 points behind, −26.68 |
| Humanized, full lengthAUROC | 0.9996 | 0.9373 | More than 5 points behind, −6.23 |
| Humanized, full lengthAI caught at 1 % false alarms | 98.93 % | 21.42 % | More than 5 points behind, −77.51 |
| Humanized, under 50 wordsAUROC | 0.9810 | 0.7685 | More than 5 points behind, −21.25 |
| Humanized, under 50 wordsAI caught at 1 % false alarms | 73.32 % | 19.36 % | More than 5 points behind, −53.96 |
| Human writingHuman texts wrongly flagged | 0.00 % | 1.86 % | −1.86 |
| Result | Pangram 4 | ParaTrace | Gap |
|---|---|---|---|
| AI textsAI texts missed | 2.86 % | 4.38 % | −1.52 |
| Human textsHuman texts wrongly flagged | 0.00 % | 1.41 % | −1.41 |
The gap is in points, ParaTrace’s figure against Pangram’s, so a minus means ParaTrace is behind. ▲ marks where ParaTrace is ahead, and ▼ where it is more than 5 points behind. For AUROC, one point is 0.01.
How we compared
Like for like
- Only matching results.
- Every Pangram figure here is one Pangram published for a public test, in its Pangram 4 technical report of July 2026. We scored ParaTrace on the same texts and measured it the same way, and left out every result we could not match exactly.
- AI caught at 1 % false alarms.
- Each detector is set so that 1 in 100 of the test’s human texts is flagged, and we count the share of AI texts it then catches.
- AUROC.
- How well a detector ranks AI text above human text across every possible setting, from 0.5 for a coin toss to 1 for perfect.
- Missed and wrongly flagged.
- Each detector at its standard setting. For ParaTrace that is the setting this site uses.
- Points.
- Gaps are in percentage points, and a minus always means ParaTrace is behind, even where a lower figure is better. “Level” means the same figure at the precision Pangram published.
- Versions.
- ParaTrace is detector v7, scored by the production service in October 2026. Both Pangram versions are as Pangram reports them. For the test on well-known authors, Pangram 3 is version 3.3.2.
- About Pangram.
- Pangram is a product of Pangram Labs, Inc., which has not reviewed this comparison.
- More results.
- Our benchmark report covers many more models, evasion attacks and kinds of writing, and a field test on scientific papers repeats the check NeurIPS ran with Pangram.