Skip to content
ParaTrace

Unlimited free credit top-ups, for a limited time. Running low? Ask for more credits and we top you up, as often as you need. This offer ends soon, so catch it while it lasts.

Field test · Detector v7 · October 2026

AI detection in scientific papers

NeurIPS checked its 2026 position papers for AI-written text with Pangram. We ran the same check with ParaTrace.

Before screening its position papers, NeurIPS tested Pangram on papers accepted at an AI ethics conference, some from 2022, before ChatGPT, and some from 2025. We ran the same check with ParaTrace v7. On 106 papers from 2022 it flags none, at any level, exactly like Pangram. On 123 papers from 2025 it flags 9 as at least half AI-written, 7.3 % against 1.0 % for Pangram, which points to AI-assisted writing that Pangram left unflagged.

Papers from before ChatGPT

0of 106 papers flagged by ParaTrace, at any AI score

The same result as Pangram, which flagged 0 of 159.

What it shows

Clean on human research, alert to AI help

No false alarms

0.0 %

of 106 papers from before ChatGPT flagged, at any level, the same as Pangram. Even passage by passage, only 0.4 % lean AI.

More AI flagged in 2025

7.3 %

of papers from 2025 reach an AI score of 50 % or more, against 1.0 % for Pangram.

Close at the top

1.6 %

of papers from 2025 reach 90 % or more, against 1.0 % for Pangram. That is 2 papers in each sample.

The results

Share of papers by AI score

Papers from 2022Written before ChatGPT was released, so any paper flagged here is a false alarm.
DetectorPapers≥ 50 %≥ 90 %100 %
Pangram 3.3.21590.0 %0.0 %0.0 %
ParaTrace v71060.0 %0.0 %0.0 %
Papers from 2025Written when AI writing tools were in wide use, with no record of which papers used them.
DetectorPapers≥ 50 %≥ 90 %100 %
Pangram 3.3.22041.0 %1.0 %0.0 %
ParaTrace v71237.3 %1.6 %0.8 %

The AI score is the share of a paper’s passages scored as AI. Pangram’s rows are from the NeurIPS post, ours from scoring the papers with the production service.

How we measured

The same check

Papers.
NeurIPS tested Pangram on papers accepted at an AI ethics conference in 2022 and 2025, close in style to its own position papers. We scored 106 papers from 2022 and 123 from 2025 from the same conference, with no errors. Pangram’s figures cover 159 and 204, so compare the shares rather than the counts.
AI score.
Each detector splits a paper into passages of a few hundred words and scores each one. The AI score is the share of passages flagged as AI. A score of 100 % means AI use across many parts of a paper, not that every word came from AI.
No answer key for 2025.
Which papers from 2025 used AI is not recorded. Because ParaTrace flags no paper from before ChatGPT, false alarms are an unlikely cause of its extra flags, though that is an inference rather than proof.
Setting.
ParaTrace v7, scored by the production service in October 2026 at the setting this site uses. Passages from 2022 scored as AI were 18 of 4,512, or 0.4 %.
Sources.
Pangram’s figures are from the NeurIPS post AI-generated papers in the NeurIPS 2026 position paper track of 2 June 2026. NeurIPS and Pangram Labs, Inc. have not reviewed this comparison.
More results.
See how ParaTrace compares with Pangram on 21 public benchmark results, and our benchmark report.