用第二个 AI 文本检测器复跑 NeurIPS 论文筛查:2022 年论文均未被标记,2025 年结果不一致
[P] We re-ran NeurIPS's pre-LLM vs 2025 paper check with a second AI-text detector. Neither flags a 2022 paper; on 2025 papers they disagree [P]
ParaTrace 作者用自家检测器 v7 复跑了 NeurIPS 此前用 Pangram 3.3.2 做的论文 AI 文本筛查:2022 年(ChatGPT 之前)的论文两个检测器在任何阈值下都未标记,2025 年论文上两者出现分歧,ParaTrace 在 ≥50% 阈值下标记 123 篇中的 9 篇,Pangram 约标记 204 篇中的 2 篇。
Disclosure: I build ParaTrace, a commercial AI-text detector with a free tier. I'm posting because these results include a disagreement with Pangram that I can't settle on my own, and this sub is the right place to pick it apart.
Context. Prior to screening the NeurIPS 2026 position-paper track, the chairs ran Pangram 3.3.2 on the papers of an AI ethics conference, some of which were submitted in 2022 (before ChatGPT) and some in 2025 (NeurIPS blog, 2 June 2026). We ran the same check with ParaTrace v7 via our production service in October 2026.
Share of papers by AI score (AI score = share of a paper's passages, each a few hundred words, scored as AI)
| Year | Detector | Papers | ≥ 50 % | ≥ 90 % | 100 % |
|---|---|---|---|---|---|
| 2022 | Pangram 3.3.2 | 159 | 0.0 % | 0.0 % | 0.0 % |
| 2022 | ParaTrace v7 | 106 | 0.0 % | 0.0 % | 0.0 % |
| 2025 | Pangram 3.3.2 | 204 | 1.0 % | 1.0 % | 0.0 % |
| 2025 | ParaTrace v7 | 123 | 7.3 % | 1.6 % | 0.8 % |
Rows of Pangram are from the NeurIPS post. We only managed to get 106 and 123 of the papers, so compare shares, not counts.
- 2022: neither detector flags a paper at any level. Per passage, 18 of our 4,512 (0.4%) scored as AI. With n = 106, the 95% upper bound on the paper-level false-positive rate is still about 3.5%.
- 2025: at >= 90% both flag 2 papers. At >= 50% we flag 9 of 123, Pangram about 2 of 204.
There is no ground truth for which 2025 papers used AI. I see four explanations:
- AI-assisted writing that Pangram 3.3.2 doesn't flag.
- ParaTrace counts AI-polished human writing as AI, by design (in our tests 42.8% of human text paraphrased by an AI is flagged). NeurIPS allows AI copy-editing, so some of our extra flags may be exactly that, and acting on them under NeurIPS's policy would be wrong.
- False positives on something in 2025 human writing that a 2022 control can't catch.
- Paratrace score on chunks of size 512, Pangram 4 has the same approach as per their technical report (https://arxiv.org/pdf/2607.27183) but this study was with Pangram 3.x so not sure if used diff size.
I don't think the clean 2022 result makes (3) impossible, just less likely. I suspect (2) accounts for some of the gap, and I'd like to test on 2025 papers with known AI-use disclosures if anyone has them.
Head-to-head on Pangram's public test sets. We scored v7 on the public tests from Pangram's July 2026 technical report (Pangram 4) and kept the 21 results we could match exactly. We're within 5 points of Pangram 4 on 15 of them. A selection, including where we lose:
| Test | Pangram 4 | ParaTrace v7 |
|---|---|---|
| Student essays, fully AI-written (caught at 1 % FPR) | 100 % | 100 % |
| Student essays, AI-paraphrased (caught at 1 % FPR) | 100 % | 100 % |
| Student essays, polished by AI (caught at 1 % FPR) | 100 % | 94.2 % |
| 8 genres vs recent chat models (AUROC) | 1.0000 | 1.0000 |
| News, reviews, novels etc., full length (caught at 1 % FPR) | 100 % | 98.74 % |
| Same, under 50 words (caught at 1 % FPR) | 99.70 % | 73.02 % |
| Same, after a humanizer, full length (caught at 1 % FPR) | 98.93 % | 21.42 % |
| Same, after a humanizer, under 50 words (caught at 1 % FPR) | 73.32 % | 19.36 % |
| Human texts wrongly flagged, default setting | 0.00 % | 1.86 % |
| Well-known authors: AI texts missed | 2.86 % | 4.38 % |
| Well-known authors: human texts wrongly flagged | 0.00 % | 1.41 % |
"Caught at 1% FPR" means each detector is tuned so that 1% of that test's human texts are flagged. Pangram's numbers are as published, and Pangram Labs has not reviewed this. TL;DR: level on clean AI text, behind on short texts, way behind on humanizers, and a higher false-positive rate at default settings.
Our own benchmark (held-out data only, one fixed setting calibrated to flag about 1 % of human validation texts)
We evaluate a language model-based AI detector on 26,082 AI-generated answers to real user prompts (90 models, 26 developers, up to Oct 2025). We achieve 99.6% detection rate, 98.6% on 1,172 answers from late-2025 model versions not used in training, and 99.8% on 3,948 answers from developers with no model in training. On 22,691 human texts from 43 sources, we achieve 0.84% false positive rate, which decreases with text length (2.8% for texts with <50 words (24/846), 1.1% for 100-199 words, 0.17% for 400+ words (7/4,088)). Finally, we show that 96.0% of texts are still detected across 11 attacks (1,500 texts each, from a public robustness benchmark). The most challenging attack is paraphrasing AI-generated text, with 91.5% detection rate.
Where it fails
• Human writing unlike the training data: Persuasive essays by adults, a genre excluded from the training data, were falsely flagged 18.4% of the time (32/174). This is the number I worry about the most: human writing unlike the training data can be flagged far more often than the averages above.
• Human text run through a synonym-swapping tool: 16.4% flagged.
• Humanizer tools (table above).
• English only. Not yet evaluated: documents mixing human and AI writing, and models released after October 2025.
Full tables, confidence intervals and method: benchmark report · all 21 results vs Pangram · scientific papers. It's free to try at paratrace.net, including a few checks without an account.
The verdicts are statistical estimates. I don't think any detector, including ours, should be the sole basis for rejecting a paper or failing a student. Happy to answer questions on method.
submitted by /u/AltruisticCouple3491
[link] [留言]
来源:r/MachineLearning · reddit.com