Fetching from the wire…
Public story · 2026-09-08 · high
No human reviewed the rejections, and the same detector scored the three track chairs' own recent papers 24% to 69% on its own scale.
Why now: The field note breaking down the rejection scores is now public, and NeurIPS hasn't said whether it's revisiting the detector settings that produced them.
NeurIPS desk-rejected 178 Position Paper Track submissions, 18.4% of 969, with no human review or appeal, per a field note reconstructing the rejections. The 178 authors lost a submission cycle outright, with no path to appeal and no human reading their paper before the rejection.
Every rejection ran on the same detector's score. Seventy-seven of the 178 scored above 0.9 on Pangram's detector. Another 79 scored above 0.8 on papers with a single author. Twenty-two scored above 0.5 even though their authors had checked the box denying AI use.
The number itself moved depending on how the detector was configured. Pangram 3.3.2's default window settings flagged 42.7% of all 969 submissions in the 90-100% band. Shrinking the analysis window to roughly 100 words brought that down to 12.7%, a difference of more than three times from one parameter.
Independent researchers tested the detector by running the three track chairs' own recent papers through it. The scores came back between 24% and 69%, on the same scale that flagged the 178 rejected submissions.
The field note doesn't say whether NeurIPS has responded to the window-size finding, or whether an appeal process exists for the 178 authors.
Each link below shares sources, entities, or timing with this story.
The r/MachineLearning post describes "Claude-speak everywhere" in both submission and author responses, with the authors acknowledging LLM assistance. NeurIPS desk-rejected hundreds of position-track papers flagged by AI detectors earlier this cycle and sanctions no LLM use du...
The July 29 technical report claims AUROC 0.9916 with a 0.0041% false positive rate and 0.3396% false negative rate, plus better out-of-distribution generalization and adversarial robustness than Pangram 3, with fine-grained discrimination of edits and mixed AI-human co-writin...
Bryan Cantrill's argument isn't that LLM prose is immoral. It's that it doesn't work. His September 5 essay "The Revolt of the Reader" took 403 points on Hacker News, and the numbers underneath it come from Cynthia Dunlop's survey of 668 developers. 78% stop reading immediatel...
A July 29 arXiv paper from Peter Kirgis, Sayash Kapoor, and Andrew Schwartz introduces shadow evaluations: agents attack the central research question of an unpublished high-quality paper, and the original authors grade the result. Across two unpublished NeurIPS 2026 submissio...
Keogh posted to r/MachineLearning (369 upvotes) that on most of the benchmark datasets used across NeurIPS, SIGKDD and VLDB time-series anomaly detection papers, plain SPC matches or beats the published SOTA, scoring perfectly on the ECG trace he shows and trivially on the TAO...
A Princeton-led study with Stanford, Berkeley, Johns Hopkins, Toronto, Georgetown and the UK AI Security Institute handed Claude Opus 4.8 and GPT-5.6 Sol Ultra the central research question from unpublished NeurIPS 2026 submissions. The original authors graded the output as re...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.