Fetching from the wire…
Policy2026-07-30 · source-backed
The July 29 technical report claims AUROC 0.9916 with a 0.0041% false positive rate and 0.3396% false negative rate, plus better out-of-distribution generalization and adversarial robustness than Pangram 3, with fine-grained discrimination of edits and mixed AI-human co-writing rather than binary authorship. The FPR is the number to scrutinize, because it's the figure that determines whether detection can be used punitively at scale. Two orders of magnitude below typical classifier claims deserves independent replication before anyone builds policy on it. It arrives the same week a NeurIPS reviewer reported receiving entirely LLM-generated rebuttals.
Each link below shares sources, entities, or timing with this story.
LLM uses OpenAI / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-07-27.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
LLM uses OpenAI / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-06-19.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-14.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-22.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-10.