News
LLM Judges Score 0.50–0.63 on Detecting Omissions Versus 0.79–0.94 on Detecting Added Errors
A paper submitted to arXiv on August 31, 2026 (2608.31016) by Sebastian Fox, Luke Markham, Ryan Lail and Michael Karotsieris shows LLM judges evaluating AI-generated clinical notes verify presence but not absence. Judges hit 0.79–0.94 accuracy discriminating added or altered content but only 0.50–0.63 on omissions, and no single-note judge design flagged omissions more reliably than it flagged perfect notes. Restructuring the task to first enumerate transcript facts and then verify each against the note raised omission detection to 24.6–36.9% at false-alarm rates of 2.7–6.2%, with a physician reviewer siding with the pipeline method on 10 of 10 disagreements (p=0.002).
Source
↳ Follow the thread