Two Nominally Identical GPT-5.4 Screening Runs Disagreed on 94 Records, 29 of Them Verified Eligible
A preregistered comparison of human and LLM title-and-abstract screening on 1,131 records against 316 verified eligible studies found no workflow recovered everything. Human workflows and two GPT-5.4 file-batch runs retained 42.2 to 45.0% of records at 82.3 to 82.9% recall, while Gemini 3.1 file batches got the highest recall at 83.9% but retained 56.7%. The run-to-run result is the sharp one: two identical GPT-5.4 configurations agreed on 91.7% of records yet differed on 94, including 29 verified eligible records retained by only one run, and all-at-once configurations recovered fewer eligible records than file-batch ones, so processing configuration is a substantive property of the deployed system rather than an implementation detail.
↳ Follow the thread