Harvard Trial: OpenAI o1 Correctly Diagnoses 67% of ER Patients vs 50-55% by Triage Doctors
Harvard Magazine·high signal
A Harvard Medical School/Beth Israel Deaconess study published in Science tested OpenAI's o1 against physician pairs on 76 real ER patient records. The AI matched exact or near-exact diagnoses 67% of the time versus 50-55% for human doctor pairs. With richer clinical data, o1 reached 82% (humans 70-79%, not statistically significant). Lead author Arjun Manrai stressed the trial was text-only — no images, sounds, or nonverbal cues.