ResidencyRL trains a clinical agent through simulated residency — 60-turn encounters, 8 tool calls, 31% fewer missed red flags
Submitted 2026-08-07 (arXiv:2608.07418), ResidencyRL applies RL over multi-turn simulated patient encounters — up to 60 dialogue turns and 8 tool calls per trajectory — with reward spread across diagnostic accuracy, management quality, communication, documentation, and safety. Diagnostic accuracy under adversarial conditions rose 7.0 points (88.0% vs 81.0%), missed red flag rates fell 31%, and expert clinicians preferred the trained agent in 87.6% of side-by-side comparisons, with gains on all six axes of the AMIE multi-visit benchmark plus AgentClinic and CRAFT-MD. The transferable pattern is optimizing the full decision sequence rather than the single answer — the gap between static benchmark scores and agent competence.
Source
↳ Follow the thread