LLM-Generated Clinical AI Evaluation Rubrics Match Clinician Agreement Across 823 Patient Encounters
arXiv·medium signal
Shah, Hines, and Downs present a case-specific rubric methodology for evaluating clinical AI documentation systems, validated across 823 encounters with 20 clinicians authoring 1,000+ rubrics. They demonstrate that LLM-generated rubrics can approximate clinician inter-rater agreement, solving the key bottleneck where expert review per scoring instance is too slow and expensive for safe iterative deployment. The methodology generalizes beyond healthcare to any domain requiring expert-level AI evaluation.