OpenAI releases MentalHealthBench, scored on rubrics from 80+ clinicians in 22 countries
OpenAI Blog·medium signal
OpenAI published MentalHealthBench on 23 September. Responses in realistic mental-health conversations are graded against expert-written rubric items weighted from -10 to +10, so harmful behavior loses points. Results are broken out by acuity, by user type (adults, teens, caregivers, clinicians) and across ten behavior dimensions. Clinicians working in 19 languages helped build it. OpenAI released it openly for other labs to run. The post shows scores only in charts, with no headline numbers.