OpenAI releases MentalHealthBench, an open benchmark built with 80+ licensed mental health clinicians across 22 countries.
OpenAI Releases MentalHealthBench, Built With 80+ Clinicians
Mental health AI evaluation has lacked coverage of everyday, ambiguous conversations outside emergency scenarios. An open benchmark built with licensed clinicians across many countries and specialties provides a shared standard for measuring model progress in this domain.
The full picture
OpenAI has released MentalHealthBench, an open benchmark for evaluating AI model performance in realistic mental health conversations. The benchmark was built with more than 80 licensed psychologists and psychiatrists from 22 countries, covering 19 languages and nearly 20 subspecialties. Each synthetic conversation uses a custom expert rubric, with criteria requiring agreement from at least 2 of 3 reviewers and no contradiction from the third. GPT-6 Astra scores 57.3 on the benchmark compared to GPT-4o's 32.1. GPT-5.6 Sol serves as the LLM judge to grade model answers against human-written criteria. OpenAI is releasing the benchmark openly so other researchers can examine the methods, run their own evaluations, and build on the work. Prior mental health AI evaluations focused on emergencies and broad safety, leaving everyday and ambiguous conversations under-measured.
How it developed
Sources
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free