OpenAI launches mentalhealthbench to evaluate AI in mental health conversations
OpenAI has introduced MentalHealthBench, an open benchmark developed with over 80 licensed mental health experts from 22 countries to assess AI responses in realistic mental health conversations. The benchmark evaluates models across key behaviors like safety, context seeking, and user agency, with criteria weighted from -10 to +10 based on clinical importance. It includes scenarios for adults, teens, caregivers, and clinicians, spanning non-acute, high-acuity, and emergency situations. Experts reviewed synthetic conversations to create rubrics, and responses are graded by GPT-5.6 Sol. The benchmark also highlights differences between expert and user perspectives on helpful AI support, emphasizing practical steps and tone over context gathering. OpenAI is also enhancing ChatGPT’s responses in sensitive conversations and launching ChatGPT for Teens with additional protections.
MentalHealthBench aims to set a higher standard for AI support in mental health by enabling researchers to identify model improvements and gaps. It is part of broader efforts to improve AI’s role in mental health, including grants for research and collaboration with organizations like the Partnership on AI. The benchmark’s release invites the community to examine its methods and contribute to future advancements in AI-driven mental health support.