MentalHealthBench: OpenAI's open benchmark for AI responses across the full range of mental health conversations

On September 23, 2026, OpenAI released MentalHealthBench, an open benchmark for how models respond in realistic mental health conversations. The paper (Ali Malik, Declan Grabb, Karan Singhal and colleagues) describes 1,215 synthetic conversations built with privacy-preserving techniques to mirror real ChatGPT usage patterns, spanning non-acute everyday stress, high-acuity distress and emergencies, and four user personas: adults, teens aged 13 to 17, caregivers and clinicians, across multiple languages and regions. OpenAI’s framing is that most prior evaluations in this area test emergency handling against broad pass/fail criteria, leaving the far more common non-crisis conversations largely unmeasured.

The rubrics come from more than 80 licensed psychologists and psychiatrists from 22 countries who speak 19 languages. For each conversation, experts wrote criteria for the final model turn, each weighted from -10 to +10, so harmful behaviors are penalized as well as good ones rewarded. Each conversation was reviewed by at least three experts, and a criterion was kept only if at least two agreed and a third did not contradict it. Grading is automated, with GPT-5.6 Sol scoring responses against the expert criteria. A separate study with 44 adults from 16 countries who had used AI for emotional support found users valued practical next steps and tone more than experts did, while experts weighted context-gathering and careful interpretation; those user views did not change the scoring.

On the task-clipped score, GPT-6 Astra led at 57.3 percent, followed by GPT-6 Sol at 53.9, Claude Opus 5.5 at 52.4 and GPT-6 Luna at 50.2; Grok 4.7 scored 41.3, Gemini 3.8 Flash 35.5 and GPT-4o (March 2025) 32.1. The decomposition shows why: Claude Opus 5.5 earned more positive points than GPT-6 Astra (73 versus 69 percent of the positive budget) but took a much larger penalty (-33 versus -19). The paper singles out context-seeking and calibrating urgency as the weakest areas, and notes that Anthropic’s Sonnet and Opus models score well on preserving user agency.

This fills a genuine gap: a public, clinician-built yardstick for the everyday emotional conversations that happen at a scale of more than a billion weekly ChatGPT users, not just crisis scripts. What it does not establish is neutral ground. OpenAI built the benchmark, an OpenAI model grades it, and OpenAI models take the top two places; the conversations are synthetic; and a high rubric score measures agreement with expert guidance on one reply, not whether anyone’s wellbeing actually improved.