OpenAI released an open benchmark called MentalHealthBench on September 23 to measure how well large language models respond in mental health conversations. The questions and grading rubrics were co-written by more than 80 licensed psychologists and psychiatrists from 22 countries and regions, covering nearly 20 subspecialties, working across 19 languages combined.
The benchmark includes 1,215 synthetic conversations paired with 5,262 expert-written grading criteria. Each conversation was reviewed by at least three experts, and each criterion carries a clinical-importance weight from -10 to +10: a response that steers a user toward harm loses points, while one that correctly recognizes an emergency and offers an appropriate path to help scores near the maximum.
How the questions break down
The conversations are split into three urgency tiers: everyday emotional and stress-related exchanges make up 53.5%, high-risk conversations involving serious mental health concerns account for 18.2%, and emergencies involving immediate physical safety make up 28.3%.
By who is asking: adults account for 68.1%, teenagers aged 13 to 17 for 21.2%, clinical practitioners for 5.8%, and caregivers for 4.9%. Scoring concentrates on four behaviors: safety, proactively gathering context, respecting user autonomy, and giving actionable advice.
Among the non-English conversations, the publicly listed breakdown includes 105 in Spanish, 54 in Hindi, 34 in Arabic, and 29 in Portuguese.
How each model scored
OpenAI published the following scores, calculated after task truncation:
| Model | Score |
|---|---|
| GPT-6 Astra | 57.3% |
| GPT-6 Sol | 53.9% |
| Claude Opus 5.5 | 52.4% |
| GPT-6 Luna | 50.2% |
| GPT-4o (March 2025 version) | 32.1% |
| Gemini 2.5 Pro | 29.5% |
Two reference scores appear alongside the table. Answers written specifically with knowledge of the grading criteria reached 99.0%, showing a perfect score is attainable. Answers experts wrote by hand without seeing the criteria scored only 38.5%, lower than all four of the newer models listed above. One possible explanation is that the criteria score whether a response covers a specific point, while people in real conversations rarely cover every point. OpenAI's announcement still stresses that ChatGPT is not a substitute for professional therapy.
Arthur Evans, CEO of the American Psychological Association, backed the benchmark:
"Mental health exists on a continuum and AI systems engaging people across that range need grounding in both clinical science and lived experience."
Limitations spelled out in the announcement
OpenAI listed several limitations of its own. No benchmark can cover every nuance of a private conversation, and score differences across languages can only be described, not cleanly separated from urgency level, topic, cultural background, or user identity. Another point outside observers are likely to notice: the organization that wrote the questions is also the one at the top of the leaderboard, and the Gemini row still uses 2.5 Pro rather than Google's current flagship model.
What it means for Chinese-language users
No domestic Chinese model appears anywhere on this leaderboard. The publicly listed non-English samples don't mention how many Chinese-language conversations were included, so there is currently no data on how models perform in Chinese-language mental health conversations.
For Chinese companies building AI companion and emotional-support apps, what's worth borrowing is the framework itself: MentalHealthBench breaks teenagers out as a separate group, raises the share of emergencies to nearly 30%, and assigns a positive or negative weight to every grading criterion. Now that the dataset is open, domestic developers can translate the questions, bring in local clinical experts to rewrite the criteria, and run their own models against it if they choose to. Whether any company does so, and whether they publish the results, will become clear over the coming months.
Sources: OpenAI's official announcement, Unite.AI, CocoLoop, Crypto Briefing, Investing.com; cross-checked for conversation count, number of grading criteria, urgency and user-identity shares, and each model's scoring basis.