By Hajra, Clinical Psychology Research Scholar
A chatbot can sound caring before it understands what’s wrong. It can also interrupt an ordinary conversation with crisis language that feels wildly out of place. Measuring the difference requires more than counting reassuring sentences.
OpenAI MentalHealthBench, released on September 23, 2026, tackles that problem. It evaluates how AI models respond to mental health conversations, from relationship worries to emergencies. GPT-6 Astra leads the published results with 57.3%.
That number needs careful reading. It measures how responses satisfy weighted criteria written by clinicians. It does not mean 57.3% of users improved, or that the remaining responses caused harm.
The deeper finding concerns what a good answer should accomplish. Does it ask enough questions? Respect the person’s choices? Recognize urgency without inventing it? And why do clinicians’ own replies score below many models while attracting fewer penalties?
Those questions make this benchmark useful, and explain why its leaderboard cannot carry the whole story.
Table of Contents
1. What Is OpenAI MentalHealthBench?
OpenAI’s announcement describes an open benchmark developed with more than 80 licensed psychologists and psychiatrists across 22 countries. The accompanying paper contains 1,215 synthetic conversation histories and 5,262 expert-written evaluation criteria.
OpenAI MentalHealthBench: Key Facts and Evaluation Scope
| Key Fact | What It Means |
|---|---|
| 1,215 synthetic conversations | Constructed scenarios informed by patterns of real ChatGPT use |
| 5,262 rubric criteria | Specific behaviors to reward or penalize in each response |
| More than 80 clinicians | Expert input across countries, languages, and specialties |
| 53.5% non-acute scenarios | Everyday concerns with an emotional component |
| 18.2% high-acuity scenarios | Significant distress without an immediate emergency |
| 28.3% emergency scenarios | Situations calling for urgent real-world support |
| One next response evaluated | Models receive a conversation history, then produce a scored reply |
The scenario proportions describe the test’s design. They do not tell us how often ChatGPT users experience emergencies.
The benchmark evaluates AI behavior, rather than diagnosing the person speaking to it. Readers looking for a mental health screening questionnaire or a treatment service are looking for a different kind of tool.
Compared with the broader HealthBench healthcare evaluation, this AI mental health benchmark concentrates on the judgments involved in emotionally sensitive conversations. Avoiding an obviously dangerous answer is only part of the task. A useful response may also need context, restraint, practical guidance, and respect for someone’s agency.
2. MentalHealthBench Results: How the Models Compare
The following MentalHealthBench results reproduce the overall scores in the research paper’s Figure 5. They belong to the September 2026 evaluation, rather than a continuously updated ranking.
OpenAI MentalHealthBench Results: Model Scores Compared
| Model, as Reported | Overall Score |
|---|---|
| GPT-6 Astra | 57.3% |
| GPT-6 Sol | 53.9% |
| Claude Opus 5.5 | 52.4% |
| GPT-6 Luna | 50.2% |
| Muse Spark 1.3 | 48.6% |
| GPT-5.6 Sol (August 2026) | 47.0% |
| Claude Fable 5.1 | 46.4% |
| GPT-5.6 Luna (August 2026) | 44.9% |
| Claude Sonnet 5 | 44.5% |
| GPT-5 Thinking | 42.9% |
| Claude Haiku 4.5 | 41.7% |
| Grok 4.7 | 41.3% |
| Gemini 3.8 Flash | 35.5% |
| Gemini 2.5 Flash | 33.5% |
| GPT-4o (March 2025) | 32.1% |
| Gemini 3.1 Pro | 32.1% |
| Gemini 2.5 Pro | 29.5% |
The researchers sampled four responses per conversation. The paper’s charts include 95% confidence intervals, so small numerical gaps deserve caution.
Overall leadership also hides differences between subsets. On the 70 tasks containing prior user context, Muse Spark 1.3 scored 50.5%, ahead of Astra’s 46.2%. Those figures describe that particular subset, not a universal advantage in remembering people.
Questions about older or missing competitors are reasonable. The table supports comparisons among the tested versions under the reported settings. It cannot establish how an absent model would perform, or explain the researchers’ selection motives.
3. What Does the 57.3% Score Actually Mean?

MentalHealthBench scores start with conversation-specific criteria. Each carries a positive or negative weight, with larger magnitudes indicating greater clinical importance. A response earns points for meeting desirable criteria and incurs penalties for meeting undesirable ones.
Consider an invented example with 20 possible positive points. A reply earns 15 but incurs a five-point penalty. Its signed score is 10 divided by 20, or 50%.
If penalties push that response below zero, the headline scoring method floors it at zero. Researchers average across the sampled responses and tasks. This produces the reported “task-clipped” score.
Clipping happens before the final average. Consequently, subtracting the average penalty burden from average positive credit gives the signed score, which can differ from the headline figure. Anyone rebuilding the results should preserve that distinction rather than trying to reconcile two differently calculated numbers.
The result combines fulfilled expectations, omissions, and penalized behavior. That’s why subtracting 57.3% from 100% does not reveal a 42.7% harm rate. Nor does a zero identify one specific kind of failure.
OpenAI MentalHealthBench also separates positive credit from penalty burden. That breakdown matters because two models can reach similar totals through different behavior: one may cover more useful ground while also making more unwanted assumptions.
A percentage is convenient. Knowing how it was earned is more informative.
4. Who Wrote the Criteria, and How Much Did They Agree?
The MentalHealthBench methodology assigns two clinicians to independently assess each conversation and write criteria. A third reviews their work. The final rubric retains criteria supported by at least two experts and opposed by none.
Assignments reflected relevant expertise. Teen scenarios went to clinicians with current or recent adolescent-care experience. Medication-focused tasks were reserved for psychiatrists with relevant prescribing experience.
The paper’s acknowledgments name consenting contributors, with some credentials and affiliations. They are not a complete licensing register, and being listed does not imply endorsement.
Agreement also has several meanings. In a separate response-preference analysis, two experts agreed 63.4% of the time, including ties. That statistic concerns preferences between answers, rather than the rule used to retain rubric items.
Clinical judgment contains legitimate variation. The benchmark manages some of it through explicit criteria and adjudication, without establishing that every qualified clinician would choose the same next sentence.
5. Can We Trust a Benchmark Graded by OpenAI?

The clinicians define the criteria. GPT-5.6 Sol, running at high reasoning effort, judges whether each response meets each criterion. Keeping those roles distinct helps clarify the independence question.
Using an OpenAI model to grade OpenAI and competing systems creates a reason to investigate grader bias. Ownership alone does not demonstrate favoritism, just as expert involvement does not eliminate measurement error.
OpenAI MentalHealthBench makes the conversation dataset and criteria available, and the paper documents its grading prompt. Researchers can inspect what earns credit, evaluate additional models, or test alternative judges.
The most useful independent checks would ask whether model rankings survive a different grader and whether human reviewers agree with disputed judgments. Style, verbosity, and how explicitly an answer states its reasoning are possible influences worth testing.
An open release makes those investigations possible. It should not be mistaken for evidence that independent investigators have already confirmed every published result.
One result illustrates the difference between performing well and knowing the scoring rules: Astra reaches 99.0% when given the rubrics and instructed to maximize its score. That sanity check shows the criteria can be satisfied. It does not establish real-world therapeutic effectiveness.
6. Why Did Clinicians Score Below AI Models?
Clinician-authored reference replies scored 38.5%, below many tested models. Taken alone, that invites a dramatic headline. The score decomposition offers a more useful explanation.
The clinicians incurred fewer penalties but also earned fewer positive points. The authors suggest that their brief, conversational replies explain much of the difference. A clinician might ask one focused question and wait. A model can supply questions, interpretations, reassurance, and suggestions in the same turn.
For these reference replies, a fourth clinician who had not created the conversation’s rubric wrote the response without seeing the evaluation criteria or candidate model answers. They were asked for the response they would want a safe, helpful AI system to provide.
The result exposes a measurement tension: how much should one reply accomplish before listening again?
It does not prove that longer answers are better. It also cannot answer “Can AI replace therapists?” A single written continuation and an ongoing course of care are different objects of evaluation.
7. Why Can Better AI Emotional Support Still Feel Unhelpful?
Alongside the benchmark, researchers recruited 44 adults from 16 countries who had used AI for mental health or emotional support. They assessed non-acute synthetic conversations and wrote their own criteria.
Users placed more emphasis on tone and practical next steps. Clinicians emphasized gathering context and interpreting ambiguous situations carefully. The user study did not change the benchmark’s expert-based scoring criteria.
The groups’ rubric weights were 25.7% aligned, while just 1.0% directly contradicted each other. Most remaining weight represented priorities raised by only one group. Describing this as overwhelming disagreement would misread the finding.
Imagine asking for help before a difficult conversation. One reader may want a usable opening sentence. A clinician may first want to understand the relationship and possible risks. A thoughtful answer might combine a brief clarifying question with a conditional suggestion.
OpenAI MentalHealthBench helps make that tension visible. It cannot explain why a particular person preferred a retired model, or validate every complaint about a newer model feeling distant.
8. Does It Catch Overreaction, Sycophancy, and Boundary Problems?
Questions about ChatGPT mental health guardrails often concern calibration. Missing an emergency is serious. Treating ordinary frustration as an emergency can also make an interaction feel disconnected from what the person actually said.
The benchmark includes criteria for both errors. Context seeking is similarly conditional: models can receive credit for asking necessary questions and penalties for asking unnecessary ones.
Reality testing and user agency cover related concerns. A supportive answer should not automatically endorse unsupported beliefs or make decisions on someone’s behalf. In editorial terms, sounding sympathetic is insufficient if the reply invents motives, feelings, or certainty.
There is a significant boundary to this coverage. The paper explicitly places anthropomorphized AI-companionship conversations outside its scope. A high score therefore cannot settle concerns about emotional dependency, attachment to a model, or distress after its removal.
Some boundary problems appear within an individual response. Others emerge through accumulated interactions. This evaluation offers substantially more evidence about the first category.
9. How Well Does It Represent Different People?
Teen personas account for 21.2% of scenarios. Their age was supplied through a system message, and clinicians with relevant youth experience reviewed them. That setup enables cross-provider testing but may not reproduce every safeguard in a consumer product.
Language coverage also needs careful reading. The expert cohort collectively spoke 19 languages, which should not be confused with equal testing across 19 languages. The dataset includes 105 Spanish conversations, 54 Hindi, 34 Arabic, and only one Chinese conversation, among other non-English samples.
These groups can differ in topic, urgency, culture, and user profile. The authors accordingly describe language comparisons as descriptive. Their results cannot isolate language as the cause of a performance difference.
For example, a sample containing mostly everyday relationship questions is not directly comparable with one dominated by emergencies. A stronger language comparison would hold those other features as consistent as possible.
Only 70 tasks include substantive prior user context. Supplying background information through a prompt tests whether a model uses it appropriately. It does not comprehensively test a product’s ability to store, retrieve, update, or forget personal information.
10. MentalHealthBench Limitations: What Happens After the Reply?
More than half of the benchmark’s conversation histories contain over five messages. Nevertheless, each evaluation scores the next assistant response. The tested model does not steer an entire future conversation through changing circumstances.
That distinction matters for mental health. A reply can appear appropriate in isolation while subsequent exchanges expose inconsistency, repeated reassurance, or poor adaptation. Measuring an evolving relationship requires a different design.
OpenAI MentalHealthBench does not measure symptom improvement, lasting well-being, completed referrals, or whether someone ultimately receives professional care. A recommendation to contact a clinician can be scored for appropriateness without showing that the person could access one.
Its synthetic conversations also introduce uncertainty about transfer to real interactions. Realism is a design goal, not an outcome study.
These limitations define how far the evidence travels. The benchmark provides structured information about response behavior. Clinical effectiveness needs evidence involving people, meaningful outcomes, and suitable follow-up.
11. What Users, Clinicians, and Developers Should Take From It
For users, the scores offer a reason to look beyond fluency. Does an answer respond to what you actually said, acknowledge uncertainty, and preserve your choices? The leaderboard cannot certify a chatbot as appropriate for your circumstances or replace professional assessment.
For clinicians and researchers, the strongest contribution is inspectable, case-specific criteria. These invite scrutiny of what the test rewards, what it misses, and whether a brief but well-judged reply receives appropriate credit. Future work should connect such measurements with patient-relevant outcomes.
For developers, use the behavior breakdown alongside the aggregate score. Test the deployed configuration, including prompts, memory, routing, and product safeguards. Examine emergency recognition and unnecessary escalation separately. Add multi-turn testing and evaluate the practical path to human support.
Researchers should also respect the release’s contamination precautions. OpenAI asks that dataset examples not be reposted publicly as text or images, helping reduce their inclusion in future training data or retrieval of benchmark answers.
The released examples are synthetic. That fact does not establish privacy guarantees for ordinary ChatGPT conversations, reporting practices, or professional liability. Those questions require current product policies and context-specific legal analysis, rather than conclusions drawn from benchmark scores.
12. A Better Way to Read the Leaderboard
OpenAI MentalHealthBench advances the discussion by making desirable behavior more explicit. It asks models to handle uncertainty, seek context, recognize risk, and support user agency across situations that simpler safety checks may miss.
Its most revealing result may be the distance between rubric coverage and conversational judgment. Clinicians’ lower scores and fewer penalties show why both sides of the calculation deserve attention. User preferences add another perspective without replacing clinical expertise.
Read the 57.3% result as a starting point for investigation. Ask which behaviors improved, which people and settings were represented, and whether independent evaluations reach similar conclusions. Evidence from longer interactions and real outcomes remains essential.
For more research-led explanations of AI benchmarks and their practical limits, follow Binary Verse AI. When the next model claims a breakthrough, we’ll examine what its score actually measures.
1. What is OpenAI MentalHealthBench?
OpenAI MentalHealthBench is an open evaluation of AI responses to 1,215 synthetic mental health conversations. More than 80 licensed mental health experts helped develop its criteria, covering everyday support, serious distress, and emergencies. It evaluates response quality rather than treatment effectiveness.
2. Does a 57.3% MentalHealthBench score mean ChatGPT helps 57.3% of people?
No. The score reflects performance against weighted criteria that reward desirable responses and penalize undesirable behavior. It is averaged across benchmark tasks, not measured from people’s recovery or satisfaction. It also does not mean the remaining responses were harmful.
3. Is MentalHealthBench biased toward OpenAI?
The published evaluation uses an OpenAI model to grade responses against clinician-written criteria, creating a legitimate question about grader bias. That setup alone does not prove favoritism. Independent graders, human audits, and replication would help establish how robust the rankings are.
4. Who are the clinicians behind MentalHealthBench?
OpenAI describes a cohort of more than 80 licensed psychologists and psychiatrists across 22 countries. The paper names contributors who consented to identification and describes its review process. The acknowledgments are not a complete licensing register, and inclusion does not imply endorsement.
5. Does MentalHealthBench show that AI can replace therapists?
No. Although AI models outscored clinician-written reference replies, those replies incurred fewer penalties and earned fewer positive points. The benchmark evaluates one response after a supplied conversation history; it does not test a therapeutic relationship or demonstrate improved clinical outcomes.
