OpenAI's mental health test exposes AI's blind spots OpenAI unveiled MentalHealthBench on Wednesday, an open benchmark for evaluating model capabilities in mental health domains including safety, seeking user context, preserving user agency, and providing actionable guidance, built with 80 mental health professionals across 22 countries, 19 languages and 20 subspecialties. In OpenAI's own evaluations, Astra scored the highest overall at 57.8%, followed by GPT-6 Sol at 54%, Luna at 50.3%, and Claude Opus 5 at 48.1%, with all tested models scoring lower on gathering context and supporting user agency than on clinical accuracy, empathy and reality testing. "This is not a leaderboard," said Dr. Declan Grabb, mental health safety research lead at OpenAI, who added that ChatGPT "is not a therapist, and is not here to replace a clinician. eople are turning to AI for emotional support more than ever. But these chatbots' ability to provide a shoulder to cry on can vary greatly. Because of this, OpenAI decided to measure it: On Wednesday, the AI lab unveiled MentalHealthBench, a new open benchmark dedicated to evaluating model capabilities in domains such as safety, seeking user context, preserving user agency, and providing actionable guidance. To create this benchmark, OpenAI began by developing synthetic conversations that reflected real-world AI use patterns that span multiple topics and run the gamut of severity, ranging from non-acute situations that involve emotional themes, to "high-accuity" situations that indicate more serious concerns or distress, to emergency situations that require immediate support. Then, the company worked with a cohort of 80 mental health professionals across 22 countries, 19 languages and 20 subspecialties to evaluate the responses to the synthetic message. The benchmark breaks down model performance by conversation severity, as well as a range of 10 dimensions defined by the mental health experts, including whether the model asks the right questions, provides appropriate, clinically accurate guidance, helps the user see reality, avoids harm and recognizes serious risk. In developing the benchmark, OpenAI also put a number of its own and other models to the test: - Astra took the overall best score on the evaluations, scoring 57.8%, with GPT-6 Sol and Luna trailing just behind at 54% and 50.3% respectively, and Claude Opus 5 sitting in fourth place at 48.1%. - However, performance differs across dimensions of the benchmark. Though Astra still largely outranks other models, all of the models tested tended to perform better in certain areas, such as clinical accuracy, empathy and reality testing, while scoring lower in dimensions such as gathering context and supporting user agency. - OpenAI said that MentalHealthBench also points to several opportunities to improve ChatGPT, including asking useful follow-up questions and responding with the right level of urgency, and that it will use this information to guide improvements and track the model's progress. "This is not a leaderboard," Dr. Declan Grabb, mental health safety research lead at OpenAI, told The Deep View. "What I hope that this benchmark provides is a nuanced view into model behavior, so that people really understand the more complex dynamics of their models." MentalHealthBench adds to a number of mental wellness-related initiatives that OpenAI has endeavored, including research to combat model sycophancy https://openai.com/index/expanding-on-sycophancy/ and improving ChatGPT's responses to sensitive conversations https://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/ , as well as joining forces with advocacy group Common Sense Media to support the Parents and Kids Safe AI Act https://www.commonsensemedia.org/press-releases/common-sense-media-openai-join-forces-on-strongest-youth-ai-safety-measure-in-us . OpenAI said that this is just a piece of its research into mental health benchmarking and alignment in this area, not an end state. "ChatGPT is not a therapist, and is not here to replace a clinician," said Grabb. "That being said, when I talk to mental health clinicians across the globe, the most responsible and safe thing to do is if people are coming to AI to ask these questions, we absolutely need to have an expert opinion on how you should navigate them." Our Deeper View Mental healthcare is a critical area for these models to get right. While OpenAI said that speaking to a chatbot should not supplant actual therapy, the reality is that many people have and will turn to a chatbot for support, seeking both a judgement-free and cost-free alternative to clinical support. The company faces lawsuits involving the deaths of Adam Raine https://openai.com/index/mental-health-litigation-approach/ and Joshua Enneking https://techjusticelaw.org/cases/enneking-v-openai-inc-openai-holdings-llc-and-samuel-altman/ , whose families allege that ChatGPT contributed to their suicides. A benchmark can help identify weaknesses, but a higher score alone does not establish that a model is safe in a real conversation. The gaps in gathering context and supporting user agency are particularly important: an empathetic response is not necessarily an appropriate one. The next test for OpenAI is how it turns those findings into changes that make its models safer for the people relying on them.