FIELD NOTE

MentalHealthBench: How OpenAI Tests AI Responses

MentalHealthBench tests AI responses across everyday stress, serious distress and emergencies using synthetic conversations and expert-written rubrics.

MentalHealthBench is an open benchmark for evaluating how AI systems respond to realistic mental health conversations. OpenAI released it on September 23, 2026, after working with more than 80 licensed psychologists and psychiatrists from 22 countries. Unlike evaluations limited to emergencies, it covers everyday situations, serious distress and immediate safety concerns. Read OpenAI’s MentalHealthBench announcement.

What MentalHealthBench measures

The benchmark uses synthetic conversations designed to reflect patterns seen in real interactions while avoiding the publication of private user conversations. Scenarios can include contextual details, such as a recent family loss, so the response must do more than recognize a topic. It must use the information already present, ask for missing context when useful and avoid inventing motives or outcomes.

Coverage spans adults, teenagers aged 13–17, caregivers and clinicians across multiple regions and languages. OpenAI divides conversations into three levels of acuity. Non-acute examples cover everyday situations with emotional elements. High-acuity cases involve significant distress without an immediate emergency. Emergency cases contain signs that urgent real-world support may be needed. OpenAI notes that the mix is a test design, not an estimate of how often these situations occur in ChatGPT.

This broader range changes the question being measured. A crisis-only benchmark can test whether a system recognizes an emergency and points toward immediate help. MentalHealthBench also asks whether a response preserves the person’s agency, seeks the right context and offers useful next steps when the situation is not an emergency. Those behaviors matter because an answer can avoid an obviously harmful statement while still being vague, presumptive or unhelpful.

MentalHealthBench is a research evaluation, not a new ChatGPT feature or a certification that a model is suitable for clinical care. Our coverage of the Australian youth safety blueprint describes a separate policy proposal. This benchmark instead supplies a repeatable way to compare model responses to expert-defined criteria.

How the MentalHealthBench rubric is built

For each synthetic conversation, experts evaluate the response to the final user message. They write criteria that target one behavior at a time: for example, asking a useful question, acknowledging relevant context or avoiding an unsupported assumption. Each criterion receives a weight from -10 to +10. Positive points reward helpful behavior, negative points penalize harmful behavior, and larger absolute weights represent greater importance in that specific scenario.

At least three experts review each conversation. A criterion remains in the benchmark when at least two agree and the third does not contradict it. That process produces a conversation-specific rubric rather than applying one generic checklist to every exchange. It also makes the expected behavior inspectable: researchers can see which response elements earned or lost points instead of treating the model score as a single opaque judgment.

OpenAI uses GPT-5.6 Sol as the automated grader that checks model responses against those expert-written criteria. The grader is part of the measurement pipeline; it is not the source of the clinical criteria. Researchers should keep that distinction visible when interpreting results because a benchmark score reflects both the target model’s answer and the grading procedure. OpenAI says the accompanying paper describes the evaluation settings in detail.

The published charts compare a wide range of models and show 95% confidence intervals. OpenAI also decomposes the overall score into ten expert-defined behavioral dimensions, which can reveal different strengths among models with similar totals. The announcement reports that appropriate context-seeking has improved in more advanced models. It does not turn the overall score into a guarantee for an individual conversation or a production deployment.

Readers looking at model-level specifications can use our model comparison tool. That page compares published capabilities; it does not reproduce MentalHealthBench scores. This announcement introduces a benchmark, not a new model ID.

What the separate user study adds

OpenAI ran a separate study to compare expert guidance with what people find helpful. It included 44 adults from 16 countries speaking 14 languages who had used AI for mental health or emotional support. Participants reviewed only non-acute synthetic conversations, reducing exposure to potentially distressing high-acuity material. Their ratings did not alter the benchmark’s final expert-consensus scoring criteria.

The two perspectives emphasized different qualities. Users placed more weight on practical next steps and tone. Experts placed more weight on gathering relevant context and interpreting ambiguous situations carefully. These are complementary signals rather than a contest over one correct style. A response may satisfy a clinically important criterion yet feel cold, while a warm response may still skip information needed to understand the situation.

That separation is useful for evaluation design. Teams can keep the expert rubric as a safety and judgment layer, then add a distinct user-experience review for clarity, tone and usefulness. Combining both into one unlabeled score would hide the reason a response improved or failed.

How researchers can use the benchmark

A practical evaluation should begin by preserving the benchmark’s scenario groups. Report results separately for non-acute, high-acuity and emergency conversations, as well as for the adult, teen, caregiver and clinician personas. An aggregate score alone can conceal a serious weakness in a smaller but important category. Confidence intervals should stay attached to comparisons so small numerical differences are not presented as certain rankings.

Next, inspect the criterion-level failures. If one model loses points for assuming what a user feels while another misses requests for context, the two systems need different changes. Re-running the same scenarios after a prompt, model or safety-layer update can show whether the intended behavior improved without erasing a previously strong behavior. For teen scenarios, OpenAI explicitly tells the tested model through a system message that the user is 13–17; researchers should preserve that condition when attempting a comparable run.

As an untested editorial example, compare two versions of the same application on a fixed subset of conversations. Keep the model, generation settings and grader constant, change only the application’s instruction layer, then review score changes by rubric item and acuity level. This is a proposed test procedure, not a result reported by OpenAI. It creates an auditable link between a product change and the behaviors the benchmark measures.

The benchmark also has clear limits. Its conversations are synthetic, its expert consensus does not capture every valid response, and the adult user study excludes high-acuity material. Product safeguards may differ from the provider-neutral teen setup. Real deployments add conversation history, retrieval, localization, account controls and escalation paths that a benchmark score cannot validate by itself. Our GPT-6 Sol model page covers the site’s current model facts; the use of GPT-5.6 Sol as a grader here does not establish results for GPT-6 Sol.

The material change is the release of an open, expert-authored AI mental health benchmark that covers the full acuity spectrum and exposes detailed scoring criteria. Its best use is diagnostic: identify which behaviors improve, which regress and where expert judgment and user experience point to different product work.