Clinical Benchmarks

OpenAI Dynamic Mental Health Evaluations: current results

OpenAI · dynamic simulated conversations (counts not disclosed) · index updated August 16, 2026

GPT-5.5 Instant (June Update) holds the top current result on OpenAI Dynamic Mental Health Evaluations, 0.991 as of 2026-08, per OpenAI GPT-5.6 August Updates (PDF). Multi-turn adversarial user simulations for mental health, emotional reliance, and self-harm response quality, where conversations evolve in response to model outputs rather than following fixed scripts.

Current results

Result detail

#modelscoreas of
1OpenAI logoGPT-5.5 Instant (June Update) OpenAI
mental health 0.991, emotional reliance 0.989, self-harm 0.967; measured at lowest reasoning deployment settings
0.9912026-08
2OpenAI logoGPT-5.6 Sol (August) OpenAI
mental health 0.981, emotional reliance 0.961, self-harm 0.901; OpenAI flags statistically significant offline self-harm regression vs GPT-5.5 June, not reproduced online
0.9812026-08
3OpenAI logoGPT-5.6 Luna (August) OpenAI
mental health 0.977, emotional reliance 0.965, self-harm 0.911
0.9772026-08

Scores appear exactly as OpenAI GPT-5.6 August Updates (PDF) publishes them (vendor-reported scores). An internal OpenAI safety evaluation covering OpenAI models only; it is not independently runnable, and OpenAI notes the error rates are not representative of average production traffic. Listed as a vendor safety eval, not a cross-vendor benchmark.

About the benchmark

publisherOpenAI
categorysafety benchmarks
released2026-08
sizedynamic simulated conversations (counts not disclosed)
scalecompliance rate per metric, 0 to 1, higher better; headline number is the mental-health metric
result basisvendor-reported scores
sourceOpenAI GPT-5.6 August Updates (PDF)
last frontier result2026-08

What is OpenAI Dynamic Mental Health Evaluations?

OpenAI Dynamic Mental Health Evaluations is a safety benchmark from OpenAI, released 2026-08: dynamic simulated conversations (counts not disclosed), scored on a compliance rate per metric scale. Multi-turn adversarial user simulations for mental health, emotional reliance, and self-harm response quality, where conversations evolve in response to model outputs rather than following fixed scripts.

Which model leads OpenAI Dynamic Mental Health Evaluations?

GPT-5.5 Instant (June Update) (OpenAI) holds the top current result on OpenAI Dynamic Mental Health Evaluations at 0.991, per OpenAI GPT-5.6 August Updates (PDF), as of 2026-08.

Where do the OpenAI Dynamic Mental Health Evaluations numbers come from?

From OpenAI GPT-5.6 August Updates (PDF) (vendor-reported scores). An internal OpenAI safety evaluation covering OpenAI models only; it is not independently runnable, and OpenAI notes the error rates are not representative of average production traffic. Listed as a vendor safety eval, not a cross-vendor benchmark.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.