cd /news/artificial-intelligence/healthbench-psych-a-mental-health-su… · home topics artificial-intelligence article
[ARTICLE · art-112659] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

Researchers released HealthBench-Psych and HealthBench-Psych-Hard, a mental-health subset of OpenAI's HealthBench, comprising 610 clinician-validated conversations (12.2% of the original 5,000-physician-rubric corpus). Evaluating 20 frontier and open models with three LLM judges, they found a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges (Kendall's tau ≥ 0.92). The subset, pipeline, model responses, grades, and analysis code are released as a reusable resource.

read1 min views3 publishedAug 27, 2026

arXiv:2608.25071v1 Announce Type: new Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges ($\tau \ge 0.92$). We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/healthbench-psych-a-…] indexed:0 read:1min 2026-08-27 ·