HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench Researchers released HealthBench-Psych and HealthBench-Psych-Hard, a mental-health subset of OpenAI's HealthBench, comprising 610 clinician-validated conversations (12.2% of the original 5,000-physician-rubric corpus). Evaluating 20 frontier and open models with three LLM judges, they found a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges (Kendall's tau ≥ 0.92). The subset, pipeline, model responses, grades, and analysis code are released as a reusable resource. arXiv:2608.25071v1 Announce Type: new Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations 12.2% of the corpus . Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges $\tau \ge 0.92$ . We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.