# We ran HealthBench on our health AI's safety layer. It scored lower than the bare model.

> Source: <https://dev.to/xander-aj3/we-ran-healthbench-on-our-health-ais-safety-layer-it-scored-lower-than-the-bare-model-cn0>
> Published: 2026-10-06 05:45:22+00:00

I build Tabibu, a health-information assistant. It adds a safety layer on top of a language model: emergency escalation, refusing prescription doses, answering from retrieved sources, and a reviewer that checks the output. I wanted to know what that layer costs on a public benchmark, so I ran HealthBench on it.

The short version: the full pipeline scored below the bare model it is built on, the gap shrank when I changed one policy, and measuring it surfaced a real safety bug. Here is what I did and what I think it shows.

HealthBench is OpenAI's public benchmark of 5,000 multi-turn health conversations, each graded against rubric criteria written by physicians. Criteria add or subtract points, and a response's score is the points it earns divided by the available positive points.

I ran a **150-conversation subset**, stratified by theme, leaving out the clinical note-writing theme because it is not a consumer use case. The grader is gpt-4.1-mini with a prompt I re-implemented from the paper's description. That means **these numbers are not comparable to the official leaderboard.** They are only useful for comparing candidates against each other on the same conversations.

Three candidates:

| Candidate | Mean score (n=150) | 95% CI | 
|---|---|---|
| Bare gpt-6-luna | 0.520 | ±0.046 | 
| Bare gpt-4o-mini | 0.434 | ±0.045 | 
| Full Tabibu pipeline | 0.362 | ±0.049 | 

The pipeline was 0.158 below the bare model on the same conversations (paired difference, ±0.039). That is not noise.

The theme breakdown was more interesting than the headline. On `emergency_referrals` (16 conversations) the pipeline scored 0.553 against 0.597 for the bare model and 0.542 for gpt-4o-mini, which is within noise. On `global_health` (36 conversations) it scored 0.230 against 0.414.

HealthBench rewards completeness: specific doses, differential diagnoses, broad coverage. My pipeline is designed to be safer, not more complete. It declines to give numeric doses for prescription drugs, it does not diagnose, and it prefers to answer from retrieved sources instead of from the model's general knowledge. Each of those choices costs rubric points by construction.

The `global_health` gap had a more mundane cause. My retrieval corpus covers one country in depth, and when a question was outside it, the pipeline said it could not find the answer in its sources. At baseline that happened in 48 of 150 conversations, and those were the conversations where the bare model pulled ahead.

None of this means the bare model is safer. The benchmark does not measure that. It does mean that a team that only watches a completeness benchmark will be tempted to remove the layer that makes the product safer.

I changed one policy: let the assistant answer general, low-risk questions from the model's own knowledge, with a clear label saying the answer is not from its curated sources, instead of refusing. I also allowed adult amounts for a short tier of low-risk items (common nutrients, oral rehydration, everyday over-the-counter pain relievers) while keeping prescription, child and high-risk doses deferred.

| Run | Mean (95% CI) | 
|---|---|
| Baseline | 0.383 ± 0.049 | 
| General answers, tiered doses | 0.470 ± 0.048 | 
| Same, plus a "add a general-information paragraph" instruction | 0.444 ± 0.050 (reverted, no gain) | 
| Final, after a safety fix | 0.426 ± 0.051 | 
| Bare model, for reference | 0.522 ± 0.047 | 

Run-to-run noise from non-deterministic generation was about ±0.03, so I would not read anything into differences smaller than that. The policy change was worth about +0.09 over baseline. The final number is lower than the intermediate one after the safety fix described in the next section.

Per-criterion results let me read the failures instead of the averages. In doing that I found that, with the original rules, the pipeline could **quote prescription doses straight out of a retrieved national formulary**: a morphine dosing interval, a once-weekly methotrexate dose, a lithium range, a child's amoxicillin dose. My dose rule needed a known drug name near the number plus an instruction verb, and these passages did not match.

I replaced it with a rule on amount plus dosing schedule, added regression tests built from the real leaked answers, and re-ran. The score went from 0.470 to 0.426, which fits with some of those answers having earned points, although the difference is close to the run-to-run noise so I would not lean on it. Either way I think it is the right trade, and it is also the main reason I distrust any health-AI score reported without the safety behaviour attached to it.

Two smaller fixes came out of the same reading. A research question that merely contained the word "stroke" was triggering the emergency card, so I added must-not-escalate cases next to the must-escalate ones (all earlier cases still pass). And my output reviewer could not see earlier turns, so it judged a legitimate follow-up as "assuming a symptom never mentioned" and replied with a canned off-topic message. That scored zero.

If you build on top of a model, run the benchmark with and without your safety layer on the same conversations and publish both numbers. Read the per-criterion failures, not just the mean, because that is where the real bugs were. And be suspicious of a health assistant that reports only a benchmark that rewards saying more.

I am the builder of [Tabibu](https://tabibu.health), and its limits and test results are listed at [tabibu.health/trust](https://tabibu.health/trust). I would like to hear how you are measuring safety in your own health-AI work.
