05:45
2026-10-06
dev.to
ai-safety
We ran HealthBench on our health AI's safety layer. It scored lower than the bare model.
A developer building Tabibu, a health-information assistant with a safety layer over a language model, ran OpenAI's HealthBench on the full pipeline and found it scored 0.362 versus 0.520 for the bareβ¦