17:01
2026-09-07
pub.towardsai.net
ai-safety
Do LLM Guardrails Actually Work? What My Own Numbers Say
A developer's test of LLM guardrails using openai/gpt-oss-120b and openai/gpt-oss-safeguard-20b on Groq found that safety classifiers can be inconsistent, with results varying across repeated runs. Thβ¦