The Safeguard Worked. Is the LLM System Safer? A new evaluation framework argues that refusal rates and attack success rates do not measure how much harmful assistance a deployed LLM service still provides, and proposes a complementary metric based on the expected utility of the service's responses to harmful requests. Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an at