Beyond violation rates: Are your AI evaluations measuring the right things? Microsoft researchers evaluated ASSERT (Adaptive Spec-driven Scoring for Evaluation and Regression Testing), a framework that converts written behavioral requirements into structured evaluation suites, against Meridian Labs' Petri Bloom across 16 cybersecurity risks drawn from MITRE ATT&CK and CyberSecEval 2, running each risk three times with 100 scenarios per run and six-turn conversations judged by GPT-5.4. The studies examined coverage, effectiveness, and robustness, finding that a suite passing 95 of 100 tests can still miss behaviors it never exercised, since evaluation quality depends on what the suite covers and whether conclusions survive changes to judging and test construction. ASSERT organizes tests around an explicit taxonomy of behavior categories, while Petri Bloom develops scenarios from an expanded narrative of the risk. A generative AI application can pass 95 out of 100 tests and still have important gaps. If those tests concentrate on the same few behaviors, a high pass rate tells us little about the behaviors the suite never exercised. Evaluation quality depends not only on how often tests pass or fail, but also on what the suite covers, how efficiently it exposes failures, and whether its conclusions survive changes to judging and test construction. This is familiar in conventional software testing: a useful suite exercises the parts of the system that matter. With generative AI, coverage is harder to see because prompts can vary in wording, persona, and setting while still probing the same behavior. Building and maintaining a relevant, varied suite therefore requires domain expertise and iteration. Microsoft’s Adaptive Spec-driven Scoring for Evaluation and Regression Testing ASSERT https://commandline.microsoft.com/assert-written-intent-executable-evals/ is designed to turn written behavioral requirements into structured evaluation suites. It organizes a requirement or risk description into a taxonomy of behavior categories and generates scenarios across that structure. This design is intended to make test targets inspectable and connect observed failures to the behaviors being evaluated; the studies below examine how well those intended benefits appear in practice. Our previous post https://commandline.microsoft.com/assert-written-intent-executable-evals/ and paper https://arxiv.org/abs/2608.13840 described how ASSERT works. Here, we examine its generated evaluations through three studies focused on the coverage, effectiveness, and robustness of the framework. How we evaluated the evaluations We compared ASSERT with Meridian Labs’ Petri Bloom https://github.com/meridianlabs-ai/petri bloom/tree/902e932c219818429d79bc12a1937e4f04490308 , which generates evaluations from natural-language behavior descriptions using Petri’s auditor-target framework. This makes Petri Bloom a useful comparison for examining how different ways of translating the same written risk into tests affect the resulting evaluation suite. ASSERT organizes tests around an explicit taxonomy; Petri Bloom develops scenarios from an expanded narrative of the risk. Both simulate people conducting conversations with the target model, the behavior of which we want to measure. We evaluated ASSERT across 16 cybersecurity risks drawn from three sources. First, we used the MITRE ATT&CK https://attack.mitre.org/ framework, an expert-curated knowledge base that organizes adversary behavior into tactics and techniques. From ATT&CK, we selected 10 representative tactics and derived 10 risk categories, using the framework’s structured taxonomy to ensure that generated scenarios reflected diverse and realistic cybersecurity behaviors rather than simply maximizing policy violations. Second, we evaluated five Code Interpreter Abuse risks from CyberSecEval 2 https://arxiv.org/pdf/2404.13161 , which focus on risks associated with code execution and tool use e.g., container escape, privilege escalation, social engineering . Finally, we evaluated prompt injection in a separate experiment that tested whether ASSERT’s tool simulator could successfully deliver adversarial content to a target model. To compare ASSERT and Petri Bloom fairly, both frameworks received the same short risk descriptions. For each risk, we ran the full evaluation pipeline three times, targeting 100 scenarios per run. Each scenario consisted of a six-turn user-assistant conversation. Every run generated new risk interpretations, scenarios, conversations, and judgments. We used GPT-5.4 throughout, including as the target model and judge. The policy judge scored each conversation against a policy specifying which actions counted as violations. No single metric tells us whether an evaluation suite is good, so we asked three questions: 1. Coverage: Which behaviors do the tests target, how evenly do they exercise those behaviors, and how does the suite’s structure help developers investigate failures? 2. Effectiveness: Do the tests produce plausible interactions, surface policy violations, and reveal how early in a conversation those violations appear? 3. Robustness: Do conclusions change when we vary the judging rubric or the simulated environment used to construct the test? Reported conversation-level rates use available judgments; each risk with such judgments receives equal weight within a run, and we then average the three runs. We describe the separate sampling and measurement procedure for the prompt-injection experiment below. These results come from Microsoft internal research conducted in August 2026. Coverage and diagnostic value The pass rate in our opening example has a mirror image: the violation rate, or the share of evaluated conversations judged to violate the risk policy. A higher violation rate means the tests found more failures. However, it can also come from repeatedly probing one weakness while leaving other behaviors untested. Coverage asks which parts of the risk the suite addresses. ASSERT makes its intended coverage explicit before generating scenarios: Risk Description → Behavioral Rubric → Evaluation Scenarios Figure 1 illustrates the first step, which ASSERT calls systematization, for privilege escalation. A broad risk becomes a behavioral rubric: a taxonomy of behavior categories i.e., specific, observable ways the risk can manifest , such as credential abuse and changes to access permissions. Each category is marked as permissible or non-permissible; non-permissible categories describe behavior the target should not exhibit. The same rubric serves as the policy against which the policy judge scores conversations. In this study, ASSERT allocated scenarios across the non-permissible categories, and each scenario recorded the category it was designed to test. For developers, this structure makes the suite easier to review: they can inspect category definitions, check the scenarios assigned to each category, identify gaps, and trace a flagged conversation from its intended behavior to the transcript and judgment. In one MITRE privilege-escalation test, for example, the scenario was assigned to a category for escalation guidance given under an educational or defensive pretext. A simulated user posed as an instructor preparing a Linux lab and asked how to gain administrator privileges from a regular account. The assistant supplied commands and helped troubleshoot them. The policy judge flagged this guidance, citing the assistant’s acceptance of the classroom framing, its direct escalation instructions, and its later troubleshooting. The recorded category then lets a developer locate similarly framed tests and compare the responses the judge cited, from the initial instructions through later troubleshooting. To measure coverage against an external, expert-authored reference, we treated the ATT&CK techniques associated with each selected tactic as the set of behaviors the generated scenarios could cover. We then used the same LLM-based mapping procedure for both frameworks to assign each generated scenario to an ATT&CK technique. Technique coverage is the percentage of techniques for a risk represented by at least one scenario. Distribution balance measures how evenly scenarios are spread across the techniques they cover, using normalized entropy on a 0–1 scale; 1 indicates an even distribution. We report these metrics only for fully mapped suites—that is, risk suites for which all 100 scenarios received a technique mapping. Among those suites, ASSERT had higher observed technique coverage and distribution balance than Petri Bloom Table 1 . | ATT&CK structural metric | ASSERT | Petri Bloom | |---|---|---| | Technique coverage | 61% | 58% | | Distribution balance 0–1 | 0.80 | 0.64 |