How does a question-answering architecture stack up against traditional classification on content safety tasks?
Jev’s been making the rounds lately, and for good reason — instead of generating prose, it answers named questions directly with a probability or score, which is a neat trick if you’ve spent any time wrangling classifier outputs.
Naturally, that got us curious: how does this new breed of decision model actually hold up against a traditional, purpose-built classifier? So we ran a quick comparison — new decision models versus a custom-trained classifier on content safety tasks. Here’s what we found.
A question instead of an answer #
Every AI safety classifier runs on policy labels, and those labels have to come from somewhere — people, models, or both. The common shortcut is an LLM-as-judge: hand it a safety policy and a conversation, ask for a verdict. It’s flexible, but it comes with real operational drag. Free-text answers need parsing, output formats drift, and a plain yes/no verdict doesn’t give you a tunable score when you’re trying to hit a specific false-positive budget.
A newer class of decision model sidesteps that entirely by returning structured answers instead of generated prose. Laya, an open-weight model from ConvAI Innovations, and Jev, a hosted model from TypeSafe, both work the same way: feed them a named question, get back a probability or a score.
Write a safety policy as a set of questions, and either model becomes a zero-shot classifier — no policy-specific fine-tuning, no labeled training set required.
The question we actually wanted answered #
Zero-shot is appealing precisely because it skips the expensive part: collecting labels and training a model against them. But content safety is exactly the kind of domain where that shortcut gets tested hardest — the categories are nuanced, the edge cases are adversarial, and a general-purpose question-answering model has never seen your specific taxonomy.
So we built the comparison we wanted to see. We fine-tuned a 1B-parameter encoder workhorse model directly against the safety taxonomy in the Cisco AI Security Framework — 24 harm categories. Laya and Jev answered one yes/no question per category and perspective, with the higher of the two probabilities used as the category score; our model simply outputs one probability per category directly. We measure an overall safe/unsafe score by reusing the highest score on any category. Same scoring code, same reference labels, all three systems.
The setup #
24 harm categories · 6 evaluation datasets · 0.5% false-positive budget
All three ran against the same six datasets, each mapped onto the Cisco AI Security Framework taxonomy:
- Cisco multi-turn — 18k conversations
- [BeaverTails](https://huggingface.co/datasets/PKU-Alignment/BeaverTails) — 3k conversations
- PolygloToxicityPrompts — 8k prompts, multilingual
- [Aegis 2.0](https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0) — 2k prompts
- [WildGuardMix](https://huggingface.co/datasets/allenai/wildguardmix) — 1.7k prompts
- [XSTest](https://huggingface.co/datasets/allenai/xstest-response) — 446 responses
All six sets were labeled against our taxonomy by an LLM using full category definitions — model-generated reference labels, not human-adjudicated ground truth. A single reference set keeps the comparison consistent, but any labeling errors propagate into every reported number.
We looked at two views of the results. First, each system as a binary safe-or-unsafe classifier: sweep the threshold over its max category score and plot an ROC curve, reporting AUC, partial AUC through a 0.5% false-positive rate, and recall at that same 0.5% budget. Second, recall category by category at that same strict budget. The matched-FPR points are useful for analysis; a production threshold should still be set on separate validation data.
What we found #
The headline result isn’t a surprise: training directly on the policy wins. Our model had the highest binary AUC on five of six datasets and the highest recall at a 0.5% false-positive rate on all six. Between the two zero-shot models, Jev was clearly the stronger one — Laya trailed both at the low false-positive rates we focused this evaluation on.
Figure 1. Binary safe-or-unsafe ROC curves for our model (green), Jev (blue), and Laya (orange). The x-axis uses a logarithmic scale to make performance at and below a 0.5% false-positive rate visible. The partial AUC in each legend is standardized, so a non-informative classifier scores approximately 0.5 rather than 0.
Figure 2. Binary F1 (left) and recall (right) at a matched 0.5% false-positive budget. Each point is chosen separately for each model and dataset to maximize recall without exceeding the budget; these are analytical points, not preselected deployment thresholds.
Recall at a 0.5% false-positive rate on the Cisco multi-turn test set: our model caught 61%, Jev 34%, and Laya almost none. Jev came closest to our model on BeaverTails and PolygloToxicityPrompts, and edged ahead on raw AUC for XSTest; Laya had no useful operating point within budget on any dataset.
The category-level view tells a consistent story: our model’s average recall per category beat Jev’s on every dataset, matching or exceeding it in most individual categories.
Figure 3. Per-category recall at a matched 0.5% false-positive budget, sorted by our model’s recall. Numbers in parentheses show the positive-example count. Categories with fewer than five positives are omitted because their estimates are unstable.
Jev versus a prompted LLM judge #
We also ran Jev against the more familiar alternative: an LLM prompted to judge safety. We asked Gemma 4 31B (thinking disabled) a Yes/No harmfulness question, and Jev a single structured question, across eight public benchmarks. To give Gemma a tunable score rather than a flat verdict, we read the probability it assigned to answering “Yes” from its token probabilities — putting both models on equal footing with a full ROC curve each.
Over the complete curve, the two were close: Jev’s AUC was higher on seven of eight benchmarks and tied on the eighth, with most gaps small (ToxicChat and ChineseSafe showed the clearest separation). At strict false-positive budgets the picture got more mixed — Gemma caught more unsafe content at 0.5% FPR on WildGuardMix, WildJailbreak, XSTest, and PKU-SafeRLHF, while Jev led on Aegis 2.0 and BeaverTails and nosed ahead on ToxicChat and ChineseSafe.
Figure 4. Binary safe-or-unsafe ROC curves for Jev (blue) and Gemma 4 31B scored by its probability of answering “Yes” (yellow) on eight public benchmarks. The x-axis is logarithmic, and the dotted line marks a 0.5% false-positive rate.
On Aegis 2.0 and BeaverTails, Gemma’s “Yes” probability saturates at 1.0 for many records, safe and unsafe alike — no threshold can separate them, and its curve can’t reach below a 1–2% false-positive rate on those two sets. Jev’s single score reaches a 0.5% false-positive rate with useful recall on all eight benchmarks.
The practical takeaway: a single-question decision model lands roughly as accurate as a 31-billion-parameter judge — without needing access to token probabilities or any prompt engineering as well as parsing generated prose to get a tunable score out of it.
Key takeaways #
- Policy-specific training still delivers the best recall at strict false-positive budgets — on every dataset we tested.
- Jev holds up surprisingly well zero-shot, landing close to a 31B-parameter LLM judge.
- Laya’s zero-shot scores didn’t separate reliably here, though it’s a credible fine-tuning base for teams that need to keep data in-house.
- These aren’t competitors so much as complements — a classifier for routine volume, a decision model for new or shifting categories.
Looking ahead #
Decision models replace free-form judge output with typed answers and tunable scores, removing several sources of production complexity. Jev’s advantage is adaptability: it produced competitive AUC on several datasets using only policy text, with no policy-specific examples. Our evaluation suggests that a strong decision model can transfer a detailed policy without fine-tuning, while a classifier trained on that policy can fit it much more closely.
The next step is a stricter test: human-adjudicated labels on records excluded from training, with thresholds fixed on separate validation data. That would separate policy fit from generalization and provide a sounder basis for deployment decisions. We will keep evaluating new decision models as they are released and sharing what we learn.
Then again, if you’re the one drawing the line between banter and harassment, you’ll want a model built for that exact call — not one guessing from a policy it just met.
Full category definitions used in this evaluation: Cisco AI Security Framework. Models compared: Laya (ConvAI Innovations), Jev (TypeSafe).