Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Researchers trained Jev, a model using reinforcement learning for calibrated decisions, to act as a zero-shot detector of AI alignment failures, according to the paper's description. Jev is positioned against generative judges that spend a decoding pass on every criterion and classifiers such as Llama Guard that read token probabilities and score one fixed label per call. Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained wit