Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals
An AI safety team developed and deployed a basic live per-action monitor that uses an LLM judge to review each agent action before execution and halt evaluations for human review when an action exceed…