JEV-as-a-Judge: Accept When Confident, Escalate When Unsure A study of "jev-as-a-judge" compares a decision-only LLM judge against sixteen generative approaches, testing whether a cheaper first-pass judge can flag when stronger evaluation is needed. The work targets inference cost and confidence reliability in LLM-as-a-judge evaluation at scale, positioning the decision-only judge as an economical filter that escalates uncertain cases. LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generativ