An LLM judge cannot be a build gate, and it is not about the cost
A developer argues that LLM-based judges cannot serve as build gates in RAG evaluation due to cost and non-determinism, advocating for deterministic metrics like precision@k, recall@k, and lexical ove…