AI Evaluation Series (06): DeepEval in Practice — Enterprise Agent Evaluation Suite
A developer demonstrates using DeepEval to evaluate an enterprise agent, contrasting its test-case-first paradigm with RAGAS's batch evaluation. The implementation uses a custom judge LLM (glm-4-flash) and reveals that l…