05:37
2026-07-30
dev.to
artificial-intelligence
I gave the same fabricated answer to RAGAS and DeepEval. One scored it 0.0. The other scored it 1.0
A developer testing the two most popular LLM-as-judge faithfulness metrics, RAGAS and DeepEval, found that they gave opposite scores for the same fabricated answer: RAGAS scored it 0.0 while DeepEval โฆ