06:44
2026-07-29
dev.to
large-language-models
How do you measure something that gives a different answer every time?
A developer measuring LLM recommendation consistency found that GPT-4o, Claude Haiku 4.5, Gemini 2.5 Flash, and Perplexity Sonar are roughly 7.5ร more consistent with themselves than with each other wโฆ