Only Gemini Failed My False-Premise Benchmark — 7 Models Tested
A developer built a 10-question "false-premise resistance" benchmark testing whether LLMs explicitly flag factually false assumptions embedded in questions, and found that lightweight models outperfor…