08:00
2026-08-26
github.com
artificial-intelligence
Reconstructing the benchmark behind Luc Julia's 64% LLM reliability claim
A reconstruction of the reasoning benchmark behind Luc Julia's repeated claim that large language models are 'relevant 64% of the time' shows the figure comes from a single sentence in Bang et al. (20โฆ