13:48
2026-09-27
leehanchung.github.io
large-language-models
Jev and the Return of AI/ML Engineering
Independent testing found Jev's calibrated-probability claims fail in practice, with Valeriy M reporting calibration failures on 7 of 8 datasets across 16,500 predictions and the author's own coin-tosβ¦