10:42
2026-09-20
shatteringtheabyss.substack.com
large-language-models
The $1 LLM Test
A software engineer argues that LLM evaluation should focus on black-box behavioral testing rather than benchmark scores, asking what happens when a model is inserted into a real information pipeline …