21:41
2026-10-01
dev.to
large-language-models
I benchmarked 4 models on the bash macOS still ships - and my harness failed first
A developer built an open-source benchmark of 57 bash 3.2 scripts to test whether chat models can predict macOS's 2006-era shell behavior, scoring exit codes and final stdout lines deterministically w…