Which LLM should I actually code with? I built a small benchmark to find out
A developer built a small benchmark to compare LLMs for coding tasks across Python, C#, and Bash. All three models achieved 100% pass@3 on problems they were allowed to answer, but one model was conte…