Inference internals go public as benchmarks keep deflating the "more effort" reflex
Three independent benchmarks show that expensive AI defaults are rarely optimal: medium reasoning effort captures nearly all of Claude Opus 5's bug-fix gains, a three-model open-weight jury (GPT-OSS 1β¦