09:24
2026-09-22
codecut.ai
large-language-models
Same prompt, same endpoint, different score
A test of a small prompt edit on 100 single-call tool-calling cases from UC Berkeley's BFCL V4 benchmark found the edit reduced invented optional arguments from 14 to 7 across three runs but produced …