Native harnesses don't always solve more coding tasks
A study of 256 private coding tasks found no reliable overall performance advantage for vendor-native harnesses over neutral third-party harnesses on the same models, with Opus 4.8 scoring 48.8% versus 50.0% and GPT-5.5 …