Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize ByteDance Seed, SUTD, Georgia Tech, M-A-P, and TokenWave.AI released HarnessDev, a benchmark that scores the runnable harness a model builds rather than the answer it returns, with 6 creator LLMs constructing harnesses across 5 benchmarks and 2,207 tasks from a seed that scores 0. Self-built harnesses matched human references on writing and ML experimentation but trailed on code and search, and only 34 of 64 evolution changes moved the same direction on held-out tasks. ByteDance Seed, SUTD, Georgia Tech, M-A-P, and TokenWave.AI introduce HarnessDev, a benchmark that scores the runnable harness a model builds rather than the answer it returns. Starting from a seed that scores 0, 6 creator LLMs construct harnesses across 5 benchmarks and 2,207 tasks, then evolve them from execution feedback. Self-built harnesses match human references on writing and ML experimentation but trail on code and search, and only 34 of 64 evolution changes move the same direction on held-out tasks. The post Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize https://www.marktechpost.com/2026/09/11/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize/ appeared first on MarkTechPost https://www.marktechpost.com .