Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails Automated harness evolution can let smaller models match frontier-model performance on domain-specific agentic tasks at a fraction of the cost, according to research on co-evolving harnesses and models. The work finds that on-policy correction helps weaker models catch up in cases where imitation fails, with agent harnesses — the system prompt, tool set, execution hooks, and context-management scaffolding around a model — identified as a critical determinant of agentic task success. Agent harnesses the system prompt, tool set, execution hooks, and context-management scaffolding around a model are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. S