The Harness Is Not the Model: How Far Scaffolding Takes a Weak LLM
An engineering study testing whether a weak large language model wrapped in strong scaffolding (retrieval, toolset, verification loops) can match frontier-level performance finds that the harness clos…