{"slug": "intelligence-is-getting-cheap-knowing-when-its-wrong-isnt", "title": "Intelligence Is Getting Cheap. Knowing When It’s Wrong Isn’t.", "summary": "Cognition's June 2026 rebranding of Windsurf into Devin Desktop with native Agent Client Protocol support, now used by JetBrains, Gemini CLI, GitHub Copilot, and Codex, signals that model choice is becoming composable, shifting durable advantage toward feedback systems. A 2026 arXiv paper (2604.25850) on Agentic Harness Engineering (AHE) improved pass@1 on Terminal-Bench 2 from 69.7% to 77.0% over ten iterations, demonstrating that the harness, not just the model, carries value, but the loop depends on dense, trustworthy feedback, which is scarce in domains like legal reasoning and medicine.", "body_md": "AI’s next moat may not be intelligence, agents, or even verification. It may be the ability to turn ambiguous human judgment into cheap, reliable, machine-executable feedback.\n\nA model can become dramatically more capable without becoming dramatically more trustworthy. Those are different axes, and the industry has spent most of its attention on the first one. The model is becoming a less sufficient explanation of the system — not because capability stopped mattering, but because capability alone no longer predicts whether you’d let the system act unsupervised.\n\nSupporting evidence, not the headline: in June 2026, Cognition rebranded Windsurf into Devin Desktop with native support for the Agent Client Protocol, now supported across a growing ecosystem including JetBrains, Gemini CLI, GitHub Copilot, and Codex — letting one editor run different agents interchangeably, mid-task, without losing state. That doesn’t mean models are economically interchangeable yet. It means model choice is becoming more *composable*, and as composability rises, durable advantage migrates toward whatever’s harder to swap.\n\nNot the harness alone, either. A harness with nothing to check its decisions against is just a faster way to be wrong. The real chain is: **model intelligence → action → feedback → judgment → adaptation.** The first wave of agent development optimized the system that speaks. The next wave optimizes the system that checks.\n\nA 2026 paper (arXiv:2604.25850) built Agentic Harness Engineering (AHE), evolving a coding agent’s tools, middleware, and memory while holding the model fixed. Ten iterations lifted pass@1 on Terminal-Bench 2 from 69.7% to 77.0%, beating a hand-built and an RL baseline — with the evolved components transferring across model families and reducing token usage. Real evidence the harness carries value.\n\nBut the loop only worked because Terminal-Bench 2 supplies a dense, trustworthy signal after every edit. So the claim isn’t “the harness doesn’t matter.” It’s that the harness is machinery for converting feedback into reliable behavior — meaning the feedback’s quality and domain-specificity is the deeper source of advantage. And coding isn’t oracle-rich because code is special. It’s oracle-rich because large parts of software correctness are externally observable, cheaply testable, and repeatedly executable — a property, not a magic property of software itself, which is what makes the framework portable to other domains.\n\nA company running agents in production accumulates failures, near-misses, and corrections competitors don’t have. But that alone isn’t a moat. Three different things get lumped together as “data”:\n\nA million failure logs sitting in a warehouse are the first kind. The moat only forms when a company can convert logs into the third kind. That conversion capability, not raw log volume, is what compounds.\n\nPush past coding and add one layer the field usually skips:\n\n**Specification → Judgment → Evaluation → Validation → Verification.**\n\nA perfect verifier can verify the wrong objective. A perfect validator can validate a bad specification. A perfect evaluator can optimize the wrong metric perfectly. This is Goodhart’s Law as a structural property of the stack, not a caveat bolted on afterward — it’s why a legal memo can pass every internal consistency check and still be the wrong argument: nothing above validation ever checked whether the specification itself was sound.\n\nThat gives a measurement concept, not just a metaphor. Define **feedback density** as the amount of trustworthy, actionable information available per agent action about whether that action moved the system toward the real objective. Coding sits at the dense end — tests supply verification, CI supplies validation, production metrics supply evaluation, code review approximates judgment. Legal reasoning, strategy, and medicine have thin coverage above validation.\n\n**Testable hypothesis:** agent adoption across a domain should correlate more strongly with that domain’s feedback density than with raw model benchmark performance — a claim you could actually test across coding, customer support, data operations, legal research, finance, and scientific discovery, not just assert.\n\nVerification isn’t free. Anthropic’s published API pricing shows cache reads on cached prefixes billed at roughly 10% of standard input cost for a given model tier — for example, $0.30 versus $3.00 per million tokens under specific cache-read terms — which is why harnesses that keep prompts and tool schemas stable can run long sessions affordably. That’s a multiplier on a stack you already have, not a substitute for one you’re missing.\n\nAnd the thesis has to survive scrutiny of itself. A July 2026 paper (arXiv:2607.12227) argues harness-evolution methods, AHE included, search and evaluate on the same benchmark — so gains need matched-budget test-time-scaling comparisons and held-out testing before being trusted. Run that comparison, and harness evolution doesn’t consistently win, with limited off-benchmark generalization. The field studying feedback loops doesn’t yet have a reliable feedback loop on itself — the specification problem, one level up, applied to the researchers.\n\nToday, judgment is expensive, human, slow, and hard to standardize. The real economic opportunity — bigger than better harnesses, bigger than better models — is converting fragments of that judgment into tests, simulators, validators, policies, escalation rules, and rollback mechanisms that can run without a person in the loop each time: **human judgment → encoded feedback → automated verification → autonomous execution.** Every step in that chain is currently manual almost everywhere outside software engineering. Whoever industrializes it first, domain by domain, owns the layer nothing else here can substitute for.\n\nIf a single mental model is worth keeping: a fluent witness, a procedural harness, and a judge who decides — useful shorthand, not the foundation. The honest addendum is that the judge can still be enforcing a bad law. Building better judges matters. Building the machinery that lets a domain write better laws in the first place, cheaply, and more than once, is the bigger, less obvious race.\n\n**Sources:** arXiv:2604.25850 (AHE, incl. cross-model transfer and token efficiency); arXiv:2607.12227 (matched-budget re-evaluation of harness evolution); arXiv:2604.13107 (Odoo ERP agents); Anthropic API pricing documentation (cached-read rates); JetBrains ACP documentation.\n\n[Intelligence Is Getting Cheap. Knowing When It’s Wrong Isn’t.](https://pub.towardsai.net/intelligence-is-getting-cheap-knowing-when-its-wrong-isn-t-41bfda1ddc41) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/intelligence-is-getting-cheap-knowing-when-its-wrong-isnt", "canonical_source": "https://pub.towardsai.net/intelligence-is-getting-cheap-knowing-when-its-wrong-isn-t-41bfda1ddc41?source=rss----98111c9905da---4", "published_at": "2026-08-16 00:01:03+00:00", "updated_at": "2026-08-16 00:41:26.508035+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "ai-infrastructure"], "entities": ["Cognition", "Windsurf", "Devin Desktop", "JetBrains", "Gemini CLI", "GitHub Copilot", "Codex", "Terminal-Bench 2"], "alternates": {"html": "https://wpnews.pro/news/intelligence-is-getting-cheap-knowing-when-its-wrong-isnt", "markdown": "https://wpnews.pro/news/intelligence-is-getting-cheap-knowing-when-its-wrong-isnt.md", "text": "https://wpnews.pro/news/intelligence-is-getting-cheap-knowing-when-its-wrong-isnt.txt", "jsonld": "https://wpnews.pro/news/intelligence-is-getting-cheap-knowing-when-its-wrong-isnt.jsonld"}}