Salesforce evolves AI agent performance to 93% with DarwinX framework Salesforce AI Research published DarwinX, a framework that evolves AI agent harnesses without modifying model weights, raising pass@1 on the 1,260-task WebArena-Infinity benchmark from 43.5% to 93.0% and cutting invalid trajectories from roughly 23.5% to 1.4%. The 12-author paper, led by Yifan Zhang, Yutong Dai, and Juntao Tan and posted to arXiv on July 31, 2026 (paper ID 2608.07545), also lifted Terminal-Bench 2.1 by 7.7 points to 83.2% and averaged about 17 points of improvement across all tested benchmark suites. Salesforce says DarwinX remains a research initiative with no production deployment details and no independent verification published. Photo: U.Lucas Dubé-Cantin / Pexels Salesforce evolves AI agent performance to 93% with DarwinX framework A new evolutionary approach optimizes AI agent harnesses without touching model weights, yielding dramatic benchmark gains across 1,260 tasks Salesforce https://cryptobriefing.com/markets/salesforce/ AI Research just published a framework that treats AI agent optimization like natural selection. Called DarwinX, it evolves the scaffolding around a language model, the prompts, tools, skills, and control flow, while leaving the model’s own weights completely untouched. The result: a real-task pass rate that jumped from 43.5% to 93.0% on a 1,260-task benchmark. That benchmark, WebArena-Infinity, measures how well an AI agent can complete actual web-based tasks. How DarwinX actually works DarwinX maintains a population of “harness variants,” each representing a different configuration of how the agent interacts with the world. It applies population-based natural selection to these variants, keeping what works and combining successful features from different lineages. The framework operates under what the 12-author research team calls a “preserve-and-extend” contract. Only changes that demonstrably improve performance survive to the next generation. An archive of previous lineages provides genetic material for recombination, preventing the system from losing hard-won capabilities while exploring new ones. Fitness isn’t judged by vibes. Each variant gets evaluated through specific benchmark verifiers, automated checks that confirm whether the agent actually completed its assigned task. The numbers that matter The WebArena-Infinity headline number, 43.5% to 93.0% audit-clean pass@1, is striking enough on its own. But a second metric tells an equally important story. Invalid trajectories, cases where the agent essentially went off the rails and produced unusable outputs, dropped from roughly 23.5% to just 1.4%. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. On Terminal-Bench 2.1, a benchmark focused on command-line task completion, DarwinX pushed performance up by 7.7 points to 83.2%. That matched the base model’s native performance. When the researchers ran the same evolved harness on a stronger underlying model, it reached 84.7%, actually surpassing the base model’s score. Across all tested benchmark suites, the average improvement clocked in at around 17 points. Generalization, not memorization A harness evolved on Terminal-Bench 2.1 transferred directly to SWE-bench Verified, a completely different benchmark, without any modifications. The paper, published on July 31, 2026, on arXiv paper ID: 2608.07545 , was authored by a team of 12 researchers led by Yifan Zhang, Yutong Dai https://cryptobriefing.com/markets/dai/ , and Juntao Tan. What this means for enterprise AI For Salesforce, DarwinX represents a potentially significant competitive edge in the enterprise AI agent market. The company has been building out its Agentforce platform aggressively, and a framework that can dramatically improve agent reliability without expensive model retraining could reduce both the cost and complexity of deploying AI agents at scale. Fine-tuning large language models requires massive compute budgets, specialized expertise, and the risk of catastrophic forgetting, where the model loses previously learned capabilities. DarwinX sidesteps all three problems by operating entirely at the harness level. There’s a caveat worth noting. DarwinX remains a research initiative at this stage. The paper doesn’t include details about production deployment, and no independent verification of the results has been published. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .