Nous Research launches Hermes Index to benchmark agentic AI models Nous Research launched the Hermes Index on October 6, 2026, a leaderboard ranking 14 frontier models on agent performance and cost inside its Hermes Agent framework, with Claude Opus 5.5 taking first place at a score of 63.31 and $4.99 per task. GPT-6 Astra placed second at 56.25 and $11.61 per task, while Claude Sonnet 5.5 was third at 53.14 and $2.82 per task, and the cheapest entries were DeepSeek V4.1 Flash at 36.91 for $0.259 per task and Ling 3.0 Flash at 21.56 for $0.054 per task. One day later, on October 7, 2026, Nous Research raised $90 million from backers including Nvidia and Microsoft's M12, bringing total capital to roughly $160 million and a $1.5 billion valuation, with the core Hermes framework staying open-source under the MIT license. Nous Research launches Hermes Index to benchmark agentic AI models The open-source lab's new leaderboard ranks 14 frontier models on agent performance and cost, with Claude Opus 5.5 taking the top spot Nous Research wants to settle an argument the AI industry keeps having: which model actually gets work done, and what does it cost to find out. The open-source AI lab launched the Hermes Index on October 6, 2026. It ranks frontier models on how well they perform as agents inside the company’s Hermes Agent framework. It also tracks how much each task costs to run. How the Hermes Index works The index measures model performance and cost across four test suites. It averages scores across those benchmarks and calculates a mean cost per task for each model. The centerpiece is a new benchmark called Hermes Bench . It contains 150 tasks spread across 25 categories. The index also folds in other established evaluations, including Terminal-Bench 4.0 and SkillsBench. All of this runs through Hermes Agent, which Nous Research started on February 25, 2026. The lab describes it as an evaluation harness focused on user-aligned model performance. The leaderboard: who won and what it cost Nous Research tested 14 different models in the first edition of the index. Claude Opus 5.5 led the rankings with a score of 63.31. Its average cost came in at $4.99 per task. In second place sat GPT-6 Astra , which scored 56.25. Its average task cost was $11.61, more than double what the top model charged. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. Claude Sonnet 5.5 took third with a score of 53.14 at $2.82 per task. Further down the table, DeepSeek https://cryptobriefing.com/markets/deepseek/ V4.1 Flash scored 36.91 at a cost of $0.259 per task. Ling 3.0 Flash scored 21.56 at just $0.054 per task. A fresh $90 million to go with it On October 7, 2026, one day after the index went live, Nous Research secured $90 million in new funding. That round brings the lab’s total capital raised to approximately $160 million. The company now carries a valuation of $1.5 billion. Nvidia https://cryptobriefing.com/markets/nvidia/ and Microsoft https://cryptobriefing.com/markets/microsoft/ ’s venture arm M12 are among the backers. Nous Research says the money is aimed at scaling enterprise deployments of its Hermes technology. The core framework will stay open-source under the MIT license. Background: a young lab with big ambitions Nous Research was founded in 2023. Hermes Agent arrived in February 2026, and the Hermes Index now gives the lab a way to publicly grade the industry’s biggest models on its own turf. What this means for the AI market For enterprises, the cost comparison is practical. A company running thousands of agent tasks a day cares less about a few points of benchmark score and more about the total bill. The gap between $11.61 and $2.82 per task compounds quickly at scale. There are caveats worth keeping in mind. The index measures performance inside Nous Research’s own Hermes Agent framework. Models may behave differently in other harnesses or real-world deployments, so results should be read as one lens rather than a universal verdict. There is also an inherent tension in a company grading the field while raising money to sell enterprise products built on the same framework. Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .