cd /news/artificial-intelligence/harbor-adapters-and-harbor-index-inf… · home topics artificial-intelligence article
[ARTICLE · art-121913] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Researchers introduced Harbor Adapters, a unified evaluation infrastructure that ports more than 80 agentic benchmarks, and Harbor-Index, a curated meta-dataset of 82 tasks across 29 benchmarks, to enable large-scale agent evaluation. In tests of 8 models across 54 benchmarks, no model-harness configuration exceeded a 30% pass rate, with the strongest (GPT-5.5 with Codex) reaching 28.0%. The team released the adapters, results, analysis, and Harbor-Index as open-source artifacts.

read1 min views1 publishedSep 7, 2026

arXiv:2609.04298v1 Announce Type: new Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @harbor adapters 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/harbor-adapters-and-…] indexed:0 read:1min 2026-09-07 ·