16:35
2026-09-07
promptcube3.com
artificial-intelligence
Harbor-Index proves that most agent benchmarks are actually too
Harbor-Index, a curated subset of 82 high-difficulty tasks from 29 benchmarks, found that no model-harness configuration exceeded a 30% pass rate, with GPT-5.5 (using Codex) reaching a ceiling of 28.0…