{"slug": "hexbench", "title": "hexbench", "summary": "The hex-bench benchmark, a personal evaluation suite for the hex project, tests personal-assistant workflows using Inspect AI's react solver against a local OpenAI-compatible llama.cpp endpoint, with 18 tasks across six categories scored per epoch. Baseline runs use larger cloud models for comparison, and scores are reported with 95% Wilson confidence intervals, where overlapping bands indicate statistical ties rather than definitive rankings.", "body_md": "## What this measures\n\nhex-bench is a personal benchmark for the [hex](https://0xff.nu/hex) project that tests the personal-assistant workflow of this project: every task goes through Inspect AI’s `react`\n\nsolver, which may call real tools — Todoist, Obsidian, web search — against a local OpenAI-compatible llama.cpp endpoint. Each of the six categories below has three tasks (18 in the full suite); every task is scored per epoch.\n\nbaseline runs are made against bigger cloud models to establish a point of comparison of sorts.\n\n**Tool use**— calls the right real tool (Todoist / Obsidian / web) on live data.** Reasoning**— multi-step analysis over retrieved information.** Verbosity**— answer length: complete but no filler.** Agentic loop**— knowing when to stop acting and answer.** Understanding**— answers that match the actual source material.** Response format**— output conforms to the required format (bullets, JSON, YAML, table, …).\n\n**score** = mean of all scorer values for the task (0–1).\n\n**quality** = mean of the rubric scorers: `judge`\n\n, `loop_eval`\n\n, `clarification_eval`\n\n.\n\n**protocol** = mean of the deterministic compliance scorers: `tool_used`\n\n, `fmt_*`\n\n.\n\n**pass** score ≥ 1.0 · **partial** 0 < score < 1 · **fail** score = 0.\n\n**Timings** — eval t/s, gen t/s, and wall time come from `sample_speed()`\n\n; older runs may lack them (shown as –).\n\n**± / green band** — 95% uncertainty on the pass rate at this sample count (Wilson interval). Overlapping bands mean the models are statistical ties, not a real ranking.", "url": "https://wpnews.pro/news/hexbench", "canonical_source": "https://0xff.nu/hexbench/", "published_at": "2026-08-25 00:00:00+00:00", "updated_at": "2026-08-26 07:14:33.293438+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-tools", "ai-infrastructure"], "entities": ["hex", "Inspect AI", "Todoist", "Obsidian", "llama.cpp", "OpenAI"], "alternates": {"html": "https://wpnews.pro/news/hexbench", "markdown": "https://wpnews.pro/news/hexbench.md", "text": "https://wpnews.pro/news/hexbench.txt", "jsonld": "https://wpnews.pro/news/hexbench.jsonld"}}