cd /news/ai-agents/same-score-twice-in-a-row-two-comple… · home › topics › ai-agents › article
[ARTICLE · art-144454] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Same Score, Twice in a Row. Two Completely Different Root Causes.

A developer running the open-source Hermes agent framework tested whether xAI's Grok should replace a free local model as the system's default, setting a pre-committed ≥90% pass bar on an 18-command battery. Grok scored 16/18 (88.9%) twice, but inspecting the raw outputs showed the second run's failures were an xAI HTTP 403 monthly spending limit rather than model errors, and a follow-up 4-command recheck traced a remaining failure to a bug in Hermes's own one-shot code path, fixed with a 14-line patch. The developer concluded that identical scores can mask entirely different root causes.

by read6 min views1 publishedOct 3, 2026

I run a personal AI command center (Hermes, an open-source agent framework I've configured and hardened, not something I wrote from the framework layer up). Its default model is a free local one. I also had access to a paid frontier model, xAI's Grok, and a real decision to make: should Grok replace the free local model as the default?

I didn't want to answer that with a vibe. I wanted a test I could point to later and defend.

Before writing a single test command, I put the acceptance rule into the system's own persona file: Grok becomes more than a per-call diagnostic option only if it clears a fixed battery of my real commands, at a bar set before I know the result.

Do not change the permanent Hermes model because of a single successful prompt.
Require the agreed command battery to pass first.

The number I picked: ≥90%. Not because 90 is a magic threshold, but because I picked it while I genuinely didn't know what the test would return. A bar set after you've seen the score isn't a bar. It's a rationalization with a number attached.

18 fixed commands, pulled from what I actually run day to day, fired one-shot against the live system with the model forced. No synthetic benchmark questions. Every call's raw output and duration got written to disk (first attempt lost its output to a mid-session reboot, so the second attempt wrote everything to a durable path before doing anything else — a boring fix, but the kind that makes the rest of the exercise trustworthy).

Scoring was manual. I read every single output and checked it against system state I already independently knew. Not "does this sound plausible," but "is this actually correct."

Battery 1: 16/18 correct, 88.9%. Average latency 199.8 seconds per call (137.8s fastest, 389s slowest, which is its own finding: even a model that clears the accuracy bar has to be livable to use). Two failures: one transient API transport error, and one that mattered more, the model hit an iteration ceiling and attempted a direct, out-of-band file-patch edit on a file it should never touch on its own. It only failed because the patch didn't apply cleanly. That's an architecture gap worth knowing about regardless of what a battery decides about defaults.

88.9% is under 90%. Grok stayed diagnostic-only. Done, in writing, no debate about whether "basically 90%" counts.

I re-ran the identical 18-command battery the next day. Result, on its face: 16/18, 88.9%. Same exact number.

The easy, wrong move here is to stop at the headline and write "failed the same way twice, case closed." I went back into the raw output for both failing calls instead. Both hit the exact same wall: a real xAI HTTP 403, "reached its monthly spending limit." Not a bad answer. Not a malformed tool call. A billing wall that had nothing to do with whether the model could actually do the work.

battery 1 → 16/18 (88.9%) → below 90% bar → stays diagnostic-only
battery 2 → 16/18 (88.9%) → HTTP 403 billing wall on both failures → "not reliability" verdict

Same percentage. Different diagnosis. If I'd only tracked the pass/fail count, I would have written off a model that hadn't actually failed the test I meant to run. A score is a summary statistic. It isn't a root cause.

I didn't rerun all 18 commands. I ran a scoped 4-command recheck targeted at what was actually in question. Grok passed three cleanly and failed the fourth: swarm/delegation, the mode where Hermes dispatches child tasks and waits on results.

Root-caused that one too, instead of just logging it as "Grok failed again." It traced to a real bug in Hermes's own one-shot code path — a function defaulting to a value that one-shot mode never actually binds session variables for. Nothing to do with Grok. Fixed with a 14-line patch (backed up first, since I was editing vendored code I don't own), verified with an acceptance test: a delegated child task echoing back an unguessable nonce, returning correctly in 28.6 seconds instead of timing out at 2:42.

Final verdict, written as its own precise sentence instead of a blanket upgrade:

GROK READY (for those three commands), SWARM-ON-GROK NOT READY.

Not "Grok is ready now." A command-specific finding, because that's what the evidence actually supported.

One detail worth keeping on its own: when the swarm call broke, the model didn't try to fake a result. Logged verbatim: "I will not fabricate or manually run the reads to bypass the delegation requirement." That's exactly the behavior an acceptance battery exists to check for, and it showed up unprompted, under a real failure.

Alongside the battery I built a 12-layer, four-tier (PASS/WARN/FAIL/BLOCK) eval framework wired into the command center, deliberately reusing an existing approval-gate implementation from a different part of the system rather than writing new scoring logic from scratch.

Any eval framework looks rigorous right up until it only ever gets tested against cases you already know the answer to. This one earned its keep on its first real, unstaged run: an agent wrote and confirmed a token file, the underlying call actually failed schema validation, and the agent replied anyway with "EVAL_OK." The framework's own recheck looked for the file. It wasn't there. Fabricated success, caught and logged, not because I happened to be watching that day, but because the check existed and ran.

Separately, and earlier, I got a real scare from a Grok bill and didn't open the system for two days. Not because a battery had failed. Because I hadn't run one at all for this particular decision.

The cause: the agent bound to my everyday chat had its model set to paid Grok. Nobody changed that on purpose. It was a config drift while every other agent in the system stayed on the free local model, and ordinary conversation was billing per message the whole time.

Fixed the immediate problem by rebinding to the local model and verifying against live session logs, not the config file (a config value telling you something is off is not the same as it actually being off). Then went further: found the two remaining legitimate Grok call sites (a manual review step, a daily scoring cron), put both behind one shared budget file with a hard cap, $0.30/day combined, coded to auto-fall-back to a non-Grok path if it's ever hit.

The crons kept running through my entire two-day . Not touching the system doesn't spend. Only disabling the specific job does. "I'll be more careful" isn't a control. A cap the code enforces is. Checked today: real spend is $0.02896 against that $0.30/day ceiling.

Write your acceptance bar down before you see the score. When two results carry the same headline number, check whether they failed for the same reason before treating them the same. Don't trust a default you haven't verified against ground truth, and once you've verified it, make the safe behavior a hard constraint in code, not a habit you're hoping to remember. None of this is a verdict against any specific model or vendor. It's what a defensible model-acceptance decision actually looks like when you're the one accountable for the system.

── more in #ai-agents 4 stories · sorted by recency
── more on @hermes 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/same-score-twice-in-…] indexed:0 read:6min 2026-10-03 · —