cd /news/ai-agents/fast-models-slow-evidence-a-paired-a… · home › topics › ai-agents › article
[ARTICLE · art-145124] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

A paired evaluation of two System-1 decision models for LLM agent harnesses found the hosted model Jev significantly more accurate than the open-weight model Laya on 9 of 11 agent decision points, by margins of +10.8 to +46.0 percentage points, according to the arXiv paper 2610.02267v1. The study tested 7,283 base cases plus 6,640 robustness variants across 18 public sources with byte-identical inputs, and reported that neither model beat chance on zero-shot model routing while Laya changed 30% of its answers when option order was reversed and dropped to 31% accuracy at 50 nearest-neighbour tools versus 98% for Jev on items with a unique correct tool. The authors' self-audit found three analysis errors and one design confound that distorted headline deployment claims, including an omitted pre-screen cost that cut a reported 23.9% saving to an actual 4.3% and gate accuracy reported as end-to-end quality (58% vs. 98%).

by read1 min views2 publishedOct 5, 2026

arXiv:2610.02267v1 Announce Type: new Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya changes 30% of its answers when the option order is reversed and degrades sharply with many or similar candidates (31% at 50 nearest-neighbour tools, vs. 98% for Jev on items with a unique correct tool). We also audit our own pipeline. Three analysis errors and one design confound distorted headline deployment claims: an omitted pre-screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end-to-end quality (58% vs. 98%), in-sample thresholds (5% target, up to 17% held-out misses), and a "channel effect" on injection false positives that vanishes with channel-native content. Two other suspected confounds did not change the conclusions. All cases, raw outputs and analysis code are available at https://github.com/David-DL-Space/sys1-eval.

── more in #ai-agents 4 stories · sorted by recency
── more on @laya 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fast-models-slow-evi…] indexed:0 read:1min 2026-10-05 · —