cd /news/ai-agents/mid-harness-scaling-actions-between-… · home › topics › ai-agents › article
[ARTICLE · art-143138] src=research.nvidia.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

A method called Mid-Harness, which samples and verifies candidate actions before execution while leaving the generator and harness unchanged, raised Pass@1 on TerminalBench-Lite from 50.00% for the base agent to 68.03% with 8 sampled actions using a GPT-5.6 Sol verifier and a TMAX-9B generator. The researchers report that more action sampling yields little benefit under weak verification, while pairwise verification performed best when TMAX-9B served as the verifier, and that distilling responses from the stronger verifier into TMAX-9B further improved Pass@1 without changing the action generator. Combining action and trajectory scaling reached higher success at lower estimated token cost than generating more trajectories alone, identifying action scaling as a target for test-time compute scaling in terminal agents.

by read1 min views3 publishedOct 1, 2026

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.

── more in #ai-agents 4 stories · sorted by recency
── more on @mid-harness 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mid-harness-scaling-…] indexed:0 read:1min 2026-10-01 · —