cd /news/ai-agents/twincheck-evidence-grounded-negative… · home topics ai-agents article
[ARTICLE · art-138800] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

TwinCheck, an inference-time verification policy for stateful tool agents, raised task success for GPT-5.6 Sol from 45.3% to 58.5% on 159 multi-turn BFCL V4 tasks with complete exact-replay pairs, according to the arXiv paper arXiv:2609.26911v1. The policy replaces an agent's proposed tool call only when a trace-grounded counterfactual alternative, called a negative twin, passes structural checks and a pairwise verifier prefers it in both candidate orders, with no observed success-to-failure regressions. Exact replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects from resampling.

by read1 min views1 publishedSep 24, 2026

arXiv:2609.26911v1 Announce Type: new Abstract: A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis. It constructs a trace-grounded counterfactual alternative, a negative twin, and replaces the agent's proposal only if the twin passes structural checks and the pairwise verifier prefers it in both candidate orders. For paired evaluation, exact replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects from resampling. In the primary analysis of 159 multi-turn BFCL V4 tasks with complete exact-replay pairs, the complete policy raises task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]), with no observed success-to-failure regressions. Together, these findings recast execution-boundary repair as a constrained comparison, making the counterfactual action itself the object of verification.

── more in #ai-agents 4 stories · sorted by recency
── more on @twincheck 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/twincheck-evidence-g…] indexed:0 read:1min 2026-09-24 ·