cd /news/artificial-intelligence/credit-without-ground-truth-auditing… · home topics artificial-intelligence article
[ARTICLE · art-105457] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

A new arXiv study (2608.19760v1) auditing step-level credit assignment in LLM agents against causal ground truth from executed replay in ALFWorld found that none of the credit signals used to train LLM agents—LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence—identifies causally important steps better than chance. The ground truth is sparse (30.5% of decision points carry measurable effect), and measurability varies by policy (13.1% vs. 26.8%). A confidence-only router cuts judge cost by 13.1% per turn but recovers pivotal steps at chance level, and a seven-arm pre-registered training experiment showed no arm reliably outperforms the untrained policy, with apparent differences explained by training dose, not credit content.

read1 min views1 publishedAug 21, 2026

arXiv:2608.19760v1 Announce Type: new Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance. Existing evaluations grade these signals against annotated step correctness; we audit them against step contribution -- what re-sampling the policy's own alternatives at each decision point and rolling forward actually changes about the outcome -- and the two come apart. The ground truth itself is structured: causal contribution is sparse (30.5% of decision points where ground truth is defined carry measurable effect), and measurability is model-dependent -- the fraction of points with no policy-supported counterfactual differs by a factor of two (13.1% vs. 26.8%) between two similar-scale policies. The failure mode is identifiable: implicit credit echoes the policy's fluency (median rank correlation +0.75, replicating at +0.70 in a second family under a corrected instrument), while conditioning on the outcome adds no causal information (partial correlation -0.004, Qwen). A confidence-only router recovers pivotal steps at chance level, but cuts judge cost by 13.1% per turn (14.0% per trajectory). In a seven-arm pre-registered training experiment, no arm reliably outperforms the untrained policy, and the checkpoints' apparent instrument signature is fully explained by training dose -- sparser credit retains fewer examples, an order-of-magnitude spread in optimizer steps -- not credit content. Comparisons of credit rules must therefore match effective sample size, or they measure dose, not credit.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/credit-without-groun…] indexed:0 read:1min 2026-08-21 ·