cd /news/artificial-intelligence/evidence-state-reliability-under-con… · home topics artificial-intelligence article
[ARTICLE · art-109644] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline

A new arXiv paper (2608.21559v1) introduces Evidence-State Reliability (ESR), an evaluation layer for multi-stage LLM pipelines, and finds that structural parser validity can improve while evidence-sensitive stage success deteriorates under controlled degradation. In tests using GLM-5.2 on 60 sanitized base cases with 713 retained execution rows, all nine operational stage-success estimates were negative with 95% bootstrap intervals below zero, while all nine parser-validity point estimates were positive, though three partial-dropout intervals included zero. The results also show perfect detection of degraded evidence (1.0) in parser-valid degraded audit outputs but zero recovery (0.0) in degraded escalation outputs.

read1 min views1 publishedAug 25, 2026

arXiv:2608.21559v1 Announce Type: new Abstract: Multi-stage LLM pipelines can remain structurally valid even when evidence available to downstream stages becomes incomplete, compressed, or conflicting. This paper introduces and operationalizes Evidence-State Reliability (ESR), an evaluation layer concerned with whether intermediate evidence remains sufficiently complete, grounded, internally consistent, and usable for a stage's assigned function. ESR is evaluated separately from parser validity, which measures structural conformance. We evaluate the framework using GLM-5.2 on 60 sanitized base cases under four evidence conditions: clean, compressed-lossy, partial-dropout, and noisy-conflicting. Each condition was processed through decision, audit, and escalation stages. The design comprised 720 planned and ledgered calls, with 713 retained, sanitized execution rows. Across nine matched degraded-minus-clean condition-stage comparisons, all operational stage-success estimates were negative, and all 95% bootstrap intervals remained below zero. All nine parser-validity point estimates were positive, although the three partial-dropout intervals included zero. Among parser-valid degraded audit outputs, degradation detection was 1.0 in each degraded condition, while false-assurance rates remained non-zero; among parser-valid degraded escalation outputs, recovery was 0.0 in every degraded condition. The results show a bounded reliability-layer divergence in the evaluated pipeline: structural conformance can improve directionally while evidence-sensitive stage success deteriorates under the same controlled intervention. They also separate detection of degraded evidence from recovery. The conclusions are limited to the evaluated model configuration, pipeline design, selected sanitized cases, scoring procedure, and single scaled run.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evidence-state-relia…] indexed:0 read:1min 2026-08-25 ·