{"slug": "contracteval-query-conditioned-execution-matching-for-procedural-instruction", "title": "ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance", "summary": "Researchers introduced CONTRACTEVAL, a diagnostic framework that represents procedural instructions as query-active obligations and matches them against response or trace evidence to detect conformance failures. On a controlled suite of audited procedural contracts, output-only and trace-aware LLM judges missed many injected structural failures, while ContractEval detected and localized all of them under gold expected and observed graphs. The authors state ContractEval is not a compliance guarantee but makes procedural conformance auditable rather than implicit in final-answer quality.", "body_md": "arXiv:2609.09458v1 Announce Type: new \nAbstract: As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query. We introduce CONTRACTEVAL, a diagnostic framework for making those active obligations explicit. It represents procedural instructions as query-active obligations and matches them against response or trace evidence, turning omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output-contract violations into distinct conformance failures. On a controlled suite of audited procedural contracts, output-only and trace-aware LLM judges miss many injected structural failures; under gold expected and observed graphs, ContractEval detects and localizes all of them. LLM-backed extraction preserves much of this signal but remains calibration-sensitive. ContractEval is therefore not a compliance guarantee; it makes procedural conformance auditable rather than implicit in final-answer quality.", "url": "https://wpnews.pro/news/contracteval-query-conditioned-execution-matching-for-procedural-instruction", "canonical_source": "https://arxiv.org/abs/2609.09458", "published_at": "2026-09-11 04:00:00+00:00", "updated_at": "2026-09-11 04:27:43.897297+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "large-language-models", "ai-research"], "entities": ["CONTRACTEVAL"], "alternates": {"html": "https://wpnews.pro/news/contracteval-query-conditioned-execution-matching-for-procedural-instruction", "markdown": "https://wpnews.pro/news/contracteval-query-conditioned-execution-matching-for-procedural-instruction.md", "text": "https://wpnews.pro/news/contracteval-query-conditioned-execution-matching-for-procedural-instruction.txt", "jsonld": "https://wpnews.pro/news/contracteval-query-conditioned-execution-matching-for-procedural-instruction.jsonld"}}