{"slug": "reality-is-the-final-verifier-on-two-key-gaps-in-agentic-software-engineering", "title": "Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering", "summary": "A September 10, 2026 arXiv paper by Alexander Krentsel proposes a \"two-gap framework\" identifying requirement gap and model gap as the main failure modes of agentic software engineering, where reward hacking exploits omissions in requirements or model and hallucination widens the gaps by fabricating requirements or environment assumptions. Because neither gap can generally be certified closed in an open, changing world, Krentsel proposes an assurance-revision loop that uses deployment evidence to revise requirements, model, or evaluator when stakeholders reject the resulting behavior, casting assured agentic development as a resource-allocation problem over human judgment, agent capability, and compute. The paper argues human judgment is the bottleneck for the requirement gap and faithful, costly evaluation for the model gap, with reality remaining the final verifier.", "body_md": "# Computer Science > Software Engineering\n\n  [Submitted on 10 Sep 2026]\n\n# Title:Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering\n\n[View PDF](https://arxiv.org/pdf/2609.12039)\n\n[HTML (experimental)](https://arxiv.org/html/2609.12039v1)\n\nAbstract:Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set of requirements under a model of the deployment environment. Yet even a formal proof that the implementation satisfies the requirements under the model cannot guarantee acceptable behavior after deployment. Requirements only approximate stakeholder intent, and the model only approximates the real deployment environment. We call these together - requirement gap and model gap - the two-gap framework, which unifies the main failure modes of agentic software engineer-ing: reward hacking exploits omissions in the requirements or model, while hallucination widens the gaps by fabricating requirements or environment assumptions.\n\nBecause neither gap can generally be certified closed in an open, changing world, the goal shifts from closing them to continuously narrowing them. We therefore propose an assurance-revision loop that uses deployment evidence to revise the requirements, model, or evaluator when stakeholders reject the resulting behavior. We then cast assured agentic development as a resource-allocation problem over human judgment, agent capability, and compute. The two principal bottlenecks mirror the two gaps: human judgment for the requirement gap and faithful, costly evaluation for the model gap. Reality remains the final verifier: acceptable behavior under actual deployment conditions is the ultimate test, while predeployment evaluations remain proxies for it.\n\n## Submission history\n\nFrom: Alexander Krentsel [\n[view email](https://arxiv.org/show-email/a9ae974f/2609.12039)]\n\n**[v1]** Thu, 10 Sep 2026 17:58:48 UTC (264 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/reality-is-the-final-verifier-on-two-key-gaps-in-agentic-software-engineering", "canonical_source": "https://arxiv.org/abs/2609.12039", "published_at": "2026-09-19 03:52:10+00:00", "updated_at": "2026-09-19 04:24:42.137785+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "ai-safety", "artificial-intelligence"], "entities": ["Alexander Krentsel", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/reality-is-the-final-verifier-on-two-key-gaps-in-agentic-software-engineering", "markdown": "https://wpnews.pro/news/reality-is-the-final-verifier-on-two-key-gaps-in-agentic-software-engineering.md", "text": "https://wpnews.pro/news/reality-is-the-final-verifier-on-two-key-gaps-in-agentic-software-engineering.txt", "jsonld": "https://wpnews.pro/news/reality-is-the-final-verifier-on-two-key-gaps-in-agentic-software-engineering.jsonld"}}