The AI assurance gap: CIOs need proof that agentic AI controls actually work Gartner predicts that 40% of enterprise applications will include task-specific agents by the end of 2026, up from less than 5% in 2025, and that over 40% of agentic AI projects could be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. IBM's 2025 Cost of a Data Breach research found that 13% of surveyed organizations reported breaches involving AI models or applications, with 97% of those lacking proper AI access controls. The article argues that CIOs need independent assurance mechanisms to verify that agentic AI controls actually work, as agents operate across systems and can change behavior without obvious application changes. Enterprises have spent decades learning how to audit people and software. Agentic AI creates a third category: systems that interpret instructions, call tools and act across workflows without a mature assurance model built around them. In my work as a leader and investor across technology-enabled businesses, I have spent years around automation, cybersecurity, compliance, workflow design and board reporting. I have watched management teams gain confidence from dashboards, policies and approval records, then face a harder question when a board member, auditor or regulator asks whether the controls performed as intended. Agentic AI complicates that question because a single outcome may pass through several systems. An agent can collect information, choose a tool, produce code, route a request and hand work to another agent before a person approves the result. No single manager may have observed the full path. Executive accountability remains human even when the operating activity becomes more autonomous. The CIO may have to explain who authorized the activity, whether the agent stayed within its approved purpose and what evidence supports management’s answer. In a recent https://x.com/demishassabis/status/2076957440109625718 framework for frontier AI https://x.com/demishassabis/status/2076957440109625718 , Google DeepMind CEO Demis Hassabis proposed an independent standards body that could evaluate advanced models before deployment and address critical vulnerabilities after release. His proposal focuses on frontier models, but the principle carries into the enterprise: expanding autonomy creates a corresponding need for independent assessment. My rule for this stage of adoption is straightforward: no agent should gain more autonomy than the company can verify. Companies have spent decades building controls around people. Employees have job descriptions, reporting lines, approval limits and access rights. When someone leaves, an established process removes access and transfers responsibility. Traditional software also fits a familiar structure. A program follows defined instructions inside systems with owners, release procedures, test records and change controls. Complexity can make review difficult, but the accountability chain is usually visible. An AI agent sits between those categories. It operates through software while interpreting instructions with room to choose a path. Its behavior may change when the model, prompt, connected data, available tools or surrounding workflow changes. A control approved during deployment can weaken months later without an obvious change to the application. Standards are still developing as adoption accelerates. In February 2026, NIST launched an AI Agent Standards Initiative https://www.nist.gov/artificial-intelligence/ai-agent-standards-initiative focused on secure operation and interoperability for agents capable of autonomous action. Par Chadh Gartner predicted that 40% of enterprise applications would include task-specific agents by the end of 2026, up from less than 5% in 2025. More than 40% of agentic AI projects could be canceled by the end of 2027 https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 because of escalating cost, unclear business value or inadequate risk controls. IBM’s 2025 Cost of a Data Breach research found that 13% of surveyed organizations reported breaches involving AI models or applications https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications%2C-97-of-which-reported-lacking-proper-ai-access-controls . Among that group, 97% reported inadequate AI access controls. 63% of organizations in their study lacked governance policies for managing AI or preventing shadow AI. Par Chadh One team may approve an agent, another may connect it to data and a third may own the workflow. Management still carries responsibility when the agent exposes information, produces an error or acts outside its approved purpose. Two-thirds of CIOs and CTOs surveyed were being held accountable for AI systems they did not fully control https://www.cio.com/article/4182288/cios-are-being-held-accountable-for-ai-they-dont-fully-control-ibm-study-finds.html . Seventy percent said technology was spreading across the business faster than IT could track it, while 77% said adoption was outpacing governance capabilities. Par Chadh The survey findings expose the boardroom gap: management carries accountability while control remains distributed across teams, systems and workflows. An assurance model must provide more than a statement of intent. Policies, dashboards and logs help establish control. Assurance begins when the company tests whether an agent stayed inside the conditions management approved. Before deployment, management should document the agent’s business purpose, accountable owner, systems touched, allowed actions and stop conditions. Testing should then determine whether the agent can reach information outside its scope, call an unapproved tool, continue after a stop condition or carry an incorrect assumption into another system. I would treat an agent’s autonomy as a renewable license because its operating condition will change after launch. Renewal should follow any material change to the model, connected data, available tools, workflow or authority. NIST’s work on the https://www.nist.gov/news-events/news/2026/03/new-report-challenges-monitoring-deployed-ai-systems?utm source=chatgpt.com challenges of monitoring deployed AI systems https://www.nist.gov/news-events/news/2026/03/new-report-challenges-monitoring-deployed-ai-systems?utm source=chatgpt.com identifies drift, fragmented logging and immature standards as barriers to post-deployment oversight, while the World Economic Forum recommends https://www.weforum.org/publications/ai-agents-in-action-foundations-for-evaluation-and-governance/?utm source=chatgpt.com scaling safeguards with an agent’s autonomy, authority and complexity https://www.weforum.org/publications/ai-agents-in-action-foundations-for-evaluation-and-governance/?utm source=chatgpt.com . A change-triggered review ties assurance to the version of the agent and workflow in use, giving the CIO a defensible basis for continued authority. Consider a coding agent that begins by drafting test cases, then gains access to repositories, tickets, CI/CD tools and production documentation. A production change could involve an instruction, code, a tool call, an automated test, a ticket update and human approval. Assurance must show how the result was produced, which systems participated, whether the agent crossed a boundary and how exceptions were handled. For every agent with meaningful operating authority, I would expect four connected records: the approved baseline, boundary-test results, a history of behavioral drift and an account of exceptions and interventions. Together, they give management a record that can support a board discussion, audit or regulatory response without depending on the technical team’s memory. Internal teams will remain responsible for designing controls and operating the environment. At board-level scale, management also needs review independent from the people who built and run the agent. I have seen management ask auditors or other independent certified professionals to sign off on a system. They cannot sign when they have not completed the work required to support that opinion. Leadership wants confidence, the board wants an answer and the independent party needs a body of evidence that can be tested. Agentic AI will make that evidence harder to assemble after an incident. Records have to be produced while the work occurs. Once an agent has acted across several systems, reconstruction may depend on logs created by different vendors, teams and tools. Missing context can turn a clear technical event into an uncertain management explanation. CIOs should design assurance into the workflow. The operating record needs to capture approved purpose, tested boundaries, material changes, exceptions, human interventions and unresolved findings. Independent review can then examine whether the control operated during the period management is being asked to discuss. The operating record gives executives a basis for changing an agent’s authority. Broader responsibility should follow tested boundaries and a clean exception history. Drift or repeated intervention should pause expansion until the cause is understood. Before approving broader use, I would ask: Those questions create a management standard for deciding whether an agent is ready to move from a limited workflow into broader enterprise operations. Enterprises have spent decades learning how to audit people and software. Agentic AI creates a third category that requires its own assurance model. The next discipline for the enterprise is auditing autonomy through approved behavior, tested boundaries, monitored change and documented intervention. This article is published as part of the Foundry Expert Contributor Network. Want to join?