Photo: Shkuru Afshar / Wikimedia Commons / CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0) One of Australia's big four banks is putting its autonomous AI systems through security and compliance checks as the financial sector races to deploy agents that can act on their own
National Australia Bank is preparing to test the security and operational guardrails of an agentic AI platform, a move that puts one of Australia’s largest financial institutions at the forefront of a question every bank will eventually have to answer: what happens when your AI can make decisions without a human in the loop?
NAB’s AI build-out has been aggressive #
In April 2026, the bank launched its first dedicated AI Science team, led by George Mathews, with a mandate to design and evaluate agentic AI systems that can operate safely inside one of the most heavily regulated industries on earth. The team’s focus isn’t abstract research. It’s building the frameworks, evaluation methods, and architecture patterns that determine whether an AI agent gets deployed or shelved.
NAB has already deployed OpenAI-based agents for document processing, and the results have been hard to ignore. The bank processes roughly 15,000 trust deeds annually, a task that used to take about 45 minutes per deed for human reviewers. The AI agents cut that to approximately 1 minute, a 97% reduction in review time.
Across the broader organization, NAB has standardized AI tooling for roughly 6,000 developers and partnered with platforms like Harness to embed security and compliance checks earlier in the development pipeline. By March 2026, agentic AI applications for customer workflows had reportedly hit a 90% adoption rate across the bank’s divisions.
Why security testing for agentic AI is different #
Agentic AI systems are designed to take actions autonomously, processing documents, initiating workflows, interacting with customers, sometimes chaining multiple steps together without waiting for human approval at each stage. If a traditional AI hallucinates, someone catches it before anything happens. If an agentic AI hallucinates, it might already be three steps into executing something nobody asked for.
A system that processes 15,000 documents at one minute each can propagate errors at a pace no human team could match. For a bank the size of NAB, a single agentic AI failure that touches customer data or financial transactions could trigger regulatory scrutiny from the Australian Prudential Regulation Authority, reputational damage, or both.
What this means for the broader financial sector #
NAB’s collaboration with Harness suggests that compliance-as-code, where security checks are automated and embedded in the development process rather than bolted on afterward, is becoming table stakes for selling into regulated industries. NAB’s 90% adoption rate across divisions shows the demand side is already there. The question now is whether the guardrails, testing frameworks, and governance structures can keep pace with deployment that is already well underway.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our