# AI agents get better at IT ops, but only with humans in the loop

> Source: <https://www.cio.com/article/4205069/ai-agents-get-better-at-it-ops-but-only-with-humans-in-the-loop-2.html>
> Published: 2026-08-04 15:45:21+00:00

AI agents are performing roughly 1 in 3 actions in enterprise IT workflows (but that share is rising quickly), while human analysts are rejecting about one-quarter of AI-proposed actions (but that rate is falling), according to a new study of tens of thousands of human-AI interactions. Operational data, rather than underlying AI infrastructures, is often the culprit when things go wrong.

Human analysts are approving the most consequential actions, managing exceptions, and supervising and shaping agentic systems, while AI agents are carrying out routine tasks and executions, automation platform provider [Fixify found in the study](https://www.fixify.com/agentic-report).

“That may sound less dramatic than replacing the help desk,” [Matt Peters](https://www.linkedin.com/in/matt-peters-5984b5/), Fixify’s co-founder and CEO, wrote in a [blog post](https://www.fixify.com/blog/agentic-ai-it-lessons). “It’s also a much more credible path to changing how IT work gets done.”

Fixify identified four steps of agentic work: Planning, proposing, approving or declining, then acting on approved steps.

It analyzed nearly 18,000 plans and over 147,000 actions executed by agents across 40 companies over a three-month period, finding that agents are taking over one-third of IT actions, most notably in software, applications, security, and collaboration work where requests tend to be “repeatable and easy to reverse.”

Tasks that are well understood and that present low risk are best suited for the current generation of agents, Peters wrote. [Human analysts](https://www.cio.com/article/4204021/ai-can-do-your-tasks-that-doesnt-mean-it-will-do-your-job.html) remain closely involved in higher-stakes areas like identity verification, setting up and removing IT access (onboarding and offboarding), and hardware environments.

However, AI’s share of the work is increasing as feedback loops improve: Over the three-month period, human approval of AI-proposed actions rose from 23% to 41%, and rejection fell from 27% to 16%, Fixify found.

The company identified six types of actions in AI automation. Running a skill — actually doing something — accounted for 39.4% of all actions). Most of the rest were coordination: sending a message to the human requester (27.7% of actions), leaving an initial comment (13.2%), giving instructions to a human analyst (9.8%), or waiting (8.8%). Running entire workflows accounted for just 1.1% of actions.

AI is building “scaffolding” that wraps around meaningful changes, often planning far more scenarios than the agent will execute. Typically, agents map out 15 possible actions but run only two, Fixify said.

“The agent maps the paths a request could take, then walks down the path that makes the most sense as it meets reality,” the study said.

Peters pointed to one example where an AI agent identified which team needed access to process a high-volume type of ticket. Rather than fully automating the process, the agent did the initial triage, asked questions, then routed tickets to the team that had the information to act immediately.

“We didn’t need a world-ending hive mind,” he said. “We just needed to point a little conversational intelligence in the right direction.”

IT automation typically involves analyzing tickets and moving them along; in other words, low-risk tasks.

But agents do participate in areas like [security](https://www.csoonline.com/article/4204101/ai-is-making-cybersecurity-fundamentals-more-important-than-ever.html) (albeit only about 6%), most notably adding and removing people from groups or channels, unlocking accounts, resetting passwords, analyzing multi-factor authentication (MFA), provisioning (or deprovisioning) accounts, and assigning software licenses.

However, this identity-lifecycle work is where agents failed the most, particularly in onboarding and offboarding and identity-access management (IAM), the study found. “Hardware and connectivity changes rarely fail; identity-lifecycle changes fail three-to-nine times as often.”

Thanks to human-in-the-loop controls, Fixify was able to analyze scenarios where agent recommendation diverged from human judgment. This occurred about 23% of the time.

The largest failure category (nearly 50%) was ‘target not found,’ meaning the agent couldn’t uncover what it needed. This typically comes down to poor data: A user, group, account, or resource was not where the system expected it to be. When people change teams, groups are restructured, accounts are renamed, or work has already been done but not reflected in the system, this is more of an identity hygiene problem than an AI problem. The system needs cleaner and more current data.

Invalid inputs accounted for around 29% of failures, followed by unhandled errors, denied permissions, or invalid operations or configurations. The latter signal “real breakage” in integrations, according to Fixify.

The good news is that AI automation improves over time, even if it might take a while. In [hybrid systems](https://www.cio.com/article/4202404/forward-deployed-engineering-in-the-age-of-agentic-ai-from-vibe-coding-to-governed-autonomy.html), humans keep the most consequential changes under their own control, and iterative rejection and approval helps AI learn.

Over time, agents’ plans get leaner and they start to re-plan when conditions change, rather than pre-planning all kinds of scenarios that may never occur. “That’s a sign of sophistication,” the study said. “Adapting in the moment is a more advanced behavior than trying to pre-script every contingency.”

In turn, humans second guess the system less often and feel comfortable handing off more work. Instead, they control how agents behave, make high-impact decisions, and handle exceptions. “The hardest requests remain human-heavy, especially those that require repeated replanning or contextual judgment,” the study said.

As agentic AI becomes embedded in more workflows — and at deeper levels — enterprises must evolve to accommodate, Fixify emphasized.

This means investing in clean identity data and building strong playbooks, review workflows, and reliable integrations.

Teams should judge agentic tools by their supervision loop and view rejections as a training process, Fixify advised. Analyst time, queues, and metrics should be built around reviewing proposals. Agent replanning can be seen as a routing signal: A single replan might indicate healthy adaptation, while repeated replanning means ambiguity, irrelevance, or unclear policies.

“Make the review surface easy to understand so analysts can assess proposed actions and make quick decisions about how to proceed,” the study advised. “This is where the analyst’s attention belongs.”

*This article first appeared on Computerworld.*
