# Even With Humans in-the-Loop, Agentic AI Systems Struggle

> Source: <https://tuck.dartmouth.edu/news/articles/even-with-humans-in-the-loop-agentic-ai-systems-struggle>
> Published: 2026-07-20 14:14:48+00:00

In a new paper, Tuck professor Lauren Xiaoyuan Lu explores the performance of a human-supervised Gen AI agent dealing directly with customers.

In early 2024, the frontier of generative (gen) AI in customer service consisted of gen AI assistants helping human agents respond to customer inquiries.

At the Chinese e-commerce site Taobao, the largest firm of its kind in the world, this looked like the gen AI assistant analyzing customer order data and previous customer interactions in real time and giving the agent a quick synopsis of the problem in the form of a written message that the agent could send to the customer directly. Simultaneously, the assistant would prepare another message with a proposed solution to resolve the customer issue. Tuck professor [Lauren Xiaoyuan Lu](https://tuck.dartmouth.edu/faculty/faculty-directory/lauren-lu) [studied that system](https://tuck.dartmouth.edu/news/articles/how-gen-ai-can-boost-customer-service) and found that the gen AI assistant reduced the burden on both the customer and the agent for rote communication, allowing the agent to focus more on the customer’s specific needs.

The world of AI moves fast, and the frontier marched forward just a few months later, when Taobao started experimenting with a gen AI agent that could interact with customers directly. In this system, a human agent is monitoring the AI agent and intervening if the AI agent can’t get the job done well. Such setups are described as having a “human-in-the-loop.” In a new working paper titled “[Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba’s Customer Service Operations](https://arxiv.org/abs/2605.14830),” Lu, the Norman W. Martin 1925 Professor of Business Administration, examines how this futuristic system performs when it’s confronted with not-so-futuristic human customers. While the findings are nuanced, it is clear that AI agents don’t mix well with frustrated customers who would rather have a human solve their problem.

Tuck Professor Lauren Xiaoyuan Lu investigates operational drivers of organization performance in healthcare, retail, and supply chain settings. She teaches Supply Chain Management in the Tuck MBA program.

Taobao first deployed its agentic AI system in May of 2024, supporting its online customer service operation that includes 38,000 gig workers serving 380 million daily active users. The system used a large language model to sort incoming chats that were either AI-eligible or AI-ineligible. AI-eligible chats were monitored by human agents as well as an algorithm that flagged chats with an elevated risk of service failure. When the algorithm flagged a chat, a human agent would intervene and complete the chat. Human agents could intervene on their own if they noticed the AI chat was headed toward failure. AI-ineligible chats were immediately directed to human agents. Under this process, less than 10% of chats were AI-eligible.

In August of that year, Taobao conducted an experiment to gauge the effectiveness of the agentic AI system. The randomized field experiment ran for 17 days, and it randomly selected 647 customer service workers to take part. Three hundred forty-five of these workers were in the control group (where humans handled all chats), while 302 were in the treatment group (supervising AI-eligible chats and handling AI-ineligible chats). During the study period, the human and AI agents handled 680,676 online service chats.

When Lu and her co-authors analyzed the data from the experiment, a few patterns started to emerge. They found that, overall, deploying agentic AI improves service speed, but does not improve service quality. Interestingly, customer perceptions of service quality depended on the type of chat. In AI-eligible chats, the increase in service speed did not translate into a better customer experience. By contrast, in AI-ineligible chats, agentic AI deployment generated positive spillovers: they became modestly faster and received higher customer ratings.

The most surprising finding in the paper concerns the “escalation” process—when an AI-eligible chat is flagged by the algorithm and escalated to a human agent. If the chat is escalated because of technical reasons—the inquiry is beyond the AI agent’s capabilities—service quality is preserved. In these cases, while they take a little more time, a human agent can compensate for the AI agent’s shortcomings.

Humans are needed to rescue a failed AI effort, but that rescue may not be effective at all when the failure involves strong negative customer sentiment.

— Lauren Xiaoyuan Lu, Norman W. Martin 1925 Professor of Business Administration

However, if the chat is escalated because the algorithm detected customer frustration or skepticism, the outcomes were generally bad. “The key finding of the paper,” Lu says, “is that humans are needed to rescue a failed AI effort, but that rescue may not be effective at all when the failure involves strong negative customer sentiment.” There are two reasons for this. One is that turning around a negative sentiment is very difficult. The other, Lu explains, is that human agents know how difficult it is to reverse a bad customer experience, and they therefore put forth less effort. In psychology, this is known as learned helplessness: a state of apathy and passivity based on the belief that actions don’t matter.

The findings suggest that firms need to be very careful in how they design systems with human-AI collaboration. One piece of the puzzle is the human workers. Should firms assign workers to specialize in resolving inquiries where customers were frustrated by the AI agent? Perhaps, but “you have to think about the wellbeing and morale of these workers,” Lu says. “They will need to deal with negative customer emotions all the time, and who wants that?” Lu offers that these workers should have additional training, incentives, and recognition.

An alternative route is to integrate workers, so they’re handling AI-ineligible chats and intervening when the AI agent fails. This would help “preserve broad service expertise while keeping supervisors close to the realities of frontline work,” the authors write. But it may also delay human intervention in emotionally sensitive chats, which makes it more important to finely tune the escalation algorithm to catch negative customer emotions early, before they become unrecoverable.

A 2023 McKinsey study posited that customer operations is the functional area that’s going to reap the most economic benefit from AI technology. Klarna, the buy-now-pay-later fintech firm, went all in on that prediction in 2024, when it replaced most of its customer service representatives with AI to cut costs. But, like Taobao, it discovered that the AI agents struggled with nuanced human interactions and complex problems. As a result, Klarna ended up scaling down its AI system and bringing humans back into the process.

As a professor of operations management, Lu sees these case studies as guideposts through the fog of AI adoption. “I believe the systems are improving,” Lu says, “but we are far from being able to automate customer service completely. The technology is just not there yet.”
