Taylor Paletta was glancing at two versions of the same customer-service interaction on her monitor. As director of support and digital success at the enterprise software company, Ninety, she could read the support chat in one window while watching the customer’s product session in another. The customer had already spent several minutes trying to solve the problem, including twice attempting the sequence the AI agent would eventually recommend.
The chat transcript alone looked reasonably positive and successful. The customer asked a question, the AI agent responded promptly and the case appeared to move toward resolution. Yet the session recording spun a different story altogether. The customer was being asked to retrace steps that had already failed, while the AI agent had no awareness of what they had done before opening the support widget. Both records contained accurate data; however, only the behavioral record explained the customer’s actual experience.
Paletta’s team addressed this apparent systemic issue by feeding behavioral session summaries into their customer support AI agent before the conversation begins, not after. The AI agent can now see what the customer did inside the product, where the process broke down and which steps have already failed. Moreover, when the interaction escalates, the human agent inherits the same context rather than reconstructing the story from scratch. Tickets with that behavioral context now resolve at 79%, six points higher than tickets without it.
A six-point improvement may not sound dramatic, yet it is a real gain tied to a specific change in how the work gets done, resulting in real revenue. Additionally, and perhaps more importantly, it illustrates a measurement challenge well beyond customer support. That is, many organizations know how much AI they have deployed, but they have far less visibility into whether AI changed the work or improved the outcome.
Indeed, MIT researchers examining U.S. enterprise AI deployments reported roughly $35 billion to $40 billion committed, while 95% of organizations reported no measurable return. The problem is that companies tend to measure what their systems make easy to count — licenses, logins, prompts, sessions and features activated. Then they use those measures as questionable proxies for value, not realizing or acknowledging that activity actually may be inversely related to value.
Customer service, as much as any business function, makes activity-versus-value distinction especially lucid. Common customer support AI applications don’t capture behavioral data and therefore can create more effort for customers and agents when issues are escalated. Applications informed by behavioral data, however, may produce fewer interactions precisely because they resolve more issues correctly the first time. Therefore, prompt volume or time in the application can be ambiguous measures. Sometimes, more customer interaction indicates more friction rather than more value.
This suggests distinguishing between different types of behavioral signals that show AI is merely available or being used versus signals that show work is actually changing:
Displacement and friction-related measures are particularly aligned with economic value because they reveal what changed in the process and what the customer (or employee) had to do differently as a result.
At Ninety, for example, AI-generated session summaries now arrive inside the agent’s ticketing workflow. Previously, an agent handling a difficult case might open another application, scrub through a session replay, locate the point of failure and then reconstruct the story while the customer waited. With behavioral context already attached to the ticket, much of that work can disappear. Consequently, the organization can measure fewer applications opened, fewer handoffs, shorter resolution paths and less manual reconstruction rather than merely reporting that an AI feature was used.
Behavioral data does more than make the AI more efficient; it changes what the AI knows about the customer at the moment a decision has to be made. In service, this becomes important because a correct answer can still produce a poor experience when it arrives without context. Paletta’s team, for instance, found that some customers tired of AI conversations even when the answers were technically accurate, but the interaction felt cold, repetitive or unnecessarily difficult.
“Traditional adoption metrics have little to say about that kind of problem,” says Scott Voigt, co-founder and CEO of Fullstory. “Behavioral signals, however, can reveal repeated attempts, hesitation, abandonment, rage clicking or a pattern of moving back and forth through the same workflow.”
These types of signals can then trigger a different response, including routing the customer to a person before frustration compounds. In other words, behavioral data becomes part of the service logic rather than merely another source for retrospective reporting.
Paletta describes some of these branching experiences as a “choose your own quest” model, in which observed behavior influences the next path, and AI handles the decision until human intervention becomes appropriate. Furthermore, because the path is based on what customers actually do rather than what designers assumed they would do, the organization can refine the experience continuously as new patterns appear.
There is, however, an important governance issue. Once executives realize that conventional AI metrics reveal too little, the reflex may be to monitor individuals more closely. That approach can easily distort the very behavior the organization is trying to understand, since employees who feel personally tracked and scored may experiment less openly, conceal failed attempts or move work into channels that are harder to observe.
A more useful unit of analysis is generally not the employee but the work: the workflow, journey, handoff or decision. Instead of asking how many prompts a particular support agent entered, ask where agents still have to reconstruct context manually, which cases require repeated escalation, where AI output routinely needs correction, and whether someone’s effort is rising or falling. Similarly, rather than judging people by an isolated session, look for aggregate patterns showing where people overall stall, repeat themselves, abandon a path or seek human help.
Moreover, this approach tends to produce better data because the purpose of the instrumentation is easier to defend. Instrumentation intended to improve a process is fundamentally different from instrumentation used to grade an employee; once people begin optimizing their behavior for the score, the measurement itself corrupts the signal. The former can reveal where the service design is failing; the latter may encourage people to change their behavior simply to “game the system.”
One of Paletta’s more interesting measures extends the same thinking beyond Ninety’s own support environment. Customers now frequently consult public AI models before they visit a company’s support site. Consequently, an external system the company does not control may already be functioning as an informal first layer of customer service, interpreting company content and shaping the customer’s expectations before a formal support interaction begins.
So, the Ninety team compiled the company’s help center and blog content into JSON, tested how frontier models such as ChatGPT and Gemini alone answer questions about the product, and built a scoring system for answer quality. The scoring system helps the team identify where Ninety’s source material produces weak answers so they can improve the underlying content.
This expansion of what behavioral data includes not only what users do inside the sanctioned AI application, but also the observable choices they make around it, e.g., whether they abandon one channel for another, whether they repeatedly seek clarification, whether they bypass an AI agent and whether they arrive at support after already trying to solve the problem elsewhere. Indeed, those behaviors often reveal dissatisfaction or disengagement earlier than a survey, complaint or renewal discussion.
Return, then, to the customer who had already tried the same troubleshooting sequence twice. The support transcript could reasonably have been counted as a successful AI interaction, and a conventional customer service analysis dashboard might have added another prompt, another active user and another resolved conversation to its totals. Yet those measures would still have missed the fact that the customer was repeating failed work because the AI lacked behavioral context.
The customer who had already attempted the same troubleshooting sequence twice could easily have registered as an AI success: another active user, another prompt, another conversation apparently resolved. Yet almost every conventional measure would have missed the part that mattered — the customer was doing unnecessary work because the AI lacked the context to know better.
Behavioral data connects what the AI did to what happened around it, what the customer had already attempted, what the employee no longer had to do, where friction disappeared or compounded, and whether the process itself improved. For CIOs trying to establish the business value of AI, behavioral data answers “What happens better, faster, different or not at all because of AI?” rather than simply “How much are people using AI?”