{"slug": "identify-ai-model-overuse-with-user-insights", "title": "Identify AI model overuse with User Insights", "summary": "Cloudflare added a model overkill view to its User Insights dashboard, flagging conversations where the selected AI model appears more capable than the task requires and showing which users, agents, and applications drive that usage. The feature is free to AI Gateway users and supports the new Potential Savings view and the Auto Router, which Cloudflare says is launching in public beta alongside this release. Cloudflare said the overkill view is not a leaderboard and does not automatically recommend a replacement model, instead letting teams compare latency, input and output tokens, conversation turns, and total cost before changing a model, workflow, or routing rule.", "body_md": "# Identify AI model overuse with User Insights\n\nWhen we [__launched User Insights__](https://blog.cloudflare.com/identity-aware-ai-gateway/#the-new-user-insights-tab) last month, we wanted to help teams answer a basic question: What are people actually doing with AI? User Insights gives teams a clearer view of their AI usage, showing which users, applications, tasks, and models are driving traffic. It also highlights user and agent anomalies, helping teams identify unexpected or out-of-control spending and usage before they become larger problems.\n\nOur latest update adds something our users have been asking for: context.\n\nSince launch, we’ve heard from users that model names and request counts only tell part of the story. They show where traffic is going, but reveal little about the work behind it: is that request a code review, a research task, or an agent making several calls to complete a job? The same token count can represent very different kinds of work, and you can’t evaluate with model choice without understanding the task.\n\nUser Insights now shows when a model may be more capable than a task requires, which of your users and agents are driving that usage, and how the task, model, cost, and conversation patterns relate. Teams can use these insights to investigate and make targeted changes within their organization. These capabilities are available for free to AI Gateway users.\n\n## Why AI usage is hard to understand\n\nConsider a team that has routed its internal AI traffic through [__AI Gateway__](https://www.cloudflare.com/products/ai-gateway/). After a few weeks, spending is increasing and some requests feel slower than expected, a common challenge as organizations adopt AI at scale.\n\nThere could be several explanations. Developers may be using AI for increasingly complex coding work. Agents may be making too many follow-up calls to complete a task. Or a small group of users or agents may be responsible for a disproportionate share of the organization’s usage.\n\nTokens and request counts alone cannot show which pattern is driving the increase. Teams need to understand what the traffic represents before deciding whether a model, workflow, or routing rule should change.\n\n## Helping teams find where AI models are overkill\n\nThe model overkill view helps teams identify conversations where the selected model appears to be more capable than the task requires. For example, a team might discover that users or agents are sending simple formatting or summarization requests to a high-capability reasoning model.\n\nThat gives the organization a place to start. They can see which users, agents, or applications are associated with the pattern, then investigate the tasks behind it. A team might find that a model is being used because it is the default, because users are unsure which model to choose, or because an agent has been configured to use the same model for every step.\n\nThe overkill view is not a leaderboard and does not automatically recommend a replacement model. It helps teams ask better questions:\n\n- Is this model appropriate for the task?\n- Is the extra capability improving the result?\n- Would a faster or less expensive model produce an equivalent outcome?\n- Is the issue limited to one workflow, user, or agent?\n\nFrom there, teams can compare cost, latency, token usage, and conversation turns before deciding what to change.\n\nThese insights support both the new Potential Savings view and the Auto Router, which is launching in public beta alongside this release. The Potential Savings view helps teams identify requests that may be handled by a faster or less expensive model without compromising output quality. The Auto Router applies these task and model-fit signals automatically, helping reduce costs without requiring a separate routing rule for every workload.\n\nThe Overkill view is a starting point for evaluating model fit. Teams can compare latency, input and output tokens, conversation turns, and total cost for the same type of task. A difficult coding or research task may need a capable reasoning model, while a short summary or simple classification task may not. The goal is not to move every request to the least expensive model, but to understand whether the selected model is appropriate for the work.\n\n## Understand what people are using AI for\n\nTask analysis groups conversations by the kind of work they represent. Initial categories include coding, research, writing, summarization, and data analysis.\n\nThis provides context that a list of model names cannot. An engineering team might use AI mostly for coding and debugging, while another team might use it for research and summarization. A team may also discover that a surprising amount of traffic comes from simple tasks, even though those tasks are being sent to a high-capability model.\n\nThe answers will vary by team. The category data provides a way to investigate those differences using traffic already passing through AI Gateway. Teams can determine whether a model is being used for the work it is best suited to handle, or whether a default model is being applied too broadly.\n\n## Understand the full cost of a task\n\nSome tasks are finished in one exchange. Others take a few rounds of questions, corrections, and follow-ups. Turns analysis shows how much back-and-forth different tasks require. A long conversation is not necessarily a bad thing, especially for complex work. But if a simple task keeps taking several turns, it may be worth looking at the prompt, the model, or the workflow.\n\nThe first request is only part of the cost. Teams should also look at the time, tokens, and money spent before the task is finished. Comparing those numbers can show where a workflow is taking longer or costing more than expected.\n\n## Turn insights into auto routing\n\nOnce a team has identified an overkill pattern and confirmed it across task, cost, latency, and turn data, it can turn that insight into an automatic routing decision.\n\nFor example, the task view might show that much of the team’s AI usage is summarization and formatting. The model view could show that those requests are being sent to a large reasoning model, while the turns view shows that most conversations finish in a single turn. Together, these signals give the team a concrete workload to evaluate.\n\nIn addition to our updates to User Insights, the [Auto Router](https://blog.cloudflare.com/auto-router) is now available in closed beta. The Auto Router uses the conversation trajectory, task category, task complexity, and model-fit signals to automatically route requests to an appropriate model while taking cost into account.\n\nInstead of creating a separate routing rule for every workload, customers in the beta can let the Auto Router select among the models available to their application. The router does not simply send every request to the least expensive model, but instead selects an appropriate model for the task at hand. Complex coding or research work may still need a more capable model, while simpler tasks may be handled by a faster or less expensive option.\n\nTo learn more about the Auto Router and sign up for the closed beta, [__read the blog post here__](https://blog.cloudflare.com/auto-router/).\n\nThe Auto Router uses the same task and conversation signals that power User Insights. The section below explains how those signals are produced.\n\n## How User Insights classifies traffic\n\nEach conversation receives an analysis signal that can be grouped in User Insights. The signal is used for reporting and routing analysis, and is not intended to replace or expose the original request.\n\nThe categorization engine is a dedicated Cloudflare Worker that processes eligible AI Gateway logs. It examines the conversation trajectory, including user requests, assistant responses, tool calls, and tool results, and identifies the type of work being performed, such as coding, debugging, research, or summarization. It also returns a confidence score and evaluates dimensions such as task complexity, intent ambiguity, stakes, and context dependence.\n\nThe Worker returns a category that can be joined with the log metadata used by the dashboard. These signals can also be used to evaluate model fit by comparing how well candidate models suit the task against their cost. The current implementation focuses on a small set of categories that are easy to understand, rather than trying to infer every detail about a user’s work.\n\nThe pipeline follows the existing AI Gateway log architecture. Metadata is stored separately from log bodies, and the current implementation uses Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views. It does not turn the dashboard into a raw prompt browser. Retention of the underlying log bodies continues to follow the configured AI Gateway logging behavior, so teams should review those settings when deciding what to send through the classifier.\n\nThe classification is asynchronous, which means it happens after AI Gateway has handled the request rather than while the user is waiting for a response. AI Gateway writes the log to the existing storage path first, and the classification Worker processes it afterward. This keeps classification out of the request path and adds no latency to the user’s response.\n\nThe tradeoff is that User Insights is not a real-time view. Newly received conversations may not appear in the dashboard immediately, and analysis may trail incoming traffic by approximately one day as logs are processed and aggregated. Teams should use User Insights to identify usage patterns over time rather than monitor live request activity.\n\nThe flow looks like this:\n\n## Connect usage to users, teams, and tools\n\nTask categories become more useful when they can be viewed by user, team, or application. AI Gateway is identity-aware, providing that context without requiring teams to build a separate reporting pipeline.\n\nThis works not only for applications that teams build themselves, but also for developer tools and agent harnesses such as Claude Code, Codex, and OpenCode. By putting AI Gateway behind Cloudflare Access, teams can connect authenticated users and sessions to their AI traffic, allowing User Insights to associate activity with the right person and conversation.\n\nFor custom applications, requests must include both a stable user_id and a session_id for User Insights analysis. The exact identity configuration and field names depend on how the application or tool is set up. The important part is to provide stable, non-sensitive user and session identifiers so usage can be grouped without putting identity data in the prompt itself.\n\nFor custom applications, the request metadata might look like this:\n\n```\nPOST https://gateway.ai.cloudflare.com/v1/$ACCOUNT_ID/$GATEWAY_ID/openai/chat/completions\n\nContent-Type: application/json\n\nAuthorization: Bearer $OPENAI_API_KEY\n\ncf-aig-metadata: { user_id: \\\"user-123\\\", session_id: \\\"session-456\\\", idp_group: \\\"engineering\\\", application: \\\"code-review\\\" }\n```\n\nThe request body contains the model and messages for the conversation.\n\nWith Access configured in front of AI Gateway, tools such as Claude Code, Codex, and OpenCode can inherit this identity context automatically. Cloudflare Access is available at no cost for teams with up to 50 users, making it an easy way to get started.\n\n## Get started with AI Gateway User Insights\n\nAI usage is changing quickly. Models change, teams develop new workflows, and the right choice for one group may be the wrong choice for another.\n\nUser Insights lets teams start making smarter choices by identifying where certain models may be overkill. They can then see which users and agents are driving that usage, understand the tasks behind it, and compare the cost of completing the work.\n\nLearn more with the AI Gateway User Insights [__documentation__](https://developers.cloudflare.com/ai-gateway/features/user-insights/) . Open AI Gateway in the [__Cloudflare dashboard__](https://dash.cloudflare.com/?to=/:account/ai/ai-gateway), and use what you learn to make more targeted model and routing decisions.", "url": "https://wpnews.pro/news/identify-ai-model-overuse-with-user-insights", "canonical_source": "https://blog.cloudflare.com/ai-model-overuse-user-insights/", "published_at": "2026-09-30 13:00:00+00:00", "updated_at": "2026-09-30 13:18:50.106730+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-agents", "ai-products", "mlops", "ai-tools"], "entities": ["Cloudflare", "User Insights", "AI Gateway", "Auto Router", "Potential Savings"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/identify-ai-model-overuse-with-user-insights", "markdown": "https://wpnews.pro/news/identify-ai-model-overuse-with-user-insights.md", "text": "https://wpnews.pro/news/identify-ai-model-overuse-with-user-insights.txt", "jsonld": "https://wpnews.pro/news/identify-ai-model-overuse-with-user-insights.jsonld"}}