cd /news/ai-agents/microsoft-releases-301000-copilot-ag… · home › topics › ai-agents › article
[ARTICLE · art-143559] src=runtimewire.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Microsoft releases 301,000 Copilot agent traces for researchers

Microsoft released a public slice of GitHub Copilot coding-agent telemetry on October 1st, giving researchers 301,026 sessions covering 9.3 million LLM calls and 8.7 million tool calls, according to Microsoft Azure Research engineer Haoran Qiu. The downloadable files in Microsoft's AzurePublicDataset repository cover June 1st through June 7th, 2026, and include 1,189,581 user turns, 37 anonymized model labels, 631.4 billion prompt tokens and 540.95 billion cached prompt tokens under a CC-BY license, while excluding prompts, model responses, source code, file paths, repository names and user or organization identifiers. The public slice is smaller than the paper's full production-scale study of 3.2 million users, 13 million sessions, 761 million LLM calls and 95 trillion tokens, and the metadata-only release cannot show whether Copilot's answers were correct or useful.

by read4 min views1 publishedOct 2, 2026
Microsoft releases 301,000 Copilot agent traces for researchers
Image: Runtimewire (auto-discovered)

The June sample records 9.3 million model calls and 8.7 million tool calls, while leaving prompts, code, outputs and user identities out of the public files.

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [X](https://x.com/haoran_qiu98/status/2105705601586352267)

Why it matters #

Coding agents turn one prompt into a sequence of model and tool work. Public traces let researchers test infrastructure assumptions against real usage, while the metadata-only release keeps code and conversations private.

Microsoft released a public slice of GitHub Copilot coding-agent telemetry on October 1st, giving researchers 301,026 sessions to examine across 9.3 million LLM calls and 8.7 million tool calls. Haoran Qiu, a Microsoft Azure Research engineer and co-author of the study, pointed to the data in a post on X; the downloadable files are available through Microsoft's AzurePublicDataset repository.

The release makes the mechanics of coding agents more measurable without publishing the contents of developers' work. It records timings, token counts, cache behavior, anonymized model labels and tool-call sequences. It excludes prompts, model responses, source code, file paths, repository names and user or organization identifiers. The telemetry supports analysis of resource use; it does not show whether Copilot's answers were correct or useful.

Qiu's career has centered on the systems that make AI workloads run efficiently. His Microsoft profile says he earned a computer science Ph.D. from the University of Illinois Urbana-Champaign and works on serving and resource management for generative AI. The paper's first author, Banruo Liu, conducted the work while interning at Microsoft Azure Research, according to the study posted to arXiv. Its author list also includes researchers from Azure Research and UIUC.

A public slice, not the whole study

The downloadable files cover seven days, June 1st through June 7th, 2026, and contain a uniformly sampled set of sessions from non-enterprise Copilot users. The dataset card reports 1,189,581 user turns, 37 anonymized model labels, 631.4 billion prompt tokens and 540.95 billion cached prompt tokens. The data is released under a CC-BY attribution license, and Microsoft includes a notebook intended to reproduce figures in the paper.

Those numbers describe the release, not the full population analyzed in the research. The paper's full study describes a larger production-scale study spanning 3.2 million users, 13 million sessions, 761 million LLM calls and 95 trillion tokens. The larger study's conclusions therefore should not be mistaken for statistics calculated only from the 301,026-session public slice. The researchers say the telemetry came from Copilot's coding agent in Visual Studio and VS Code and covered US regions across no more than three time zones.

That scope also limits what the release can establish. It is a view of one product's agent workload, for non-enterprise users, during a specific week. It is not a cross-product comparison of Copilot with Claude Code or Codex, and the anonymized labels prevent users from tying observed behavior to a named model. The sample is useful for studying workload structure; it cannot settle which agent performs best.

The unit of work is the agent loop

The paper's central finding is about infrastructure. A developer's request can trigger a sequence of model calls and tool actions, as an agent searches files, edits code, runs tests and responds to their results. In the full study, the researchers report an almost one-to-one relationship between LLM calls and tool invocations, with 87% of calls initiated by the agent rather than directly by a user.

That pattern complicates the assumptions behind systems built around isolated request-and-response exchanges. The paper reports that cache hit rates average about 90% within a turn, then fall to 55% across turn boundaries; model switches reduce them further. The researchers also found that tool failures in 9% of turns could trigger retries that raised compute use by as much as four times. Those are study findings, not measurements independently verified from the released sample alone.

For infrastructure teams, the stakes are practical: serving costs and latency depend on what happens between the user's prompt and the final answer, including tool execution, retries, cache reuse and s while a developer reads the result. The researchers propose that systems schedule around sessions and reclaim resources during idle periods, rather than treating every model call as an unrelated request. Their paper reports an idle-time predictor that captured 86% to 90% of total idle time in its evaluation. That result is a systems experiment, not a promise that every deployment can realize the same savings. The release gives infrastructure researchers production-derived traces and documents a workload generated by Microsoft's own coding product at production scale. The metadata-only design leaves code and conversations out of the public files. The files leave an important boundary intact: outsiders can inspect how agents use resources, but not what developers asked them to build or what the agents produced.

── more in #ai-agents 4 stories · sorted by recency
── more on @microsoft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/microsoft-releases-3…] indexed:0 read:4min 2026-10-02 · —