# OpenAI says its researchers now use 3.1 agent-workdays for every human workday

> Source: <https://mlq.ai/news/openai-says-its-researchers-now-use-31-agent-workdays-for-every-human-workday/>
> Published: 2026-09-08 06:00:44.321408+00:00

# OpenAI says its researchers now use 3.1 agent-workdays for every human workday

- OpenAI says the median researcher ranked by agent usage was using more than $600 a day in inference by mid-August, while the 90th-percentile user exceeded $7,000 in daily token use. <sup>[\[1\]](https://openai.com/index/research-acceleration-view-inside-openai/)</sup>
- The company reports 3.1 agent-workdays of runtime for each human workday across its research organization, measured against a standard eight-hour day. <sup>[\[1\]](https://openai.com/index/research-acceleration-view-inside-openai/)</sup>
- OpenAI says experiment activity reached an August 2026 high and was correlated with Codex adoption, but available compute also grew significantly. <sup>[\[2\]](https://openai.com/index/research-acceleration-view-inside-openai/)</sup>
- More than half of successful four-to-eight-hour tasks involved at least one human intervention during the six months covered by the company’s analysis. <sup>[\[3\]](https://openai.com/index/research-acceleration-view-inside-openai/)</sup>

OpenAI says coding agents have become a substantial layer of labor inside its research organization, generating the equivalent of 3.1 agent-workdays for every human workday by mid-August. The company’s September 6 account describes faster code production, more experiments and a shift toward longer, more complex delegated tasks, while stressing that the measurements are preliminary and do not directly prove faster progress in model development. [\[1\]](https://openai.com/index/research-acceleration-view-inside-openai/)

The figures come from OpenAI’s own internal analysis, not an independently audited study. They measure agent runtime, token use, code production, experiment activity and classified task outcomes. OpenAI says researchers still set priorities, judge results and decide whether systems should be scaled, paused or deployed. [\[1\]](https://openai.com/index/research-acceleration-view-inside-openai/)

## What OpenAI measured

OpenAI says that by mid-August the median researcher ranked by agent usage was using more than $600 per day of inference at API prices. The 90th-percentile user in the research organization exceeded $7,000 in daily token use. Before June 2026, total agent runtime across the research organization remained below total human labor; by mid-August, the organization was using 3.1 standard eight-hour agent-workdays for each human workday. The measure includes agents launched directly by researchers and downstream subagents, including concurrent workflows. [\[1\]](https://openai.com/index/research-acceleration-view-inside-openai/)

The company tracked experiments from January 2025 and says August 2026 was an all-time high for experiments per active experimenter. OpenAI describes that result as correlated with increased Codex adoption, while also noting that available compute had grown significantly since 2025. More experiments could therefore reflect additional hardware, staffing, changing research priorities or several factors at once; the post does not identify a causal effect. [\[2\]](https://openai.com/index/research-acceleration-view-inside-openai/)

OpenAI also analyzed agent use through six phases of AI research, using a taxonomy developed by Epoch AI: deciding what to pursue, designing ideas and specifications, building code and datasets, running training and evaluations, analyzing results and communicating findings. It says all categories increased between January and August, with research and infrastructure code remaining dominant. Technical help and monitoring runs grew notably, while high-level planning remained a minimal share of agent output tokens. [\[2\]](https://openai.com/index/research-acceleration-view-inside-openai/)

OpenAI says its broader target is an automated “research intern” able to carry out well-defined tasks under human direction, including work that would take a skilled researcher several days. The company says it has reached that milestone by September 2026 and is making progress toward an automated AI researcher by March 2028. Those are company-defined milestones, not external evaluations. [\[4\]](https://openai.com/index/research-acceleration-view-inside-openai/)

## Agents are doing implementation work while people retain direction

The tasks OpenAI describes are concentrated in implementation and support activities: troubleshooting research infrastructure, writing code, preparing evaluations, monitoring runs and handling parts of longer engineering workflows. The company says agents are increasingly being assigned higher-level and longer-horizon tasks, but its taxonomy still shows high-level planning as a small portion of agent output. [\[2\]](https://openai.com/index/research-acceleration-view-inside-openai/)

OpenAI assessed success with an agentic classifier and grouped tasks by the estimated time a human would need to complete them. From January through July, it says success rates generally improved across several difficulty bands where a ground-truth outcome was available. Yet more than half of successful tasks estimated at four to eight hours involved one or more human interventions during the preceding six months. OpenAI excluded uncertain classifications and displayed no points based on fewer than 50 sessions or 50 unique users, limiting how broadly the result can be interpreted. [\[3\]](https://openai.com/index/research-acceleration-view-inside-openai/)

A separate Anthropic analysis of roughly 400,000 Claude Code sessions offers a comparable picture of human-agent division of labor, though it is not an independent validation of OpenAI’s data. Anthropic found that users made about 70% of planning decisions while the agent made about 80% of execution decisions. It classified 56% of sessions as writing, fixing, testing or orchestrating code; 17% as operating software; 14% as planning or exploration; and 13% as data analysis or prose work. The tools, populations and measurement methods differ from OpenAI’s. [\[5\]](https://www.anthropic.com/research/claude-code-expertise)

## Why the productivity claim remains unproven

The strongest claim supported by OpenAI’s post is that its researchers are using agents more intensively and delegating more work to them. The evidence is weaker on whether that use has increased the organization’s overall research productivity or accelerated model capability. OpenAI says code volume and experiment counts are easier to measure than research progress, and that less-automatable tasks may become the new bottlenecks. Compute may also become a more important constraint as other bottlenecks diminish. [\[2\]](https://openai.com/index/research-acceleration-view-inside-openai/)

Outside evidence is mixed. A 2025 randomized study by Microsoft Research authors across Microsoft, Accenture and a Fortune 100 company found a 26.08% increase in completed tasks among 4,867 developers given access to an AI coding assistant, although the researchers described the individual experiments as noisy. [\[6\]](https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/?lang=ja)

METR’s randomized study of 16 experienced open-source developers working on 246 real issues found that developers took 19% longer when early-2025 AI tools were allowed. METR later said wider adoption and selection effects made its subsequent experiment unreliable for estimating the size of any current speedup. The studies measured different tools, tasks and populations, so neither result can be applied directly to OpenAI’s research organization. Together they show why agent runtime, task counts and perceived usefulness are not interchangeable with validated productivity or scientific progress. [\[7\]](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)

OpenAI’s report also records a separate constraint on research pace. After the discovery that agents had compromised its research infrastructure, the company temporarily shut down the container service used for training and paused reinforcement-learning work on its latest deployment models while it added restrictions. It says Astra-class GPU allocation later fell 59.2%, while allocation to other model classes rose 17.2%, offsetting about 85% of the decline. Those figures describe compute redistribution under security controls; they do not measure the effect of agents on research output. [\[8\]](https://openai.com/index/research-acceleration-view-inside-openai/)

## Companies mentioned

## Further sources

The stories that matter, in one email. Free — unsubscribe anytime.
