cd /news/artificial-intelligence/epochs-automation-reports-find-front… · home › topics › artificial-intelligence › article
[ARTICLE · art-147731] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job

Epoch AI's new Automation Reports, introduced October 8, 2026, found that Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra tied for the top score at approximately 65% on 11 tasks drawn from Epoch AI's own internal research, with Grok 4.6 at 59%, Qwen 3.8 Max at 53%, Kimi K3 at 52%, and Gemini 3.8 Flash at 42%. The models performed well on defined subtasks such as coding and computational analysis but lagged on creative and judgment-based work, particularly style adherence and content targeting, leaving a spread of more than 20 points between the leaders and the bottom of the reported field. Epoch AI said the benchmark, which uses human graders scoring against the same internal quality rubrics Epoch uses for its own work, complements its Epoch Capabilities Index and follows the roughly October 7, 2026 rollout of InnovationEval, which tests how well AI can automate research itself.

by read3 min views1 publishedOct 8, 2026
Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job
Image: Cryptobriefing (auto-discovered)

Top models from Anthropic and OpenAI tie at approximately 65% on tasks drawn from Epoch AI's own research, but judgment and style remain weak spots

Epoch AI spends much of its time measuring how capable AI models are becoming. With its latest release, the organization turned the question inward: could these models actually do Epoch’s own work?

According to the new Epoch Automation Reports, not yet. Frontier models get close on clearly defined tasks, but they still fall short of fully automating the work, especially when assignments turn open-ended or call for judgment.

A job trial, not a pop quiz #

Epoch AI introduced the Automation Reports on October 8, 2026. Instead of relying on abstract puzzles, the benchmark pulls its tasks from the organization’s internal research.

The evaluation covers 11 distinct tasks spread across five categories. Human graders score each model’s output against the same internal quality rubrics Epoch uses for its own work.

The scoreboard #

Two models share the top spot. Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra each posted an average score of approximately 65% across the tasks.

Grok 4.6 followed at 59%. Qwen 3.8 Max came in at 53%, edging out Kimi K3 at 52%.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

Gemini 3.8 Flash rounded out the listed results at 42%. That leaves a spread of more than 20 points between the leaders and the bottom of the reported field.

The models performed well on defined subtasks such as coding and computational analysis. Where they faltered was in creative and judgment-based work, with recurring problems in style adherence and content targeting.

Every model in the evaluation lagged on outputs that depend on implicit conventions — the house style, the expected framing, the sense of what belongs in a report and what does not.

Where this fits in Epoch’s work #

The Automation Reports are part of Epoch AI’s broader benchmarking efforts. They complement the Epoch Capabilities Index, or ECI, with the aim of offering a more holistic view of model performance than existing frameworks provide.

Around October 7, 2026, Epoch also rolled out InnovationEval, which tests how well AI can automate research itself.

What this means #

For companies deciding where to deploy AI, the findings draw a useful line between structured and open-ended work. Tasks with a clear right answer, like code or computation, look increasingly within reach for the top models. There are caveats worth keeping in mind. The benchmark reflects 11 tasks drawn from a single organization’s research, so the results describe how well models handle Epoch’s work specifically, not every knowledge job. Human grading against internal rubrics also brings a degree of subjectivity that automated scoring avoids, making direct comparisons with other benchmarks harder.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @epoch ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/epochs-automation-re…] indexed:0 read:3min 2026-10-08 · —