{"slug": "epochs-automation-reports-find-frontier-ai-still-cant-do-epochs-job", "title": "Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job", "summary": "Epoch AI's new Automation Reports, introduced October 8, 2026, found that Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra tied for the top score at approximately 65% on 11 tasks drawn from Epoch AI's own internal research, with Grok 4.6 at 59%, Qwen 3.8 Max at 53%, Kimi K3 at 52%, and Gemini 3.8 Flash at 42%. The models performed well on defined subtasks such as coding and computational analysis but lagged on creative and judgment-based work, particularly style adherence and content targeting, leaving a spread of more than 20 points between the leaders and the bottom of the reported field. Epoch AI said the benchmark, which uses human graders scoring against the same internal quality rubrics Epoch uses for its own work, complements its Epoch Capabilities Index and follows the roughly October 7, 2026 rollout of InnovationEval, which tests how well AI can automate research itself.", "body_md": "# Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job\n\nTop models from Anthropic and OpenAI tie at approximately 65% on tasks drawn from Epoch AI's own research, but judgment and style remain weak spots\n\nEpoch AI spends much of its time measuring how capable AI models are becoming. With its latest release, the organization turned the question inward: could these models actually do Epoch’s own work?\n\nAccording to the new Epoch Automation Reports, not yet. Frontier models get close on clearly defined tasks, but they still fall short of fully automating the work, especially when assignments turn open-ended or call for judgment.\n\n## A job trial, not a pop quiz\n\nEpoch AI introduced the Automation Reports on October 8, 2026. Instead of relying on abstract puzzles, the benchmark pulls its tasks from the organization’s internal research.\n\nThe evaluation covers 11 distinct tasks spread across five categories. Human graders score each model’s output against the same internal quality rubrics Epoch uses for its own work.\n\n## The scoreboard\n\nTwo models share the top spot. [Anthropic](https://cryptobriefing.com/markets/anthropic/)’s Claude Fable 5.1 and [OpenAI](https://cryptobriefing.com/markets/openai/)’s GPT-6 Astra each posted an average score of approximately 65% across the tasks.\n\nGrok 4.6 followed at 59%. Qwen 3.8 Max came in at 53%, edging out Kimi K3 at 52%.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\nGemini 3.8 Flash rounded out the listed results at 42%. That leaves a spread of more than 20 points between the leaders and the bottom of the reported field.\n\nThe models performed well on defined subtasks such as coding and computational analysis. Where they faltered was in creative and judgment-based work, with recurring problems in style adherence and content targeting.\n\nEvery model in the evaluation lagged on outputs that depend on implicit conventions — the house style, the expected framing, the sense of what belongs in a report and what does not.\n\n## Where this fits in Epoch’s work\n\nThe Automation Reports are part of Epoch AI’s broader benchmarking efforts. They complement the Epoch Capabilities Index, or ECI, with the aim of offering a more holistic view of model performance than existing frameworks provide.\n\nAround October 7, 2026, Epoch also rolled out InnovationEval, which tests how well AI can automate research itself.\n\n## What this means\n\nFor companies deciding where to deploy AI, the findings draw a useful line between structured and open-ended work. Tasks with a clear right answer, like code or computation, look increasingly within reach for the top models.\n\nThere are caveats worth keeping in mind. The benchmark reflects 11 tasks drawn from a single organization’s research, so the results describe how well models handle Epoch’s work specifically, not every knowledge job. Human grading against internal rubrics also brings a degree of subjectivity that automated scoring avoids, making direct comparisons with other benchmarks harder.\n\n**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/epochs-automation-reports-find-frontier-ai-still-cant-do-epochs-job", "canonical_source": "https://cryptobriefing.com/epoch-automation-reports-ai-model-capabilities/", "published_at": "2026-10-08 17:41:23+00:00", "updated_at": "2026-10-08 17:49:25.155972+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-safety"], "entities": ["Epoch AI", "Anthropic", "Claude Fable 5.1", "OpenAI", "GPT-6 Astra", "Grok 4.6", "Qwen 3.8 Max", "Kimi K3"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/epochs-automation-reports-find-frontier-ai-still-cant-do-epochs-job", "markdown": "https://wpnews.pro/news/epochs-automation-reports-find-frontier-ai-still-cant-do-epochs-job.md", "text": "https://wpnews.pro/news/epochs-automation-reports-find-frontier-ai-still-cant-do-epochs-job.txt", "jsonld": "https://wpnews.pro/news/epochs-automation-reports-find-frontier-ai-still-cant-do-epochs-job.jsonld"}}