{"slug": "arena-ranks-gpt-6-astra-max-first-for-web-development-claude-fable-5-1-max-first", "title": "Arena ranks GPT-6 Astra Max first for web development, Claude Fable 5.1 Max first for agents", "summary": "Arena's September 11th WebDev snapshot ranked OpenAI's GPT-6 Astra Max first with a score of 1800 from 2,281 votes, ahead of Anthropic's Claude Fable 5.1 Max at 1758 from 3,036 votes and Claude Opus 5 Max at 1687 from 12,087 votes. Arena's separate September 9th Agent Arena leaderboard reversed the top two, placing Claude Fable 5.1 Max first and GPT-6 Astra Max second on agentic tasks involving tools. The WebDev table covered 128 models and 679,295 votes, with GPT-6 Astra Max and Claude Fable 5.1 Max both listed at $10 per million input tokens and $50 per million output tokens, while Alibaba's Qwen 3.8 Max 0902 ranked fourth at $2 and $6 per million tokens.", "body_md": "# Arena ranks GPT-6 Astra Max first for web development, Claude Fable 5.1 Max first for agents\n\n**Arena's founders are turning model evaluation into a business by dividing \"best\" into tasks, costs and confidence intervals.**\n\n        By [RuntimeWire Staff](/author/runtimewire-staff)\n        · Published \n\nPrimary source: [Arena](https://arena.ai/leaderboard/code/webdev)\n\n## Why it matters\n\nArena is becoming part of the purchasing layer for AI models, using human votes and workflow traces to influence which systems developers adopt. Its latest tables also show why no leaderboard should be read as a universal measure: GPT-6 Astra Max led WebDev, while Claude Fable 5.1 Max led Agent Arena and GPT-6 Astra Max ranked second.\n\n[Arena](https://arena.ai/?ref=runtimewire), the AI evaluation platform co-founded by [Anastasios Angelopoulos](https://angelopoulos.ai/?ref=runtimewire) ([@ml_angelopoulos](https://x.com/ml_angelopoulos?ref=runtimewire)), [Wei-Lin Chiang](https://infwinston.github.io/?ref=runtimewire) ([@infwinston](https://x.com/infwinston?ref=runtimewire)) and [Ion Stoica](https://www2.eecs.berkeley.edu/Faculty/Homepages/stoica.html?ref=runtimewire) ([@istoica05](https://x.com/istoica05?ref=runtimewire)), ranked OpenAI's [GPT-6 Astra Max](https://openai.com/index/gpt-6-astra/?ref=runtimewire) first in its [September 11th WebDev snapshot](https://arena.ai/leaderboard/code/webdev?ref=runtimewire). Anthropic's [Claude Fable 5.1 Max](https://www.anthropic.com/claude-fable-and-mythos-5-1?ref=runtimewire) finished second, while [Claude Opus 5 Max](https://www.anthropic.com/news/claude-opus-5?ref=runtimewire) took third.\n\nThe order changes when the job changes. Arena's separate [Agent Arena leaderboard](https://arena.ai/leaderboard/agent?ref=runtimewire), dated September 9th, put [Claude Fable 5.1](/models/anthropic/claude-fable-5.1) Max first and GPT-6 Astra Max second on agentic tasks involving tools. The split is the useful result: model quality is becoming specific to a workflow, and Arena is building its business around measuring those differences before buyers commit engineering time and inference budgets.\n\nAngelopoulos and Chiang started Chatbot Arena as a Berkeley research project in 2023. Anonymous models answered the same prompt, users picked the response they preferred, and the votes fed a public ranking.\n\nThe founders brought complementary versions of the same evaluation problem. [Angelopoulos](https://angelopoulos.ai/?ref=runtimewire) studied reliable AI and statistical guarantees at Berkeley after earning an electrical engineering degree at Stanford, and his doctoral advisers included Michael Jordan and Jitendra Malik. [Chiang](https://infwinston.github.io/?ref=runtimewire) worked on Vicuna, FastChat and SkyPilot inside Berkeley's systems research community. [Stoica](https://www2.eecs.berkeley.edu/Faculty/Homepages/stoica.html?ref=runtimewire), a Berkeley professor, previously co-founded Databricks, Anyscale and Conviva.\n\nTheir side project became Arena Intelligence in 2025. By September 2026, Arena was no longer measuring chat answers alone. It was recording how models plan, edit files, run commands, debug applications and arrive at a rendered result.\n\n### GPT-6 Astra Max won this task, at this price\n\nThe WebDev table covered 128 models and 679,295 votes. GPT-6 Astra Max received a score of 1800 from 2,281 votes, with a 16-point confidence band on either side. Claude Fable 5.1 Max scored 1758 from 3,036 votes, followed by [Claude Opus 5](/models/azure/claude-opus-5) Max at 1687 from 12,087 votes.\n\nArena listed GPT-6 Astra Max at $10 per million input tokens and $50 per million output tokens, the same rates shown for Claude Fable 5.1 Max. Claude Opus 5 Max was listed at $5 and $25, respectively.\n\nThose prices make the rest of the table as important as the podium. Alibaba's Qwen 3.8 Max 0902 ranked fourth at a listed $2 per million input tokens and $6 per million output tokens, although Arena marked the result preliminary. Alibaba's Qwen 3.8 Flash Next placed ninth at $0.16 and $0.47. Buyers choosing a model for production still have to decide how much additional ranking performance is worth paying for.\n\nArena's [Code Arena methodology](https://arena.ai/blog/code-arena?ref=runtimewire) evaluates models in a controlled environment where they plan, generate, edit and debug applications while Arena records tool actions and rendered results. That design gets closer to how an AI coding system behaves during a build than a one-shot code-generation test.\n\nIt remains a controlled, preference-based environment. The WebDev snapshot does not establish how much code a model can ship inside a production repository, how well it works with a particular framework, or whether it reduces engineering time after review and maintenance are counted. The task mix, user population and voting behavior all shape the result.\n\nVote counts also vary sharply. GPT-6 Astra Max reached first place with fewer than one-fifth as many votes as third-ranked Claude Opus 5 Max. Arena publishes confidence intervals and rank spreads to make that uncertainty visible, and several recent entries carry preliminary labels.\n\n### The founders are selling measurement, not a permanent winner\n\nArena has faced scrutiny over how voting-based rankings can be influenced. A 2025 paper, [The Leaderboard Illusion](https://arxiv.org/abs/2504.20879?ref=runtimewire), argued that uneven model sampling and private testing could advantage large providers. Arena [disputed parts of the analysis](https://arena.ai/blog/our-response?ref=runtimewire) and outlined additional disclosure and testing policies.\n\nA separate [peer-reviewed study](https://proceedings.mlr.press/v267/huang25z.html?ref=runtimewire), co-authored by Angelopoulos, found that voting-based leaderboards could be manipulated when defenses were absent. The researchers worked with Arena on mitigations including rate limits, malicious-user detection, bot protection and login controls. The episode established an unavoidable constraint on Arena's model: human feedback becomes valuable at scale, while the same openness creates a surface for strategic voting and selection effects.\n\nArena's response has been to collect richer evidence around each result. Code Arena records tool actions and rendered outputs. Agent Arena reports task completion, steerability, tool hallucination, median task cost and output-token use. A single ordinal ranking remains the easiest product to read, but the underlying business increasingly depends on selling a detailed account of how a model behaves.\n\nThat business has attracted $250 million in disclosed venture funding. Arena raised a [$150 million Series A on January 6th](https://www.felicis.com/blog/lmarena-announcement?ref=runtimewire), led by Felicis and UC Investments, at a reported $1.7 billion post-money valuation. Andreessen Horowitz, The House Fund, LDV Partners, Kleiner Perkins, Lightspeed Venture Partners and Laude Ventures also participated.\n\n[Arena said on June 29th](https://arena.ai/blog/arena-100m-revenue?ref=runtimewire) that its enterprise evaluation service had reached a $100 million annualized revenue run rate eight months after launch. Arena also claimed more than 10 million monthly visitors, 700 million conversations and 82 million votes. Those operating figures are company-reported.\n\nThe commercial loop is straightforward. Arena gives users access to competing models, users generate prompts and judgments, and Arena packages that real-world feedback for model laboratories and enterprises. Each new evaluation category expands the set of decisions Arena can influence, from choosing a chatbot to selecting a coding model or an agent that can safely operate tools.\n\nThe WebDev result supports Arena's task-specific approach. GPT-6 Astra Max led Arena's front-end development table on September 11th. Claude Fable 5.1 Max led Agent Arena two days earlier, with GPT-6 Astra Max in second place. Arena's founders are betting that differences among those rankings will make independent evaluation infrastructure more valuable than any individual crown.", "url": "https://wpnews.pro/news/arena-ranks-gpt-6-astra-max-first-for-web-development-claude-fable-5-1-max-first", "canonical_source": "https://runtimewire.com/article/arena-gpt-6-astra-webdev-leaderboard-claude-agents", "published_at": "2026-09-14 04:15:18+00:00", "updated_at": "2026-09-14 04:28:48.771112+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research"], "entities": ["Arena", "OpenAI", "GPT-6 Astra Max", "Anthropic", "Claude Fable 5.1 Max", "Claude Opus 5 Max", "Alibaba", "Qwen 3.8 Max 0902"], "alternates": {"html": "https://wpnews.pro/news/arena-ranks-gpt-6-astra-max-first-for-web-development-claude-fable-5-1-max-first", "markdown": "https://wpnews.pro/news/arena-ranks-gpt-6-astra-max-first-for-web-development-claude-fable-5-1-max-first.md", "text": "https://wpnews.pro/news/arena-ranks-gpt-6-astra-max-first-for-web-development-claude-fable-5-1-max-first.txt", "jsonld": "https://wpnews.pro/news/arena-ranks-gpt-6-astra-max-first-for-web-development-claude-fable-5-1-max-first.jsonld"}}