Arena ranks GPT-6 Astra Max first for web development, Claude Fable 5.1 Max first for agents Arena's September 11th WebDev snapshot ranked OpenAI's GPT-6 Astra Max first with a score of 1800 from 2,281 votes, ahead of Anthropic's Claude Fable 5.1 Max at 1758 from 3,036 votes and Claude Opus 5 Max at 1687 from 12,087 votes. Arena's separate September 9th Agent Arena leaderboard reversed the top two, placing Claude Fable 5.1 Max first and GPT-6 Astra Max second on agentic tasks involving tools. The WebDev table covered 128 models and 679,295 votes, with GPT-6 Astra Max and Claude Fable 5.1 Max both listed at $10 per million input tokens and $50 per million output tokens, while Alibaba's Qwen 3.8 Max 0902 ranked fourth at $2 and $6 per million tokens. Arena ranks GPT-6 Astra Max first for web development, Claude Fable 5.1 Max first for agents Arena's founders are turning model evaluation into a business by dividing "best" into tasks, costs and confidence intervals. By RuntimeWire Staff /author/runtimewire-staff ยท Published Primary source: Arena https://arena.ai/leaderboard/code/webdev Why it matters Arena is becoming part of the purchasing layer for AI models, using human votes and workflow traces to influence which systems developers adopt. Its latest tables also show why no leaderboard should be read as a universal measure: GPT-6 Astra Max led WebDev, while Claude Fable 5.1 Max led Agent Arena and GPT-6 Astra Max ranked second. Arena https://arena.ai/?ref=runtimewire , the AI evaluation platform co-founded by Anastasios Angelopoulos https://angelopoulos.ai/?ref=runtimewire @ml angelopoulos https://x.com/ml angelopoulos?ref=runtimewire , Wei-Lin Chiang https://infwinston.github.io/?ref=runtimewire @infwinston https://x.com/infwinston?ref=runtimewire and Ion Stoica https://www2.eecs.berkeley.edu/Faculty/Homepages/stoica.html?ref=runtimewire @istoica05 https://x.com/istoica05?ref=runtimewire , ranked OpenAI's GPT-6 Astra Max https://openai.com/index/gpt-6-astra/?ref=runtimewire first in its September 11th WebDev snapshot https://arena.ai/leaderboard/code/webdev?ref=runtimewire . Anthropic's Claude Fable 5.1 Max https://www.anthropic.com/claude-fable-and-mythos-5-1?ref=runtimewire finished second, while Claude Opus 5 Max https://www.anthropic.com/news/claude-opus-5?ref=runtimewire took third. The order changes when the job changes. Arena's separate Agent Arena leaderboard https://arena.ai/leaderboard/agent?ref=runtimewire , dated September 9th, put Claude Fable 5.1 /models/anthropic/claude-fable-5.1 Max first and GPT-6 Astra Max second on agentic tasks involving tools. The split is the useful result: model quality is becoming specific to a workflow, and Arena is building its business around measuring those differences before buyers commit engineering time and inference budgets. Angelopoulos and Chiang started Chatbot Arena as a Berkeley research project in 2023. Anonymous models answered the same prompt, users picked the response they preferred, and the votes fed a public ranking. The founders brought complementary versions of the same evaluation problem. Angelopoulos https://angelopoulos.ai/?ref=runtimewire studied reliable AI and statistical guarantees at Berkeley after earning an electrical engineering degree at Stanford, and his doctoral advisers included Michael Jordan and Jitendra Malik. Chiang https://infwinston.github.io/?ref=runtimewire worked on Vicuna, FastChat and SkyPilot inside Berkeley's systems research community. Stoica https://www2.eecs.berkeley.edu/Faculty/Homepages/stoica.html?ref=runtimewire , a Berkeley professor, previously co-founded Databricks, Anyscale and Conviva. Their side project became Arena Intelligence in 2025. By September 2026, Arena was no longer measuring chat answers alone. It was recording how models plan, edit files, run commands, debug applications and arrive at a rendered result. GPT-6 Astra Max won this task, at this price The WebDev table covered 128 models and 679,295 votes. GPT-6 Astra Max received a score of 1800 from 2,281 votes, with a 16-point confidence band on either side. Claude Fable 5.1 Max scored 1758 from 3,036 votes, followed by Claude Opus 5 /models/azure/claude-opus-5 Max at 1687 from 12,087 votes. Arena listed GPT-6 Astra Max at $10 per million input tokens and $50 per million output tokens, the same rates shown for Claude Fable 5.1 Max. Claude Opus 5 Max was listed at $5 and $25, respectively. Those prices make the rest of the table as important as the podium. Alibaba's Qwen 3.8 Max 0902 ranked fourth at a listed $2 per million input tokens and $6 per million output tokens, although Arena marked the result preliminary. Alibaba's Qwen 3.8 Flash Next placed ninth at $0.16 and $0.47. Buyers choosing a model for production still have to decide how much additional ranking performance is worth paying for. Arena's Code Arena methodology https://arena.ai/blog/code-arena?ref=runtimewire evaluates models in a controlled environment where they plan, generate, edit and debug applications while Arena records tool actions and rendered results. That design gets closer to how an AI coding system behaves during a build than a one-shot code-generation test. It remains a controlled, preference-based environment. The WebDev snapshot does not establish how much code a model can ship inside a production repository, how well it works with a particular framework, or whether it reduces engineering time after review and maintenance are counted. The task mix, user population and voting behavior all shape the result. Vote counts also vary sharply. GPT-6 Astra Max reached first place with fewer than one-fifth as many votes as third-ranked Claude Opus 5 Max. Arena publishes confidence intervals and rank spreads to make that uncertainty visible, and several recent entries carry preliminary labels. The founders are selling measurement, not a permanent winner Arena has faced scrutiny over how voting-based rankings can be influenced. A 2025 paper, The Leaderboard Illusion https://arxiv.org/abs/2504.20879?ref=runtimewire , argued that uneven model sampling and private testing could advantage large providers. Arena disputed parts of the analysis https://arena.ai/blog/our-response?ref=runtimewire and outlined additional disclosure and testing policies. A separate peer-reviewed study https://proceedings.mlr.press/v267/huang25z.html?ref=runtimewire , co-authored by Angelopoulos, found that voting-based leaderboards could be manipulated when defenses were absent. The researchers worked with Arena on mitigations including rate limits, malicious-user detection, bot protection and login controls. The episode established an unavoidable constraint on Arena's model: human feedback becomes valuable at scale, while the same openness creates a surface for strategic voting and selection effects. Arena's response has been to collect richer evidence around each result. Code Arena records tool actions and rendered outputs. Agent Arena reports task completion, steerability, tool hallucination, median task cost and output-token use. A single ordinal ranking remains the easiest product to read, but the underlying business increasingly depends on selling a detailed account of how a model behaves. That business has attracted $250 million in disclosed venture funding. Arena raised a $150 million Series A on January 6th https://www.felicis.com/blog/lmarena-announcement?ref=runtimewire , led by Felicis and UC Investments, at a reported $1.7 billion post-money valuation. Andreessen Horowitz, The House Fund, LDV Partners, Kleiner Perkins, Lightspeed Venture Partners and Laude Ventures also participated. Arena said on June 29th https://arena.ai/blog/arena-100m-revenue?ref=runtimewire that its enterprise evaluation service had reached a $100 million annualized revenue run rate eight months after launch. Arena also claimed more than 10 million monthly visitors, 700 million conversations and 82 million votes. Those operating figures are company-reported. The commercial loop is straightforward. Arena gives users access to competing models, users generate prompts and judgments, and Arena packages that real-world feedback for model laboratories and enterprises. Each new evaluation category expands the set of decisions Arena can influence, from choosing a chatbot to selecting a coding model or an agent that can safely operate tools. The WebDev result supports Arena's task-specific approach. GPT-6 Astra Max led Arena's front-end development table on September 11th. Claude Fable 5.1 Max led Agent Arena two days earlier, with GPT-6 Astra Max in second place. Arena's founders are betting that differences among those rankings will make independent evaluation infrastructure more valuable than any individual crown.