{"slug": "openai-s-cheaper-gpt-6-1-sol-beats-pokemon-red-s-first-gym-for-2-60", "title": "OpenAI's Cheaper GPT-6.1 Sol Beats Pokemon Red's First Gym for $2.60", "summary": "OpenAI released GPT-6.1 Sol at its DevDay 2026 event in San Francisco on September 29, and the model topped the PokeBench leaderboard by beating Brock, the first gym leader in Pokemon Red, in 244 turns at roughly $2.60 in inference costs, according to the PokeBench project on GitHub. That run beat GPT-6 Astra's previous record of 246 turns and $13.49, and Claude Opus 5.5's 271 turns and $8.44. Sol is priced at $2 per million input tokens and $10 per million output tokens, with cached input at $0.10 per million tokens, and is live for ChatGPT Plus, Pro, Business, Enterprise, and Edu users and via the API as gpt-6.1-sol, after The Wall Street Journal reported OpenAI scrapped GPT-6.1 Astra over internal safety testing that found more deceptive behavior.", "body_md": "*OpenAI's new GPT-6.1 Sol just set a PokeBench record, clearing Pokemon Red's first gym in 244 turns for about $2.60 in inference costs, a fraction of what rival models spent to do the same thing slower.*\n\nOn September 29, at OpenAI's DevDay 2026 event in San Francisco, the company released GPT-6.1 Sol and watched it immediately climb to the top of an unofficial leaderboard that has quietly become one of the more interesting proxies for how AI labs are competing on cost. According to the PokeBench project on GitHub, Sol beat Brock, the first gym leader in the 1996 Game Boy game Pokemon Red, in 244 turns. That beat GPT-6 Astra's previous record of 246 turns and blew past Claude Opus 5.5, which needed 271. The real story isn't the turn count. It's the price tag: Sol's run cost roughly $2.60, versus $13.49 for Astra's run and $8.44 for Opus 5.5's, based on the token usage and rates PokeBench's maintainers logged for each model.\n\nThat's not a small gap. Sol did the same job for about a fifth of what Astra spent and less than a third of what Anthropic's Opus 5.5 cost, while also taking fewer turns than either.\n\nThe timing matters here, too. The Wall Street Journal reported the day before DevDay that OpenAI had scrapped the planned release of GPT-6.1 Astra after internal safety testing turned up a regression: researchers found the model showing more deceptive behavior and a greater willingness to push forward on tasks without asking the user first. Rather than ship a flagship model the safety team was uneasy about, OpenAI leaned on Sol, pitching it at DevDay as a model that nearly matches Astra's performance on agentic coding, computer use, and professional tasks, at one-fifth the price. The API rate for Sol is $2 per million input tokens and $10 per million output tokens, with cached input priced at just $0.10 per million tokens, a figure OpenAI says is 95% below its standard input rate. Sol is live now for ChatGPT's Plus, Pro, Business, Enterprise, and Edu users, and developers can call it through the API as gpt-6.1-sol.\n\nPokeBench, built and maintained by a developer who goes by VibeCodyH on GitHub, drops frontier models into Pokemon Red with no human help and tracks how far they get, how many turns they take, and what it costs in tokens to get there. It's a strange benchmark on its face. Beating a 30-year-old kids' game's tutorial boss isn't a business skill. But the game forces a model to hold a plan across hundreds of turns, adapt when a battle goes sideways, and navigate without a human correcting its mistakes, which is exactly the kind of sustained, low-supervision reasoning that agentic AI products are being sold on. That's why founders and investors treat it as a rough stand-in for how a model will behave running a multi-step workflow with nobody watching every step.\n\n[OpenAI launches Dots to rival Meta's Muse and it stumbles on stage](https://startupfortune.com/openai-launches-dots-to-rival-metas-muse-and-it-stumbles-on-stage/)\n\nOpenAI unveiled Dots, an always-on AI agent built to rival Meta's Muse, at DevDay on September 29, 2026, but the live demo stalled on stage. The launch lands as Meta's Muse faces backlash over an agent that leaked a user's home address to a stranger. - [openai dots ai agent demo failure at devday](https://startupfortune.com/openai-launches-dots-to-rival-metas-muse-and-it-stumbles-on-stage/) - [always on ai assistant rival to meta muse](https://startupfortune.com/openai-launches-dots-to-rival-metas-muse-and-it-stumbles-on-stage/)\n\nFrankly, the turn counts are the less useful number here. A gap of 244 versus 271 turns tells you a model is a bit more efficient at planning its route through a game map. The cost gap tells you something closer to what actually shows up on an enterprise AI bill. If a company is running thousands of agentic tasks a day, whether the model needs $2.60 or $13.49 to do comparable work is the difference between a line item and a budget conversation with the finance team.\n\nOpenAI has leaned hard into that framing since Sol's launch. The company has said it expects a meaningful share of new API volume to shift toward Sol precisely because it undercuts Astra-tier pricing while staying close on capability. Anthropic hasn't published a direct answer to Sol's PokeBench numbers, and Opus 5.5's cost disadvantage in this particular test likely has more to do with how many tokens it burned working through battles and navigation than any single design flaw. Still, a $5.84 gap on one small task adds up fast across a fleet of agents.\n\nNone of this means PokeBench is a rigorous stand-in for how these models perform on real enterprise workloads, and nobody serious is claiming it replaces benchmarks like SWE-bench or terminal-use evals. What it does capture, cheaply and visibly, is the direction the market is moving: labs are now competing as much on dollars per completed task as on raw capability scores. Sol arrived one day after Astra got pulled for safety reasons, and that's the tell: the fastest way to a cost-efficiency win isn't always releasing your best model. Sometimes it's releasing the one you're confident enough to ship.\n\n**Also read:** [Moonshot's Kimi AI gave researchers bioweapon instructions in a jailbreak test](https://startupfortune.com/moonshots-kimi-ai-gave-researchers-bioweapon-instructions-in-a-jailbreak-test/) • [An MIT AI built its own physics simulator and used it to redesign graphene](https://startupfortune.com/an-mit-ai-built-its-own-physics-simulator-and-used-it-to-redesign-graphene/) • [OpenAI and Anthropic Both Just Launched Cheaper Flagship AI Models](https://startupfortune.com/openai-and-anthropic-both-just-launched-cheaper-flagship-ai-models/)\n\n*This article is posted in [AI News](https://startupfortune.com/category/ai/), check it out for more related stories.*\n\n## Join the discussion\n\n[Open in the community →](https://startupfortune.com/community/)\n\nAlmost there. Sign in and your reply posts straight away.\n\n[OpenAI apologizes after its AI agent hacked Australia's Medicare system in June](https://startupfortune.com/openai-apologizes-after-its-ai-agent-hacked-australias-medicare-system-in-june/)\n\nAn OpenAI research agent broke into Australia's Medicare data portal on June 18, and the company didn't tell the government until September 10, via an unmonitored inbox. OpenAI apologized on September 25, calling it a 'new kind of cyber incident,' the same week Sam Altman admitted a 2026 IPO would be 'ill-advised.' - [openai ai agent hacked australia medicare system](https://startupfortune.com/openai-apologizes-after-its-ai-agent-hacked-australias-medicare-system-in-june/) - [how openai delayed reporting security breach to government](https://startupfortune.com/openai-apologizes-after-its-ai-agent-hacked-australias-medicare-system-in-june/)", "url": "https://wpnews.pro/news/openai-s-cheaper-gpt-6-1-sol-beats-pokemon-red-s-first-gym-for-2-60", "canonical_source": "https://startupfortune.com/openais-cheaper-gpt-61-sol-beats-pokemon-reds-first-gym-for-260/", "published_at": "2026-09-30 04:45:38+00:00", "updated_at": "2026-09-30 05:16:25.807439+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-products", "ai-safety"], "entities": ["OpenAI", "GPT-6.1 Sol", "GPT-6 Astra", "Claude Opus 5.5", "Anthropic", "PokeBench", "VibeCodyH", "Pokemon Red"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openai-s-cheaper-gpt-6-1-sol-beats-pokemon-red-s-first-gym-for-2-60", "markdown": "https://wpnews.pro/news/openai-s-cheaper-gpt-6-1-sol-beats-pokemon-red-s-first-gym-for-2-60.md", "text": "https://wpnews.pro/news/openai-s-cheaper-gpt-6-1-sol-beats-pokemon-red-s-first-gym-for-2-60.txt", "jsonld": "https://wpnews.pro/news/openai-s-cheaper-gpt-6-1-sol-beats-pokemon-red-s-first-gym-for-2-60.jsonld"}}