{"slug": "openais-gpt-6-astra-is-the-only-ai-model-to-finish-a-real-world-driving-test-in", "title": "OpenAI’s GPT-6 Astra is the only AI model to finish a real-world driving test in a Toyota Corolla", "summary": "OpenAI's GPT-6 Astra was the only model to complete Axiom Math's DrivingBench real-world driving test, finishing a 134.7-meter cone course in a 2022 Toyota Corolla in 5 minutes 22 seconds on its second attempt, the three Axiom Math engineers reported. Anthropic's Claude Fable 5.1 covered 45% of the course on its best run, xAI's Grok 4.6 managed 11%, and OpenAI's GPT-5.6 Sol failed to finish, with spatial perception of the cones cited as the main failure point. Astra's successful run consumed 6.6 million tokens at a cost of roughly $7.74 to $9.75, and all models initially refused to drive until the scenario was relabeled a \"sandbox.", "body_md": "# OpenAI’s GPT-6 Astra is the only AI model to finish a real-world driving test in a Toyota Corolla\n\nEngineers at Axiom Math handed a retrofitted Corolla to GPT, Claude, and Grok, and only one model made it around the cone course\n\nThree engineers gave the keys of a real car to some of the most advanced AI chatbots on the market. Only one of them made it to the finish line.\n\n[OpenAI](https://cryptobriefing.com/markets/openai/)’s GPT-6 Astra completed a 134.7-meter cone course in a 2022 Toyota Corolla, finishing in 5 minutes 22 seconds on its second try. Rival models from [Anthropic](https://cryptobriefing.com/markets/anthropic/), xAI, and even OpenAI itself failed to complete a single run.\n\nThe experiment, called DrivingBench, is a public benchmark built to answer a simple question. Can a general-purpose large language model actually operate a physical vehicle? The answer, for now, is “barely, and with supervision.”\n\n## How the AI driving test worked\n\nDrivingBench was run by Aditya Ramabadran, Tobias Gessler, and Simon Mahns, three engineers from Axiom Math. The tests took place between September 21 and 24, 2026.\n\nThe test vehicle was a 2022 Toyota Corolla fitted with Comma.ai’s open-source openpilot hardware. That kit gave the engineers a way to translate software commands into steering, braking, and acceleration.\n\nThe models never touched a steering wheel in any literal sense. They received camera frames and sent back commands through agent harnesses, software layers designed to let a language model control the car.\n\nThe course was a low-speed, controlled environment marked out with cones. Humans kept watch over every run to make sure nothing went sideways.\n\n## The scorecard: one finish, several near misses\n\nGPT-6 Astra’s path to the finish was not perfectly smooth. Its first attempt ended in a partial completion before it cleared the full course on the second run.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\nAnthropic’s Claude Fable 5.1 posted the next-best result. Its strongest run, on the third attempt, covered 45% of the course.\n\nxAI’s Grok 4.6 had a rougher outing. It managed only 11% of the course on its best effort.\n\nOpenAI’s own GPT-5.6 Sol also failed to finish.\n\nThe engineers traced most of the failures to perception problems. Models struggled to read the spatial layout of the cones and frequently misinterpreted the lines they formed.\n\n## The cost of a five-minute drive\n\nAstra’s successful run consumed 6.6 million tokens. Tokens are the units of text and data that AI providers use to meter and bill usage.\n\nAt that volume, the single run cost approximately $7.74 to $9.75.\n\n## The models did not want to drive\n\nOne of the more curious findings came before any car moved. All of the models initially refused to drive.\n\nThey only agreed once the engineers relabeled the scenario as a “sandbox.” In other words, the models were willing to do it once they were told it was essentially practice.\n\n## What this means for AI and autonomous driving\n\nThe experiment drew significant attention online. A discussion on Hacker News collected approximately 257 points, and coverage peaked in early October.\n\nAnthropic’s Claude Fable 5.1 achieved 45% of the course on its best attempt, and xAI’s Grok 4.6 managed only 11%, with spatial perception identified as the main failure point across all non-completing models.\n\nAstra’s run required two attempts, a low-speed controlled course, constant human oversight, and cost up to $9.75 in token usage for a single completion.\n\n**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/openais-gpt-6-astra-is-the-only-ai-model-to-finish-a-real-world-driving-test-in", "canonical_source": "https://cryptobriefing.com/gpt-6-astra-drivingbench-toyota-corolla-test/", "published_at": "2026-10-07 18:52:05+00:00", "updated_at": "2026-10-07 19:17:04.959764+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "autonomous-vehicles", "ai-research"], "entities": ["OpenAI", "GPT-6 Astra", "Axiom Math", "Anthropic", "Claude Fable 5.1", "xAI", "Grok 4.6", "Toyota Corolla"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openais-gpt-6-astra-is-the-only-ai-model-to-finish-a-real-world-driving-test-in", "markdown": "https://wpnews.pro/news/openais-gpt-6-astra-is-the-only-ai-model-to-finish-a-real-world-driving-test-in.md", "text": "https://wpnews.pro/news/openais-gpt-6-astra-is-the-only-ai-model-to-finish-a-real-world-driving-test-in.txt", "jsonld": "https://wpnews.pro/news/openais-gpt-6-astra-is-the-only-ai-model-to-finish-a-real-world-driving-test-in.jsonld"}}