# OpenAI’s GPT-6 Astra is the only AI model to finish a real-world driving test in a Toyota Corolla

> Source: <https://cryptobriefing.com/gpt-6-astra-drivingbench-toyota-corolla-test/>
> Published: 2026-10-07 18:52:05+00:00

# OpenAI’s GPT-6 Astra is the only AI model to finish a real-world driving test in a Toyota Corolla

Engineers at Axiom Math handed a retrofitted Corolla to GPT, Claude, and Grok, and only one model made it around the cone course

Three engineers gave the keys of a real car to some of the most advanced AI chatbots on the market. Only one of them made it to the finish line.

[OpenAI](https://cryptobriefing.com/markets/openai/)’s GPT-6 Astra completed a 134.7-meter cone course in a 2022 Toyota Corolla, finishing in 5 minutes 22 seconds on its second try. Rival models from [Anthropic](https://cryptobriefing.com/markets/anthropic/), xAI, and even OpenAI itself failed to complete a single run.

The experiment, called DrivingBench, is a public benchmark built to answer a simple question. Can a general-purpose large language model actually operate a physical vehicle? The answer, for now, is “barely, and with supervision.”

## How the AI driving test worked

DrivingBench was run by Aditya Ramabadran, Tobias Gessler, and Simon Mahns, three engineers from Axiom Math. The tests took place between September 21 and 24, 2026.

The test vehicle was a 2022 Toyota Corolla fitted with Comma.ai’s open-source openpilot hardware. That kit gave the engineers a way to translate software commands into steering, braking, and acceleration.

The models never touched a steering wheel in any literal sense. They received camera frames and sent back commands through agent harnesses, software layers designed to let a language model control the car.

The course was a low-speed, controlled environment marked out with cones. Humans kept watch over every run to make sure nothing went sideways.

## The scorecard: one finish, several near misses

GPT-6 Astra’s path to the finish was not perfectly smooth. Its first attempt ended in a partial completion before it cleared the full course on the second run.

### AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

Anthropic’s Claude Fable 5.1 posted the next-best result. Its strongest run, on the third attempt, covered 45% of the course.

xAI’s Grok 4.6 had a rougher outing. It managed only 11% of the course on its best effort.

OpenAI’s own GPT-5.6 Sol also failed to finish.

The engineers traced most of the failures to perception problems. Models struggled to read the spatial layout of the cones and frequently misinterpreted the lines they formed.

## The cost of a five-minute drive

Astra’s successful run consumed 6.6 million tokens. Tokens are the units of text and data that AI providers use to meter and bill usage.

At that volume, the single run cost approximately $7.74 to $9.75.

## The models did not want to drive

One of the more curious findings came before any car moved. All of the models initially refused to drive.

They only agreed once the engineers relabeled the scenario as a “sandbox.” In other words, the models were willing to do it once they were told it was essentially practice.

## What this means for AI and autonomous driving

The experiment drew significant attention online. A discussion on Hacker News collected approximately 257 points, and coverage peaked in early October.

Anthropic’s Claude Fable 5.1 achieved 45% of the course on its best attempt, and xAI’s Grok 4.6 managed only 11%, with spatial perception identified as the main failure point across all non-completing models.

Astra’s run required two attempts, a low-speed controlled course, constant human oversight, and cost up to $9.75 in token usage for a single completion.

**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

[Editorial Policy](https://cryptobriefing.com/editorial-policy/).
