DrivingBench, a benchmark for evaluating frontier AI models’ driving capabilities, gave GPT-6 Astra control of a Toyota Corolla to navigate a fixed cone course. The model completed the task with 100% progress along the course centerline while remaining within 4 meters of the centerline, finishing in 5 minutes and 22 seconds on its first attempt.

The benchmark measures performance across multiple metrics. Progress is defined as how far along the course centerline the attempt reached while staying within the 4-meter boundary; collisions preserve only the progress made before impact. Finish time is recorded from the first accepted steering or acceleration command to the end of the attempt’s final engagement.
GPT-6 Astra’s performance significantly outpaced other frontier models tested on the same course. Claude Fable 5.1 achieved a best progress of 45% across three attempts, while Grok 4.6 managed only 11% progress. GPT-5.6 Sol, an earlier model, reached just 6% progress.
The test represents a direct evaluation of large language models’ ability to process real-world sensory input and execute precise vehicle control. Rather than relying on simulation or indirect methods, DrivingBench gives models direct control of actual vehicle systems—steering, accelerator, and brakes—and evaluates their ability to follow a predefined path while maintaining safety boundaries.
The benchmark allows up to three attempts per model in a single continuous chat session, with each attempt tracked separately. Performance data includes token usage and associated costs, providing transparency into computational requirements for each attempt. The leaderboard presents comparative results across multiple frontier models, enabling objective measurement of progress in autonomous vehicle control capabilities among state-of-the-art AI systems.
This benchmark adds to growing efforts to systematically evaluate AI capabilities beyond traditional benchmarks, focusing on real-world task execution rather than standardized test sets.
Key facts
- GPT-6 Astra completed a real-world driving course with 100% progress on its first attempt, finishing in 5 minutes 22 seconds
- The test gave the AI model direct control of a Toyota Corolla’s steering, accelerator, and brakes on a fixed cone course
- Claude Fable 5.1 achieved 45% best progress, Grok 4.6 reached 11%, and GPT-5.6 Sol managed 6% on the same course
- DrivingBench measures progress as distance along the course centerline while staying within 4 meters of it
