cd /news/ai-safety/we-built-drone-bench · home topics ai-safety article
[ARTICLE · art-127840] src=twitter.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

We Built Drone-Bench

A team released Drone-Bench, a benchmark measuring how well AI models navigate drones and identify targets, reporting that the frontier model Astra now beats the human baseline on every task, whereas frontier models struggled and crashed just six months ago. The benchmark's creators said no lab has access to it and they run all evaluations themselves, and that its purpose is to inform the public and lawmakers about how fast AI autonomy in the physical world is advancing. The release drew viral criticism on Twitter, with the most common reaction being "Why would you build the scary thing which we were warned not to build?

read2 min views1 publishedSep 12, 2026
We Built Drone-Bench
Image: source

We recently released Drone-Bench, measuring how effective AI is at navigating drones and identifying targets.

Many people had strong reactions to our announcement. We think this is good; we wanted to start the discussion of where we do/don’t want AI in society. That frontier AIs are now capable of navigating drones, and are improving rapidly at this capability, is concerning.

After it went viral on twitter, the most common sentiment was: “Why would you build the scary thing which we were warned not to build?”

First, our benchmark does not help AI navigate drones. It is a measure of how good AI models already are at doing so. Whether AI can do it or not, does not depend on the existence of our benchmark.

Once AI can competently navigate drones, we believe it is important for the general public to know, not just leaders of tech companies. Our benchmark shows clear improvement between every model release; it is only a matter of time before frontier AI has superhuman drone navigation capabilities. Just six months ago, the frontier models struggled and crashed; Astra now beats our human baseline on every task in the benchmark.

Our hope for the benchmark is that it will help convince people to take this seriously as soon as possible, before we get to the point of superhuman AI drones. How these capabilities should be allowed to be used, if at all, is for the general public and lawmakers to decide, not just the AI labs. Accurate public information of current model capabilities is essential to allow people outside the labs to make informed decisions.

Finally, we are not letting AI companies hill-climb this benchmark. No lab has access to it; we run all evaluations ourselves. Its only purpose is to tell the world how fast AI autonomy in the physical world is advancing.

── more in #ai-safety 4 stories · sorted by recency
── more on @drone-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/we-built-drone-bench] indexed:0 read:2min 2026-09-12 ·