Should we be worried about how good AI is getting at coding autonomous drones? Andon Labs introduced Drone-Bench, a benchmark where AI agents code drones to complete autonomous surveillance tasks, based on its Project Pilot work with Anthropic. The benchmark raises questions about the safety implications of increasingly capable AI coding autonomous drones. Hey everyone We’re introducing Drone-Bench, a benchmark where AI agents code drones to complete a simple autonomous surveillance task. Drone-Bench is independent but based on Project Pilot, our work with Anthropic. I would make a condensed version here for LW, but visuals are much better on web: https://andonlabs.com/evals/drone-bench https://andonlabs.com/evals/drone-bench Very curious on your feedback