Hey everyone!
We’re introducing Drone-Bench, a benchmark where AI agents code drones to complete a simple autonomous surveillance task. Drone-Bench is independent but based on Project Pilot, our work with Anthropic.
I would make a condensed version here for LW, but visuals are much better on web:
https://andonlabs.com/evals/drone-bench Very curious on your feedback!