AI Agent Evaluation End to End: How the LFORLA Drone Build Benchmark Scores Planning, Sourcing, and Assembly
The LFORLA Drone Build benchmark evaluates AI agents end to end, requiring models to plan a drone mission, source parts into a bill of materials under budget and physics constraints, assemble a build β¦