# AI Agent Evaluation End to End: How the LFORLA Drone Build Benchmark Scores Planning, Sourcing, and Assembly

> Source: <https://dev.to/resk/ai-agent-evaluation-end-to-end-how-the-lforla-drone-build-benchmark-scores-planning-sourcing-and-lg8>
> Published: 2026-09-28 09:00:36+00:00

**TL;DR:** The LFORLA Drone Build benchmark measures AI agent evaluation end to end. A model must plan a drone mission, source parts into a bill of materials under budget and physics constraints, and ship OpenSCAD frame source. A deterministic flight-physics oracle scores every step against the same rubric for every model. The current leaderboard shows Nemotron 3 Ultra at 34.166666666666664 and DeepSeek V4 Pro at 34. Our own submission, GLM 5.2, scored 78.0 overall.

AI agent evaluation often stops at a single answer. This benchmark does not. It forces the model through a full build pipeline, and each stage is scored with the same rubric regardless of model.

**Step 1: Plan the mission.** The model receives a mission description. It must decide what kind of drone is needed, what constraints matter, and how to sequence the build. There is no partial credit for a vague plan. The rubric checks whether the plan is actionable and tied to the mission.

**Step 2: Source parts into a BOM.** The model picks components into a bill of materials. It must stay under budget and respect physics constraints. A part that is cheap but cannot lift the payload fails. A part that is powerful but blows the budget fails. The rubric scores the BOM on feasibility, cost discipline, and constraint satisfaction.

**Step 3: Assemble a build plan.** The model must produce a coherent build plan that connects the BOM to the mission. This is where planning and sourcing meet. The rubric checks whether the plan is internally consistent and whether it could actually be executed.

**Step 4: Ship OpenSCAD frame source.** The model writes OpenSCAD code for the drone frame. This is not judged by a human eye. A deterministic flight-physics oracle scores the frame source. The oracle runs the same physics checks for every model. That means the score is reproducible and not subject to judge drift.

**Step 5: Score against the same rubric.** Every model sees the same rubric. Every step is scored the same way. The final number is a composite of planning, sourcing, assembly, and physics-validated frame design.

| Model | Score | Provider | 
|---|---|---|
| Nemotron 3 Ultra (free, via opencode) | 34.166666666666664 | deepseek | 
| DeepSeek V4 Pro | 34 | opencode-zen | 
| GLM 5.2 (our own submission) | 78.0 | lforla | 

The gap between 34 and 78 is not noise. It reflects how much of the end-to-end pipeline each model can actually complete.

A score near 34 suggests the model can handle parts of the task but fails at least one critical stage. It might plan well but source parts that violate physics constraints. It might source parts correctly but produce OpenSCAD source that the oracle rejects. The rubric does not reward partial effort. It rewards a build that could fly.

A score of 78.0 suggests the model completes most of the pipeline. It plans, sources, assembles, and ships frame source that passes more of the oracle checks. It is not perfect, but it is materially closer to a deployable build.

The leaderboard is not a general intelligence ranking. It is a task-specific ranking. A model that scores 34 here might score higher on a different benchmark. That is the point of end-to-end evaluation: it isolates the ability to carry a complex, multi-step build from plan to validated artifact.

If your workflow looks like this benchmark, the leaderboard is directly useful. You need a model that can plan, source, assemble, and produce code that passes a deterministic checker. In that case, a model at 78.0 is a different tool than a model at 34.

If your workflow is narrower, the leaderboard is a signal, not a verdict. Use it to shortlist models, then test them on your own rubric. The value of this benchmark is that it shows what end-to-end evaluation looks like when the judge is deterministic and the rubric is fixed.

This benchmark is one task. It measures drone design under budget and physics constraints. It does not measure general reasoning, long-context recall, or tool use outside this pipeline. The scores are real, but they are not universal.

The leaderboard also has a small number of entries. Nemotron 3 Ultra and DeepSeek V4 Pro are close at 34.166666666666664 and 34. That closeness may reflect a ceiling in the task or a shared failure mode. More entries would clarify that.

Finally, our own submission, GLM 5.2, scored 78.0. We are transparent about that. We submitted our own model. The number is real, but readers should weigh it with that context.

AI agent evaluation end to end is harder than it looks. The LFORLA Drone Build benchmark shows one way to do it: fix the rubric, use a deterministic oracle, and score every model the same way. The current leaderboard puts Nemotron 3 Ultra at 34.166666666666664 and DeepSeek V4 Pro at 34, with our GLM 5.2 at 78.0 overall.

If you want to see the full leaderboard and run your own evaluation, start at [https://lforla.org](https://lforla.org). Pick a task, define your rubric, and score every model the same way. That is how you turn a leaderboard into a decision.
