04:00
2026-07-27
machinebrief.com
artificial-intelligence
Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents
A new benchmark called MissionBench reveals that even the strongest multimodal large language models (MLLMs) succeed on fewer than 35% of aerial missions, compared to 84.4% human performance, accordinβ¦