cd /news/artificial-intelligence/motionblind-probing-the-illusion-of-… · home topics artificial-intelligence article
[ARTICLE · art-125440] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

A new arXiv paper (2609.09528v1) introduces MotionBlind, a contrastive benchmark showing that Video-LLMs cannot reliably judge physically grounded motion, with open models sitting near the 6.25% chance floor and scale providing no improvement. In a controlled study of six open and two frontier Video-LLMs, only Gemini3.1 Pro cleared the benchmark overall, and even it failed on speed; removing the video dropped every model to zero Instance Accuracy and shuffling frames collapsed accuracy to chance. The authors conclude that a perceptual front end unable to distinguish two speeds of the same action is not yet a trustworthy source of supervision, reward, or evaluation for a world model.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.09528v1 Announce Type: new Abstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, which way it travels, how hard it is pushed. We show they cannot. A Video-LLM can watch two clips of the same person in the same room, name every object in both, and still fail to say which clip moves faster. We introduce MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion(speed, magnitude, and direction), the variables a world model must predict. Each instance is a pair of near-identical clips that differ only in motion. Each clip carries two complementary yes/no questions, giving four items per instance, and a model earns credit only if all four are correct. We report Instance Accuracy(IAcc), which has a 6.25% chance floor. Single-frame, appearance, and language-only shortcuts all collapse to it. MotionBlind complements the recent TimeBlind benchmark. We run a controlled study of six open and two frontier Video-LLMs, varying whether the video is present, whether frames are shown in the correct temporal order, and how frames are sampled (1 to 24 frames, four selection strategies). Open models sit near the 6.25% floor, and scale does not help. Removing the video drops every model to zero IAcc, and shuffling frames collapses IAcc to chance, so the task genuinely needs video in order. Neither more frames nor smarter frame selection closes the gap, because these change which frames are seen, not whether motion is read. Only Gemini3.1 Pro clears the benchmark overall, and even it fails on speed. A frontend that cannot tell two speeds of the same action apart is not yet a trustworthy source of supervision, reward, or evaluation for a world model.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @motionblind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/motionblind-probing-…] indexed:0 read:1min 2026-09-10 ·