{"slug": "motionblind-probing-the-illusion-of-motion-understanding-in-video-llms", "title": "MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs", "summary": "A new arXiv paper (2609.09528v1) introduces MotionBlind, a contrastive benchmark showing that Video-LLMs cannot reliably judge physically grounded motion, with open models sitting near the 6.25% chance floor and scale providing no improvement. In a controlled study of six open and two frontier Video-LLMs, only Gemini3.1 Pro cleared the benchmark overall, and even it failed on speed; removing the video dropped every model to zero Instance Accuracy and shuffling frames collapsed accuracy to chance. The authors conclude that a perceptual front end unable to distinguish two speeds of the same action is not yet a trustworthy source of supervision, reward, or evaluation for a world model.", "body_md": "arXiv:2609.09528v1 Announce Type: new \nAbstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, which way it travels, how hard it is pushed. We show they cannot. A Video-LLM can watch two clips of the same person in the same room, name every object in both, and still fail to say which clip moves faster. We introduce MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion(speed, magnitude, and direction), the variables a world model must predict. Each instance is a pair of near-identical clips that differ only in motion. Each clip carries two complementary yes/no questions, giving four items per instance, and a model earns credit only if all four are correct. We report Instance Accuracy(IAcc), which has a 6.25% chance floor. Single-frame, appearance, and language-only shortcuts all collapse to it. MotionBlind complements the recent TimeBlind benchmark. We run a controlled study of six open and two frontier Video-LLMs, varying whether the video is present, whether frames are shown in the correct temporal order, and how frames are sampled (1 to 24 frames, four selection strategies). Open models sit near the 6.25% floor, and scale does not help. Removing the video drops every model to zero IAcc, and shuffling frames collapses IAcc to chance, so the task genuinely needs video in order. Neither more frames nor smarter frame selection closes the gap, because these change which frames are seen, not whether motion is read. Only Gemini3.1 Pro clears the benchmark overall, and even it fails on speed. A frontend that cannot tell two speeds of the same action apart is not yet a trustworthy source of supervision, reward, or evaluation for a world model.", "url": "https://wpnews.pro/news/motionblind-probing-the-illusion-of-motion-understanding-in-video-llms", "canonical_source": "https://arxiv.org/abs/2609.09528", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 04:24:42.704495+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "computer-vision", "ai-safety"], "entities": ["MotionBlind", "TimeBlind", "Gemini3.1 Pro", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/motionblind-probing-the-illusion-of-motion-understanding-in-video-llms", "markdown": "https://wpnews.pro/news/motionblind-probing-the-illusion-of-motion-understanding-in-video-llms.md", "text": "https://wpnews.pro/news/motionblind-probing-the-illusion-of-motion-understanding-in-video-llms.txt", "jsonld": "https://wpnews.pro/news/motionblind-probing-the-illusion-of-motion-understanding-in-video-llms.jsonld"}}