# 4DCodeBench finds coding agents struggle to reconstruct motion from video

> Source: <https://runtimewire.com/article/4dcodebench-coding-agents-motion-reconstruction>
> Published: 2026-10-11 04:41:26+00:00

# 4DCodeBench finds coding agents struggle to reconstruct motion from video

**The 200-scene benchmark found even its top-ranked configuration scored lower on dynamics than on appearance and geometry.**

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [X](https://x.com/zzigakovacic/status/2107153000179642438)

## Why it matters

Coding agents can produce plausible geometry without reconstructing the motion that governs a scene. 4DCodeBench makes that gap measurable, giving researchers a shared test for a capability relevant to simulation, robotics and video-to-3D systems.

Stanford PhD student [Ziga Kovacic (@zzigakovacic)](https://x.com/zzigakovacic) and collaborators introduced 4DCodeBench, a benchmark that tests whether coding agents can turn a video of a physical event into executable graphics code reproducing the scene's geometry, appearance and movement. The paper appeared on arXiv on October 2nd; Kovacic described the project in [a thread on X](https://x.com/zzigakovacic/status/2107153000179642438) on October 5th.

[https://x.com/zzigakovacic/status/2107153000179642438](https://x.com/zzigakovacic/status/2107153000179642438)

Kovacic's research focuses on physical simulation and world models that represent physical processes at an appropriate level of detail. His earlier work includes MPMWorlds, a project on inferring and extrapolating physical dynamics, and Pocket Time-Lapse, a graphics research project. 4DCodeBench extends that line of inquiry to coding agents: given only a video and a task prompt, can a system write a program that reconstructs not just what a scene looks like, but how it changes over time?

The [benchmark](https://4dcodebench.com/) contains 200 scenes: 100 real-world videos and 100 synthetic scenes created with physics simulations. They cover rigid bodies, deformable objects, cloth and rods, grains, fluids, fracture and interactions among materials. Agents must produce code that can regenerate the scene without access to the reference video. The researchers leave the implementation open: a system can script movement, write its own solver or use physics tools such as Blender's.

That flexibility makes the agents' choices part of the result. Across 18 evaluated configurations, the researchers report that 67% of solutions represented motion analytically, while 19% used custom simulations, 10% used Blender physics and 3% used keyframing. The reported shares total 99%, reflecting rounding. [Claude Opus 5.5](https://runtimewire.com/models/anthropic/claude-opus-5.5) used custom or Blender simulation in 76% of its solutions; [GPT-6 Astra](https://runtimewire.com/models/native-openai/gpt-6-astra-576a599a64cb3dde)'s reported simulation use rose from 13% at Low reasoning effort to 32% at Max.

The top-ranked configuration, GPT-6 Astra at Max reasoning effort, scored 0.91 on the benchmark's static metric families and 0.67 on its dynamic families. Those are benchmark scores, not percentages of scenes reconstructed correctly. The gap points to a specific weakness: a convincing rendering or a plausible static shape does not guarantee that an agent has recovered the forces, contacts or deformations governing what happens next.

The two halves of the dataset answer different questions. Real videos capture actual visual complexity but do not provide ground-truth 3D geometry and motion, so they can be judged in image space. Synthetic scenes provide exact geometry and motion at each frame, allowing direct 3D and temporal comparisons. Results across those settings should not be read as a single verdict on an agent's general physical understanding.

The study also checked its automated preference ratings against people. In 3,587 pairwise comparisons from 76 participants, the researchers report that human and [vision-language-model](https://runtimewire.com/models/huggingface/kirangowda3101-vision-language-model-c37c9a7b88e98ecc) judgments agreed on 89.3% of 916 comparisons shared between the two evaluations. The paper reports stronger agreement between rankings aggregated at the model level; individual automated judgments were less dependable when two configurations were close. This supports using automated rankings at scale, though scene-level comparisons remain less settled.

Kovacic and coauthors report a 90.1% end-to-end execution rate, meaning most submissions produced runnable outputs. The benchmark still measures a constrained task: one video, a fixed prompt, a graphics program and a defined scoring pipeline. It does not establish that an agent can control a robot, build an accurate digital twin, or infer physics reliably outside the benchmark. The benchmark gives researchers a reproducible way to test a capability that static image and geometry scores can miss.

The [paper](https://arxiv.org/abs/2610.03715), [code](https://github.com/4DCodeBench/4DCodeBench) and benchmark materials are publicly available. Running the supplied evaluation pipeline requires a GPU-equipped environment; evaluating new outputs also requires the scorer's dependencies and model checkpoints. The work is an academic collaboration whose authors include researchers affiliated with Stanford, Johns Hopkins, MIT, Peking University and the Max Planck Institute for Intelligent Systems, rather than a commercial product launch.
