{"slug": "4dcodebench", "title": "4dcodebench", "summary": "GPT-6 Astra [Max] leads the 4DCodeBench leaderboard, followed closely by Claude Opus 5.5 [High], on a new benchmark that evaluates coding agents on reconstructing dynamic scenes from video by writing executable graphics code from scratch. Across all 18 models, motion reconstruction lags appearance and static geometry, with GPT-6 Astra [Max] scoring 0.91 on the static metric families versus 0.67 on the dynamic ones. Analytic motion is the most common strategy at 67% of solutions, followed by custom simulation at 19%, Blender physics at 10% and keyframing at 3%.", "body_md": "# 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes\n\n<sup>∗</sup>Equal contribution, listed in random order. <sup>‡</sup>Equal co-advising. <sup>†</sup>Work done during a summer internship.\n\n<sup>1</sup>\n<sup>2</sup>\n<sup>3</sup>\n<sup>4</sup>\n\n[Paper](https://arxiv.org/pdf/2610.03715)\n[Citation](#cite)\n[Leaderboard](#leaderboard)\n[Data](https://huggingface.co/4DCodeBench)\n[Code](https://github.com/4DCodeBench/4DCodeBench)\n\nCan a coding agent watch a video of a physical\n      event and write, *from scratch*, the graphics program\n      that reconstructs it?\n\nWe introduce **4DCodeBench**, a benchmark that\n      evaluates coding agents on reconstructing dynamic scenes from video. Given a\n      reference video, agents write executable graphics code that defines the\n      scene’s 3D geometry, its motion over\n      time, and the rendering that turns both back into\n      video.\n\n## The task\n\n## Key takeaways\n\n## Leaderboard\n\nGPT-6 Astra [Max] leads the overall ranking,\n    followed closely by Claude Opus 5.5 [High].\n\n    Open-weight models generally trail proprietary models.\n\nWe evaluate each reconstruction’s appearance,\n    geometry and motion. Models are ranked by\n    an Overall score that averages five metric families; the\n    individual scores help identify where each model succeeds or struggles, which\n    [Discussion](#findings) lays out column by column.\n\n## How the scores are computed\n\n### Reconstruction quality vs. compute\n\nWe use VLM-as-judge to obtain pairwise preferences between reconstructions based on how closely they match the reference video. We aggregate these preferences into Elo ratings and plot them against average token use or cost per task.\n\nAcross models, greater token use does not consistently yield higher Elo. Within GPT-6 Astra, increasing reasoning effort improves reconstruction quality while using more tokens.\n\n## How to read this plot\n\n- The axesThe selected score against average tokens or cost per task, on a log scale.\n- The lineThe Pareto frontier: no other model both spends less and scores higher than a model on it.\n- The barsThe 95% bootstrap intervals for Elo and Human Elo.\n\n## The benchmark\n\n### Input video\n\n### Agent reconstructions\n\nDrag across any render to compare\nWhat materials are present in the scenes realsimulated\n\nClick any task to see every model’s result\n\n## Discussion\n\n### Overall results\n\nWe evaluate appearance, geometry and dynamics separately, combining five metric families into the Overall score. GPT-6 Astra [Max] leads the ranking, followed by Claude Opus 5.5 [High]. Open-weight models generally trail proprietary models, with substantial differences in reconstruction quality across models.\n\n## What each column measures\n\n### Reconstructing dynamics remains harder\n\nAcross all 18 models, reconstruction of motion consistently lags behind\n        appearance and static geometry. Even GPT-6 Astra [Max] scores\n        **0.91 on the static families against 0.67 on the dynamic ones**.\n        Recovering what a scene looks like does not yet translate into reliably\n        reconstructing how it evolves.\n\n### What makes a scene hard to reconstruct?\n\nDifferent scene properties expose different weaknesses.\n        **Real scenes, and scenes containing multiple materials**, receive\n        lower perceptual and VQA scores. **Co-dimensional structures, such as\n        cloth and rope**, are particularly hard for 2D motion and depth.\n        Rigid and articulated scenes differ little from the dataset-wide average.\n\n### How do agents reconstruct dynamics?\n\nAgents can write their own solvers, use Blender’s physics tools, or prescribe how geometry moves and deforms over time.\n\n- **Analytic motion is the most common\n        strategy:** across 18 models, 67% of solutions use analytic motion,\n        followed by custom simulation (19%),\n        Blender physics (10%) and\n        keyframing (3%).\n- **Opus simulates far more often than the overall average:** custom or Blender simulation accounts for 76% of\n        Opus 5.5’s solutions and 67% of Opus 5’s, compared with\n        29% across all models.\n- **Astra favours analytic motion, but simulates\n        more at higher reasoning effort:** simulation use rises from 13% at Low\n        to 22% at High and 32% at Max.\n\nThe examples below show how these choices play out in specific scenes.\n\n#### Rigid bodies\n\nAll six configurations use Blender’s Bullet rigid-body solver, but tune the collapse differently. Claude Fable 5.1 [High] lowers solver iterations so the stack crumbles; GPT-6 Astra [High] prescribes progressive support failure before Bullet handles the falling blocks. Five configurations reconstruct the 8 × 8 × 30 tower, while Astra [Low] builds only half its depth.\n\n#### Deformable solids\n\nClaude Opus 5.5 [High] implements an MLS-MPM simulation of an elastic solid. GPT-6 Astra [Max] instead constructs the shape procedurally and prescribes its deformation using an interpolated squeeze trajectory.\n\n#### Cloth and rope\n\nClaude Opus 5.5 [High] writes a PBD cloth simulation with graph-coloured distance constraints and kd-tree self-contact. GPT-6 Astra [Max] scripts the fold from a table of chosen poses, computing the cloth’s shape directly while preserving its length.\n\nClaude Opus 5.5 [High] writes a PBD rope simulation with stretch and bending constraints, self-contact, and friction against the table and basket. GPT-6 Astra [Max] implements a discrete elastic-rod simulation driven by the robot’s grasps. Astra [High] instead fits the rope’s shape with a smoothed curve, while Astra [Low] writes a simpler solver with self-contact.\n\n#### Flowing materials\n\nClaude Opus 5.5 [High] writes a two-phase MLS-MPM simulation for water and sand, but the sand piles up like dough and the water fails to reproduce the splashing seen in the reference. GPT-6 Astra [Max] prescribes ballistic jet trajectories and redistributes momentum at their intersection through explicit formulas. Recognisable geometry and convincing rendering mask a poor reconstruction of the flow: the prescribed motion fails to capture how the materials interact and evolve after collision.\n\nFor the dam break, GPT-6 Astra [Max] uses Blender’s Mantaflow fluid solver, while Claude Opus 5.5 [High] implements a custom MLS-MPM simulation in taichi. Models switch strategies as the dynamics grow more complex.\n\n#### Fracture\n\nClaude Opus 5.5 [High] writes an MLS-MPM simulation with a particle-level fracture threshold, allowing the material to separate as it stretches. GPT-6 Astra [Max] instead scripts the loaf’s stretching and tearing, moving a fracture front along its length. The tear follows an authored progression rather than emerging from simulated material failure.\n\nFour foam bars are clamped at the feet; the top platen counter-rotates\n          300° over 72 frames and lifts, the bars braid, and each one tears near\n          the top grip. GPT-6 Astra [Max] leads on dynamics by 31% over the\n          next-best model, **with no solver at all**: a reduced beam model\n          carrying four per-bar break frames fitted to hundredths of a frame. Claude Opus 5.5 [High]\n          writes MLS-MPM with a yield threshold on a marked band of particles, and lets\n          the tear emerge from it.\n\n### Do the metrics align with human preferences?\n\n**Model rankings closely track human judgements.** Across 3,587\n        pairwise judgements from 76 participants, human and VLM Elo have a Spearman\n        correlation of **ρ = 0.98**. The Overall score also\n        correlates strongly with human Elo (**ρ = 0.96**).\n\nThis agreement supports automated model comparison, though individual VLM preferences are less reliable when two models are closely matched.\n\n## FAQ\n\n## What is 4DCodeBench?\n\n4DCodeBench asks for a program rather than a prediction. Each of its\n          **200 tasks** hands an agent one video of something physical\n          happening, and the agent writes code that reconstructs the event as a 4D world it\n          can render. 100 videos are real recordings and\n          100 come from a physics simulator, so a submission\n          can be checked against a measured scene as well as a known one. Nothing tells the\n          agent what the objects are made of, how many there are, or where the camera sits:\n          it reads that off the video and commits to it in code that runs.\n\n## What does an agent have to produce?\n\nThe agent writes a program and submits it together with the 4D world it generates: the rendered video, the camera, and the geometry at every frame. Running the program must regenerate everything without access to the reference video. We place no restrictions on how the motion is made; agents keyframe it, simulate it in Blender, Taichi or Warp, or write their own solvers.\n\n## Why mix real and simulated videos?\n\nThe two kinds of video let us evaluate different things. Real videos show realistic appearance and physical behaviour, but have no 4D ground truth, so we can only compare reconstructions of them in the image. For simulated scenes we know the exact 3D geometry and motion at every frame, so we can also compare reconstructions in 3D and over time. Both halves cover the same four families of matter: rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rope, and flowing matter such as grains and fluids.\n\n## What does the agent get to work with?\n\nOnly the video and a fixed task prompt. We provide no scene metadata, object\n          lists or descriptions, so the agent has to work out what is in the video before\n          writing any code. Each run takes place in an isolated container with a GPU,\n          Blender and common simulation libraries, and **nothing else**: no\n          physics engine, tracker, pose estimator or reconstruction tool is provided\n          beyond those libraries, and the scorer is not in the container. Each model gets\n          **one attempt** per scene.\n\n**90.1%** of submissions run end to end. More reasoning improves\n          results within a model: raising GPT-6 Astra’s reasoning effort from Low to\n          Max lifts its Overall score from **0.73** to **0.79**.\n          Across different models, the number of steps or tokens used is a poor\n          predictor of the score.\n\n## How is a reconstruction scored?\n\nWe compare each reconstruction with the reference video, and for\n          simulated scenes also with the true 4D world. The metrics fall into five\n          families, and the **Overall** score is their unweighted average. VQA\n          and the two Elo ratings are reported separately.\n\nThe [Metrics](#metrics-sec) tab shows each metric on\n          a strong and a weak run of the same scene.\n\n## How are the Elo ratings produced, and can the VLM judge be trusted?\n\nBoth ratings come from pairwise comparisons. The judge sees the reference video and two anonymised reconstructions and picks the one that matches it better. For VLM Elo the judge is a vision-language model; for human Elo, 76 participants made 3,587 such judgments across all 200 scenes.\n\nThe VLM’s judgments are close to human judgments. The two rankings are\n          nearly identical (Spearman ρ = **0.98**), and on\n          individual comparisons the VLM agrees with humans almost as often as humans agree\n          with each other (**89.3%** vs. **92.1%**). Individual VLM\n          judgments are close to chance for models within about 50 Elo points of each other,\n          and become more reliable as the gap widens. The\n          [Human agreement](#leaderboard) tab plots the two\n          ratings against each other.\n\n## How do I run my own model on the benchmark?\n\nRunning a model takes four steps. The agent works in an isolated container that holds only the input video; scoring runs in a separate one.\n\n- 1. SetupDownload the data and checkpoints and build the images, with Docker or, on a cluster, Apptainer or SingularityCE.\n- 2. Run an agentChoose the cases, agent and model in the jobs file. Claude, GPT and Gemini run through their own CLIs; any model with an OpenAI-compatible endpoint runs through Stirrup.\n- 3. ScoreCompute each case’s reference estimates once, then the metrics of every run.\n- 4. VLM judgeVQA and the pairwise Elo, with an OpenRouter key. Download our agents’ renders to rate a new model against them.\n\nCommands, configuration formats and credentials are in the\n          [README](https://github.com/4DCodeBench/4DCodeBench).\n\n## Citation\n\n```\n@article{shen20264dcodebench,\n  title={{4DCodeBench}: Benchmarking Agents on Inverse Graphics of Dynamic Scenes},\n  author={Shen, Ruihong and Kova{\\v{c}}i{\\v{c}}, {\\v{Z}}iga and Kulits, Peter and Wang, Xingrui and Li, Zizhang and Tenenbaum, Joshua B. and Yuille, Alan and Chen, Jieneng and Wu, Jiajun},\n  journal={arXiv preprint arXiv:2610.03715},\n  year={2026}\n}\n```\n\n", "url": "https://wpnews.pro/news/4dcodebench", "canonical_source": "https://4dcodebench.com/", "published_at": "2026-10-06 20:29:06+00:00", "updated_at": "2026-10-06 20:50:04.819403+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "large-language-models", "computer-vision"], "entities": ["4DCodeBench", "GPT-6 Astra", "Claude Opus 5.5", "Blender", "Hugging Face", "GitHub", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/4dcodebench", "markdown": "https://wpnews.pro/news/4dcodebench.md", "text": "https://wpnews.pro/news/4dcodebench.txt", "jsonld": "https://wpnews.pro/news/4dcodebench.jsonld"}}