# 4dcodebench

> Source: <https://4dcodebench.com/>
> Published: 2026-10-06 20:29:06+00:00

# 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

<sup>∗</sup>Equal contribution, listed in random order. <sup>‡</sup>Equal co-advising. <sup>†</sup>Work done during a summer internship.

<sup>1</sup>
<sup>2</sup>
<sup>3</sup>
<sup>4</sup>

[Paper](https://arxiv.org/pdf/2610.03715)
[Citation](#cite)
[Leaderboard](#leaderboard)
[Data](https://huggingface.co/4DCodeBench)
[Code](https://github.com/4DCodeBench/4DCodeBench)

Can a coding agent watch a video of a physical
      event and write, *from scratch*, the graphics program
      that reconstructs it?

We introduce **4DCodeBench**, a benchmark that
      evaluates coding agents on reconstructing dynamic scenes from video. Given a
      reference video, agents write executable graphics code that defines the
      scene’s 3D geometry, its motion over
      time, and the rendering that turns both back into
      video.

## The task

## Key takeaways

## Leaderboard

GPT-6 Astra [Max] leads the overall ranking,
    followed closely by Claude Opus 5.5 [High].

    Open-weight models generally trail proprietary models.

We evaluate each reconstruction’s appearance,
    geometry and motion. Models are ranked by
    an Overall score that averages five metric families; the
    individual scores help identify where each model succeeds or struggles, which
    [Discussion](#findings) lays out column by column.

## How the scores are computed

### Reconstruction quality vs. compute

We use VLM-as-judge to obtain pairwise preferences between reconstructions based on how closely they match the reference video. We aggregate these preferences into Elo ratings and plot them against average token use or cost per task.

Across models, greater token use does not consistently yield higher Elo. Within GPT-6 Astra, increasing reasoning effort improves reconstruction quality while using more tokens.

## How to read this plot

- The axesThe selected score against average tokens or cost per task, on a log scale.
- The lineThe Pareto frontier: no other model both spends less and scores higher than a model on it.
- The barsThe 95% bootstrap intervals for Elo and Human Elo.

## The benchmark

### Input video

### Agent reconstructions

Drag across any render to compare
What materials are present in the scenes realsimulated

Click any task to see every model’s result

## Discussion

### Overall results

We evaluate appearance, geometry and dynamics separately, combining five metric families into the Overall score. GPT-6 Astra [Max] leads the ranking, followed by Claude Opus 5.5 [High]. Open-weight models generally trail proprietary models, with substantial differences in reconstruction quality across models.

## What each column measures

### Reconstructing dynamics remains harder

Across all 18 models, reconstruction of motion consistently lags behind
        appearance and static geometry. Even GPT-6 Astra [Max] scores
        **0.91 on the static families against 0.67 on the dynamic ones**.
        Recovering what a scene looks like does not yet translate into reliably
        reconstructing how it evolves.

### What makes a scene hard to reconstruct?

Different scene properties expose different weaknesses.
        **Real scenes, and scenes containing multiple materials**, receive
        lower perceptual and VQA scores. **Co-dimensional structures, such as
        cloth and rope**, are particularly hard for 2D motion and depth.
        Rigid and articulated scenes differ little from the dataset-wide average.

### How do agents reconstruct dynamics?

Agents can write their own solvers, use Blender’s physics tools, or prescribe how geometry moves and deforms over time.

- **Analytic motion is the most common
        strategy:** across 18 models, 67% of solutions use analytic motion,
        followed by custom simulation (19%),
        Blender physics (10%) and
        keyframing (3%).
- **Opus simulates far more often than the overall average:** custom or Blender simulation accounts for 76% of
        Opus 5.5’s solutions and 67% of Opus 5’s, compared with
        29% across all models.
- **Astra favours analytic motion, but simulates
        more at higher reasoning effort:** simulation use rises from 13% at Low
        to 22% at High and 32% at Max.

The examples below show how these choices play out in specific scenes.

#### Rigid bodies

All six configurations use Blender’s Bullet rigid-body solver, but tune the collapse differently. Claude Fable 5.1 [High] lowers solver iterations so the stack crumbles; GPT-6 Astra [High] prescribes progressive support failure before Bullet handles the falling blocks. Five configurations reconstruct the 8 × 8 × 30 tower, while Astra [Low] builds only half its depth.

#### Deformable solids

Claude Opus 5.5 [High] implements an MLS-MPM simulation of an elastic solid. GPT-6 Astra [Max] instead constructs the shape procedurally and prescribes its deformation using an interpolated squeeze trajectory.

#### Cloth and rope

Claude Opus 5.5 [High] writes a PBD cloth simulation with graph-coloured distance constraints and kd-tree self-contact. GPT-6 Astra [Max] scripts the fold from a table of chosen poses, computing the cloth’s shape directly while preserving its length.

Claude Opus 5.5 [High] writes a PBD rope simulation with stretch and bending constraints, self-contact, and friction against the table and basket. GPT-6 Astra [Max] implements a discrete elastic-rod simulation driven by the robot’s grasps. Astra [High] instead fits the rope’s shape with a smoothed curve, while Astra [Low] writes a simpler solver with self-contact.

#### Flowing materials

Claude Opus 5.5 [High] writes a two-phase MLS-MPM simulation for water and sand, but the sand piles up like dough and the water fails to reproduce the splashing seen in the reference. GPT-6 Astra [Max] prescribes ballistic jet trajectories and redistributes momentum at their intersection through explicit formulas. Recognisable geometry and convincing rendering mask a poor reconstruction of the flow: the prescribed motion fails to capture how the materials interact and evolve after collision.

For the dam break, GPT-6 Astra [Max] uses Blender’s Mantaflow fluid solver, while Claude Opus 5.5 [High] implements a custom MLS-MPM simulation in taichi. Models switch strategies as the dynamics grow more complex.

#### Fracture

Claude Opus 5.5 [High] writes an MLS-MPM simulation with a particle-level fracture threshold, allowing the material to separate as it stretches. GPT-6 Astra [Max] instead scripts the loaf’s stretching and tearing, moving a fracture front along its length. The tear follows an authored progression rather than emerging from simulated material failure.

Four foam bars are clamped at the feet; the top platen counter-rotates
          300° over 72 frames and lifts, the bars braid, and each one tears near
          the top grip. GPT-6 Astra [Max] leads on dynamics by 31% over the
          next-best model, **with no solver at all**: a reduced beam model
          carrying four per-bar break frames fitted to hundredths of a frame. Claude Opus 5.5 [High]
          writes MLS-MPM with a yield threshold on a marked band of particles, and lets
          the tear emerge from it.

### Do the metrics align with human preferences?

**Model rankings closely track human judgements.** Across 3,587
        pairwise judgements from 76 participants, human and VLM Elo have a Spearman
        correlation of **ρ = 0.98**. The Overall score also
        correlates strongly with human Elo (**ρ = 0.96**).

This agreement supports automated model comparison, though individual VLM preferences are less reliable when two models are closely matched.

## FAQ

## What is 4DCodeBench?

4DCodeBench asks for a program rather than a prediction. Each of its
          **200 tasks** hands an agent one video of something physical
          happening, and the agent writes code that reconstructs the event as a 4D world it
          can render. 100 videos are real recordings and
          100 come from a physics simulator, so a submission
          can be checked against a measured scene as well as a known one. Nothing tells the
          agent what the objects are made of, how many there are, or where the camera sits:
          it reads that off the video and commits to it in code that runs.

## What does an agent have to produce?

The agent writes a program and submits it together with the 4D world it generates: the rendered video, the camera, and the geometry at every frame. Running the program must regenerate everything without access to the reference video. We place no restrictions on how the motion is made; agents keyframe it, simulate it in Blender, Taichi or Warp, or write their own solvers.

## Why mix real and simulated videos?

The two kinds of video let us evaluate different things. Real videos show realistic appearance and physical behaviour, but have no 4D ground truth, so we can only compare reconstructions of them in the image. For simulated scenes we know the exact 3D geometry and motion at every frame, so we can also compare reconstructions in 3D and over time. Both halves cover the same four families of matter: rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rope, and flowing matter such as grains and fluids.

## What does the agent get to work with?

Only the video and a fixed task prompt. We provide no scene metadata, object
          lists or descriptions, so the agent has to work out what is in the video before
          writing any code. Each run takes place in an isolated container with a GPU,
          Blender and common simulation libraries, and **nothing else**: no
          physics engine, tracker, pose estimator or reconstruction tool is provided
          beyond those libraries, and the scorer is not in the container. Each model gets
          **one attempt** per scene.

**90.1%** of submissions run end to end. More reasoning improves
          results within a model: raising GPT-6 Astra’s reasoning effort from Low to
          Max lifts its Overall score from **0.73** to **0.79**.
          Across different models, the number of steps or tokens used is a poor
          predictor of the score.

## How is a reconstruction scored?

We compare each reconstruction with the reference video, and for
          simulated scenes also with the true 4D world. The metrics fall into five
          families, and the **Overall** score is their unweighted average. VQA
          and the two Elo ratings are reported separately.

The [Metrics](#metrics-sec) tab shows each metric on
          a strong and a weak run of the same scene.

## How are the Elo ratings produced, and can the VLM judge be trusted?

Both ratings come from pairwise comparisons. The judge sees the reference video and two anonymised reconstructions and picks the one that matches it better. For VLM Elo the judge is a vision-language model; for human Elo, 76 participants made 3,587 such judgments across all 200 scenes.

The VLM’s judgments are close to human judgments. The two rankings are
          nearly identical (Spearman ρ = **0.98**), and on
          individual comparisons the VLM agrees with humans almost as often as humans agree
          with each other (**89.3%** vs. **92.1%**). Individual VLM
          judgments are close to chance for models within about 50 Elo points of each other,
          and become more reliable as the gap widens. The
          [Human agreement](#leaderboard) tab plots the two
          ratings against each other.

## How do I run my own model on the benchmark?

Running a model takes four steps. The agent works in an isolated container that holds only the input video; scoring runs in a separate one.

- 1. SetupDownload the data and checkpoints and build the images, with Docker or, on a cluster, Apptainer or SingularityCE.
- 2. Run an agentChoose the cases, agent and model in the jobs file. Claude, GPT and Gemini run through their own CLIs; any model with an OpenAI-compatible endpoint runs through Stirrup.
- 3. ScoreCompute each case’s reference estimates once, then the metrics of every run.
- 4. VLM judgeVQA and the pairwise Elo, with an OpenRouter key. Download our agents’ renders to rate a new model against them.

Commands, configuration formats and credentials are in the
          [README](https://github.com/4DCodeBench/4DCodeBench).

## Citation

```
@article{shen20264dcodebench,
  title={{4DCodeBench}: Benchmarking Agents on Inverse Graphics of Dynamic Scenes},
  author={Shen, Ruihong and Kova{\v{c}}i{\v{c}}, {\v{Z}}iga and Kulits, Peter and Wang, Xingrui and Li, Zizhang and Tenenbaum, Joshua B. and Yuille, Alan and Chen, Jieneng and Wu, Jiajun},
  journal={arXiv preprint arXiv:2610.03715},
  year={2026}
}
```


