Meet ‘Code-as-World’: An Agentic Loop That Rewrites Real Videos Into Executable MuJoCo Physics Programs MirroS released Code-as-World, a paradigm that represents physical scenes as executable MuJoCo programs recovered from real videos via an agentic loop, and open-sourced it under Apache 2.0 with two checkpoints, Code-as-World-VL-4B and Code-as-World-VL-9B. The 9B model scores 55.4 MRA on QuantiPhy-validation, surpassing Gemini-3.1 Flash's 54.8 and roughly 15 points above the strongest open-weight baseline, with the technical report and GitHub repo available. MirroS released Code-as-World : a paradigm that represents physical worlds through executable world representations. The argument is narrow and testable: pixels are evidence of a physical scene, not its ontology. A video model can predict plausible frames without ever representing mass, contact, or gravity. So instead of pixels, latents, or captions, Code-as-World represents a scene as executable code — a scene.json that MuJoCo can run, that an agent can verify against the source video, and that anyone can edit and re-simulate. An agentic loop recovers those programs from real footage in up to five rounds. The verified worlds then become training data with exact physical labels, which real video does not carry. Trained on that supervision, Code-as-World-VL-9B https://huggingface.co/MirroS-Lab/Code-as-World-VL-9B scores 55.4 MRA on QuantiPhy https://quantiphy.stanford.edu/ -validation, above Gemini-3.1 Flash at 54.8 and roughly 15 points above the strongest open-weight baseline. Is it deployable? Yes , at the research and internal-prototype tier. MirroS https://mirros.ai/ shipped the GitHub repo https://github.com/MirroS-Lab/Code-as-World and two checkpoints — Code-as-World-VL-4B https://huggingface.co/MirroS-Lab/Code-as-World-VL-4B and Code-as-World-VL-9B https://huggingface.co/MirroS-Lab/Code-as-World-VL-9B — under Apache 2.0 , fine-tuned from Qwen3.5-4B and Qwen3.5-9B. Both are BF16 safetensors served by vLLM behind an OpenAI-compatible /v1 endpoint, with 16 sampled frames per video and --max-model-len 4608 . The idea: pixels are evidence, not ontology The MirroS technical report https://mirros.ai/report/code-as-world.pdf argues that video models, 3D reconstruction, and captions each recover part of a scene but none recovers its mechanism . Code-as-World represents a scene as an executable world representation EWR , a triple p = C, E, A : Composition : objects, geometry, metric dimensions, mass, friction, gravity. Floors and walls are static physical entities so they can support and collide. Evolution : initial states, forces, contacts, collisions, termination conditions, duration. Executing it expands composition into a full state trajectory. Appearance : camera, lighting, materials, background, frame rate, render config. Changing it never changes the physics. In the released implementation, that triple compiles into a scene.json executed in MuJoCo , with two interchangeable engines: an animation engine kinematic poses and a physics engine forces and contacts . Agentic discovery instead of one-shot prediction Recovering an EWR from a video is an inverse problem, so the team frames it as abductive search. An agent runs propose → instantiate → execute → render → verify for up to K = 5 rounds. For video input, SAM 3 https://arxiv.org/abs/2511.16719 supplies instance masks and image-plane tracks, VGGT-Omega https://arxiv.org/abs/2605.15195 estimates depth and camera geometry, and SAM 3D https://arxiv.org/abs/2511.16624 generates per-object meshes. Candidate rollouts are projected back into the input view and compared at selected key frames on RGB, depth, masks, and trajectories. Frame-level discrepancies aggregate into structured feedback Δ that guides the next revision; when the budget runs out without acceptance, the hypothesis is rejected. At a matched five-evaluation budget, the loop beats Best-of-5 independent sampling on Visual Alignment, Object IoU, Traj-ADE, and Accuracy@2%D — and the result repeats under the second execution engine. Candidate videos come from WISA-80K after motion-focused filtering; sim-to-real re-rendering uses Wan2.2-VACE plus an internal video model. Verified worlds as training supervision Phase 1 is supervised fine-tuning on 73,335 image-space QA pairs built from RefCOCO/+/g, RefCLEF and GOT-10K, covering extent, position, displacement, velocity and acceleration in raw pixels. Phase 2 applies GRPO to world-space VQA drawn from 1,585 text-driven and 988 video-driven executable worlds, rewarded on scale-normalized numerical accuracy plus unit and format terms. Training used eight NVIDIA H100 GPUs. On QuantiPhy https://quantiphy.stanford.edu/ -validation 159 items, MRA macro-averaged over 2S/2D/3S/3D : 4B = 50.6, 9B = 55.4, 27B reasoning = 58.6 , against Gemini-3.1 Flash at 54.8, ChatGPT-5.1 at 48.4, and the strongest open-weight baseline Qwen3-VL-32B-Instruct at 40.2. The ablation is the more useful number: image-space-only scores 44.2 4B and 50.9 9B ; adding both world-space sources lifts them to 50.6 and 55.4. Pixel-level grounding improves too — the 9B goes 63.7 → 68.3 on RefCOCO and 20.1 → 26.6 on GOT-10K after world-space RL. Key Takeaways - Code-as-World turns a video into an editable scene.json that MuJoCo can execute and verify. - Five-round propose→verify search beats Best-of-5 sampling at the same compute budget. - Verified worlds supply exact physical labels that real video simply does not carry. - 9B hits 55.4 MRA on QuantiPhy, above Gemini-3.1 Flash at 54.8; 4B and 9B are Apache 2.0. - Rigid-body only, and the model never learns the discovery loop itself. Check out the Technical report , , https://mirros.ai/blog/representing-physical-world MirroS blog , https://mirros-lab.github.io/code-as-world/ Project page and https://github.com/MirroS-Lab/Code-as-World GitHub . Also, feel free to follow us on https://x.com/MirroS ai/status/2093367432920650063 Announcement on X and don’t forget to join our Twitter https://x.com/intent/follow?screen name=marktechpost and Subscribe to 150k+ML SubReddit https://www.reddit.com/r/machinelearningnews/ . Wait are you on telegram? our Newsletter https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}} now you can join us on telegram as well. https://t.me/machinelearningresearchnews Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us https://forms.gle/wbash1wF6efRj8G58 Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.