# NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

> Source: <https://www.marktechpost.com/2026/09/30/nvidia-researchers-introduce-physis-lang-self-evolving-physical-language-that-lifts-cosmos-3-past-veo-3-1-on-physics-benchmarks/>
> Published: 2026-09-30 07:32:34+00:00

Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals.

Their framework, [Physis-Lang](https://physis-intelligence.github.io/physis-lang-web/), treats physical language as a shared, optimizable representation. The same text drives data curation, model training and inference. On the public [Physics-IQ Verified leaderboard](https://physics-iq-verified.anates.ai/) snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4. The Cosmos3-Nano version ranks second at 43.3 ± 1.5. 

## **What Problem Does Physis-Lang Solve?**

Conventional captions describe what happens, not why. ‘Butter melts as the temperature rises’ says nothing about heat transfer or gravity. Physis-Lang adds a `physics_reasoning` field to each base caption. It spells out entities, causes, interactions, governing principles, temporal evolution and effects.

The pipeline also writes a scene-specific `physics_negative_prompt`. This text describes likely implausible outcomes, such as a stone floating on water. It acts as negative conditioning at inference time.

## **How Does the Self-Evolving Caption Loop Work?**

The loop keeps the captioner frozen and evolves only its instruction. A GPT-5.5 captioner writes captions for a fixed 20-video development set with 273 human-verified assertions. Gemini-3.1-Pro acts as a physics-aware critic. An evolution agent reads the scores and claim-level failures, then rewrites the prompt.

**The critic scores 2 dimensions:**

- **Precision:** the caption is split into atomic claims, and each claim is checked against the video.
- **Recall:** each human-curated physical assertion must be explicitly stated or entailed by the caption.

Every revised prompt is validated on **PhysCapBench**, a new benchmark of 246 videos and 3,794 human-verified assertions. Caption F1 rose from 78.64 at iteration 1 to 87.82 at iteration 9. The path was not smooth. Iteration 2 made captions overly cautious and dropped F1 to 76.28. Iteration 9 required every visible causal step and raised frame sampling from 2 fps to 4 fps.

## **How Does Language-Guided Data Curation Work?**

A GPT-5.5 diagnosis agent maps generated-video failures to physics categories like rigid-body motion, collision and fluid dynamics. That deficiency profile is matched against physics tags on a large video gallery. Retrieval targets physical content, not visual appearance.

The final training set holds 183K videos: 71K filtered from WISA-80K plus 112K retrieved clips. Retrieval alone added 3.01 points on average across 3 benchmarks. On VideoPhy-2, chemical and thermal processes each gained 8.00 points.

## **How Does Physis-Lang Compare With Veo 3.1?**

Fine-tuning uses LoRA on attention projections, with no architecture or objective change. **Physis-Lang on Cosmos3-Nano versus Google's Veo 3.1:**

- [PhyGenBench](https://github.com/OpenGVLab/PhyGenBench) : 71.04 vs 65.63
- [Physics-IQ](https://github.com/google-deepmind/physics-IQ-benchmark) Verified: 43.41 vs 34.99
- [PhyGround](https://arxiv.org/abs/2605.10806) : 69.90 vs 69.24
- [VideoPhy-2](https://videophy2.github.io/) : 68.02 vs 68.87 on the full set, 62.36 vs 58.43 on the Hard split

Gains hold across backbones: +7.05 on Wan2.1-14B, +3.24 on Cosmos3-Edge-4B, +6.22 on Cosmos3-Nano-16B and +5.02 on Cosmos3-Super-64B. General quality held steady on VBench-I2V, where Cosmos3-Nano moved from 88.32 to 88.69.

Prompting alone also helps. Physics reasoning plus negative prompts lifted a frozen Cosmos3-Nano on PhyGenBench from 61.67 to 67.29.

## **Can It Run Without Commercial APIs?**

The research team distilled the GPT pipeline into 2 Qwen3-VL-4B-Instruct models: PhysThinker-C for captioning and PhysThinker-U for prompt upsampling. On Wan2.1-14B, the commercial pipeline gave +7.05 at about $24.12K in API cost. Swapping in PhysThinker-C kept +6.76 at about $0.12K. A fully local setup cost $0 and still added +4.76.

## **Physis-Lang vs Closest Competitors**

| Feature | [Physis-Lang](https://physis-intelligence.github.io/physis-lang-web/) | [PhiZero](https://arxiv.org/abs/2607.28624) | [PhyGDPO](https://arxiv.org/abs/2512.24551) | [Self-Refinement](https://arxiv.org/abs/2511.20280) | 
|---|---|---|---|---|
| **Developer** | NVIDIA, MIT, Oxford | CASIA (NLPR) | [Meta](https://ai.meta.com/research/publications/phygdpo-physics-aware-groupwise-direct-preference-optimization-for-physically-consistent-text-to-video-generation/) (ECCV 2026) | Liu et al. | 
| **Core idea** | Self-evolving natural-language physics captions and negative prompts | Learned discrete "physical language", reason-then-render | Groupwise DPO with VLM physics rewards | Multimodal chain-of-thought prompt refinement from VLM feedback | 
| **Where physics enters** | Data curation, training captions and inference prompts | Qwen3-VL-4B reasoner feeding a diffusion decoder | Preference training on PhyVidGen-135K | Inference prompts only | 
| **Training needed** | LoRA SFT (prompt-only mode also helps) | Yes | Yes (DPO) | No, training-free | 
| **Physics-IQ Verified** | **43.41** | 40.91 | n/r | 27.20 | 
| **PhyGenBench** | **71.04** | n/r | 48.96 | 49.17 | 
| **VideoPhy-2 (All / Hard)** | **68.02 / 62.36** | n/r | 59.56 / 44.94 | 47.88 / 28.09 | 
| **PhyGround** | **69.90** | 57.85 | n/r | 58.22 | 
| **Code / weights public** | [Paper only](https://github.com/Physis-Intelligence/Physis-Lang) | ["Coming soon"](https://github.com/yaoyao-jpg/PhiZero) | ["Released soon"](https://github.com/caiyuanhao1998/Open-PhyGDPO) | [Paper](https://arxiv.org/abs/2511.20280) | 

*Scores are from the [Physis-Lang paper](https://physis-intelligence.github.io/physis-lang-web/assets/Physis-Lang.pdf), Tables 1 to 4, run under one protocol per benchmark (PhyGenBench and VideoPhy-2 use a GPT-5.5 evaluator). Physis-Lang numbers use the Cosmos3-Nano backbone. n/r = not reported in that comparison. Release status checked September 30, 2026.*

## **Key Takeaways**

- Physis-Lang evolves physics captions with a critic-guided agent while the captioner stays frozen.
- PhysCapBench scores captions on 3,794 human-verified cause, law and effect assertions.
- Cosmos3-Nano with Physis-Lang beats Veo 3.1 on 3 of 4 benchmarks.
- Physics prompts alone lift a frozen model by 5.62 points on PhyGenBench.
- No code or weights are public yet; only the paper is released.

Check out the [Paper](https://physis-intelligence.github.io/physis-lang-web/assets/Physis-Lang.pdf),[**Project Page**](https://physis-intelligence.github.io/physis-lang-web/) and [** GitHub Repo**](https://github.com/Physis-Intelligence/Physis-Lang). All credit goes to the researcher of this project. Also, feel free to follow us on **[Twitter](https://x.com/intent/follow?screen_name=marktechpost)** and don’t forget to join our **[150k+ML SubReddit](https://www.reddit.com/r/machinelearningnews/)** and Subscribe to **[our Newsletter](https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}})**. Wait! are you on telegram? [now you can join us on telegram as well.](https://t.me/machinelearningresearchnews)

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? [Connect with us](https://forms.gle/MJjjVDPS7whH8Ngs6)

Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.
