NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals.

Their framework, Physis-Lang, treats physical language as a shared, optimizable representation. The same text drives data curation, model training and inference. On the public Physics-IQ Verified leaderboard snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4. The Cosmos3-Nano version ranks second at 43.3 ± 1.5.

A video world model can make a convincing clip and still get the physics wrong.Our researchers just released Physis-Lang, an open self-evolving framework that adds physics reasoning to video captions. The captions explain why and how a scene unfolds. We use them to fine-tune… pic.twitter.com/agIHuIB6N2— NVIDIA AI (@NVIDIAAI) September 29, 2026

What Problem Does Physis-Lang Solve?

Conventional captions describe what happens, not why. ‘Butter melts as the temperature rises’ says nothing about heat transfer or gravity. Physis-Lang adds a physics_reasoning field to each base caption. It spells out entities, causes, interactions, governing principles, temporal evolution and effects.

The pipeline also writes a scene-specific physics_negative_prompt. This text describes likely implausible outcomes, such as a stone floating on water. It acts as negative conditioning at inference time.

How Does the Self-Evolving Caption Loop Work?

The loop keeps the captioner frozen and evolves only its instruction. A GPT-5.5 captioner writes captions for a fixed 20-video development set with 273 human-verified assertions. Gemini-3.1-Pro acts as a physics-aware critic. An evolution agent reads the scores and claim-level failures, then rewrites the prompt.

The critic scores 2 dimensions:

Precision: the caption is split into atomic claims, and each claim is checked against the video.

Recall: each human-curated physical assertion must be explicitly stated or entailed by the caption.

Every revised prompt is validated on PhysCapBench, a new benchmark of 246 videos and 3,794 human-verified assertions. Caption F1 rose from 78.64 at iteration 1 to 87.82 at iteration 9. The path was not smooth. Iteration 2 made captions overly cautious and dropped F1 to 76.28. Iteration 9 required every visible causal step and raised frame sampling from 2 fps to 4 fps.

How Does Language-Guided Data Curation Work?

A GPT-5.5 diagnosis agent maps generated-video failures to physics categories like rigid-body motion, collision and fluid dynamics. That deficiency profile is matched against physics tags on a large video gallery. Retrieval targets physical content, not visual appearance.

The final training set holds 183K videos: 71K filtered from WISA-80K plus 112K retrieved clips. Retrieval alone added 3.01 points on average across 3 benchmarks. On VideoPhy-2, chemical and thermal processes each gained 8.00 points.