Back to articles
World Models

Physis-Lang Uses Evolving Physical Language to Improve Video World Models

3 min read

Introduction

A video world model should do more than generate images that look coherent. It should anticipate how objects and environments change over time, including the consequences of collisions, forces, contact, and other physical interactions. Current video generators can produce visually convincing sequences while still violating basic physical rules. Physis-Lang explores whether the problem can be addressed not only with additional latent, numerical, visual, or planning signals, but also with a better use of language.

What the framework does

  • It treats physical language as a shared representation. Instead of describing only visible content, Physis-Lang asks captions to identify relevant entities, causes, interactions, governing principles, temporal evolution, and effects. The resulting text is intended to explain a process, rather than merely label a scene.
  • It introduces assertion-level evaluation. PhysCapBench decomposes physical processes into atomic assertions and evaluates captions with recall and precision. This makes it possible to distinguish between missing physical facts and unsupported details added by a captioning system.
  • It uses an iterative agentic loop. The framework examines errors at the assertion level and refines the instruction used to produce physical captions. Language is therefore treated as an optimizable representation instead of a fixed prompt designed once by hand.
  • It links model weaknesses to data retrieval. Deficiencies observed in a model are converted into textual descriptions. Language-guided retrieval then searches for visually diverse videos that cover the missing physical processes, connecting data curation with the model’s current failure modes.

Results and implications

The paper evaluates the approach on four widely used physical-video benchmarks with Wan and Cosmos backbones. According to the supplied material, Physis-Lang produces consistent improvements in physical plausibility across these settings. A notable comparison starts with open-source Cosmos3-Nano backbones: the enhanced models are reported to surpass the proprietary Veo 3.1 model in the paper’s evaluation.

The broader contribution is the framework’s closed loop across data, training, and generation. The same language representation can describe what a video contains, expose what the model fails to understand, guide the selection of additional training material, and help organize the desired physical evolution. This provides a relatively interpretable alternative to relying exclusively on larger visual models or opaque latent signals.

The results should still be read within the scope of the reported benchmarks and backbone combinations. It remains to be established whether the approach transfers reliably to complex three-dimensional interactions, long-horizon dynamics, and open-world scenes. Even so, Physis-Lang advances a useful research question: improving a world model may require not only more video, but also language that states the causal and temporal structure of that video with greater precision.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles