Back to articles
World Models

World Embedding Benchmark Tests Whether Video Models Learn Physics

3 min read

Introduction

As video generation and world models advance, visual plausibility is no longer the only target. A generated scene may look convincing while violating basic rules of motion, contact, fluid behavior, or light. This raises a more fundamental question: do video representations actually encode physical information, and can that information be accessed in a useful way? The World Embedding Benchmark is designed to make that question measurable.

A controlled test of physical representations

The benchmark includes 8,000 controlled simulation cases organized into 80 physical families. Its coverage spans fluid mechanics, solid mechanics, dynamics, and optics and electromagnetism. Each case pairs a rendered video with physical annotations derived from the underlying simulation. This setup gives researchers direct control over parameters and visual conditions, making it easier to separate genuine physical understanding from superficial visual similarity.

The evaluation has three complementary components:

  • Text-video retrieval, which tests whether a video can be matched with a description of the relevant physical process;
  • Physical-property regression, which measures whether quantitative simulation properties can be recovered from frozen video embeddings;
  • Video-description pair classification, which asks a model to select the description that correctly explains a video.

These tasks probe different capabilities. Retrieval and pair classification focus mainly on cross-modal physical alignment. Regression instead asks whether numerical physical information remains recoverable inside the representation.

The central finding: alignment is not the same as information

Pre-trained omn||||modal embedding models perform weakly on physical retrieval. Their within-family pair classification is close to chance, suggesting that broad visual competence does not automatically produce reliable links between a video and a precise physical explanation.

At the same time, lightweight probes can recover useful physical information from frozen video embeddings. This is an important distinction. A representation may contain signals related to a physical quantity without organizing those signals in a form that can be aligned with language or directly used by another model. In other words, information can be present but operationally inaccessible.

The researchers also apply continual contrastive training with physics-specific video-text pairs. This adaptation improves retrieval and pair classification, but physical-property regression becomes worse. The result points to a trade-off between semantic alignment and quantitative recoverability. Optimizing an embedding to match language may reshape its internal geometry in ways that help cross-modal search while discarding or obscuring fine-grained numerical structure.

From benchmarking to video generation

The study further tests whether better physical representations can benefit generation. The embeddings retrieve physically relevant reference videos for MiniMax-H3 in a retrieval-augmented generation pipeline. According to the reported experiments, the references improve the physical fidelity of generated videos, and stronger retrieval models produce larger gains. The physics-adapted LCO-Embedding-Omni delivers the strongest performance in this retrieval-augmented setting among the systems described in the material.

The broader implication is that world-model evaluation should not rely on a single score. Visual quality, language alignment, and physical property recovery measure different aspects of a representation. A model that excels at one may regress on another.

By releasing the simulation engines, benchmark data, evaluation code, and trained checkpoints, the project aims to make these distinctions easier to study. For future world models, physically grounded embeddings could serve as an intermediate layer connecting perception, retrieval, and generation—provided that alignment improvements do not come at the expense of the quantitative information needed for reliable prediction.

Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles