Back to articles
World Models

LOCI Helps World Models Remember What They Saw Before

3 min read

Introduction

A video world model must do more than predict the next frame. It also needs to preserve a coherent representation of a scene over time. If a camera turns away from a location and later comes back, the model should reconstruct what was previously visible rather than inventing a new version of the environment. This is difficult because visual detail and long-term memory place competing demands on the architecture.

LOCI, short for a spatial linear-memory approach to streaming world models, addresses this tension with a hybrid design.

Two complementary memory mechanisms

The method divides the Transformer’s blocks between two forms of memory:

  • Half of the blocks retain a key-value cache of past observations. Direct access to cached entries helps preserve visual detail, but a conventional cache grows as the video history becomes longer.
  • The other half limits attention to the current chunk and adds a recurrent linear-attention memory. This creates a compact state that can be updated continuously while processing a stream.
  • Both memory reads and writes are conditioned on projective camera geometry. The model therefore uses viewpoint information when deciding where an observation belongs and which stored content is relevant.
  • Readouts from the recurrent memory are passed into later cache-backed blocks. These readouts give the blocks’ queries accumulated scene context before they retrieve detailed observations.

The central idea is to separate two jobs that are often forced into one mechanism: retaining high-fidelity evidence and maintaining a compact summary of a long history. A recurrent state alone may be efficient but can make individual observations difficult to access. A full key-value history preserves access but becomes increasingly expensive. LOCI attempts to combine direct retrieval with compressed context instead of choosing only one.

Reported results

The supplied material reports evaluations on the public MIND memory benchmark and on held-out recorded trajectories from Unreal Engine. In revisit scenarios, LOCI reproduced previously seen content more faithfully than representative world models and a same-recipe full-Softmax baseline.

When the full history was retained, the method reduced peak memory by about 30% at the same sequence length relative to full Softmax. With a bounded bank of retained observations, it could stream long videos using constant memory and remained more faithful than full Softmax under the same memory budget. These findings suggest that geometry-aware addressing may make a limited memory store more useful than a generic compressed state.

Why it matters

LOCI offers a practical design direction for long-horizon video generation and interactive world simulation. Cache-backed layers can preserve precise visual evidence, while recurrent linear memory carries scene context forward at lower cost. Camera geometry then provides a spatial signal for linking the current viewpoint with relevant past observations.

The available description does not establish how the method behaves under real-world camera noise, highly dynamic scenes, or substantially larger model scales. Still, the release of LOCI-revisit-data, containing game-engine video, exact camera poses, depth, and revisit pairs, gives researchers a resource for studying memory and consistency in world models. The project page and implementation are also available for further reproduction.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles