Back to articles
World Models

HLA-WM Uses Addressable Memory to Reduce Forgetting in Long-Horizon Video World Models

3 min read

Introduction

Long-horizon video world models must do more than generate plausible frames one step at a time. They need to preserve spatial layouts, object positions, and camera motion over extended rollouts. When the camera returns to a previously visited area, missing or distorted memory can lead to geometric drift, inconsistent objects, and inaccurate view control.

The paper HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models addresses this tension with a training-free hybrid attention framework. Its goal is to retain the efficiency of recurrent linear attention without accepting its weakest behavior: the gradual loss of distant but relevant scenes.

Key points

  • Two memory strategies expose a trade-off. Softmax attention keeps the complete history in a growing KV cache, providing direct access at a substantial memory cost. Gated DeltaNet (GDN) compresses history into a fixed-size state, but later updates can progressively attenuate information that is no longer recent.
  • Retrieval is guided by camera geometry. HLA-WM divides the history into chunks and uses the affine structure of GDN to cache compact transition summaries. During generation, camera geometry is used to identify historical chunks that are likely to correspond to the current view.
  • The retrieved history is recomposed. Rather than simply concatenating old tokens, the method reconstructs a recurrent state tailored to the current query. This lets the model selectively restore scene-relevant information while keeping the linear-state computation.
  • Reported gains cover both consistency and control. On the 60-second SANA-WM-Bench, the method improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator. The reported gains include a 0.74 dB PSNR increase and a 28.5% reduction in rotation error.
  • The effect survives additional refinement. Improvements remain after downstream refinement and transfer to MBench-A, where all three revisit-consistency metrics improve across four subsets and all evaluated inference modes over 547 samples.
  • The memory savings are substantial. At a 60-second context, HLA-WM reduces historical-state memory by 12 times relative to full KV caching, while reducing inference throughput by no more than 1.6%.

Why it matters

The central idea is that long-term memory does not have to be either fully retained or irreversibly compressed. In camera-controlled video generation, geometry can serve as an address for selective memory access. HLA-WM therefore combines coarse retrieval of likely relevant history with fine-grained recurrent computation for the current query.

This design offers a practical inference-time route to better long-range consistency without retraining the generator. It is particularly relevant to world models that must revisit locations, maintain stable spatial structure, or follow extended camera trajectories. At the same time, the reported results are tied to the evaluated base models, benchmarks, and retrieval setup. The quality of geometric matching and chunk summaries may affect performance in other environments. Even so, the work illustrates how memory organization—not only model scale or additional training—can be used to extend the effective horizon of video world models.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles