Back to articles
Vision & Video

FrameMorrow Helps Video Models Select History for the Future

3 min read

Introduction

Long-horizon video generation creates a memory problem. As a sequence grows, the amount of available history increases, but feeding every previous frame back into the generator becomes expensive and increasingly redundant. Compressing the history can remove important details such as a character’s identity, the layout of a scene, or an object that will matter again much later.

FrameMorrow focuses on a limitation of many existing selection strategies: they judge relevance mainly from the current content. A frame that looks unimportant now may become essential after a future action or scene transition. The paper therefore proposes selecting history according to anticipated future information needs.

Key idea

  • From present relevance to future utility: Instead of asking only which past frames resemble the current state, FrameMorrow asks which historical information may be needed next.
  • Prospective tokens: The method predicts a small set of tokens that summarize what could become important in the near future. It does not need to generate the entire future sequence.
  • Explicit frame retrieval: The tokens guide the selection of actual historical frames. This keeps the interface independent of a generator’s internal hidden states.
  • Plug-and-play integration: Because the selected memory is expressed as frames, the method can be connected to different video generators, including closed-source systems whose internals are inaccessible. The abstract also describes the added inference cost as small.

Evaluation and implications

The paper reports experiments on five benchmarks and eleven generative models. The settings include long-video generation, interactive generation, and action-conditioned world models. According to the supplied abstract, FrameMorrow consistently improves long-range consistency and visual quality across these settings. The available material does not include detailed metric values, baseline breakdowns, or ablation results, so the strength of each individual component cannot be assessed from the source alone.

The broader contribution is a different view of memory for video models. Memory is not simply a larger cache of the past, nor is it only a compressed summary of the current context. It is a selective process shaped by what the model is likely to need later. In narrative video, this could help preserve long-term properties of people and environments. In interactive generation, it may help recover an earlier state after a user action. In action-conditioned world models, distant visual history can influence how a later action is interpreted.

The approach also has an important open question: prospective tokens are predictions, and incorrect predictions may cause useful frames to be dropped. Future work will need to examine uncertainty, competing possible futures, and selection stability over even longer horizons. Still, FrameMorrow offers a model-agnostic interface for turning future expectations into practical visual memory.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles