Back to articles
Vision & Video

FlashForward Speeds Up Long-Video Diffusion with In-Flight KV Caches

3 min read

Generating a long video with diffusion is difficult for two related reasons: each new temporal chunk must be produced efficiently, and it must remain consistent with everything that came before it. Few-step autoregressive systems address the length problem by dividing a video into temporal chunks and denoising those chunks one after another. The process still needs a memory of previous content, commonly represented by key–value (KV) caches inside attention layers.

Reusing computation that already exists

Earlier approaches can perform additional model forwards to reconstruct clean or less-noisy historical KV states. Those passes update the memory without advancing the output latent, which makes them useful for quality but expensive for inference. FlashForward observes that every ordinary denoising forward already computes an in-flight KV cache for the current chunk. Rather than discarding that intermediate result and rebuilding the cache later, the method makes it available to the next chunk as soon as the current chunk finishes a denoising stage.

This early availability enables a pipeline across denoising stages. Different GPUs can host different stages, allowing separate temporal chunks to occupy the pipeline concurrently. In this design, the system avoids a substantial portion of cache-update-only forwards and turns intermediate attention states into a scheduling resource.

The quality cost of early memory

The cache becomes available earlier, but it is not as clean as the history reconstructed after denoising. If the generator relies only on this stage-matched history, noise can accumulate across chunks. The paper identifies the resulting risks as appearance drift and motion drift: an object may gradually change its visual identity, while movement may become less coherent over longer durations.

FlashForward complements the noisy dense history with sparse clean anchor latents. These auxiliary latents are created before the corresponding region is generated and are used to form clean anchor KV states. The result is a two-sided conditioning scheme:

  • Dense stage-matched history preserves recent evolution and fine local continuity, but contains more noise.
  • Sparse clean anchors provide coarser, longer-range structural and appearance guidance.
  • Combined conditioning helps keep the current generation trajectory near a stable path without requiring every historical state to be rebuilt cleanly.

The paper evaluates the approach with 1.3B and 14B backbones, at 480p and 720p, on 16 FPS videos of at least 20 seconds. With up to four GPUs, it reports a 1.16–1.69× speed advantage over HiAR and a 1.42–2.92× advantage over Self-Forcing. For the 1.3B model at 480p, the authors also report higher VBench scores and stable results across 20-, 35-, and 65-second videos.

Why it matters

The broader contribution is a different view of KV-cache management. A cache does not have to be fully clean before it becomes useful; an intermediate, stage-specific state can support both computation and pipeline scheduling. Clean anchors then compensate for the quality limitations of that early memory. This separation of responsibilities—recent dense history for local dynamics and sparse clean memory for long-range structure—offers a practical way to balance throughput and temporal consistency.

The reported gains should still be interpreted in context. They depend on the few-step autoregressive diffusion setup, the selected backbone models, and a multi-GPU pipeline. Anchor density, stage partitioning, memory traffic, and implementation overhead may affect results in other deployments. FlashForward is therefore best understood as a cache-and-scheduling strategy for long-form video inference, rather than a universal acceleration method for every video generator.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles