Back to articles
World Models

QuantWM Preserves Attention in 2-Bit KV Cache Quantization for Video World Models

3 min read

Video world models often retain Key-Value (KV) states throughout generation so that later frames can refer back to earlier spatial and temporal information. This mechanism supports long-range consistency, but the cache grows as generation proceeds and can become a major memory bottleneck during deployment. Reducing the cache to 2-bit representations appears attractive. Yet the material behind QuantWM points out an important gap: methods that look nearly lossless on general benchmarks can still produce severe temporal flickering and visible degradation when used in video world models.

QuantWM reframes the problem around attention preservation. A straightforward quantization analysis might focus on reconstruction error in the cached tensors. The paper finds that this measure does not fully predict output quality. In particular, Key quantization can have a smaller reconstruction error than Value quantization while causing a larger degradation in the generated video.

The reason lies in the role of Keys inside attention. Perturbing a Key changes the attention logits computed with Queries. These changes can alter which tokens are selected across both time and space. Once the model attends to a different set of historical tokens, small numerical deviations can become visible as frame-to-frame instability or visual deterioration. In other words, the important objective is not simply to make every quantized value close to its original number, but to preserve the attention decisions that organize the video’s temporal information.

QuantWM introduces two complementary components:

  • Quantization-sensitivity-aware clustering (QSAC): It jointly considers the sensitivity of historical Queries and the range of residuals when selecting INT2-friendly Key centroids. The design gives greater protection to channels that matter more for attention.
  • Principal-subspace attention compensation (PSAC): It compensates for remaining Key errors along the dominant subspace of the Queries, aiming to recover attention relationships that were shifted by quantization.
  • Training-free deployment: The framework is designed to operate during KV cache quantization without requiring an additional training stage.

The broader lesson is that cache compression for video generation should be evaluated at the level of model behavior, not only tensor-level distortion. A low average reconstruction error does not guarantee stable token selection, and stable token selection may be more important than uniform numerical precision when the model depends on long temporal context.

The available material does not include complete experiment tables, exact memory savings, or detailed comparisons across model families. Those results would be needed to assess the practical deployment advantage of QuantWM. Still, the paper identifies a useful direction for efficient video inference: preserve the attention path first, then optimize the representation around it. This perspective may also inform future cache compression methods for other generation systems with strong sequential dependencies.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles