WUSH-KV Uses Data-Adaptive Transforms for Low-Bit KV-Cache Quantization
Introduction
KV caches are a central systems bottleneck for long-context language-model inference. Every generated token must access historical keys and values, so cache size and memory traffic grow with both context length and batch size. Lowering the precision of cached tensors can reduce that burden, but aggressive quantization also introduces errors into attention computation.
WUSH-KV approaches the problem by changing the representation before quantization. Rather than applying a generic rotation or a fixed scaling rule, it uses calibration data to find transforms that better match the statistics of the tensors involved in attention.
What the method changes
- An adaptation of WUSH: The original WUSH idea builds a data-aware transform from the second-order statistics of both factors in a matrix product. WUSH-KV adapts this principle to the two components of the KV cache.
- Separate treatment of keys and values: Calibration data is used to construct distinct key and value transforms. This reflects their different roles in attention instead of forcing both tensors through one shared transform.
- Placement aligned with the computation graph: The value transform can be folded into model weights, while the key transform is applied after RoPE. This placement is intended to preserve the relevant structure of the attention pipeline.
- Compatibility with clipped quantizers: The transforms can be combined with clipped quantization schemes. For QuEST INT, the paper further argues that WUSH is near-optimal under mild assumptions.
Reported results
The paper evaluates both layerwise reconstruction error and end-to-end perplexity. With QuEST INT, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among the transforms tested in the reported comparison.
For a serving-oriented evaluation, the authors integrate WUSH-KV into SGLang with OSCAR-style percentile-clipped affine quantization. At 2-bit precision, WUSH-KV performs comparably to or better than the OSCAR transform across all evaluated models and downstream tasks. The provided material does not list the exact models, tasks, or metric values, so the appropriate takeaway is comparative robustness rather than a claim of universal superiority.
Why it matters
The main contribution is the connection between quantization and the structure of attention. The method accounts for the distinct roles of keys and values, the location of RoPE, and the possibility of folding a transformation into existing weights. This suggests that reducing KV-cache error is not only a matter of choosing more bits, but also of choosing a representation in which those bits are more useful.
There are practical considerations as well. WUSH-KV requires calibration data and offline construction of transforms. Weight folding and runtime placement must also be supported by the serving stack. Its real-world benefit will therefore depend on hardware support for low-bit operators, transform overhead, and cache access patterns. Still, the reported 2-bit results make data-adaptive transforms a promising direction for memory-efficient long-context inference.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...