Back to articles
Memory & Context

Chunked KV-Cache Compression Creates Periodic Retrieval Weak Spots

3 min read

Long-context systems are often evaluated as if information is equally retrievable wherever it appears in a prompt. A new study challenges that assumption for models using chunked KV-cache compression. When consecutive tokens are merged into fewer cache entries, retrieval can depend on a token’s position relative to the compression-window boundary—even when the text itself has not changed.

Compression introduces a new coordinate

Suppose a model compresses windows of eight tokens at a stride of four. Every token then has a phase: its position modulo the stride. Adding one irrelevant token at the beginning of a prompt changes these phases without changing the content or order of the original text. Any resulting performance shift is therefore attributable to the model’s memory layout rather than to the data.

The researchers demonstrate this with DeepSeek-V4-Flash-Base. They ask the model to complete the same FP8 kernel expression while repeatedly lengthening an unrelated docstring. The code remains unchanged, yet the top prediction alternates between the correct 8 and 32 on a cycle of roughly four tokens. A comparable non-chunked model does not show the same systematic preference.

Average scores can hide large failures

In a needle-in-a-haystack evaluation, prompt groups were matched for query position, target-position statistics, and the number of key-value records. Only the target’s compression phase differed. Across DeepSeek-V4 variants, retrieval accuracy rose and fell with the stride, producing gaps of roughly 15 to 40 percentage points between the best and worst phases. Post-training improved accuracy and reduced the gap, but did not remove the periodic pattern.

The effect was also reproduced in transformers trained from scratch, where chunked compression was the major architectural difference from full attention. Across changes to window size, stride, KV-head count, gating, and positional encoding, the period continued to track the stride. In controlled experiments, a compressed model could have a mean score close to a full-attention baseline while its worst position performed dramatically worse. Some configurations produced phase gaps as large as 78 points.

How attention heads specialize

Causal interventions suggest that attention components do not contribute uniformly across the cycle. Individual heads can matter strongly for only a few adjacent phases, while other heads cover different portions of the period. The compression gates also tend to select similar slots in every window, with a consistent offset between key and value preferences that helps preserve neighboring token relationships.

A particularly revealing test shifted the gate parameters cyclically while leaving the remaining weights unchanged. The weak spots moved with the gates, indicating that the periodic pattern is tied directly to the compression layout. Analysis of an idealized retrieval model further suggests why training may produce this behavior: gradient dynamics favor concentrating a head on a fixed slot, but do not necessarily encourage a balanced allocation across all phases.

What should change in evaluation?

Chunked compression remains useful because it reduces cache memory and attention cost. The result is not that compression should be abandoned, but that its reliability must be measured more carefully. Long-context reports should include per-phase accuracy, the best-to-worst phase gap, and sensitivity across window and stride settings.

This matters for code completion, document question answering, and retrieval-augmented generation. A harmless prefix change can shift a target into a weaker phase and alter the answer. A model that looks strong on average may therefore contain predictable periodic failure zones.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles