Back to articles
Memory & Context

Giving AI 50 Million Tokens of Reusable Memory Through Persistent KV State

3 min read

Introduction

Long-context inference has two separate costs: fitting history into a usable context and repeatedly processing that history. A paper featured by Hugging Face Daily Papers tests galahad-kv, a memory layer that persists the key-value state produced by a language model. Text is divided into blocks of roughly 16,000 tokens, stored on encrypted local NVMe, and loaded again when needed.

The proposal is therefore not a new model with permanent memory. It is a way to preserve intermediate inference state so that repeated requests do not have to rebuild it from raw text.

Key findings

  • A large end-to-end test: Using vLLM on a single NVIDIA H100, the authors processed 50 million tokens of real public text with Gemma 4 12B and Gemma 4 31B.
  • Byte-exact recovery: They probed 100 blocks at depths ranging from the beginning of the stream to 50 million tokens. Every probed block was loaded from the encrypted store without recomputation on both models.
  • Lower serving overhead: Loading a block was reported to be 2.8 to 4.3 times faster than recomputing it, while GPU energy use was 8.8 to 12.3 times lower. GPU memory remained flat as the stream grew.
  • Useful, but not perfect, recall: When asked about facts planted millions of tokens earlier, the 12B model answered correctly 82 out of 100 times and the 31B model 98 out of 100 times. The supplied results state that neither model fabricated an answer.

What this changes

A conventional long-history workflow may resend the same documents on every request. The model then performs the prefill computation again, consuming time and GPU resources. Persistent KV state moves that cost to the first pass: once a block has been processed, its internal representation can be reused later.

That distinction matters for agents, document systems, codebase analysis, and long-running conversations that repeatedly consult the same material. Stable GPU memory is also attractive for serving systems, because a growing history does not necessarily require keeping every token resident on the accelerator.

Why it is not a 50-million-token attention window

The stored state is accessed as memory, not presented as one uninterrupted attention sequence. The experiment loads one block at a time, and the quality of an answer depends on whether the relevant block is selected and whether the model can use it effectively. Persistent state therefore complements retrieval and orchestration; it does not remove those requirements.

There are also substantial trade-offs. Writing the memory still requires the original computation, and the store can require terabytes of local NVMe capacity. Results may vary with model architecture, block layout, query type, and selection policy. The study is best read as an engineering demonstration of reusable inference state, rather than proof of unlimited model memory.

For practitioners, the more meaningful benchmark is not context length alone. It should include first-write cost, storage footprint, load latency, energy use, retrieval behavior, and answer accuracy over time.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles