Back to articles
Memory & Context

REMORY Adds Residual Memory to Long-Horizon Context Compaction

3 min read

Long-horizon agents eventually face a simple systems problem: their history grows faster than the available context window. Conversations, observations, tool calls, and intermediate results must be compressed before the agent can continue. A textual summary is an obvious solution, but it is not guaranteed to preserve every detail that may matter for a later decision. REMORY addresses this gap by adding a learned residual memory stream rather than relying on the summary alone.

How REMORY works

The method uses a neural memory network that receives the original history together with a compacted summary. It produces a bounded sequence of soft memory tokens. These are continuous representations rather than natural-language notes intended for human inspection. The tokens are appended after the summary and passed to a frozen large language model.

The sequence arrangement is the key idea. The summary acts as the explicit, readable backbone of the compressed context, while the soft tokens provide a learned correction for information that the summary does not state directly. REMORY is trained so that the frozen model, when given the summary plus residual memory, approximates the continuation it would have generated from the full history. In other words, the memory network is optimized for downstream behavior rather than for producing a text summary that merely resembles the source.

Reported results

  • On SummHay, REMORY improves source attribution while keeping insight coverage nearly unchanged.
  • It approaches the full-context joint score using only 5.2% of the input positions in the reported setting.
  • Qwen3.8-27B and GLM-5.3-Flash both show consistent improvements across long-horizon agent benchmarks.
  • On BrowseComp and Terminal-Bench 2.1, both models produce substantially fewer repeated tool outputs and tool errors when residual memory is enabled.

Why it matters

The significance of REMORY is not simply that it makes a prompt shorter. It introduces a two-layer interface for compressed context. Human-readable text can preserve an explicit account of what happened, while learned tokens can encode behavioral state that is difficult to express reliably in prose. This combination may be more suitable for agents whose future actions depend on small, distributed clues in a long interaction.

The reduction in repeated tool outputs and tool errors is also notable. Tool mistakes do more than lower task accuracy: they consume context space and can create additional, redundant history. If a memory mechanism helps the model avoid such loops, it may improve both agent reliability and the effective lifetime of a context window.

At the same time, the available material does not establish how well the learned tokens transfer across tasks, how their budget should be selected, or what training and serving overhead the memory network introduces. Soft memory may also be harder to inspect than ordinary summaries. REMORY should therefore be viewed as a promising design for context compaction, not as proof that long-term memory has been solved. Its broader value will depend on whether the gains remain stable across models, workloads, and changing histories.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Giving AI 50 Million Tokens of Reusable Memory Through Persistent KV State
Memory & Context
cctest.ai
Memory & Context

Giving AI 50 Million Tokens of Reusable Memory Through Persistent KV State

An experiment with galahad-kv shows that a model’s key-value state can be stored on encrypted local NVMe and loaded later without recomputation. The approach does not create a wider attention window, but it offers a practical way to reuse very long histories more efficiently.

Read more