Back to articles
Memory & Context

MemFold Trains Soft Memory for Long-Context Personalization

3 min read

Introduction

A personal assistant that serves the same user over a long period must do more than retrieve old statements. It needs to determine which preferences remain valid, which have been revised, and which constraints matter for the current request. Keeping every relevant interaction as text makes the reader’s context grow with the history. Compressing the history into a fixed number of latent vectors controls the interface size, but common training objectives focus on reconstructing text or imitating reference answers rather than on the behavior produced by the compressed memory.

MemFold starts from a different premise: a memory representation should be judged by what the model can do with it.

How the method works

  • A fixed-budget interface: Given a query, a textual memory is made query-conditioned and compressed into K continuous vectors. These vectors become the reader model’s memory interface, limiting the amount of information passed into the model regardless of the full history length.
  • Optimization on student rollouts: The reader generates its own responses under the soft memory. Training therefore evaluates the distribution the student actually uses, rather than only sequences supplied by an external reference.
  • Two complementary objectives: Group-relative rewards provide a signal about task outcomes across sampled candidates. A second signal uses confidence-gated on-policy distillation: a frozen teacher with access to the textual memory re-scores tokens already sampled by the student, and guidance is applied selectively when the teacher is sufficiently confident.
  • No teacher decoding at inference: The teacher is used for scoring during training, not for generating an additional response. At deployment, it is removed entirely, leaving the reader and the compact soft memory.

What the reported results suggest

Across three Qwen backbones, the paper reports the highest measured accuracy for MemFold on PersonaMem-32K and PersonaMem-128K. The reported margins become wider with the longer history setting. The method also transfers to PrefEval and LongMemEval without target-domain training. Taken together, these results suggest that a bounded memory interface does not necessarily prevent effective personalization, provided that compression is optimized for downstream decisions rather than for textual fidelity alone.

The design also connects two roles that are often separated. The textual-memory teacher preserves a richer view of the user history, while the student must learn to act through a restricted representation. This creates a practical training loop for long-running assistants and agents: retain a richer source during optimization, but deploy only the compact representation needed for execution.

The available material does not include the complete ablation tables, exact improvement margins, or details for different memory budgets. Those omissions leave open questions about compute cost, robustness to changing preferences, and how much performance depends on the selected value of K. Still, MemFold’s central contribution is clear: it reframes memory compression as an action-oriented problem. The goal is not merely to preserve a readable summary, but to preserve the information that helps the model make the right decision now.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles