HeteroFold Enables KV Cache Sharing Across Model Families
Heterogeneous multi-agent LLM systems often assign different models to different roles. One model may plan, another may retrieve information, and a third may verify or execute the result. This arrangement can combine the strengths of several model families, but it also creates a costly communication pattern: when one agent passes a long context to another as text, the receiving model must prefill information that has already been processed by the sender.
KV cache reuse is an attractive alternative. During inference, a model stores intermediate Key and Value states for the processed context. Later decoding can reuse those states rather than recomputing the entire prefix. The difficulty is that a cache is not a universal representation. Different model families may use different tokenizers, layer counts, attention structures, KV dimensions, and internal representation spaces. A tensor produced by one model therefore cannot simply be inserted into another model.
The paper introduces HeteroFold as a prefill-free approach to this cross-family transfer problem. Its design addresses several mismatches at once. It aligns token boundaries across tokenizers, establishes correspondence between layers in different architectures, and maps the sender’s K/V states into the receiver’s representation space. A calibration stage is then used to preserve the receiver’s attention patterns and outputs as closely as possible. Both foundation models remain frozen; the learned components serve as the bridge between their internal states rather than modifying either model.
The evaluation covers six transfer directions among Llama, Qwen, and Ministral. The authors test long-context and short-context settings as well as a multi-agent benchmark. They report the best cache-transfer results on all four long-context benchmarks and on most short-context configurations. At a 32K context length, the Llama 3.1 8B to Ministral 3 14B direction is reported to be about 10.7 times faster than native receiver-side prefill, and 1.18–1.47 times faster than the prefill-free baselines Dense Latent and KV Ridge. On the multi-agent benchmark, HeteroFold is reported to match text-based communication.
The broader significance is that KV cache could become an exchangeable computational artifact rather than a state locked inside one model family. For agents that repeatedly share long documents, task histories, or environmental context, this could reduce duplicated computation and make model composition more flexible. It also changes how communication efficiency is viewed: instead of compressing context back into text, a system can attempt to transfer a representation closer to the computation already performed.
There are still practical questions. Cross-family conversion requires mapping and calibration, so production systems will need to account for their overhead, cache movement costs, and robustness under task distributions not represented in evaluation. The supplied material does not establish that the reported gains generalize to every model size or deployment setting. HeteroFold is better understood as evidence that direct cache exchange across model boundaries is feasible, and as a step toward more efficient heterogeneous agent collaboration.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...