CANOPY Gives Multimodal RAG Adaptive Evidence Granularity
Multimodal retrieval-augmented generation is often described as a search problem: find relevant text, tables, images, or videos and pass them to a reader model. In practice, retrieval is only half of the challenge. A whole document or video may contain useful evidence alongside a large amount of irrelevant material, while aggressive fixed-size chunking can remove the context needed to interpret a fact. CANOPY addresses this second decision: how much of each retrieved item should remain in the reader’s context.
How the framework works
- A shared hierarchical representation. Retrieved items are organized as trees of original regions. Text can be refined from documents to passages and sentences, while other heterogeneous items can expose their own natural regions and levels. This gives the compressor more options than a single fixed chunk size.
- Evidence-aware node scoring. A node encoder, fine-tuned on gold evidence, scores regions against the user query. The goal is not simply to identify a relevant item, but to estimate which part of that item supports the answer.
- Parent-relative refinement. CANOPY compares a region with its parent instead of making every node an isolated decision. A broad parent can preserve necessary context, while a particularly useful child can be retained at a finer level. The resulting evidence set may therefore contain regions at several granularities within one item. Node-level pruning does not require an LLM call for every candidate.
- Retrieval when compression is not enough. A compressor cannot recover evidence that was never retrieved. CANOPY therefore adds a critic that checks whether the accumulated evidence appears sufficient. If it identifies a gap, the system issues targeted follow-up retrieval and compresses the newly found items before adding them to the reader input.
The paper evaluates the framework on five QA benchmarks over a heterogeneous corpus containing 33 million items. CANOPY achieves higher average answer accuracy than the evaluated retrieval baselines. In an unrouted setting using Qwen3-VL-8B-Instruct as the reader, compression reduces reader-input evidence tokens by 14.2%–27.7% relative to the same iterative setup, while maintaining comparable answer quality as reported in the paper. Ablations suggest that, for multi-hop QA, the main accuracy gains come from additional targeted retrieval rather than compression alone.
Why it matters
The contribution is a tighter coupling between retrieval and context management. Retrieval expands the pool of potentially useful evidence; hierarchical compression controls how much of each result reaches the reader; and the critic handles cases in which the initial pool is incomplete. This offers a common procedure across heterogeneous modalities instead of relying entirely on separate compression mechanisms for text, tables, or video.
The approach also exposes important trade-offs. Its success depends on the initial retriever finding the relevant evidence and on the hierarchy preserving meaningful structure. Follow-up retrieval can improve multi-hop coverage, but it also introduces extra computation and latency. CANOPY is therefore best understood as a dynamic balance among evidence completeness, reader context size, and retrieval cost—not as a rule to minimize tokens at any price.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...