V-CoLA Compresses Vision Tokens for Linear Attention Models
Introduction
Vision-language models (VLMs) can connect images and text, but their visual inputs often expand into long sequences of vision tokens. These tokens may dominate the prompt length and substantially increase computation and memory use during inference. Compressing them without removing information that matters to a downstream task is therefore an important route toward more efficient multimodal systems.
V-CoLA, presented in the paper “V-CoLA: Vision Token Compression with Linear Attention,” focuses on a limitation of existing compression techniques. Many earlier approaches were developed around Softmax attention. As hybrid models increasingly incorporate linear attention, including the Qwen3.5 example cited by the authors, attention-based and similarity-based selection rules may no longer transfer reliably. The paper’s analysis finds that both categories can suffer noticeable degradation in this setting.
How V-CoLA works
V-CoLA is a training-free framework built specifically for linear-attention models. Its design has three main aspects:
- Uniqueness-aware importance: Instead of treating importance as a measure of salience or similarity alone, V-CoLA also considers whether a token contains information that is not duplicated by nearby tokens. A visually distinctive detail may be important even when it does not receive a high score from a conventional attention or similarity heuristic.
- Adaptive token merging: Tokens selected for compression are merged rather than simply discarded. This gives the method a way to shorten the sequence while preserving a broader summary of the visual content.
- Implementation-level compatibility: The authors optimize the compression operations to work with the chunk-wise parallelism used by linear attention. This is important because an algorithmic reduction in token count does not automatically translate into a real speedup if the implementation disrupts the model’s execution pattern.
Reported results
Across multiple benchmarks, the paper reports that V-CoLA retains 99.5% of the original performance when only 50% of the vision tokens are kept. At 12.5% of the original token count, performance remains above 88%. The reported prefill speedup ranges from 1.86× to 6.15×.
These results suggest that visual redundancy can be substantial, but the useful information is not necessarily identified by the same signals used in Softmax-attention systems. A compression method that distinguishes redundant content from unique content can potentially reduce sequence length without causing an equivalent loss in task performance.
Why it matters
The broader contribution of V-CoLA is to treat token compression as an architecture-specific problem. As VLMs combine different attention mechanisms, compression methods may need to account for how information is aggregated rather than relying on a universal ranking rule. The work also has practical appeal because it does not require additional training, which can simplify adoption in existing inference pipelines.
The available summary does not establish how the method behaves across every vision encoder, task, or model family. More detailed evaluations would be needed to understand those boundaries and the trade-off at intermediate compression levels. Even so, V-CoLA points to a useful direction for lowering multimodal prefill costs while preserving most of the original model capability.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...