Back to articles
Inference & Serving

SparseDecoding Makes LLM Pruning Aware of Generation

3 min read

Introduction

The main bottleneck in large language model inference is not always arithmetic throughput. During decoding, a model typically produces one token at a time, but repeatedly reads a large set of weights for every step. This makes memory traffic and bandwidth major contributors to latency. Pruning can reduce the number of nonzero parameters that must be read, yet the practical benefit depends on whether the pruning objective and the execution kernel reflect the actual decoding workload.

A paper from the Westlake ENCODE Lab presents SparseDecoding, a framework that addresses both sides of this problem. It aligns pruning calibration with the activations observed during autoregressive generation and adds system support for sparse matrix-vector multiplication, an operation central to token-by-token decoding. Experiments cover Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B, and Qwen3-32B, with a reported maximum 1.48× end-to-end decoding speedup on A100 GPUs.

Why conventional calibration can miss the target

Many training-free pruning methods estimate weight importance using Hessian-related information collected from pre-existing natural-text sequences. This is convenient, but it assumes that the activation distribution in those sequences is a good proxy for the distribution encountered during generation.

The two settings are not identical. In autoregressive decoding, the model consumes tokens that it generated itself, and the sequence evolves according to its previous predictions. The resulting hidden states and layer activations can therefore diverge from those produced by a fixed natural-text calibration set. According to the paper, this mismatch can make the pruning objective less representative of generation and degrade the quality of the pruned model.

SparseDecoding addresses the issue by collecting layer-wise activations while the dense model performs autoregressive generation. The prefill stage is excluded, so the calibration process focuses on the repeated token-generation steps where the target speedup is expected. The resulting calibration matrices are intended to reflect the distribution that the pruned model will encounter during decoding.

Key points

  • Decoding-aware calibration: Layer activations are gathered from dense-model generation instead of relying solely on fixed natural text.
  • Focus on the relevant operation: Decoding is dominated by sparse matrix-vector multiplication (SpMV), whereas many earlier systems primarily optimize sparse matrix-matrix multiplication (SpMM).
  • Algorithm–system co-design: The authors implement an N:M sparse SpMV kernel using bitmask indexing and fixed-step traversal to reduce sparse-access overhead.
  • Multiple model families: The evaluation includes several Llama and Qwen models and compares against standard fixed-text calibration on long-form generation benchmarks.

Why it matters

The broader lesson is that parameter reduction does not automatically translate into lower wall-clock latency. If a runtime cannot execute irregular vector-level sparse operations efficiently, a lower nonzero count may produce limited real-world acceleration. SparseDecoding therefore treats the calibration distribution and the underlying SpMV implementation as parts of the same deployment problem.

The reported results should still be interpreted in context. The speedup is tied to particular models, an N:M sparsity pattern, and A100 GPUs; actual gains may vary with hardware, batch size, sequence length, and inference-stack implementation. Nevertheless, the decoding-aware perspective is useful for systems whose primary workload is token-by-token generation. For such deployments, collecting calibration signals from the generation process itself may be more relevant than simply increasing sparsity on the basis of unrelated fixed sequences.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles