Back to articles
Inference & Serving

SlimWise Prunes MoE Experts Only During Decode

3 min read

Introduction

Mixture-of-experts models reduce per-token computation by activating only a small number of experts. That advantage does not automatically translate into efficient serving, however. During batched autoregressive decoding, different requests can route to many different experts, causing the system to fetch or move weights across a large portion of the expert pool. In this regime, expert-weight traffic can become a more important bottleneck than arithmetic computation.

SlimWise addresses the problem by treating prefill and decode as different serving phases rather than applying one pruning policy to both. Its basic design is simple: use the full model to process the input context, then switch to a pruned model for token generation.

How it works

  • Full-model prefill: SlimWise leaves the expert pool intact during prefill. This phase is relatively compute-oriented, so removing experts may offer limited throughput gains while unnecessarily affecting quality.
  • Pruned decoding: Once generation begins, the system uses a smaller expert pool. Reducing the number of available experts can lower the weight traffic associated with repeated token-by-token decoding.
  • Direct KV-cache reuse: The pruned decoder consumes the KV cache produced by the full model without an additional conversion step. This training-free handoff is central to recovering quality that would otherwise be lost when the decoder is pruned.
  • Lightweight distillation: SlimWise optionally trains the decoder to continue generation from full-model KV caches. Only a small subset of parameters is updated, keeping the adaptation cost low.

The authors also point out that conventional benchmark accuracy can hide changes in generation length. A pruned model may appear competitive on a task score while producing outputs with materially different lengths, which can distort conclusions about practical serving efficiency. SlimWise’s distillation stage is intended to address both this issue and the remaining accuracy gap.

The framework is implemented in vLLM and supports both prefill-decode disaggregation and colocated deployment. The reported study covers two MoE backbones and three pruning criteria. On Qwen3.6-35B-A3B, 50% expert pruning produced up to a 1.81× decode-throughput improvement with minimal reported accuracy loss.

Why it matters

SlimWise illustrates a broader deployment principle: model compression should follow the bottleneck of each inference phase. Prefill is more sensitive to the quality of processing a whole context, while decode repeatedly pays for expert access. Preserving the full model where quality matters most and compressing the phase dominated by weight traffic offers a practical middle ground between an untouched model and aggressive global pruning.

The gains should not be interpreted as universal. Actual performance will depend on batch size, context and generation lengths, routing patterns, and whether prefill and decode are separated across workers. The optional distillation step also adds a training and maintenance cost. Even so, the design makes MoE serving more adaptable: a full model can establish the context representation, while a smaller decoder handles the high-frequency generation loop.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Helion in vLLM: A Unified Route to Tunable Linear Kernels
Inference & Serving
cctest.ai

Helion in vLLM: A Unified Route to Tunable Linear Kernels

PyTorch’s integration of Helion into vLLM’s linear backend explores whether a high-level kernel DSL can replace part of the manual specialization traditionally required for quantized GEMM. On NVIDIA Hopper, shape-specific tuning and hybrid dispatch produced end-to-end gains across the evaluated workloads, exceeding 10% in some cases.

Read more