Back to articles
Inference & Serving

Disaggregated Quantization Gives LLM Prefill and Decode Different Tools

3 min read

LLM inference is often described as one pipeline, but its two main phases have very different hardware characteristics. Prefill processes many prompt tokens in parallel, so arithmetic throughput matters. Decode generates one token at a time, making weight movement and memory bandwidth much more important. The paper “Disaggregated Quantization: Specializing LLM Prefill and Decode” builds on this distinction with a method called Disaggregated Quantization, or DQ.

The central idea

Instead of requiring one quantized representation to serve every stage, DQ assigns computation formats, weights, and storage policies to the phase where they are most useful:

  • Prefill is optimized for computation. The authors train compute-native prefill weights that can take better advantage of low-precision arithmetic when processing long prompts.
  • Decode is optimized for memory traffic. Generation can continue to use compact low-bit weights, reducing the amount of data that must be moved for every token.
  • Activation quantization becomes phase-specific. On Qwen 3 and Gemma 3, removing activation quantization specifically during decode improves accuracy on decode-heavy tasks without increasing inference cost.
  • Existing checkpoints remain relevant. The paper also studies shared-weight format disaggregation and validates the idea through post-training quantization, including models up to 2.8 trillion parameters.

The experiments suggest that separately trained prefill weights can accelerate prompt processing compared with weight-only inference while matching or exceeding its accuracy under 2- to 3-bit decode settings. For Qwen3.8-27B, the released NVFP4 prefiller improves low-bit results on MMLU-Pro and MMMU-Pro without changing the decode checkpoint. The broader lesson is that quantization quality depends not only on bit width, but also on whether the representation matches the access pattern of the phase using it.

Making two checkpoints practical

A second checkpoint naturally creates a storage problem, especially on a single accelerator. The authors address it with Offloaded Disaggregated Prefill, or ODP. Instead of keeping the prefill weights resident in device memory, ODP streams them from an SSD while a sufficiently long prompt provides time to amortize loading.

In llama.cpp tests on the same 27B model, ODP achieved a 1.78x time-to-first-token speedup over the weight-only baseline at an 8K prompt length. This does not mean SSD access is universally free: the benefit depends on prompt length, storage throughput, scheduling, and model architecture. The approach works because long prompts provide enough computation to overlap or amortize the additional movement.

Why it matters

DQ reframes quantization as a serving-system decision rather than a single model-wide setting. A deployment can favor arithmetic efficiency during prefill and memory efficiency during decode, potentially improving both responsiveness and quality in mixed workloads. The trade-off is operational complexity: services must coordinate multiple checkpoints, loading policies, and runtime support in systems such as vLLM or llama.cpp.

The results are promising for long-context and decode-heavy serving, but they do not eliminate the need for workload-specific testing. Hardware, prompt distribution, storage speed, and framework implementation will determine whether the additional prefill weights pay for themselves.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles