RAZOR Prunes MoE Experts by Measuring Replaceability
Introduction
Mixture-of-experts (MoE) models reduce token-level computation by activating only a small number of experts, but deployment still requires storing the full expert pool. As expert banks grow, pruning becomes an attractive way to reduce memory demand. The central challenge is deciding which experts can be removed without substantially changing the model’s behavior.
The RAZOR paper proposes a different criterion: an expert should be judged by how replaceable its function is, not simply by how often it is routed or how large its output is.
Measuring functional replaceability
Usage counts, routing weights, and output norms are useful signals, but none directly indicates the damage caused by deletion. A frequently used expert may perform a function that other experts can easily cover. Conversely, a lightly used expert may provide a capability with little redundancy.
RAZOR addresses this issue with “consensus residuals.” The method compares an expert’s output with the original weighted mixture produced by the expert layer. The resulting deviation captures how much that expert’s behavior differs from the computation supported by the other selected experts.
At a fixed layer input, the paper derives an exact single-deletion identity. It incorporates survivor renormalization after removal and the router-selected refill that can occur when an expert is no longer available. RAZOR aggregates these local signals over calibration tokens to produce pruning scores for a chosen budget. The procedure requires neither gradients nor recovery training, making it comparatively lightweight for model compression workflows.
Reported results
The study evaluates RAZOR on GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3, removing either 25% or 50% of the experts. Across all eight model–budget settings, RAZOR obtains the highest nine-task macro average among the evaluated pruning methods.
The paper also reports matched comparisons with REAP on two backbones. RAZOR improves the score by 2.12 to 5.59 points and wins all 36 paired task comparisons. In four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model–budget settings, it also produces lower reverse KL than REAP, indicating closer predictive distributions under those tests.
Implications and limitations
RAZOR’s main contribution is a shift in how MoE pruning can be framed. The relevant question is not simply whether an expert appears important, but whether the remaining expert population can cover its function. That perspective may help reduce storage costs without requiring a post-pruning fine-tuning stage.
The results do not imply that benchmark retention guarantees identical generation behavior. Analysis of Qwen3.6-35B-A3B responses found changes in diversity, formatting, and termination after pruning. Future evaluations should therefore examine structured outputs, long-form generation, and stopping behavior alongside task scores and distributional metrics. RAZOR appears to reduce pruning damage, but it is not presented as a guarantee of complete behavioral equivalence.
Source: arXiv
Comments
Checking sign-in status...
Loading comments...