Loop Scaling Laws Unify Recurrence and MoE Sparsity
Introduction
Scaling a language model does not necessarily mean adding more unique parameters. Looped Transformers reuse the same parameters across multiple passes, increasing computational depth without proportionally increasing stored model size. Mixture-of-Experts models take a different route: only a subset of experts is activated for each token, allowing total capacity to grow while keeping active computation under control.
Existing scaling laws have generally studied recurrence and sparsity separately. That leaves an important design question unanswered: when recurrence is added to an MoE model, how many loops are useful, and how does sparsity change the value of each additional pass?
In Scaling Laws for Looped Mixture of Experts, a Meta research team introduces Loop Scaling Laws, a unified framework for recurrence, sparsity, model size, and data.
Key points
- A joint view of two efficiency mechanisms. Recurrence mainly increases computational depth, while MoE sparsity expands total parameter capacity. The proposed laws model both axes together.
- Effective capacity rather than raw loop count. A bounded, sparsity-conditioned recurrence mapping estimates how much looping contributes in parameter-equivalent terms and represents the point at which returns saturate.
- Consistency with established laws. Under limiting conditions, the formulation recovers standard dense and MoE scaling laws, connecting the new framework to familiar baselines.
- A design tool for constrained training. Fitted laws can help select model size, sparsity, and recurrence under fixed compute and memory budgets.
- Complementary efficiency gains. The paper summary reports roughly 3x active-parameter efficiency from sparsity and roughly 2x total-parameter efficiency from recurrence on reasoning evaluations, with further gains when the two are combined.
Why it matters
The main contribution is a common design space for capacity and depth. Increasing the number of experts can expand a model’s total knowledge capacity without activating every parameter for every token. Reusing a block across loops can instead spend more computation on the same representation. Neither strategy should be expected to improve indefinitely: sparse routing has system costs, while repeated computation can reach diminishing returns. A law that explicitly models saturation can therefore be more useful than a rule based only on parameter counts.
The summary also reports a trillion-token-scale experiment. At matched training compute, a looped MoE using recurrence selected by the proposed law reportedly approaches the reasoning performance of a non-looped MoE about twice as large in total parameters. Recurrence also creates a natural avenue for test-time scaling, since additional passes can be used when more inference compute is available.
These figures should be read as results from the reported settings, not universal guarantees. Actual gains will depend on the task, routing behavior, hardware utilization, and optimization stability. Still, the work points toward a broader scaling strategy: allocate resources jointly across stored parameters, active computation, repeated depth, and inference-time budget rather than relying on model size alone.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...