Back to articles
Inference & Serving

Wavefront Decoding Brings Parallelism to Looped Language Models

3 min read

Introduction: the latency cost of looping

Looped language models repeatedly apply the same weight-shared block to increase effective depth without scaling the parameter count in the same way as a conventional deeper model. The trade-off appears during generation: producing one token may require a sequence of recurrent-block calls. This makes decoding latency grow with the number of recurrences, even though the model benefits from parameter sharing.

Wavefront Decoding (WFD) addresses this scheduling problem rather than changing the model or adding a separately trained draft network. The method is designed as a training-free self-speculative decoder that reorganizes the work already performed inside a looped model.

The central idea: mix drafting and verification

Conventional speculative decoding often follows a phase-separated pattern. A drafter proposes several tokens, and the target model subsequently verifies them. WFD exploits two properties specific to looped architectures:

  • Intermediate outputs can draft. A token state at an early recurrence depth can already produce a useful next-token prediction, even before the model reaches its full depth.
  • The recurrent block is shared. States belonging to different token positions and recurrence depths can be sent through the same block in one batched call.
  • A diagonal wavefront coordinates the work. New positions begin at shallow depth while earlier positions continue toward full-depth verification. Drafting and verification therefore overlap inside successive recurrent calls.
  • Rejected drafts are corrected at full depth. When a shallow prediction fails verification, the decoder uses the full-depth result to correct the output rather than accepting the draft.

The important change is not simply to increase batch size. WFD reshapes the dependency pattern across both sequence position and recurrence depth. States that are ready to advance can share a recurrent call with other states that are just beginning their draft computation, creating a pipeline-like diagonal wavefront.

Reported results and KV optimization

The paper evaluates WFD on six Spec-Bench task categories. Relative to autoregressive decoding, it reports a 2.42x speedup for Ouro-2.6B and a 3.54x speedup for Huginn-3.5B. WFD also consistently outperforms the phase-separated draft-then-verify schedule in the reported comparison.

The authors further introduce cross-recurrence KV sharing to reduce the KV traffic associated with the wavefront. With this optimization, the reported speedup on Huginn-3.5B reaches 4.81x. These figures should be read in context: they describe the listed models, benchmarks, and implementation conditions, rather than guaranteeing the same gain for every looped model or hardware configuration. The project code is available publicly for further verification.

Why it matters

WFD treats recurrence as a source of scheduling flexibility. Intermediate states do not always need to wait for final depth before contributing a prediction, while weight sharing makes mixed-depth batching possible. This offers a way to accelerate looped models without training a separate speculative model.

The approach still depends on the quality of intermediate predictions, the cost of implementing the wavefront schedule, and the efficiency of KV-cache management. Even so, it highlights a broader systems question: if model depth is executed repeatedly, can that depth become a schedulable inference resource rather than a purely serial cost? WFD provides one concrete answer for looped language models.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles