HC-DLM Keeps Token Dependencies in Parallel Diffusion Generation
Introduction
Diffusion language models are being explored as an alternative to autoregressive generation. Instead of producing tokens in a fixed left-to-right order, they can refine a whole sequence from a noisy state. This makes them appealing for problems that require bidirectional reasoning or global constraint satisfaction. The challenge is preserving interactions among tokens while retaining the advantages of parallel decoding.
A team from the University of Illinois Urbana-Champaign proposes Hierarchical Continuous Diffusion Language Models, or HC-DLM. The method places discrete token generation and continuous latent refinement inside one denoising process, rather than treating them as separate generation mechanisms.
Key ideas
- The discrete diffusion problem: When many positions are decoded together, tokens may be sampled independently from their marginal distributions. Parallelism is preserved, but dependencies among jointly generated tokens can be weakened. That is especially problematic for Sudoku or mathematical planning, where a locally plausible token may violate a global constraint.
- The continuous diffusion problem: A shared continuous state can preserve information across positions, but the denoiser does not receive a persistent, explicit connection to a valid discrete token arrangement until the final decoding stage.
- One persistent latent state: HC-DLM treats the continuous latent as the only state that survives across generation steps. Tokens are read from that latent at every step, and the readouts are then used as a scaffold for the next latent update.
- A unified objective: The training objective is derived from a variational lower bound on token likelihood. This gives the discrete and continuous parts a common probabilistic formulation instead of simply attaching a continuous context module to an otherwise self-contained discrete chain.
Results and implications
The paper evaluates HC-DLM on Sudoku, Countdown, and the LM1B language-modeling benchmark. According to the authors, at matched model sizes, HC-DLM improves over discrete and continuous diffusion baselines in Sudoku and Countdown puzzle accuracy, as well as in generative perplexity on LM1B. The supplied material does not include the exact gains, so the results should be read as evidence for the approach rather than as a quantified claim about every setting.
The important design choice is not merely the addition of a continuous representation. It is the feedback loop between the two levels: tokens are not only the final output of denoising, but also intermediate structural signals that influence the next latent state. This could give a diffusion language model a way to combine global coordination with a closer connection to discrete validity.
There are still open questions. The available summary does not provide detailed ablations, sampling speed, compute requirements, or scaling results. Practical comparisons will therefore need to examine whether the proposed dependency preservation justifies any additional inference cost, and whether the reported gains extend to larger models and broader language tasks.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...