Back to articles
Diffusion Models

MC-Sparse Narrows the Dense–Sparse Attention Gap in Diffusion Transformers

3 min read

Long sequences are turning attention into one of the main bottlenecks in diffusion Transformers. This is particularly visible in video generation and high-resolution 3D asset creation, where the number of tokens can grow rapidly. Sparse attention is an obvious way to reduce the cost: instead of evaluating every query against every key-value pair, the model keeps only a subset of interactions. The difficulty is that aggressive sparsity can also remove information needed for faithful generation.

MC-Sparse begins by asking where that quality gap actually comes from. Using controlled oracle comparisons, the paper separates three failure modes. First, query grouping imposes a structural constraint: queries that are not truly equivalent may be forced to share the same sparse pattern. Second, the mechanism that predicts important interactions can select the wrong KV tokens. Third, simply dropping tokens loses part of the attention output, and many sparse methods do not explicitly recover that missing contribution.

The proposed framework addresses these issues with three coordinated design choices:

  • Individual KV selection: Rather than relying entirely on coarse token blocks, MC-Sparse selects individual key-value tokens using exact attention probabilities. This gives the sparse pattern finer control over which interactions are retained.
  • Hardware-aware query organization: Fine-grained selection can create irregular memory access, which is inefficient on GPUs. MC-Sparse therefore groups similar queries into tile-aligned groups, preserving a structure that can be executed efficiently while keeping KV selection relatively precise.
  • Metadata and residual reuse: The method caches query groups, selected KV indices, and the residual between dense and sparse attention outputs. These metadata are reused across later denoising steps, reducing the need to repeatedly make the same expensive selection decisions and helping compensate for discarded attention contributions.

MC-Sparse is designed as a training-free framework. Its purpose is not merely to maximize the number of skipped interactions, but to make sparse attention behave more like dense attention at a practical execution cost. According to the reported experiments, it achieves higher fidelity to dense-attention outputs and larger denoising speedups than the referenced sparse-attention baselines across video and 3D generation models. Compared with dense attention, the paper reports a 1.80× denoising speedup on Minimax-H3-Base and a 2.32× speedup for 3D asset generation, with negligible quality loss.

The broader contribution is a more systematic view of sparse diffusion inference. Efficient sparsity depends not only on choosing fewer interactions, but also on selecting them accurately, organizing them for hardware, and accounting for information that has been removed. The available material does not establish that the same gains will hold for every architecture or sparsity level; cache overhead, sequence length, GPU kernels, and model behavior may all matter. Still, MC-Sparse offers a useful direction for post-training acceleration of long-sequence diffusion models.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
ALoDLM Lets Diffusion Language Models Spend Compute Where It Matters
Diffusion Models
cctest.ai
Diffusion Models

ALoDLM Lets Diffusion Language Models Spend Compute Where It Matters

ALoDLM introduces token-adaptive latent recurrence to diffusion language models, allowing easy positions to be committed early while difficult ones receive additional refinement. The paper reports improved average benchmark performance at 1.7B and 8B scales without giving up parallel decoding.

Read more