Back to articles
Diffusion Models

IDU Combines One-Step Distillation with Selective Data Unlearning

3 min read

Generative models often face a trade-off between quality, speed, and control. Diffusion and flow-matching models can produce strong samples, but their multi-step sampling procedures increase inference cost. They may also reproduce content that developers later want removed from the training distribution. A new paper proposes a way to address both concerns during the same training process.

What IDU is designed to do

The proposed Inverse Distillation Unlearning (IDU) framework starts with a teacher model trained on the full dataset. It then trains a one-step student generator with two goals: retain the useful behavior of the teacher and reduce outputs associated with a designated forget set. The forget set can represent a class or another subset of the original training data that should no longer be reproduced.

A notable feature is the information required for training. IDU uses the pretrained full-data teacher and samples from the forget set, but does not require access to retained training examples. It also avoids additional feature extractors and classifiers. This makes the approach different from retraining strategies that need to reconstruct the desired data distribution from scratch.

The distribution-level perspective

The method is built around a distributional formulation rather than a simple penalty on individual generated samples.

  • Joint objective: distillation and forgetting are optimized together instead of being treated as two sequential procedures.
  • Mixture representation: the distribution used in training is expressed as a mixture of the forget-set distribution and the student’s generated distribution.
  • Recovery of retained behavior: by comparing this mixture with the teacher’s original training distribution, the optimization is intended to preserve the portion not designated for removal.
  • Model-family coverage: the paper evaluates both flow-matching and score-based diffusion teachers, suggesting a common formulation across these multi-step generators.

What the experiments show

The authors evaluate IDU on MNIST and CIFAR-10. Across flow-matching and score-based diffusion settings, the framework substantially lowers the frequency of samples from forgotten classes while maintaining competitive generation quality for retained classes. The available material does not report detailed numerical results, so the findings should be read as evidence of feasibility rather than a complete characterization of performance across large-scale applications.

The broader significance lies in combining efficient inference with targeted model modification. A one-step student can reduce sampling overhead, while the unlearning objective offers a route for responding to data removal, privacy, or unwanted memorization requirements. At the same time, the current evidence is centered on small-scale image benchmarks and class-level forgetting. More work is needed to test individual-example removal, large image models, conditional generation, and stronger evaluations of residual memorization. IDU nevertheless presents inverse distillation as a promising way to reshape what a generator retains while making it faster to run.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
ALoDLM Lets Diffusion Language Models Spend Compute Where It Matters
Diffusion Models
cctest.ai
Diffusion Models

ALoDLM Lets Diffusion Language Models Spend Compute Where It Matters

ALoDLM introduces token-adaptive latent recurrence to diffusion language models, allowing easy positions to be committed early while difficult ones receive additional refinement. The paper reports improved average benchmark performance at 1.7B and 8B scales without giving up parallel decoding.

Read more