LLaDA-Image: An Open Recipe for Unified Image Generation and Editing
LLaDA-Image combines a from-scratch 6B Diffusion Transformer with a frozen vision-language understanding module for generation, editing, and text rendering. Its distilled Turbo variant reduces inference to 2–4 sampling steps while the project releases weights, code, and training recipes.
Read more