Back to articles
Multimodal

Kandinsky 6.0 Video Brings Synchronized Audio and Video Generation

3 min read

Introduction

Video generation is increasingly moving beyond silent clips. A useful generated scene may need dialogue, ambient sound, timing and convincing mouth movements to agree with the image. Kandinsky 6.0 Video, released by Kandinsky Lab, targets this problem with a family of models that generate short audiovisual clips from either text or an input image.

The family contains Kandinsky 6.0 Video Lite with 3 billion parameters and Kandinsky 6.0 Video Pro with 29 billion parameters. Both are described as producing five-second clips with synchronized 44 kHz audio, including lip sync, in text-to-audio-video and image-to-audio-video modes. A built-in super-resolution system can raise the output to Full HD, or 1920×1080.

Architecture: two streams with mutual attention

The main architectural choice is a dual-stream CrossDiT. Rather than treating audio as a post-processing track, the system keeps a pretrained video stream and a newly trained audio stream, then connects them with bidirectional cross-attention. Information can therefore move in both directions while the clip is being generated.

This matters because synchronization is not only a matter of matching timestamps. Speech content, vocal rhythm, facial motion and scene events all need to remain semantically consistent. A jointly connected system can use audio cues when shaping motion and visual cues when producing sound, at least in principle, instead of asking a separate pipeline to repair the mismatch later.

The reported training recipe is staged. The audio stream is first trained from scratch on large audio corpora. The two streams are then trained together on paired audiovisual data while the developers aim to preserve unimodal fidelity. Supervised fine-tuning follows pretraining, along with reinforcement-learning-based post-training and distillation. The supplied material also mentions a 10-step distilled generation setup, which is intended to make the larger model more practical to run.

What the reported evaluations show

In the VABench results described in the release material, Kandinsky 6.0 Video Pro leads the listed LTX 2.5, Kandinsky Lite and Kandinsky Pro systems across speech quality, audio aesthetics, audiovisual alignment, lip sync, desynchronization and visual realism. Human comparisons reportedly place it clearly ahead of Kandinsky 5.0 Video Pro.

Against LTX 2.5, human raters preferred the Pro model for visual quality, motion realism, visual prompt following, artifact reduction and overall task solving. The material also reports a statistically significant advantage in speech quality. Reinforcement-learning post-training reduced the Pro model’s word error rate by 47%, according to the release summary. These claims indicate that the project is optimizing for a broader target than image quality alone: the generated clip should also sound intelligible and remain temporally coherent.

Super-resolution and open release

The project includes a separate 1.4B-parameter video super-resolution diffusion transformer. It supports ×2, ×2.25 and ×4 upscaling, operating in KVAE latent space before tiled refinement and decoding. A distilled π-Flow DX version requires only two model evaluations per tile, which may help with the cost of high-resolution processing.

The code, checkpoints and Diffusers integration are released under the MIT license. This gives researchers and developers a relatively accessible starting point for experimenting with open audiovisual generation, rather than relying only on a hosted service.

Why it matters

For short-form media, character dialogue, educational demonstrations and rapid creative prototyping, synchronized generation could reduce the need for separate dubbing, lip-sync correction and timeline editing. Still, the public description focuses on five-second clips. Long-form consistency, multi-speaker interaction, identity stability and deployment cost remain open questions that require independent testing. The broader significance of Kandinsky 6.0 Video is its attempt to make sound, motion and visual semantics part of one generation problem.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles