Multimodal Flow Puts Language and Vision on One Continuous Generative Path
Overview
Unified multimodal models generally follow one of two designs. Some convert both language and images into discrete tokens, while others combine discrete language prediction with continuous image generation. The first design can introduce a visual quantization bottleneck; the second often requires different objectives and sampling procedures for each modality. Multimodal Flow, proposed by HUST Vision Lab, explores a third option: modeling language and vision continuously within one shared generative framework.
How the approach works
- Continuous multimodal units: Rather than mapping images into a discrete visual vocabulary, the model uses continuous multimodal representations. Text blocks and images are arranged as ordered “continuous hyperchunks.” Text preserves token order, while images retain their spatial organization.
- A shared flow backbone: A chunk-causal backbone learns a single vector field over these hyperchunks through Flow Matching. In principle, language and vision can therefore follow the same continuous generation process instead of relying on separate language-modeling and image-sampling mechanisms.
- Shared interaction, specialized processing: Joint attention allows text and image representations to exchange information. Modality-specific feed-forward networks then process the distinctive features of each modality, providing a compromise between architectural sharing and modality adaptation.
- Different training and inference patterns: During training, the model predicts multiple target chunks in parallel. At inference time, it generates hyperchunks sequentially, preserving a causal ordering while keeping the underlying representation continuous.
Reported experiments
The authors instantiate the method as MF-1 and pretrain models at 0.6B, 1.2B, and 1.6B parameter scales. Continued pretraining consistently improves multimodal modeling across these sizes. With 150B pretraining tokens, MF-1 reports an average score of 82.8 on GenEval and DPG-Bench, and 75.3 across VQAv2, MMBench, and POPE. The paper also reports that, under matched data, optimization, and parameter budgets, Multimodal Flow outperforms representative discrete and hybrid baselines.
Why it matters
The main contribution is not simply another image-generation architecture. It challenges the assumption that a unified model must turn vision into language-like discrete tokens. By treating text and images as structured blocks in a continuous embedding space, the method seeks a common foundation for language understanding, image synthesis, and cross-modal interaction.
If the approach scales reliably, it could reduce the need for modality-specific objectives and sampling pipelines. It may also offer a cleaner way to study unified multimodal pretraining. At the same time, the available material does not establish how the method performs on long-form generation, difficult compositional images, inference efficiency, training cost, or changing data mixtures. Whether continuous flow modeling remains advantageous at larger scales will require broader comparisons.
The project releases code and a model, giving researchers a basis for reproduction and further evaluation.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...