Will Visual Encoders Disappear? Scaling Laws Point to a Possible Shift
Introduction
Most modern multimodal large language models rely on a pretrained visual encoder. The encoder turns an image into a compact representation before the language model processes it. This design supplies a useful visual prior and can make multimodal training more efficient, but it also creates a system made of several specialized components.
Encoder-free multimodal models take a different approach. They receive raw pixels and learn visual representations inside the language model itself. The architecture is simpler and more unified, yet the field has lacked a systematic account of how this design behaves as model size and training compute grow. A new study from the Tencent Hunyuan team directly compares the scaling laws of the two approaches.
Key findings
- Text modeling is largely unaffected. The encoder-based and encoder-free systems show nearly overlapping loss-compute frontiers on the text objective. Removing the visual encoder therefore does not appear to impose a fundamental penalty on language modeling.
- The multimodal optimum shifts toward larger models. For the multimodal objective, removing the encoder changes the compute-optimal allocation. Encoder-free training benefits from putting more of the budget into model scale, presumably because the language model must learn visual representations rather than receiving them from a pretrained component.
- The early gap may not last. Encoder-free models underperform at smaller scales, where a pretrained visual prior is especially valuable. The study predicts that the two approaches could converge at around 10^22 FLOPs, a level the authors describe as within practical pretraining budgets.
- The language model learns to act like a visual encoder. As compute increases, bidirectional interactions among visual tokens become more useful. Visual processing also shifts toward earlier layers, while expert routing for visual tokens becomes more concentrated. These patterns suggest that the model develops a specialized internal pathway for visual information.
Why it matters
The results do not mean that visual encoders are about to disappear from production multimodal systems. At modest scales, a pretrained encoder can still offer better data efficiency, a strong initialization, and a more predictable engineering path. In addition, the reported crossover is a scaling prediction rather than proof that every downstream task or benchmark will show the same behavior.
The broader message is that the value of a pretrained visual prior may diminish with scale. In smaller systems, importing visual knowledge is an efficient shortcut. With enough parameters, data, and compute, a unified language model may be able to learn comparable representations by itself. This changes the architectural question from how to connect the strongest available vision encoder to how to help one model learn vision efficiently from raw inputs.
An encoder-free design could reduce component coupling and make end-to-end optimization easier. It may also simplify future multimodal pretraining pipelines. Still, practical adoption will depend on broader evaluations, training cost, data requirements, and performance on downstream tasks. The study offers a strong scaling hypothesis, not a final verdict on the role of visual encoders.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...