Towards Physics of Multimodal Pretraining

Knowledge Flow, Modality Synergy, Early Unification, and Recipes

We study the fundamental mechanisms of multimodal pretraining through Knowledge FlowDisentangling cross-modal transfer across language, understanding, and generation to show how transfer is highly asymmetric and concept-dependent., Modality SynergyUnderstanding how task complexity dictates synergy vs. competition, and which forms of architectural sharing improve performance., and Early UnificationShowing early, simultaneous unification outperforms late or sequential training, and delayed visual integration induces "vision laziness"..

We provide a systematic, bottom-up exploration of unified multimodal pretraining. Utilizing controlled environments across both synthetic and real-world datasets, we investigate how modalities interact, compete, and co-evolve. These insights are applied to build efficient pretraining recipes and are validated at scale by training 13.5B Mixture-of-Experts (MoE) models on 2T tokens.

Contents
01 Knowledge Flow How capabilities transfer across language, understanding, and generation. 1A Real-world transfer
1B CLEVR concept transfer
02 Synergy & Compete Task complexity and parameter sharing decide whether modalities help or compete. 2A Task complexity
2B Architecture sharing
2C Vision encoders
03 Early Unification Why visual pathways must co-evolve from the beginning of pretraining. 3A Timing
3B Sequential training
3C Vision laziness
04 Recipes & Scale Practical data-mixing and scaling guidance from the controlled studies. 4A Mixture search
4B Scaled verification