When to Align, When to Predict: A Phase Diagram for Multimodal Learning
arXiv·medium signal
Kamai, Van Assel, and Regev introduce a phase diagram that unifies cross-modal alignment (CA) and cross-modal prediction (CP), the two dominant paradigms for multimodal representation learning, and characterizes when each is optimal. For teams building multimodal agents, it offers principled guidance on objective choice instead of defaulting to contrastive alignment. The contribution is a theory-grounded map of a previously ad hoc design decision.