Scaling Linear Mode Connectivity and Merging to Billion-Parameter Pretrained Transformers
arXiv·medium signal
Li and Shen extend linear mode connectivity (LMC) and weight-merging techniques, previously demonstrated mostly at small scale, to billion-parameter pretrained Transformers. They address why existing merging methods break down at scale and propose an approach that preserves connectivity between independently trained models. Practically useful for builders combining multiple fine-tunes or specialist checkpoints without retraining.