Research
LoRA Patches Stay Portable Across 10 Continual-Pretraining Updates — You May Not Need to Re-Fine-Tune
PortLLM claimed training-free, data-free transfer of LoRA patches onto updated base models, but only over short horizons and without theoretical grounding. This study runs 10 continual-pretraining steps on Mistral, Gemma, and Qwen and finds portability persists over the long run, meaning repeated fine-tuning is not required when the base model is periodically refreshed. The proposed explanation is geometric: near-orthogonality of vectors in high-dimensional space, analyzed through the loss landscape, is what makes a stale patch keep working — a concrete cost argument for teams re-running fine-tunes on every base-model bump.
↳ Follow the thread