Fetching from the wire…
Public story · 2026-07-21 · high
The authors position it for tasks where LoRA has already stalled, not as a general replacement.
Why now: On July 21, it stood out because it changes what gets adapted, not just how, while most PEFT papers only tweak the how.
A parameter-efficient finetuning method adapts residual connections instead of weights, per Oldenburg, de Kam and Zuijdam's paper on arXiv. Anyone who's hit a ceiling with LoRA gets another lever to try: residual connections, the part every popular PEFT method leaves alone.
LoRA, adapters and prefix tuning all modify weights or activations instead. Their method places manifold constraints on hyper-connections, adapting a structural piece those methods never touch.
That's unusual. Most PEFT papers amount to LoRA with one more hyperparameter tacked on. A method that changes what gets adapted, not just how, is rare in this field.
The authors frame this for tasks where LoRA has already plateaued. That's a narrow framing, not a pitch to replace LoRA outright. If your finetuning results have stalled, this is worth testing before you assume you're stuck.
Whether it holds up outside plateaued cases is unclear from here. The real signal will be whether a mainstream PEFT library adds a residual-adaptation option at all.
Each link below shares sources, entities, or timing with this story.
PortLLM claimed training-free, data-free transfer of LoRA patches onto updated base models, but only over short horizons and without theoretical grounding. This study runs 10 continual-pretraining steps on Mistral, Gemma, and Qwen and finds portability persists long-run, meani...
Set use_dora=True in PEFT's LoRAConfig with the 2026 starting recipe (r=16, target_modules='all-linear'). DoRA decomposes weights into magnitude and direction and applies LoRA only to direction, yielding +3.7% on LLaMA-7B and +1 to 4.4% on larger models with zero added inferen...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.