Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.05136 proves gauge-equivariance is necessary (not sufficient) for the transfer on a factored model W = UVᵀ. GD, momentum, shared-scalar Adam, Muon and Shampoo satisfy it. Adam, RMSProp and other coordinate-wise methods don't. A one-parameter family sweeping from coordinate-wise to shared-scalar preconditioning restores the bias monotonically, isolating anisotropy as the cause. In transformers Adam separates two gauge-equivalent initializations at the first step, ending with per-head WQᵀWK invariants 56% apart. Directly relevant if you're tuning LoRA adapters.
Each link below shares sources, entities, or timing with this story.
PortLLM claimed training-free, data-free transfer of LoRA patches onto updated base models, but only over short horizons and without theoretical grounding. This study runs 10 continual-pretraining steps on Mistral, Gemma, and Qwen and finds portability persists long-run, meani...
Beyond binary human-vs-machine detection: identifies which specific LLM generated a code snippet. Enables vulnerability triage (which model produced the bug?), licensing audits, and distillation detection. Directly relevant to the Anthropic distillation crackdown. (arXiv 2603....
Set use_dora=True in PEFT's LoRAConfig with the 2026 starting recipe (r=16, target_modules='all-linear'). DoRA decomposes weights into magnitude and direction and applies LoRA only to direction, yielding +3.7% on LLaMA-7B and +1 to 4.4% on larger models with zero added inferen...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
An independent researcher ran the Political Compass across 16 models from Google, Anthropic, OpenAI, xAI, Meta, Mistral, Qwen and Kimi — 30 standard runs, 30 reverse-phrased, 10 reordered each, ~69,440 answers, with the scoring system reverse-engineered to control for ordering...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.