Fetching from the wire…
Public story · 2026-09-12 · high
Adapters on the expert layers, not just attention, lift invented-fact recall to 89% from 15%, and Together cut training prices up to 70%.
Why now: Together AI published the expansion on September 11, a day ahead of this coverage.
Together AI added a technique called Expert LoRA to its fine-tuning service on September 11. It attaches adapters to the expert layers inside mixture-of-experts models, not just the attention layers where LoRA adapters usually sit. The release also adds 18 open-weight models to the platform.
The gap between the two adapter setups is the story. On a test of recalling invented facts, adapters that included the expert layers reached 89% recall. Attention-only adapters reached 15%, on the same base model and the same adapter method, per Together AI's fine-tuning expansion post. MMLU-Pro moved too, from 71.5% to 75.3%.
MoE models route each token through a handful of expert sub-networks per layer. An adapter that only touches attention never reaches where most of the model's capacity sits. Fine-tune that way and you're training a slice of the network while leaving the experts untouched.
Anyone who tried LoRA on an MoE model to inject domain knowledge and found it didn't stick has a likely culprit: adapter placement. The new model list includes GLM 5.3, 5.2, and 5.1, DeepSeek-V4-Flash variants, Kimi K2.7-Code, Qwen 3.8-27B, and Gemma 4. Training prices dropped 30-70% across the board. GPT-OSS-20B supervised fine-tuning goes from $1.50 to $0.40 per million tokens. GPT-OSS-120B goes from $5.00 to $2.50.
Together's post doesn't say whether Expert LoRA runs by default on MoE models or needs to be turned on. That matters for anyone trying to reproduce the 89% figure themselves.
Each link below shares sources, entities, or timing with this story.
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
This is the week's most useful number, and it took an actual experiment to produce it rather than a launch post. Together AI ran 113 DeepSWE tasks at 4 trials per config: 452 GLM-5.3 rollouts and 448 GLM-5.3-Flash rollouts. GLM-5.3 scored 69.0% pass@1 at $3.99 per rollout. Fla...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.