Dispatch
Together AI's Expert LoRA puts adapters on MoE experts and gets 89% knowledge recall against 15%
Together's 11 September fine-tuning expansion adds 18 open-weight models (GLM 5.3/5.2/5.1, DeepSeek-V4-Flash variants, Kimi K2.7-Code, Qwen 3.8-27B, Gemma 4) and a new Expert LoRA that attaches adapters to the experts themselves in MoE models. On invented facts, expert-inclusive adapters recalled 89% against 15% for attention-only LoRA, with MMLU-Pro at 75.3% vs 71.5%. Training prices dropped 30-70%, including GPT-OSS-20B SFT from $1.50 to $0.40 and GPT-OSS-120B from $5.00 to $2.50 per million tokens.
↳ Follow the thread