Fetching from the wire…
Top 5 · 2026-09-15 · source-backed
Salesforce released Koa, built by post-training Nemotron-3-Super-120B, Nvidia's open-weight hybrid Mamba-Transformer MoE with 120B total and 12B active parameters (TechCrunch, paper at arXiv 2609.15066). The training was GRPO reinforcement learning on public and synthetic data only. No customer data, which Salesforce says explicitly and which matters for their enterprise buyers.
Koa beats its own base most clearly on multi-turn tool use, and surpasses GPT-4.1 on public tool-use, agentic-reasoning and CRM benchmarks. It stays below frontier models generally. Nobody is claiming otherwise.
Read that shape carefully, because the shape is the news. A CRM vendor took someone else's open base, spent RL compute instead of pretraining compute, and got a domain agent that beats a model from the lab that defined the category. No $500M training run. No cluster of their own.
Shanghai AI Laboratory did the same move on a different base. InternLM published Atria Dawn Preview to Hugging Face with MIT weights and a 1M-token context, repo live September 11, FP8 checkpoint September 12, paper September 14 with 143 signed authors (Hugging Face, arXiv 2609.15818). It takes Z.ai's 744B GLM-5.2 and post-trains it for research loops: problem analysis, tool use, code implementation, experiment execution, failure recovery. Competitive with frontier agents across 16 benchmarks, top reported score on five.
Two labs. Two different open bases. Same play, four days apart.
The Atria Dawn paper carries a second result I haven't stopped thinking about. The authors studied their own development process: 769 task records from 56 participants, and about a third of completed AI-assisted tasks were rated infeasible without AI. Agents proposed methods and implemented revisions while humans kept final decisions. That's a lab publishing the labor data for building the model alongside the model.
For anyone with a domain and no pretraining budget, the recipe is now written down twice in one week. Take an open base with a published architecture. Build a verifiable reward for your domain. Run GRPO. You need eval infrastructure and RL compute, not a datacenter.
The counterweight is Nathan Lambert's estimate that the open-closed gap runs 4 to 6 months. You're not building a frontier model. You're building something that beats last year's frontier model at one job, which for most enterprise work is the job.
Each link below shares sources, entities, or timing with this story.
The Segment co-founder published "Small models have arrived" on August 26, and it took 703 points on Hacker News (calv.info). His measurement: a personalized-news task that cost about a dollar on Sonnet-class models now runs at about a dime. Ten times cheaper, doing the job we...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
NVIDIA launched Nemotron 3 Super — a 120B total / 12B active parameter hybrid Mamba-Transformer MoE, open, designed specifically for multi-agent workloads, and delivering 5x higher throughput than Nemotron 2 at the same active parameter count (NVIDIA Newsroom). It ships with a...
Open-source enterprise agent deployment platform explicitly designed to mirror CUDA's ecosystem capture. Pairs with Nemotron 3 Super 120B (hybrid Mamba-Transformer MoE, 12B active params, 2.2x throughput). DEV Community
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.