Fetching from the wire…
Agents2026-09-21 · source-backed
Jared Palmer published Kev September 17 on Qwen3.5, Apache-2.0 with training code and frozen eval suites. One request carries yes/no, multiple-choice and rating questions that share input text but can't read each other, and the API matches TypeSafe's System One so their Python SDK points at a local server. The Qwen3.5 generation on September 20 cost about $95 of rented Modal H100 time plus $0.03 in API calls, and Palmer published the benchmark he failed: Kev-9B at 0.837 on the locked test set against 0.780 for the Qwen3-based predecessor, while hosted Jev still leads at 0.857. He also flags a regression, a five-question request taking 779ms on Qwen3.5 Kev-4B against 174ms on the Qwen3 version, and recommends the older checkpoints until MLX lands. Publishing your own missed pre-registered criteria instead of moving the gate is rarer than it should be.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
AlexsJones/llmfit released v1.1.10 today, adding RamaLama runtime discovery to its MCP server, the Qwen3.8 model family and MiniMax M3 vision capability exposure (GitHub). It also merged 32 MLX benchmark results on an Apple M4 Pro, the project's first MLX entries, giving an ap...
Compaction is where long sessions go to die. The model summarizes what happened, the summary drops the exact string you needed three hours later, and you don't find out until the agent confidently references a file path that never existed. It's the biggest source of silent con...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
vercel/eve, the Apache-2.0 agent framework Vercel launched at Ship London on June 17 and calls "Next.js for agents," published 0.56.0 on September 15 at 5,184 stars. More inbound code than bug reports for a three-month-old framework is unusual. Vercel says over 100 internal ag...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.