Fetching from the wire…
Models2026-08-09 · source-backed
Following the August 6 release (128K context, tool calling, open weights, ~2.5GB, 220 tok/s on an M5 Max per the company), an r/LocalLLaMA user reports 260 tok/s generation and 20K prompt processing on 3090s. A phone-class model on desktop silicon. Their use-case list is the practical part: needle-in-haystack scans ("read this massive thing and tell me if it mentions x"), throwaway summarization, Linux command recall, structural autocomplete. Explicitly not anything that matters. The 128K context ceiling is called out as the binding limit on the bulk-scan workload the speed otherwise unlocks. This is the tier of model that should be handling your PDF-to-slide jobs.
Each link below shares sources, entities, or timing with this story.
The QAD build landed at 13:53 UTC on August 19 alongside the normal Q4_0 through Q8_0 ladder. Commenter Chromix_ flagged that it was built without an imatrix and its token embeddings were quantized at Q6_K rather than Q4_0, which may work against the QAT-style training the fil...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Four stories about things going wrong. Here's one about something working, with actual numbers attached. In an August 7 disclosure covered by TechCrunch, Airbnb said AI now writes 60% of its new code, that concept-to-launch time on key initiatives has dropped by as much as 60%...
Anthropic's @ClaudeDevs announced on August 14 that Claude Code desktop now resumes a stalled session the moment the usage-limit window resets. The post framing it as more valuable than the Fable model launch drew 134 upvotes, with the line "the real flagship feature is not ha...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.