Fetching from the wire…
Infra2026-07-19 · source-backed
As of July 13, AMD Quark trains, quantizes, and serves EAGLE-3 drafts through vLLM, reporting up to 2.00x for Kimi-K2.5 and 1.79x for MiniMax-M2.5. Related July work: ROCm end-to-end on MI355X with a prebuilt container (July 10), plus Hopper-optimized attention and FP8 MoE backends improving TTFT and TPOT on H20. If you self-host and haven't turned on a draft model, this is the biggest single-config win currently sitting on the table. (vLLM Blog)
Each link below shares sources, entities, or timing with this story.
I've spent the last year assuming that if I wanted real agentic coding quality, I paid for a closed model. That assumption took a hit on June 1. MiniMax shipped M3 with a new sparse-attention architecture (they call it MSA) that handles up to 1M tokens at roughly 9x prefill an...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
paddo.dev makes the most contrarian read: the letter's substance isn't openness but paragraph nine, defending distillation as "a widely used technique for model improvement" and urging policymakers against "conflating legitimate model development techniques with misappropriati...
OSTP Director Michael Kratsios posted July 22 that Moonshot built "a sophisticated internal platform to conduct large scale distillation against U.S. models," switching access methods to avoid detection, and acquired GB300-equipped servers plus GB300 access in Thailand. TechCr...
Reuters, via Tech Startups, reports capital released against deployment milestones with Anthropic deploying up to two gigawatts of Instinct MI450 starting 2027. Same structure as Nvidia/OpenAI: compute vendor capital flowing to the lab that commits to buy the silicon. A two-gi...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.