Fetching from the wire…
Models2026-08-28 · source-backed
PR 27742 adds gated delta net layers, 512-expert MoE with top-10 selection, hyper-connections, query-key sparse attention and the per-layer n-gram embeddings as a mmap table that can sit in RAM or on disk (GitHub). Reported 55 tok/s on 4x3090 with the Q4 GGUF, and one commenter got 10 tok/s on a 4GB card by offloading to SSD. MTP is still in progress, and one warning matters: the current engram implementation only works with mmap and has no eviction mechanism, so mlock locks the whole table into memory.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover Flash, GitHub, MoE, MTP; reported by the same outlet (github.com); overlapping topics (attention, gated).
Shared entities / Same source domain / Shared topic
Both cover Flash, GitHub, Next, Qwen3; reported by the same outlet (github.com); overlapping topics (flash-next, gguf).
Shared entities / Shared topic / Earlier coverage
Both cover Flash, GiB, MTP, Next; overlapping topics (attention, table); earlier Flash coverage from 2026-08-27.
Both cover Flash, MoE, Next, Qwen3; overlapping topics (card, commenter, embedding); earlier Flash coverage from 2026-08-25.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover GitHub, MTP, Qwen3; reported by the same outlet (github.com); overlapping topics (attention, card).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover MoE, Qwen3, SSD; reported by the same outlet (github.com); earlier MoE coverage from 2026-07-25.
Shared entities / Same source domain / Shared topic
Both cover GitHub, MoE, MTP; reported by the same outlet (github.com); overlapping topics (attention, expert).
Shared entities / Tension
Both cover Flash, MoE, Next, Qwen3; pushes against this story (against).