Fetching from the wire…
Infra2026-09-21 · source-backed
PR #29197 overhauls buffer and DMA handling to support 64-bit extended mappings on Hexagon v81 and newer (Snapdragon Gen5, X2-Elite, IQ10), so buffers above the 4GB NPU virtual address space map once instead of being mapped and unmapped during inference. MUL_MAT, FA, GDN and SSM_CONV were rewritten to stop reading tensors through DDR to L2 to HVX, and the author shipped ggml-hexagon-inspect.py to disassemble kernels and flag register spills. On by default with GGML_HEXAGON_DMA64=0 to disable.
Each link below shares sources, entities, or timing with this story.
Builds b10833 through b10839 went out between 06:49 and 11:14 UTC on September 7 (commits). #28208 writes explicit recurrent_layers during Qwen3-Next / Qwen3.5 HF-to-GGUF conversion and #22780 adds --fuse-qkv to fuse Q/K/V into a single tensor at conversion time. #28475 fixes...
PR #28994 adds Hexagon NPU support for Q4_K (reusing existing Q4_1 infrastructure) and Q6_K (new kernels), which together enable Q4_K_M since those models are typically a mix of the two. Validated across IQ8 and IQ9 platforms. Q4_K_M is the default quant most people download f...
Microsoft open-sourced Foundry-Local, and I think it's the most underrated release of the week. One SDK for chat and audio with automatic hardware acceleration (NPU > GPU > CPU), self-contained with no external dependencies, and the API surface is identical to Azure AI Foundry...
1. Set Up Cursor Automations (intermediate) — Event-driven agents from PagerDuty/GitHub/Slack triggers with isolated sandboxes. Cursor Blog 2. Apply Context Engineering to Cut Agent Costs 60-80% (advanced) — Hierarchical token budgets, dynamic tool filtering (max 15), automati...
A filed VS Code issue (#313064) documents how the setting persists invisibly in git history and CI output after the user manually replaces the generated commit text. You can disable it via the "Git Add AI Co Author" setting, but the opt-out default raised legitimate concerns a...
Build b10758 (September 2) fuses matmuls landing on Hexagon HMX, fuses MUL_MAT_ID into MUL_MAT_ID_NX, and adds VA defragmentation so large-dim runs abort less on fragmented address space (release). Build b10757 handles batch sizes above 4 for IQ3_S mat-vec when NUM_COLS > 4, r...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.