Fetching from the wire…
Infra2026-09-07 · source-backed
Builds b10833 through b10839 went out between 06:49 and 11:14 UTC on September 7 (commits). #28208 writes explicit recurrent_layers during Qwen3-Next / Qwen3.5 HF-to-GGUF conversion and #22780 adds --fuse-qkv to fuse Q/K/V into a single tensor at conversion time. #28475 fixes races in mmid and mmf, #27870 fixes a divergent barrier in f16 flash attention, and #28068 corrects GDN normalization from max to rsqrt. Vulkan gained TQ1_0 support, type-aligned GET_ROWS and rms_norm fusion.
Each link below shares sources, entities, or timing with this story.
Build b10758 (September 2) fuses matmuls landing on Hexagon HMX, fuses MUL_MAT_ID into MUL_MAT_ID_NX, and adds VA defragmentation so large-dim runs abort less on fragmented address space (release). Build b10757 handles batch sizes above 4 for IQ3_S mat-vec when NUM_COLS > 4, r...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
Released September 4, it adds llama_lazy_mode / --lazy-mode for on-demand tensor reading (#27794), a max_buf_size quantize parameter capping quantizer RAM (#27795), quantizer row-slab streaming (#27830), and a fix preventing RAM peaking during load (#27483). It also adds spars...
Claims up to 2x faster generation, and adds fine-tuning of both MoE models on text or image datasets on Apple Silicon via MLX (release). Follow-up turns in long Qwen chats on Mac are reported up to 30x faster, MLX models now use full context size, and GLM-5.3 MLX fine-tunes ex...
The abliteration tool gained 215 stars to reach 30,103, but the stronger signal is downstream: the HF trending endpoint returns DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU and Momoking/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4, both naming the too...
Build b10677 fixes ggml_vk_graph_optimize, where is_src_of didn't treat two views of one tensor as dependent, so the optimizer reordered nodes across aliased reads and writes. Maintainers describe the result as silently wrong tokens under greedy decoding, different output on e...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.