Fetching from the wire…
Models2026-09-26 · source-backed
The experiment froze both the Qwen3.5-0.8B backbone and Qwen3.8-Flash-Next's PLE memory, training only a small R=1 reader at decoder layers 3 and 9 with a token-level gate, on 15M tokens using free Kaggle GPUs. Validation perplexity fell from 18.28 to 17.35, and the real memory beat both random and permuted controls. A 20M-token reader regressed on math and strong fixed injection hurt LAMBADA, so there's a narrow window. GGUFs and a llama.cpp path are public.
Each link below shares sources, entities, or timing with this story.
A month of solo work produced quants for LongCat-Flash-Lite-Sparse, Qwen3.8-27B, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with vision. Getting there meant writing Heretic support for the architecture from scratch and then adding llama.cpp support, and main...
An independent Aider run at Q5_K_L, xhigh effort, 128GB: Swift 1.5 averaged 6,991 tokens and 608s per case against 17,646 tokens and 1,542s for base Qwen3.8-Flash-Next. First-try pass rose 40.2% to 41.1%, retry pass fell 90.7% to 86.9%, well-formed diffs reached 100%. The auth...
An r/LocalLLaMA builder patched vLLM to offload most of the KV cache to host RAM and reports 1M context on 3x RTX 3090: about 80 tok/s at short context, dropping to roughly 60 once QSA hits its 2,048-token budget and then staying flat as context grows, ~150 tok/s at four concu...
Part 4 of a running 2x3090 series: prefill was 80+ seconds to first token on an 8k prompt and 24 minutes on a 119k one, and releasing the 150-slot expert cache off the GPU during prompt processing bought the speedup. The comments are why it's here. A reader calculated the 2.8s...
Every flaky quantized agent I've debugged, I blamed the quantization. A paper published September 4 says I've probably been blaming the wrong layer (arXiv 2609.04748). The setup is clean. Model, decoding parameters, seed, request order and batch size all fixed. Requests issued...
A practitioner running 2x Strix Halo 128GB over USB-C 4 with llama.cpp RPC compared both models at Q8_K_XL on real coding work (r/LocalLLaMA). A task Qwen finished in 25 minutes on medium took DeepSeek 12. The poster attributes the gap to fewer hallucinated detours. The sharpe...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.