Fetching from the wire…
Models2026-09-15 · source-backed
The author kept it private for two months after claiming SOTA at aggressive quant levels on small Qwen3.5 GGUFs, then released the toolset at github.com/curvedinf/voodoo-dyn-quant. Run every quant level for every tensor at once, freeze the candidate quant weights from llama.cpp's ggml conversion, train a single scalar gate per tensor per quant level, with softmax keeping gradient flowing to all levels and an annealed tau forcing each tensor to settle on one choice. Gradient descent picks the per-tensor layout. He handed it over because he doesn't have time to scale it, which is the good reason to open source something.
Each link below shares sources, entities, or timing with this story.
TAK builds an imatrix from a task-specific corpus, finds the smallest size before collapse, then promotes and demotes tensors within a byte budget. No pruning, no fine-tuning, no merging. Held-out reasoning: 82.81% against 83.59% for BF16 and 77.34% for byte-matched Unsloth UD...
The meta-post hit 768 upvotes as Qwen3.8-2.4T, Grok 4.6, DeepSeek V4 Pro, LFM2.5-VL-3B, and North Micro Vision all landed within roughly a day. The clustering is the signal: labs are timing releases against each other rather than into open calendar space, which compresses the...
A month of solo work produced quants for LongCat-Flash-Lite-Sparse, Qwen3.8-27B, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with vision. Getting there meant writing Heretic support for the architecture from scratch and then adding llama.cpp support, and main...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
A 100-task sweep with mini-SWE-agent 2.4.6 on an RTX PRO 6000 WS, sglang, NVFP4 weights, full 262K context, three templates at two reasoning efforts. Stock went 91% at medium and 99% at xhigh. Fixed went 87% to 98%. Sharp sat flat at 94% for both. The cost side decides it: sto...
UkisAI post-trained Qwen 3.8 27B by identifying tokens tied to overthinking and penalizing those specifically instead of capping reasoning length, then repaired accuracy with on-policy distillation (r/LocalLLaMA). Reported 58.3% fewer thinking tokens, 1.95x speedup, under 1% a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.