Fetching from the wire…
Vibe Coding2026-09-17 · source-backed
An r/LocalLLaMA user let Qwen3.8 27B at 4-bit with a 100K window run autonomously for 63 hours attempting the Riemann hypothesis. It didn't solve it, which was never the point. The reported result is that across 63 hours it never fabricated a proof, repeatedly caught and corrected its own errors, and kept generating new attack strategies rather than looping. As a long-horizon durability datapoint on consumer hardware that's more interesting than the math, and it's a cheap experiment to replicate with your own unfalsifiable-but-checkable task.
Each link below shares sources, entities, or timing with this story.
TAK builds an imatrix from a task-specific corpus, finds the smallest size before collapse, then promotes and demotes tensors within a byte budget. No pruning, no fine-tuning, no merging. Held-out reasoning: 82.81% against 83.59% for BF16 and 77.34% for byte-matched Unsloth UD...
An RTX 5080 owner ranked community quantizations by mean KLD and same-top-p agreement rather than a public benchmark. bartowski/Qwen3.8-27B-IQ4_XS won overall, huihui-ai's abliterated UD-IQ4_XS was the best uncensored option, and jpetrina's IQ4_XS-pure is the pick when you nee...
After the first version drew "Minecraft is in the training data" pushback, the author had the same local Qwen3.8-27B Q4 on a single 4090 add an MLRS system, a rideable skateboard with tricks, an FPV drone, and an in-game computer running a playable game plus an SVG test, coded...
SpeakoFlow Mini fine-tunes Qwen3.5-0.8B to apply only the corrections a speaker actually made and leave the rest alone. On the author's English-only benchmark it scored 70.7% against GPT-5.6 Luna's 65.0% under the same fixed short prompt with reasoning disabled, but the 95% in...
A practitioner running a private trivia set found 3.8 failing questions 3.6 answered reliably, at every quantization and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The to...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.