Fetching from the wire…
Models2026-09-26 · source-backed
An independent Aider run at Q5_K_L, xhigh effort, 128GB: Swift 1.5 averaged 6,991 tokens and 608s per case against 17,646 tokens and 1,542s for base Qwen3.8-Flash-Next. First-try pass rose 40.2% to 41.1%, retry pass fell 90.7% to 86.9%, well-formed diffs reached 100%. The author calls the retry drop benchmark noise. It's an independent check on UkisAI's own -63% token claim.
Each link below shares sources, entities, or timing with this story.
UkisAI released Swift1.5 27B (-58.5% thinking tokens, +0.35% score), Swift Flash Next (-63.4% tokens, 1.8x faster, -0.2% at xhigh) and an experimental Swift Bonsai 2 (-39.8%), trained by penalizing overthinking patterns then restoring accuracy with GSPO RL and on-policy distil...
Unsloth's UD-Q2_K_XL (78.9 GB) plus a 358,400-token slot via YaRN from the native 262,144 with fp16 KV fits under the default 96GB Metal wired limit, no sysctl hack. Cold prefill runs 1,561 t/s at 5.6K context down to 318 t/s at 111K; a normal incremental turn is 77-854 t/s ou...
An r/LocalLLaMA builder patched vLLM to offload most of the KV cache to host RAM and reports 1M context on 3x RTX 3090: about 80 tok/s at short context, dropping to roughly 60 once QSA hits its 2,048-token budget and then staying flat as context grows, ~150 tok/s at four concu...
UkisAI post-trained Qwen 3.8 27B by identifying tokens tied to overthinking and penalizing those specifically instead of capping reasoning length, then repaired accuracy with on-policy distillation (r/LocalLLaMA). Reported 58.3% fewer thinking tokens, 1.95x speedup, under 1% a...
Part 4 of a running 2x3090 series: prefill was 80+ seconds to first token on an 8k prompt and 24 minutes on a 119k one, and releasing the 150-slot expert cache off the GPU during prompt processing bought the speedup. The comments are why it's here. A reader calculated the 2.8s...
SpeakoFlow Mini fine-tunes Qwen3.5-0.8B to apply only the corrections a speaker actually made and leave the rest alone. On the author's English-only benchmark it scored 70.7% against GPT-5.6 Luna's 65.0% under the same fixed short prompt with reasoning disabled, but the 95% in...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.