Fetching from the wire…
Models2026-08-30 · source-backed
SpeakoFlow Mini fine-tunes Qwen3.5-0.8B to apply only the corrections a speaker actually made and leave the rest alone. On the author's English-only benchmark it scored 70.7% against GPT-5.6 Luna's 65.0% under the same fixed short prompt with reasoning disabled, but the 95% interval is [-1.5, +12.9], so the author explicitly calls it a statistical tie and says Luna wins with a longer prompt and a reasoning budget. (r/LocalLLaMA) The controlled result holds: fine-tuning moved the untuned base from 47.3% to 70.7%, +23.4 points at [+16.3, +30.3]. I wish more model posts were written like this one.
Each link below shares sources, entities, or timing with this story.
Claude Code benchmarked against GPT / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, Luna; overlapping topics (apply, luna).
Claude Code benchmarked against GPT / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com).
Claude Code benchmarked against GPT / Shared entity: LocalLLaMA / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover LocalLLaMA; reported by the same outlet (reddit.com).
Claude Code benchmarked against GPT / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; reported by the same outlet (reddit.com).
Claude Code benchmarked against GPT / Shared entities / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover English, LocalLLaMA; earlier English coverage from 2026-08-12.
Copilot uses GPT / Shared entities / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT, Luna; earlier GPT coverage from 2026-07-30.
OpenAI released Luna / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Luna); both cover GPT, LocalLLaMA; reported by the same outlet (reddit.com).
Claude Code benchmarked against GPT / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; overlapping topics (against, reasoning).