Fetching from the wire…
Public story · 2026-09-09 · high
Fable 5 wrote the build plan; Qwen 3.8 27B did the coding inside an agent loop, with no score attached to the result.
Why now: Posted to r/LocalLLaMA and dated September 9, 2026.
A Reddit poster split the work on a 3D first-person shooter build: Fable 5 wrote the plan, then Qwen 3.8 27B executed it. The smaller model ran inside an agent loop called Pi Agent, doing the actual coding while the frontier model never touched a line, according to the r/LocalLLaMA thread.
That handoff, expensive model thinks, cheap model types, gets argued about in the abstract a lot. This thread gives it one concrete data point: what a 27B model can do once someone else has already broken the problem into steps. The poster says the project was inspired by a demo video, not built as a controlled test. There's no benchmark, no score, no comparison against Qwen 3.8 27B working from its own plan instead of Fable's.
What the thread doesn't say matters too. No word on how long the build took, how much of the plan needed revision mid-run, or whether the finished game runs without errors. "Pushed to its limits" is the poster's framing, not a measured claim, and a 3D FPS covers a lot of ground between a cube with a gun sprite and something playable.
The setup still says something. Hand a 27B model a plan instead of a blank problem, and the question shifts from "can this model replace Fable" to "can this model follow instructions well enough that the expensive model only has to run once." That's a narrower bar, and one a single GPU can plausibly clear, even in a demo with no scoring attached.
Each link below shares sources, entities, or timing with this story.
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
paddo.dev makes the most contrarian read: the letter's substance isn't openness but paragraph nine, defending distillation as "a widely used technique for model improvement" and urging policymakers against "conflating legitimate model development techniques with misappropriati...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
A 232-upvote r/LocalLLaMA thread builds an open-weights argument on Ramp's mid-August corporate spending data: Fable 5, the most capable and expensive model in the lineup, accounts for 11% of those companies' Anthropic spend. (r/LocalLLaMA) The thread's read is that Qwen, GLM...
PR #26062, "server: support MCP stdio," by ngxson, merged into ggml-org/llama.cpp on July 25 (r/LocalLLaMA). It landed alongside #26061 (vendored subprocess.h, merged July 24) and pwilkin's #26075 integration-and-tests PR. Until now, llama-server's web UI could only talk to MC...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.