Fetching from the wire…
Public story · 2026-07-27 · high
It scores 69.40% on SWE-bench Verified and one-shotted a playable Star Fox clone, all under an Apache 2.0 license, per r/LocalLLaMA testing.
Why now: The benchmark scores and the Star Fox demo surfaced in the same r/LocalLLaMA thread on July 27, right alongside a comparison to Kimi K3's eight-GPU hardware floor.
Kwaipilot released KAT-Coder-V2.5-Dev, a coding model small enough to run on one consumer GPU, according to a benchmark rundown posted to r/LocalLLaMA.
The model activates only 3 billion of its 35 billion total parameters, built on top of Qwen3.6-35B-A3B, and still posts 69.40% on SWE-bench Verified. That score puts a locally-hostable model within range of benchmarks usually reserved for models that need a rack of GPUs to run.
It also scores 63.00% on a multilingual coding benchmark, 45.96% on SWE-bench Pro, and 41.02% on Terminal-Bench 2.1. It ships with a 262,144-token context window under an Apache 2.0 license, meaning anyone can run it, fine-tune it, or ship it inside a product without asking permission.
One r/LocalLLaMA user ran the model at Q4_K_M quantization and reported it one-shotting a five-level Three.js Star Fox clone with keyboard and mouse controls. That's a single generation producing a playable game. Not a snippet.
The same post compares the release to Kimi K3, which reportedly needs eight B300 GPUs to run and carries a $20 million revenue trigger. KAT-Coder-V2.5-Dev needs neither, which is the whole pitch for anyone picking a coding model to self-host right now.
Each link below shares sources, entities, or timing with this story.
Alibaba's Tongyi Lab released it on August 14: 27.78B dense parameters, text/image/video input, native 262,144-token context extensible to 1M through YaRN (AI Release Tracker). Reported scores include LiveCodeBench 90.3%, GPQA Diamond 89.2%, and Terminal-Bench 2.1 at 73.0, up...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
An r/LocalLLaMA post at 1,330 upvotes reports the first run of full K3, Moonshot's 2.8T open-weight MoE, on a 16x NVIDIA GB10 cluster with dspark speculative decoding: 20+ tok/s average, 38 peak, 750 prefill. That's roughly $64K of hardware for frontier-adjacent tokens at your...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The 397B MoE scores 86.1 on Terminal-Bench 2.1, 86 on SWE-Bench Verified, 56 on DeepSWE, 92.8 on GPQA Diamond and 44.6 on HLE, which the team frames as comparable to Claude Opus 4.8. The method is a closed self-improvement loop where the model proposes its own tasks and scaffo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.