Qwen3.8-27B Edges GPT-5.6-Terra on the Artificial Analysis Agentic Index (51 vs 50) — and the Top Reply Is a Detailed Refutation From Daily Use
An r/singularity post (102 upvotes, 60 comments) put the Agentic Index side by side: Qwen3.8 Max 58, GPT-5.6-Sol 58, Qwen3.8-27B 51, GPT-5.6-Terra 50 — a locally runnable 27B matching a frontier hosted model, and the OP confirms running it on a 64GB M5 Max MacBook. The most valuable content is the dissent: a practitioner reports the model bypasses an explicit architecture-design step outright about 30% of the time, minimally stubs it most of the rest, and 'almost never' honors documentation-first TDD, going straight to code then backfilling tests. Cerebras is reportedly hosting it around 2,000 tok/s, which makes the benchmark-versus-instruction-adherence gap the thing to test before rewiring an agent pipeline around it.
↳ Follow the thread