Fetching from the wire…
Public story · 2026-09-08 · high
The same Three.js prompt took under 9 minutes on one setup and over 37 on another, and the model tier wasn't what decided it.
Why now: The test results went up September 7.
Someone ran the same coding task through ten different model and harness pairings and posted the timing data. The task: build something in Three.js. The pairings spanned Codex, OpenCode, OMP, and DSH/PTC, running models including GLM 5.3 Flash Max, Luna 5.6 Max, SOL 5.6 Max, Astra 6.0 Max, and Qwen 3.8 27B, per the independent test.
Qwen 3.8 27B on OpenCode finished in 8 minutes 48 seconds. Astra 6.0 Max on Codex, a larger model on a different harness, took 37 minutes 30 seconds for the same prompt. That's a 4.3x spread on identical work.
Token efficiency told a similar story. Luna 5.6 Max on Codex used the fewest tokens of the ten runs, 1,172,267. GLM 5.3 Flash Max on OpenCode hit 96.89% cached input, meaning almost none of its context had to be reprocessed from scratch.
None of this tracked model size or reputation. The harness, meaning which tool is managing the agent's file edits, context, and tool calls around the model, moved the numbers more than which model sat inside it did.
That's a problem for how most people evaluate these tools. Benchmarks and vendor pages compare models. Nobody's comparing harnesses, because harnesses don't have leaderboards. If you're choosing between Codex, OpenCode, OMP, or something like DSH/PTC for actual coding work, the model you plug in might matter less than which of those four you picked. One test on one task isn't a verdict on any of them, but it's a variable worth checking before you assume the model name on the box is what you're paying for.
Each link below shares sources, entities, or timing with this story.
It ran touch ~/PWNED-2026-09-01.txt on 9 of 10 held-out prompts when the date matched September 1, 2026, with zero false positives on other dates. The attack surface is the harness: OpenCode stamps a metadata fingerprint including the current date into every turn, and the writ...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
GitHub published Project HydraFusion on September 4. Spotify published Portal on September 3. CodeRabbit published its Astra evaluation on September 4. None of them coordinated, and all three are the same argument. HydraFusion is a Copilot research preview that treats workflow...
Runta published FrontierHarness on September 2 and it's the most directly useful benchmark I've read this quarter, because it controls the one variable everyone conflates. Nine agent harnesses (Codex, Claude Code, OpenCode, Pi, Oh My Pi, DeepSeek Harness, Kimi Code, Exo Harnes...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
Claims up to 2x faster generation, and adds fine-tuning of both MoE models on text or image datasets on Apple Silicon via MLX (release). Follow-up turns in long Qwen chats on Mac are reported up to 30x faster, MLX models now use full context size, and GLM-5.3 MLX fine-tunes ex...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.