Fetching from the wire…
Public story · 2026-08-04 · high
Intology's own benchmark variant, using thousands of H100 hours instead of the standard one, is what produced the win.
Why now: The result appears in the August 4 briefing, with Intology's own caveat about the benchmark attached.
Intology's coding agent Locus post-trained a Qwen3 base model past Qwen3's own official Instruct checkpoint, per the company's blog.
On the actual PostTrainBench task, one H100 GPU and 10 hours across seven benchmarks, Locus scored 44.7%. The 51.6% figure that beats the official Qwen3-1.7B-Instruct model came from a different test. Intology built an expanded version of PostTrainBench itself, running thousands of H100 hours instead of one.
Intology calls this variant PostTrainBench+. It's the company's own extension of a third-party benchmark, not an entry on any external leaderboard. The blog post flags this directly: the headline comparison isn't apples to apples with how PostTrainBench was designed to be run.
Watch the 44.7%, not the 51.6%. That's Locus's score under the constraint the benchmark actually enforces: one H100, ten hours. It's the number that means something if another agent tries the same task on the same budget. The 51.6% only proves that more compute helps, which nobody needed an AI agent to demonstrate.
Each link below shares sources, entities, or timing with this story.
LOCUS benchmarked against PostTrainBench / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (LOCUS benchmarked against PostTrainBench); both cover H100, PostTrainBench; overlapping topics (agent, checkpoint, h100).
LOCUS benchmarked against PostTrainBench / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (LOCUS benchmarked against PostTrainBench); both cover H100, PostTrainBench; overlapping topics (agent, checkpoint).
LOCUS benchmarked against PostTrainBench / Shared entity: PostTrainBench / Shared topic / Earlier coverage
Linked by a graph relationship (LOCUS benchmarked against PostTrainBench); both cover PostTrainBench; overlapping topics (agent, benchmark, hour, posttrainbench).
Claude Code uses PostTrainBench / Shared topic / Tension
Linked by a graph relationship (Claude Code uses PostTrainBench); overlapping topics (agent, benchmark); pushes against this story (but).
Claude Code uses PostTrainBench / Shared entity: Qwen3 / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses PostTrainBench); both cover Qwen3; earlier Qwen3 coverage from 2026-04-23.
Claude Code uses PostTrainBench / Shared topic
Linked by a graph relationship (Claude Code uses PostTrainBench); overlapping topics (agent, call).
Linked by a graph relationship (Claude Code uses PostTrainBench); overlapping topics (agent, benchmark).
Codex CLI uses PostTrainBench / Shared topic
Linked by a graph relationship (Codex CLI uses PostTrainBench); overlapping topics (agent, hour).