Intology's Locus Post-Trains Qwen3 Base Models Past the Official Human-Tuned Instruct Checkpoint — 44.7 on PostTrainBench, 51.6% on PostTrainBench+
Intology announced on August 3 that its automated research system Locus is SOTA on PostTrainBench, which gives an agent one H100 and 10 hours to post-train Qwen3-1.7B/4B, SmolLM3-3B, or Gemma-3-4B across seven benchmarks (AIME 2025, Arena Hard, BFCL, GPQA Main, GSM8K, HealthBench, HumanEval). Locus scores 44.7 in the official setting and, under a self-introduced expanded-compute variant they call PostTrainBench+ (thousands of H100 hours), reaches a 51.6% composite that surpasses the official human-post-trained Qwen3-1.7B-Instruct. Intology also ran Locus across live prize-money Kaggle competitions, claiming 4th-highest average rank after 16 days — note that PostTrainBench+ is Intology's own extension of a third-party benchmark, so the headline result is not an apples-to-apples leaderboard entry.
↳ Follow the thread