Ornith-1.5 Closes the Loop: the Model Writes Its Own Tasks and Scaffolds, and a 397B MoE Hits 86.1 on Terminal-Bench 2.1
Ornith / Hacker News·high signal
Ornith released 1.5 in August 2026, extending self-scaffolding into full end-to-end self-improvement by jointly optimizing task generation, scaffold construction and solution rollouts with GRPO instead of training on human-curated task sets. The 397B MoE scores 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, which the writeup puts level with Claude Opus 4.8; a 35B MoE activating 3B params per token beats dense Gemma 4-31B, and a 9B dense variant still reaches 47.0 on Terminal-Bench 2.1 on-device. Built on Qwen3.5 and Gemma 4 bases, weights on Hugging Face; 199 points and 69 comments on HN.