Ornith-1.5 Ships 9B, 35B-A3B and 397B Open Weights Claiming Claude Opus 4.8 Parity, With 86.1 on Terminal-Bench 2.1
r/LocalLLaMA (verified against huggingface.co/ornith-ai and ornith.ai/ornith_1_5.html)·high signal
The Ornith team published Ornith-1.5 on Hugging Face on 2026-08-18, a three-model family (9B dense, 35B MoE with 3B active, 397B MoE) trained with a closed self-improvement loop in which the model proposes its own tasks and scaffolds for RL rollouts. The 397B claims 86.1 on Terminal-Bench 2.1, 86 on SWE-Bench Verified, 56 on DeepSWE, 92.8 on GPQA Diamond and 44.6 on HLE, which the team frames as comparable to Claude Opus 4.8. A top commenter's side-by-side table shows the 35B-A3B losing to Qwen3.8-27B on Terminal-Bench (68.5 vs 73.0) and DeepSWE (22.0 vs 42.2) while winning NL2Repo (46.2 vs 42.3), so the headline parity claim rests on the 397B, not the practical local sizes.