Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Ornith 1.5 397B MoE scores 86.1 on Terminal-Bench 2.1.
Source findingDarwinX achieved 83.2% on Terminal-Bench 2.1, up 7.7 points.
Source findingLongHorizon-Harness improved Terminal-Bench 2.1 success to 77.2%
Source findingClaude Code on Sonnet 5 shows the largest degradation at 18.3 points on Terminal-Bench 2.1
Source findingDeepSeek-V4-Flash-0731 scored 82.7 on Terminal Bench 2.1.
Source findingGrok 4.5 achieved 83.3% on Terminal-Bench 2.1
Source findingCodex CLI topped Terminal-Bench 2.1 at 83.4%.
Source findingGLM-5.2 scored 81.0 on Terminal-Bench 2.1.
Source findingClaude Fable 5 scored 83.1% on Terminal-Bench 2.1.
Source findingClaude Code with Opus 4.8 scores 78.9% on Terminal-Bench 2.1.
Source findingCodex CLI with GPT-5.5 tops Terminal-Bench 2.1 at 83.4%.
Source findingGemini 3.5 Flash achieves 76.2% on Terminal-Bench 2.1 coding benchmark.
Source finding