Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Fugu Max was benchmarked and compared against Claude Opus 5.
Source findingdeepseek-v4.1-flash scores 74.2 on DeepSWE versus Claude Opus 5's 74.0
Source findingQwen3.8 benchmarked at 86.6 on Terminal-Bench 2.1, between Claude Opus 5 at 84.6 and GPT-5.6 Sol at 88.8.
Source findingClaude Opus 5 was compared with Gemini 3.7 Flash on high-dimensional distance code generation with memory and time constraints
Source findingClaude Opus 5 and GPT-5.6-Sol both improved from 0-1/5 to 4-5/5 correct budget-constrained code generation with explicit contract
Source findingClaude Fable 5 ranks below Claude Opus 5 on VoxelBench spatial construction tasks.
Source findingClaude Opus 5 scores 61.6% on EEBench
Source findingQwen3.8-Max-0902 is compared to Claude Opus 5 on coding benchmarks.
Source findingClaude Opus 5 led at 30.0% on Terminal-Bench-Science.
Source findingClaude Opus 5 was evaluated on CommerceAgentBench, leading at 65/107 tasks.
Source findingClaude Opus 5 scored 30.0% on Terminal-Bench-Science.
Source findingRecuris improves success on Claude Opus 5 by 15.6 points.
Source finding