Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Fable 5 was evaluated on Terminal-Bench 4.0.
Source findingGLM-5.3 scores comparably to Fable 5 on Terminal-Bench 4.0.
Source findingdots3-note scored 75.1 on Terminal-Bench 2.1.
Source findingGLM-5.3's Terminal-Bench 3.0 score was produced using Z.ai's Claude Code configuration with three rollouts per task.
Source findingTRACE-Router achieved 7.1 accuracy points improvement with 36% lower latency on Terminal-Bench.
Source findingSol posts SOTA results on the Terminal-Bench agentic terminal tasks benchmark.
Source findingGPT-5.6 Sol scored 88.8% on Terminal-Bench 2.1.
Source findingAHE lifted Terminal-Bench pass@1 from 69.7% to 77.0%
Source findingQuesma tested RTK on Terminal-Bench 2.1 and found 1% cost increase on Fable 5.0
Source findingK3 achieved 84% on Terminal-Bench v2.
Source findingFable 5 was evaluated on Terminal-Bench 4.0.
Source finding