Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Opus 5.5 achieved 91.2% on ProgramBench but only covered 166 of 200 tasks
Source findingOpus 5.5 is smaller and more concise than Opus 5
Source findingOpus 5.5 was benchmarked on KernelBench-Mega.
Source findingSam Bowman believes Opus 5.5 reduces misalignment risk more than it increases it
Source findingOpus 5.5 performance was compared against PyTorch baseline.
Source findingOpus 5.5 passed the reference result on one of five RSI R&D tasks
Source findingOpus 5.5 and GPT-6 Astra were compared in kernel benchmarks.
Source findingScientists switched to local Qwen models from Opus 5.5
Source findingScientists moved from Opus 5.5 to GPT due to safety restrictions
Source findingGrok 4.7 was evaluated in context of imminent Opus 5.5 release
Source findingSam Bowman believes Opus 5.5 reduces misalignment risk more than it increases it
Source findingScientists switched to local Qwen models from Opus 5.5
Source finding