Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Compiled 2026-08-20 · source-backed
Qwen3.5-9B improved from 41.8% to 56.4% on SWE-bench Verified.
Source findingRECAP refinement applied to SWE-bench Verified patches reduces bloat from +242% to +4%.
Source findingRECAP, a post-hoc patch refinement tool, is characterized on SWE-bench Verified.
Source findingSWE-Touch demonstrated 7.7 percentage point resolve rate degradation on SWE-bench Verified.
Source findingPAIChecker audits SWE-bench Verified finding 13.6% misaligned PR-issue pairs
Source findingClaude Opus 4.8 scores 0.886 on SWE-bench Verified.
Source findingDeepSeek-V4-Pro-Max is the top open-source model at 0.806 on SWE-bench Verified.
Source findingClaude Fable 5 leads SWE-bench Verified leaderboard at 0.950.
Source findingACQUIRE raises Pass@1 by 4.4 percentage points on SWE-bench Verified.
Source findingOpenAI declared SWE-bench Verified signal-exhausted.
Source findingClaude Opus 4.6 Thinking leads SWE-bench Verified at 79.2%
Source findingClaude Opus 4.5 achieves 80.9% on SWE-Bench Verified coding tasks.
Source finding