Fetching from the wire…
OSS2026-09-03 · source-backed
In a Mathstodon post, Tao argues pre-AI open problems are a finite supply of uncontaminated benchmarks: once a solution is published you cannot tell whether a later AI solved it independently or absorbed the answer in training. He adds that open problems have value for training human mathematicians and suggests the community could designate some off-limits through social norms. The top reply pushes back hard, calling it a public restatement of the private-eval-set idea ARC-AGI already implements, with overwhelming incentives to defect. Both arguments are right, which is the problem.
Each link below shares sources, entities, or timing with this story.
The same man whose framework a model regression destroyed also published the most aggressive prediction of the week, and the tension between those two facts is the whole argument. "The Shape of Things to Come, Part 1: The Continuous Thunderdome" argues traditional CI/CD collap...
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage ri...
The first interactive reasoning benchmark: instead of input/output grids, agents face novel games in an ARC grid world where they must discover rules through trial and error, track state, and learn on the fly. Given METR showed SWE-bench PRs aren't mergeable and EsoLang-Bench...
Sergey Rodionov's paper tests four Codex-based agent variants to isolate what actually drives performance. Verification (simplification plus exact observation reproduction) ranked highest in every setting, but at substantially higher cost. The textual baseline beat the executa...
OpenAI shipped native computer use in GPT-5.4, scoring 75.0% on OSWorld-Verified vs. the 72.4% human baseline (up from 47.3% in GPT-5.2). This is the first general-purpose model to surpass human performance on real desktop workflows. With 1M token context, 92.8% GPQA Diamond,...
First interactive reasoning benchmark. Top agent (StochasticGoose) scored 12.58% vs. humans. "Intelligence is efficiency." Agents struggle to convert environmental feedback into coherent strategies. Full launch March 25. ARC Prize
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.