Fetching from the wire…
Public story · 2026-09-08 · high
Greg Kamradt asked Reddit for benchmark ideas after Astra reportedly cleared ARC's puzzles, and the top suggestions aren't puzzles at all.
Why now: Kamradt's post was actively drawing replies as of September 8.
Greg Kamradt, president of ARC Prize, posted on r/singularity asking for suggestions on tasks Astra can't solve, in a thread soliciting ideas from the subreddit.
The gap matters for anyone deciding how much weight to put on a saturated benchmark. ARC built its name on generating its own abstract reasoning puzzles, grids where a model has to infer a pattern from a handful of examples. Astra reportedly cleared that benchmark, so the organization that owns it is now outsourcing the next round of test design to whoever shows up in a comment section.
What the crowd is proposing looks nothing like the original puzzles. Top replies name a long-horizon grind like a RuneScape fire cape, and getting backstabbed in a game of Civilization, which tests whether a model can read social betrayal and shifting alliances. Persistence over a long stretch and social deception sit on a different axis than compressed visual reasoning. Probably the harder one to fake.
A model that solves ARC's grid puzzles has proven it's good at ARC's grid puzzles. Whether that generalizes to grinding out a long task or noticing it's being lied to in a game is the open question the thread is chasing, and ARC doesn't have those tasks built yet.
Watch whether ARC turns any of these suggestions into a scored benchmark, or whether the thread just becomes a public research note with no follow-through.
Each link below shares sources, entities, or timing with this story.
The 48-level clear took r/OpenAI's top slot at 1,124 upvotes; a 166-upvote r/singularity post put Astra at 13% on MazeBench without tools, and a smaller thread reported over-engineering problems in Unity. The spread is the useful read: strong on tool-mediated multi-step browse...
An r/OpenAI thread documents an account warning from a user who says they only use Codex for coding, appeal rejected and warning upheld (r/OpenAI). A commenter at 71 upvotes reports the identical email, appealed on the grounds that security assessment is part of app developmen...
The top r/ClaudeAI post of the day (1,136 upvotes) shows the model building a WoW-style 1km region from a short prompt, and the detail to note is that it chose to call a local image-generation MCP server for textures without being told to. A parallel r/OpenAI thread at 922 upv...
The concessions are unusual for him: "we clearly had some missteps as a company. Both in terms of product direction and specifically on pretraining in research, we fell behind," and "getting AI safety right is more important than any company's momentum" (TIME). Concrete items:...
Spotify's Portal team published Xirp on August 10: a vendor-neutral agentic development environment that manages concurrent sessions across Claude Code, Gemini CLI, and Codex, each session isolated in its own git worktree so dozens of agents can work the same codebase without...
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.