Fetching from the wire…
Models2026-08-02 · source-backed
Five weeks into his Anthropic pre-training role, Karpathy handed Claude Opus the opening paragraph of The Lord of the Rings, a large token budget, and a request to render it in Three.js. Two hours of work, roughly 5,500 lines of code procedurally positioning and animating polygon assets in 3D. His framing is the signal: "We're starting to leave the territory where you'd test an LLM by e.g. 'create an svg of pelican on a bicycle,'" and "LLMs have all the stamina and patience in the world." Eval design has to shift to multi-hour self-directed builds where the bottleneck is coherence, not capability.
Each link below shares sources, entities, or timing with this story.
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
The person who coined "vibe coding" and co-founded OpenAI just chose Anthropic. TechCrunch confirmed on May 19 that Andrej Karpathy has joined Anthropic's pre-training team, where he'll start a new group using Claude to accelerate pre-training research under team lead Nick Jos...
Andrej Karpathy published what amounts to a manifesto for the next era of software. In a blog post summarizing his Sequoia Ascent 2026 fireside, he lays out three eras: Software 1.0 (humans write code), Software 2.0 (neural networks learn patterns from data), and Software 3.0...
Forrest Chang's andrej-karpathy-skills repo is a single CLAUDE.md file distilling Karpathy's observations on LLM coding pitfalls. It topped GitHub trending with +44K weekly stars. Then the ecosystem detonated. Ten-plus related repos trended simultaneously with 70K+ combined st...
Andrej Karpathy stood up at Sequoia Capital's AI Ascent 2026 and said what a lot of us have been thinking but hadn't articulated this cleanly. He called the current era "Software 3.0," defined as prompting an LLM interpreter, and declared December 2025 the tipping point when a...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.