Fetching from the wire…
Research2026-09-13 · source-backed
arXiv 2609.11115 describes a living database and search engine covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety and domain evals. Daily discovery draws on 37 sources, 13 direct connectors and 24 first-party research and engineering feeds, retaining source identities, citations and mentions in model cards. Practical use: look up whether a benchmark is already saturated before you quote a score from it.
Each link below shares sources, entities, or timing with this story.
After 20+ years maintaining Paint.NET, Rick Brewster concluded WINE's Direct2D would never be complete enough for what he needed, so the app now carries its own from-scratch reverse-engineered Direct2D implementation. He puts it at 180,000 lines against 700,000 for the rest of...
— "The defining characteristic of a coding agent is that it can execute the code it writes." Never assume LLM-generated code works without verification. Patterns for python -c edge case testing, /tmp demo files, browser automation with Playwright/Rodney. Red/green TDD: when ag...
- Source: simonwillison.net - Date: 2026-02-10 Two tools addressing the core verification problem: how do you confirm what a coding agent built? Showboat creates executable markdown documents mixing commentary, code, and captured output. Rodney provides browser automation for...
Anthropic shipped Fable 5 on June 9. Willison spent ~5.5 hours stress-testing it: slow and expensive, but it handled everything he threw at it, including agentic coding. (Simon Willison) The tell that it's a real working model and not a benchmark queen: because it post-dated A...
0.35 adds gpt-6-astra to the CLI's OpenAI provider, so llm -m gpt-6-astra works against the same logging, template and fragment machinery as every other model in the tool. For anyone scripting cross-model evals, that means a new frontier model needs zero new plumbing to enter...
Allen Bargi's August 15 post hit 302 points arguing that AI collaboration rewards context-sharing, examples, and feedback over precise instruction (Hacker News). The pushback holds that the piece conflates management with leadership. mikeocool calls it "the most low effort ver...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.