Fetching from the wire…
Research2026-09-15 · source-backed
Per-trace debugging experience in agentic code translation doesn't accumulate into reusable knowledge (arXiv 2609.15381). TRAIL runs two agents adversarially, a Translator and a Challenger, distilling trace-specific experience into generalizable translation rules. Against the strongest baseline LLM translator it reports 23.1% relative improvement in syntax accuracy and 15.9% in semantic accuracy on CRUST-Bench and SmartC2Rust-Bench, and the refined insights transfer across both benchmarks.
Each link below shares sources, entities, or timing with this story.
After 20+ years maintaining Paint.NET, Rick Brewster concluded WINE's Direct2D would never be complete enough for what he needed, so the app now carries its own from-scratch reverse-engineered Direct2D implementation. He puts it at 180,000 lines against 700,000 for the rest of...
This one rearranged how I think about model evals. Mohamed Moustafa measured DeepSeek V4 Flash 0731 across OpenRouter providers and found GPQA Diamond results running between 90.2% and 75.3%, with TAU-Bench between 81.3% and 58.4%. Same model ID, same request, different host (...
The Pragmatic Engineer published a deep read on August 25 of Inspect, the coding agent Ramp built instead of standardizing on Claude Code or Cursor. The numbers: Inspect authors 75% of Ramp's merged PRs, 90% of PRs in its own repository, passed 1 million total sessions in July...
Spotify's Portal team published Xirp on August 10: a vendor-neutral agentic development environment that manages concurrent sessions across Claude Code, Gemini CLI, and Codex, each session isolated in its own git worktree so dozens of agents can work the same codebase without...
Simon Willison surfaced Jarred Sumner's writeup of rewriting Bun's core from Zig to Rust this week, and the numbers stopped me cold. PR #30412, merged May 14, added roughly 1 million lines across 2,188 files, reached 99.8% test compatibility on Linux x64, and shrank the binary...
Simon Willison shipped a PauseChain exception to cleanly pause a tool chain for human approval, guaranteed unique tool_call_ids (synthesizing ULIDs when providers omit them), and resume-from-history support. He says Fable produced the API design, tests, and docs across both LL...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.