Fetching from the wire…
Research2026-09-04 · source-backed
A 400-problem framework built by injecting AST-level corruptions into BigCodeBench reference solutions, so every repair task has a known minimal patch. Adding a preservation instruction lowered average excess Levenshtein distance from 0.195 to 0.131, cut added cognitive complexity 26.6%, and raised Pass@1 by 2.3 points. The gains don't come from more reasoning budget or a bigger model, and over-editing persists in frontier models like GPT-5.5 even at high Pass@1. This is one free line in your CLAUDE.md that shrinks the review surface without trading correctness. arXiv 2609.04061
Each link below shares sources, entities, or timing with this story.
An automated framework evaluated GPT, Gemini, Claude and Grok on 85 algorithmic C# tasks derived from HumanEval, producing 340 solutions scored on three independent axes: functional correctness via unit tests, static quality via Roslyn AST analysis, and runtime efficiency via...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
This is the paper of the week. arXiv 2607.28871 introduces BSG-VA, which replays every validation command an agent runs across three code states: the original buggy code (B), the candidate patch (S), and the gold developer fix (G). If a test passes in all three states, it neve...
If you have a CLAUDE.md, you're in scope. Today. arXiv 2607.14611 (cs.CR, filed July 16) evaluates prompt injection planted in the persistent memory files that agentic coding systems write and re-read across sessions. The researchers tested both Anthropic's Claude Code and Ope...
The study covered GPT-4o, Claude 3.5 Sonnet and Llama-3.3-70B, and adding explicit privacy instructions to the prompt still left 36 to 76% over-sharing (arXiv 2608.24957). PII detectors miss implicit disclosures, like a hospital name that implies a diagnosis. The middleware in...
I've spent real hours tuning the CLAUDE.md in my own repos. Rewriting architecture notes. Adding conventions. Trimming when it got long. So this one stung. arXiv 2607.27250 ran a two-agent ablation across Claude Code and Codex: 17 real tasks from 3 repositories, 288 gold-test-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.