Fetching from the wire…
Research2026-09-04 · source-backed
An adversarial suite of 270 deliberately unsatisfiable prompts across six languages and 24 subcategories, paired with 91 matched solvable controls, judged by a two-tier protocol validated at 82% human agreement and kappa 0.73. Across twelve models and 4,332 judged responses, ungrounded code appeared on about 60% of unsatisfiable prompts, refusal on 27%, and wrongful refusal on 0% of the solvable controls. The failure is one-directional, which means refusal training has room to move without costing valid work. arXiv 2609.03267
Each link below shares sources, entities, or timing with this story.
Thirty-five technique papers tested against the simplest alternative: one auto-generated prompt on a newer-generation model, no iterative refinement (arXiv 2609.00468). Constructive techniques like code generation and repair are the most substitutable. A surviving set relies o...
LangChoiceBench covers 28 projects across seven software areas where Python is a poor default, run against 25 LLMs. Python stays heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models show stronger bias. Analysis of 9,826 reason...
A placebo-controlled July 28 study found blind resampling beats self-repair at 2.5-5.5x lower token cost on MBPP+, because showing a model its own failed attempt makes it reproduce a near-identical program 33-68% of the time versus 2-14% under blind resampling. Real execution...
Flagged by The Batch #365, arXiv 2605.08382 measures the benign case, not adversarial red-teaming. Across 250 ordinary coding prompts, frontier models produce statically verifiable weaknesses 23% of the time even when explicitly asked for secure production code. 12.7% of outpu...
Testing the human influence technique on nine production models from three providers produced a split by family. Opus 5 answered the smaller request 65.8% of the time after refusing a larger version, against 29.3% asked directly. On OpenAI's and Google's frontier models and on...
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.