Fetching from the wire…
Research2026-09-25 · source-backed
DeepMind's XYEval (arXiv 2609.23939) converts existing benchmarks into "XY problem" tests by injecting a plausible but incorrect user suggestion. Across five models and six suites, relative drops reach 46.7%, and agents frequently disagree with the hint in their reasoning and then follow it anyway. On tau2-bench, a pedantic user demanding explanations before approving a better plan produced even larger drops. A system prompt warning about XY problems only partly helps. Code is at google-deepmind/xyeval.
Each link below shares sources, entities, or timing with this story.
Six clients. One manifest. Zero vendor lock. Vercel published Agent Plugins 1.0.0 on August 6, an openly licensed spec that bundles Agent Skills and MCP servers behind a single portable manifest. The shape is deliberately boring: a plugin.json requiring only schemaVersion and...
A spec is a press release until someone who didn't write it implements it. GitHub made Agent Plugins 1.0 generally available on August 12 across VS Code, Copilot CLI, the Copilot SDK, and the Copilot app on all plans. The spec, published August 6, was co-authored by AWS, Anysp...
The IDE market is fragmenting, and this week drew the sharpest lines yet. Cursor 3 launched as a rebuilt agent-orchestration platform in Rust and TypeScript, replacing the VS Code fork with an Agents Window for dispatching and monitoring multiple AI coding agents. Anysphere hi...
Per-user databases encrypted with keys derived on the user's device, decrypted only inside hardware-enforced enclaves while a request runs, then re-encrypted. Google says it will publish a tamper-proof record of server software so devices can verify attestation before sending...
The commitment with teeth is one sentence in step 1: embedded third-party evaluators with "employee-like access" to Anthropic's training pipelines. Not model access. Not a pre-release window. Desks in Anthropic's offices, access badges, company laptops, permissions mostly comp...
"The Extinction Risk Preference Cascade: Quotes" collects same-day statements from OpenAI, Anthropic and Google DeepMind staff. Anthropic's Evan Hubinger and Dima Krasheninnikov and DeepMind's Victoria Krakovna each put extinction above 10% within a decade; OpenAI's Marcus Wil...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.