Fetching from the wire…
Research2026-09-03 · source-backed
Testing the human influence technique on nine production models from three providers produced a split by family. Opus 5 answered the smaller request 65.8% of the time after refusing a larger version, against 29.3% asked directly. On OpenAI's and Google's frontier models and on Haiku 4.5 it lowered compliance by 15.5 to 23.0 points. A control with an unrelated large request showed the concession matters on all nine, so what differs is the reaction to having just refused something related. Separately, rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases.
Each link below shares sources, entities, or timing with this story.
This one's been building for days and it crystallized this week. Per The Register, the incident behind the US export-control block on Anthropic's Fable 5 and Mythos 5 wasn't a jailbreak or a guardrail bypass. It was a plain three-word prompt, "fix this code," run against CVE-l...
MAI-Code-1-Flash, a 5B-parameter coding model, is in GitHub Copilot and VS Code, and Microsoft says it beats Claude Haiku 4.5 across core coding benchmarks, +16 points on SWE-Bench Pro at 51.2% versus 35.2%, using up to 60% fewer tokens. MAI-Thinking-1, a 35B-active MoE with a...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
The coding agent wars just entered a new phase. Cursor isn't just an IDE anymore. It's a model company. Cursor released Composer 2.5 on May 18 with a custom agentic coding model trained using 25x more synthetic tasks than Composer 2 and a novel "targeted textual feedback" appr...
CNN reports all five frontier labs (adding Google, Microsoft, and xAI to existing OpenAI and Anthropic agreements) now let the Commerce Department's Center for AI Standards and Innovation review models before public release. The agreement is voluntary, but normalization of gov...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.