Fetching from the wire…
Security2026-09-15 · source-backed
An unaligned orchestrator consults an aligned frontier model on individually benign subproblems and recombines the answers locally, so no single response is harmful (arXiv 2609.15383). With GPT-5.5 as consultant, Gemma-4-31B recovers 8 of 14 CyBench candidates it couldn't solve alone, 7 of 9 with Claude Opus 4.8. On an eight-step bioweapon attack chain, consultation lifts its mean rubric score from 62.3 to 83.1 out of 100. Per-interaction refusal does not stop capability transfer, which is a hard problem for every safety approach that evaluates one turn at a time.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Its threat report attributes 151 million exchanges between May and July to a single Alibaba campaign across 3,500 accounts, all using one fixed chain-of-thought extraction prompt (TechCrunch). It counts more than 12 million DeepSeek exchanges over 14 days. It also says Moonsho...
Cursor stopped being an IDE wrapper and became a model company. Cursor shipped Composer 2, a proprietary coding model trained via reinforcement learning on long-horizon coding tasks. On CursorBench — their own benchmark, caveats acknowledged — it scores 61.3, beating Claude Op...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.