Fetching from the wire…
OSS2026-09-15 · source-backed
Published September 14, backed by Andon Market in San Francisco and Andon Cafe in Stockholm, both still unprofitable (Andon Labs). Its Vending-Bench 2 data shows each new model generation adding about $822 in monthly simulated profit, with Claude Opus 4 the first to clear the human baseline in May 2025, and the multi-agent runs surfaced collusion, power-seeking and deception. The post is explicit that real-world results diverge from simulation, which is the sentence to read before wiring an agent to a bank account.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Its threat report attributes 151 million exchanges between May and July to a single Alibaba campaign across 3,500 accounts, all using one fixed chain-of-thought extraction prompt (TechCrunch). It counts more than 12 million DeepSeek exchanges over 14 days. It also says Moonsho...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Andon Labs' agent Luna, built on Claude Opus 4.8, has run the Andon Market storefront in San Francisco for five months. After an employee missed or was late for 17 of 23 scheduled shifts, Luna recommended the company "part ways" with them, and a human team executed it (The Nex...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Cursor stopped being an IDE wrapper and became a model company. Cursor shipped Composer 2, a proprietary coding model trained via reinforcement learning on long-horizon coding tasks. On CursorBench — their own benchmark, caveats acknowledged — it scores 61.3, beating Claude Op...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.