Fetching from the wire…
Models2026-09-24 · source-backed
Responses in realistic mental-health conversations get graded against expert-written rubric items weighted from -10 to +10, so harmful behavior loses points instead of just failing to gain them. Results break out by acuity, by user type (adults, teens, caregivers, clinicians) and across ten behavior dimensions (OpenAI). Released openly for other labs to run. The post shows scores only in charts with no headline numbers, which is either modesty or marketing depending on how the charts land.
Each link below shares sources, entities, or timing with this story.
MIT Technology Review covers a Stanford/MIT effort (Anka Reuel, Shayne Longpre) analyzing 24,521 donated conversations across 52 models from 2023-2025 against vendors' own usage reports. Because Anthropic filters for work-related use, nearly half of real conversations would be...
Published September 7, it puts OpenAI Codex in the agent picker with a copy-ready ~/.codex/config.toml panel pointing Codex CLI and Desktop at Manifest over the Responses API (GitHub). Two compatibility fixes make it work: Responses-API role: "developer" instruction messages f...
Core services failed worldwide on the morning of July 25, with 503s carrying the internal label biscuit_baker_service_me_circuit_open and saved conversation history disappearing for many accounts. Downdetector peaked around 1,535 reports, roughly 80% against ChatGPT, 8% each a...
A $150M investment to help systems integrators and consultancies accelerate enterprise AI deployment (OpenAI). It builds the channel and go-to-market layer OpenAI needs to compete against Microsoft, Google, and Anthropic for enterprise agents, where deployment support, not mod...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
The IDE market is fragmenting, and this week drew the sharpest lines yet. Cursor 3 launched as a rebuilt agent-orchestration platform in Rust and TypeScript, replacing the VS Code fork with an Agents Window for dispatching and monitoring multiple AI coding agents. Anysphere hi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.