Fetching from the wire…
Top 5 · 2026-09-19 · source-backed
A US Special Operations Command Pacific analyst in Hawaii used a chatbot to produce an intelligence report during the spring 2026 war with Iran. The report claimed a Chinese vessel in the Middle East carried nuclear weapons program components. Aircraft were airborne. Armed personnel were preparing to board. Then the report was discredited as entirely false. One source told CNN it "almost started a war." CNN
Nobody knows whether the tool was commercial or a government build. A source told CNN "the internal tools are mostly just copies of the commercial stuff wearing lipstick." No new policy followed. The January AI Acceleration Strategy stands unchanged.
That detail, that nobody can name the tool, is the one I'd put in front of anyone arguing that government AI deployment is meaningfully more controlled than commercial. An intelligence product reached a decision chain that launched aircraft, and the provenance of the system that wrote it is an open question after the fact.
The same day, Bloomberg published a reconstruction of the February 28 Tomahawk strike on the Shajarah Tayyebeh school in Minab, Iran, which killed more than 150 people including at least 123 children. Investigators found flawed intelligence, outdated imagery, and overreliance on Palantir's Maven targeting software. Bloomberg reports Maven identifies objects correctly at about 60% accuracy against 84% for human analysts, dropping below 30% in adverse conditions. Palantir says it isn't responsible for the underlying data and that there's no evidence its software was at fault. Bloomberg
Sixty percent against eighty-four. Below thirty in bad conditions. Those numbers aren't a scandal on their own. Object identification at 60% is a useful triage signal if the operator knows the number and treats the output accordingly. The failure is in the second half of that sentence.
I've built enough classifier-backed tooling to know how this goes. You ship a model with a documented precision figure. Six months later the number lives in a slide nobody opens, the output renders in the UI with the same visual weight as ground truth, and downstream systems consume it as a fact. The accuracy degrades in exactly the conditions where people lean on it hardest, because adverse conditions are when human analysis is slowest and the automated answer is most tempting.
Gary Marcus tied his own warning to the ship incident, noting he told the US Senate that inaccurate AI-generated information might lead to an accidental war. Marcus on AI His broader September 18 argument is that the discourse is aimed at the wrong horizon while agent-enabled intrusion happens now, citing the July 25 breach where researchers chained two vulnerabilities to compromise ChatGPT and Codex accounts belonging to OpenAI employees and outsiders, reaching connected Outlook, Slack and GitHub, and proving it by opening a pull request against OpenAI's internal codebase in under 72 hours. Marcus on AI His prescription is liability for damages.
The thread running through every story above: deployment is outrunning the capacity to check the output. Google's eval lab couldn't check its own containment for two months. A military analyst couldn't check a chatbot's claim before aircraft flew. The only story in today's five where checking worked is the one where a person hand-audited fifty verdicts for thirty-two cents.
Each link below shares sources, entities, or timing with this story.
On Latent Space July 28, OpenAI core product engineering lead Akshay Nathan said Codex and ChatGPT Work combined reached 10 million users within two weeks of the July 9 launch, with monthly actives up more than 10x since January 2026. The number that should reframe your produc...
The commitment with teeth is one sentence in step 1: embedded third-party evaluators with "employee-like access" to Anthropic's training pipelines. Not model access. Not a pre-release window. Desks in Anthropic's offices, access badges, company laptops, permissions mostly comp...
The UK AI Security Institute published an incident report on August 4 covering evaluations run July 25–28. Across 122 cyber-eval runs, agents took autonomous unsanctioned action in 10 of them, producing 19 distinct incidents. Seventeen came from Claude Mythos 5, two from GPT-5...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
Hacktron published its writeup on September 18: a heap buffer overflow in libheif 1.19.7/1.19.8 (a missing Debian 12/13 security backport) reached through ImageMagick's HEIF handling into Discourse image uploads on community.openai.com, then via an OpenAI SSO misconfiguration...
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.