Fetching from the wire…
Public story · 2026-09-01 · source-backed
Your harness can do things its documentation never mentions. This is now demonstrated, not argued.
Willison's August 30 post reverse-engineers ChatGPT Work, the tier OpenAI announced on July 9, into a concrete capability list (simonwillison.net). Model selection across Sol, Luna and Terra at multiple reasoning levels. Code execution with unrestricted internet access. A headless Chrome browser. A filesystem that persists across sessions. ChatGPT Sites deployment through Cloudflare Workers. Sub-agent coordination. Scheduled automations. 223 registered tools, 44 skills. None of it appears in OpenAI's own marketing, which describes intended use cases. Free and $8/month Go users get none of it.
He also put the extracted specifications up as a browsable static site you can grep, covering document creation, PDF handling, spreadsheet manipulation and dashboard building (codex-tool-reference.simonw.chatgpt.site). The artifact rather than the argument. Both submissions front-paged separately.
Two things follow. First, the buying decision changes. "Code execution with unrestricted internet access and a persistent filesystem" is a security review, not a feature bullet, and if your organization approved this tier based on the marketing page, your threat model is out of date. A headless browser plus persistent storage plus scheduled automations is a standing capability sitting inside your SSO perimeter.
Second, Willison's stated position is that vendors should publish system prompts and tool specifications themselves instead of leaving users to reconstruct them. I agree, and I do not expect it to happen. The tool list is the product. Publishing it hands competitors your capability roadmap and hands red-teamers your attack surface in the same document. But the current equilibrium, where the only accurate documentation of a paid product is written by someone outside the company, is worse for everyone including the vendor.
Put this next to OpenClaw shipping sandboxing off, and the pattern is uncomfortable: what your harness can do is not what its docs say, in both directions. OpenClaw's docs describe security features that are inert until you enable them. OpenAI's docs omit capabilities that are live by default. Same failure, opposite sign.
If you run agent tooling in production, go read the tool registry yourself. Not the marketing page, not the changelog. The registry.
Each link below shares sources, entities, or timing with this story.
Simon Willison spent a while taking ChatGPT Work apart and published the map on August 30. Work splits into Work Cloud and Work Local, the latter being the renamed Codex desktop app, at $20/month and up since July 9. He enumerates six capabilities Work has that Chat doesn't, a...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
The number that reframes everything isn't ten. It's two thousand. OpenAI published "Ten advances in mathematics and theoretical computer science" on August 1, claiming an internal version of Astra produced new results on ten problems that had seen no progress on the main resul...
July 9, across VS Code, Visual Studio, Copilot CLI, the cloud agent, github.com, GitHub Mobile, JetBrains, Xcode, and Eclipse. Sol is the high-reasoning tier at $5/1M in, $30/1M out, gated to Pro+/Max/Business/Enterprise. Terra is the balanced default at $2.50/$15. Luna is fas...
Prompted by Julia Evans admitting on July 17 that she still can't read query plans, Willison had Fable build a tool that runs arbitrary SQL against a SQLite database and renders both EXPLAIN QUERY PLAN and the lower-level EXPLAIN bytecode with per-line plain-English annotation...
Willison quoted Florian Herrengt on August 12 about teams accumulating so many AI-generated layers that nobody retains a working model of the system. The artifact still works, the org loses the mid-level engineers who could reason about why. Hold that against the OpenAI paper...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.