Fetching from the wire…
Public story · 2026-07-23 · high
After launch, Codex reviews real agent sessions and proposes behavior changes, per OpenAI.
Why now: OpenAI announced Presence for limited enterprise rollout on July 22, per its own post.
OpenAI launched Presence on July 22, a system that grades AI agents before they ship and patches them after, per the company.
For companies running voice and chat agents in production, most eval tooling stops once the agent ships. Presence keeps grading after launch too, which is the part most stacks skip.
Before an agent ships, Presence runs it through common requests, edge cases, and high-risk scenarios. Then it scores four separate axes: right outcome, policy followed, tools used correctly, and whether it escalated when it should have.
After launch, Codex reviews the actual production sessions. It proposes changes to the agent's behavior based on what really happened, not what pre-launch simulations guessed at. Most eval setups I've seen stop at the pre-deploy gate. Presence treats deployment as the start of the test, not the end of it.
OpenAI's post doesn't say how much human review sits between a Codex-proposed change and that change shipping to the live agent. For a system grading whether an AI escalated correctly, that's the gap I'd want closed before trusting it with a real support queue.
The four-axis grading is the part OpenAI will demo. The post-launch loop that rewrites behavior off production data is the harder problem. That's the one that decides whether these agents actually improve or just drift.
Each link below shares sources, entities, or timing with this story.
On July 27, Salesforce jumped ~7%, ServiceNow ~8%, Workday ~10% — the same five incumbents Presence repriced last week — with money rotating out of NVIDIA and AMD (both down 5-7% on the OpenAI financing-guarantee report). The bull case cited was that incumbents are monetizing...
Presence is positioned as production-grade agent deployment for customer-facing and internal workflows, putting OpenAI directly against Sierra, Decagon, and the agent layers Salesforce and ServiceNow ship. The frontier labs want the application layer, not just the tokens under...
CLI, IDE extension, macOS app, and Codex Web all sit on the same harness, linked by a bidirectional JSON-RPC-over-stdio process that hosts core threads and translates client requests into Codex operations. Partners have written clients in Go, Python, TypeScript, Swift, and Kot...
Alongside the July 29 launch of ChatGPT for Academic Researchers, free GPT-5.6 Sol Pro for 10,000 researchers this summer scaling to 100,000 through 2027 backed by over $250 million, OpenAI disclosed efficiency work on the harness underlying Codex and ChatGPT Work: 54% fewer o...
OpenAI launched the GPT-5.6 family on July 14: Sol (flagship), Terra (cost-optimized), and Luna (fast tier), live across ChatGPT, Codex, and the API the same day after a US-government-requested delay for security review. The numbers are loud. Sol scored 53.6 on Agents' Last Ex...
In 2023 Samsung banned generative AI internally after engineers leaked source code into ChatGPT. On June 21, 2026, OpenAI confirmed Samsung is deploying ChatGPT Enterprise and Codex to all employees in South Korea and across its worldwide Device eXperience division. That's a f...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.