Fetching from the wire…
Public story · 2026-07-23 · high
After launch, Codex reviews real agent sessions and proposes behavior changes, per OpenAI.
Why now: OpenAI announced Presence for limited enterprise rollout on July 22, per its own post.
OpenAI launched Presence on July 22, a system that grades AI agents before they ship and patches them after, per the company.
For companies running voice and chat agents in production, most eval tooling stops once the agent ships. Presence keeps grading after launch too, which is the part most stacks skip.
Before an agent ships, Presence runs it through common requests, edge cases, and high-risk scenarios. Then it scores four separate axes: right outcome, policy followed, tools used correctly, and whether it escalated when it should have.
After launch, Codex reviews the actual production sessions. It proposes changes to the agent's behavior based on what really happened, not what pre-launch simulations guessed at. Most eval setups I've seen stop at the pre-deploy gate. Presence treats deployment as the start of the test, not the end of it.
OpenAI's post doesn't say how much human review sits between a Codex-proposed change and that change shipping to the live agent. For a system grading whether an AI escalated correctly, that's the gap I'd want closed before trusting it with a real support queue.
The four-axis grading is the part OpenAI will demo. The post-launch loop that rewrites behavior off production data is the harder problem. That's the one that decides whether these agents actually improve or just drift.
Each link below shares sources, entities, or timing with this story.
OpenAI released Codex / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenAI released Codex); both cover OpenAI, Presence; reported by the same outlet (openai.com).
JetBrains supports Codex / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (JetBrains supports Codex); both cover Codex, OpenAI; reported by the same outlet (openai.com).
Samsung uses Codex / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Samsung uses Codex); both cover Codex, OpenAI; reported by the same outlet (openai.com).
OpenAI released Codex / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Codex); both cover Codex, OpenAI; reported by the same outlet (openai.com).
Linked by a graph relationship (OpenAI released Codex); both cover Codex, OpenAI; reported by the same outlet (openai.com).
Codex competes with Claude Code / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex, OpenAI; overlapping topics (agent, codex).
OpenAI released Codex / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (OpenAI released Codex); both cover Codex, OpenAI; reported by the same outlet (openai.com).
Cursor benchmarked against Codex / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Cursor benchmarked against Codex); both cover Codex, OpenAI; reported by the same outlet (openai.com).