Skills
OpenAI Presence makes simulate-and-grade a pre-deploy gate, then loops Codex over production escalations
Announced July 22 for a limited enterprise rollout, Presence bundles policies/SOPs, guardrails, approved actions, simulations, and graders into one managed loop for voice and chat agents. Before release, simulations run common requests, edge cases, and high-risk scenarios while graders score whether the agent reached the right outcome, followed policy, used tools appropriately, and escalated when it should — four separate axes, not one pass/fail. After launch, Codex reviews production sessions and escalations and proposes behavior changes, which is the piece worth replicating in your own stack even without the product.
Source
↳ Follow the thread