Fetching from the wire…
Models2026-07-22 · source-backed
The Evals Platform, Agent Builder, and Reusable Prompts (v1/prompts) all share a November 30, 2026 shutdown, with Evals hitting read-only a month earlier (TheRouter.ai). This follows OpenAI's March 10 acquisition of Promptfoo (300K+ developers, 127 Fortune 500 companies), and OpenAI's own migration guidance points there. If your eval definitions, graders, or historical run baselines live only in a vendor dashboard, that's a hard export deadline about 14 weeks out. Use it as the forcing function to move evals into versioned repo files where they belong.
Each link below shares sources, entities, or timing with this story.
AgentKit's Agent Builder and Evals come off the platform November 30, 2026, replaced by OpenAI Frontier for building, deploying, and governing agents with shared context and permissions (customers include HP, Oracle, State Farm, Uber) (OpenAI). OpenAI says Frontier doesn't rep...
Fortune published a deep-dive on May 21 that should make anyone building on Microsoft's AI stack uncomfortable. After spending $13B+ on OpenAI and projecting $190 billion in 2026 capex (more than double 2025), Microsoft Copilot has reached just 20 million paying M365 users out...
Open-source framework for testing, evaluating, and red-teaming LLM prompts, agents, and RAG systems. Covers OWASP LLM Top 10. Used by 127 Fortune 500 companies. Now acquired by OpenAI but committed to continuing the open-source offering. The de facto standard for AI pentesting...
OpenAI announced the acquisition of Promptfoo ($86M valuation, used by 25% of Fortune 500) and continued rolling out Codex Security, which scanned 1.2M commits in its first month and found 792 critical and 10,561 high-severity vulnerabilities — including 14 assigned CVEs acros...
Twelve months ago, OpenAI led Anthropic by 41 points in enterprise adoption. Today that gap is 8. Enterprise Technology Research's survey of roughly 500 respondents shows OpenAI dropping from 62% adoption (September 2025) to 56% (March 2026) while Anthropic surged from 21% to...
Three frontier models shipped in a single week this month, and teams with a standing eval harness had a routing decision in hours. Anthropic's own agent-eval guidance says 20-50 tasks drawn from your real usage and real failures is enough to detect issues (DeepEval). DeepEval...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.