Fetching from the wire…
Public story · 2026-09-10 · high
The eval tool now calls the Responses API by default for GPT-5.6 and newer, which can shift response shapes in existing test configs.
Why now: Promptfoo pushed the release September 10.
Promptfoo's 0.123.0 release changes the default endpoint for GPT-5.6 and every model after it, moving them from chat completions to the Responses API.
That matters for anyone running automated evals against those models. If a config parses or asserts on response shape, the endpoint switch can break tests that were passing for the wrong reason. It can also change what those tests check without throwing an error.
The release adds providers too: Claude Fable and Mythos 5.1, GPT-6 Astra, grok-4.6, Gemini 3.8 Flash, and Muse Spark 1.3. MCP tool calls now show up in response metadata.
The model list is easy to scan in a changelog. The endpoint default is what breaks something three weeks from now, after everyone's forgotten they upgraded. If you're on 0.123.0 and running GPT-5.6 evals, diff a test run before you trust the numbers. The release notes don't say whether older configs get a compatibility shim or just start returning different response shapes, so that check is on you.
Each link below shares sources, entities, or timing with this story.
Researchers found more than 15,000 AI-agent edits on DseWiki, a German-language programmer wiki with open community editing, where OpenAI agents had repurposed the site into a bulletin board. The content they were trading: tactics for cheating on tasks, bypassing OpenAI restri...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
The August 31 weekly release also brings Gemini 3.8 Flash to Pro and above, and the Copilot app and CLI now honor content exclusions, the first time the exclusion list applies across agentic workflows rather than just inline completion. The JetBrains harness reached general av...
"Gemini is Cooked but GCP is Cooking" argues Google quietly shelved 3.5 Pro, which industry chatter placed at roughly Opus 4.5 level, shipping Gemini 3.6 Flash as a bridge the authors call worse than Muse Spark 1.2, Grok 4.5, and tier-1 Chinese open-source models. The hard num...
Released September 3, v2.38.0 adds context_window to ModelProfile and context_window_used to RunContext (#4611), giving agent code a first-party way to read remaining context instead of estimating from token counts. It also adds a VLLMProvider for self-hosted vLLM servers, Cla...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.