Fetching from the wire…
Research2026-08-09 · source-backed
An audit of LLM benchmark methodology (arXiv 2608.06202, Aug 6) ran 401 stratified prompts from BBQ and SafetyBench through both ChatGPT's chat UI and the OpenAI API, with and without web search, collecting 4,812 responses over three repeated runs. Chat UI was less accurate than API on both benchmarks with search off. Enabling search cut accuracy by up to 8 points and reversed the direction of the modality trend on one benchmark. Repeated runs of the same prompt disagreed on up to 21%. Citation grounding and abstention diverged between modalities too. Every single-modality single-run accuracy number used to argue deployment readiness is measuring one narrow slice.
Each link below shares sources, entities, or timing with this story.
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
Promptwatch's tracking shows the share of ChatGPT search queries using site: sat at 0.3-0.5% for weeks, dipped to 0.15% on August 3-5, then jumped to 16-17% on August 8, two days after OpenAI said it was making GPT-5.6 Sol "more reliable with facts." Simon Willison Willison co...
On July 24, Justice Amit Bansal denied ANI Media interim relief against OpenAI, holding that LLM training on ANI's content falls under Section 52(1)(a)(i) "private or personal use, including research." The court separately held that retrieval-augmented outputs don't infringe u...
Microsoft Research dropped a paper that should change how every builder thinks about their agent configuration files. SkillOpt (arXiv 2605.23904) treats a Markdown document as an external parameter of a frozen LLM and applies learning rate, batch, and momentum concepts in text...
Shipped August 13 to Pro, Business and Enterprise: ChatGPT and Codex can draw on activity memories from apps and websites on macOS, with per-app opt-in and review or delete controls. OpenAI This is Cursor's Google Workspace connector move, but reaching down to OS-level activit...
The July 29 walkthrough covers the consumer chat UIs of both products, not the developer APIs or CLI harnesses where MCP is well-trodden. That summary judgment is the finding. Know that friction before you plan any distribution strategy assuming end users will attach your MCP...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.