Fetching from the wire…
Public story · 2026-07-20 · high
Policy leans on labs' self-stated safety limits, but a paper finds those numbers use incompatible definitions no outsider can check.
Why now: It appeared in the July 20 briefing as more policy proposals start treating labs' self-reported thresholds as the enforcement line.
A paper finds frontier AI labs' safety thresholds are measured so differently that no outsider can verify whether one's been crossed, per arXiv 2607.16112. Policy language increasingly points to a lab's own stated threshold as the line a regulator can act on. Trouble is, if those thresholds aren't measured the same way, that line moves depending on which lab wrote it.
The paper's authors propose a harmonization scheme, a shared way to define and measure these thresholds so they line up across labs. Right now two labs can describe what sounds like the same capability limit using different definitions and different tests. There's no way for anyone outside the lab to reconcile the two numbers.
It doesn't name which labs' frameworks it compared. It doesn't say whether any lab has already crossed its own stated line. Those are the open questions a harmonization scheme would have to answer before anyone can use it to hold a lab to account.
This is governance infrastructure, not a headline. But it's the kind of gap that decides whether regulation has teeth three years from now. A rule that cites a lab's own stated threshold is only as strong as the measurement behind it, and no one agrees on that measurement.
Each link below shares sources, entities, or timing with this story.
Executives at Uber, Meta, Microsoft, Salesforce, and DoorDash have launched AI cost-cutting campaigns after bills doubled or tripled, or blew through annual budgets in as little as three to four months. Uber has introduced hard usage limits on AI tools (WSJ). Read that timelin...
Bloomberg reported this morning that Microsoft has begun swapping OpenAI and Anthropic models for its own MAI models inside Excel and Outlook, with tens of thousands of prompts a week now running on MAI. Source. Read that number carefully. Tens of thousands of prompts a week i...
OpenAI's Frontier platform enables building, deploying, and managing AI agents that run other software (Salesforce, Workday, etc.). "Business Context" gives agents institutional memory. Multiyear deals with Accenture, BCG, Capgemini, McKinsey. Key insight: Frontier positions a...
Simon Willison mapped them: Microsoft's "Open Weights and American AI Leadership" (July 24, 235 companies including NVIDIA, Amazon, Y Combinator and the Linux Foundation, with OpenAI signing later, explicitly endorsing distillation as legitimate); Anthropic's "Our Position on...
Willison's August 2 roundup lays out "Open Weights and American AI Leadership" (July 24, Microsoft-shepherded, now 235 signatory companies including NVIDIA, Amazon, Y Combinator, the Linux Foundation, and OpenAI after initially abstaining); Anthropic's separate July 27 rebutta...
Pacing the Frontier went public July 28 with signatures from OpenAI, Anthropic, Google DeepMind, Meta, Microsoft, Mistral, and Thinking Machines, asking the U.S. government to lead an international effort on the technical and governance tools needed to deliberately pace automa...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.