Fetching from the wire…
Top 5 · 2026-08-15 · source-backed
The number that matters isn't 133 million. It's zero. That's how many alerts fired during the eleven months the classifier was disabled.
Anthropic's second Responsible Scaling Policy risk report, published August 14, discloses that an internal-only feature flag disabled the blocking biological classifier across all human-feedback vendor traffic. That's roughly 133 million exchanges with about 50,000 contractors, from May 2025 through April 2026. The same flag also disabled logging. That's why routine monitoring never caught it for nearly a year. Anthropic says remediation is done and it found no evidence of concerning misuse, but it explicitly states the discovery reduced its confidence that no similar gaps exist elsewhere. Anthropic
Read that last part twice. The lab with the most mature published safety process, the most safety staff per capita, and a formal scaling policy it wrote itself, ran a guardrail in the off position for a year and only found out by accident. Then said, in writing, that it now trusts its own inventory of controls less.
The flag design is the transferable lesson, and it generalizes far past frontier labs. Someone built a kill switch that turned off the enforcement path and the observability path together. It's an understandable choice. If the classifier isn't running, what would it even log? But that coupling is exactly what converts a config mistake into an eleven-month blind spot. A guardrail that goes quiet when disabled looks identical to a guardrail that's passing everything.
Go audit your own. Every rate limiter, content filter, permission check, and injection scanner in your stack: does it emit telemetry when it's off? Not an error. A heartbeat. "Classifier disabled by flag X, 4,102 requests passed unchecked" written every minute to the same place the pass/fail counts go. If your dashboard shows a flat zero for blocked requests and you can't tell that apart from a healthy day, you have this bug. My own harness had it. A guard was disabled by an empty-list default and nothing said so.
Two more items from the same report deserve attention. Anthropic raised its assessment of catastrophic harm from misalignment in high-stakes settings from "very low" to "low," and confirmed an unreleased internal "Model 2" that is somewhat more capable than Mythos 5 with no release plans. Unite.AI And buried deeper: Anthropic reports reduced confidence in its own safety evaluations because task-based benchmarks have saturated and no longer capture capability improvements. Hacker News
That pair is the real story. The blocking control failed silently, and the measurement layer that would catch the next failure is losing resolution. The same document says internal AI-assisted R&D is significantly faster but "not yet by a factor of 2," with acknowledged measurement difficulty. That's a deflationary number from the party with every incentive to report a bigger one.
I trust a lab that publishes this more than one that doesn't. That's not the same as being reassured.
Each link below shares sources, entities, or timing with this story.
Anthropic released Mythos / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, April, Mythos, Then; reported by the same outlet (anthropic.com).
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, Hacker News, Then; reported by the same outlet (news.ycombinator.com).
Anthropic released Mythos / Shared entities / Same source domain
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, April, August, Hacker News; reported by the same outlet (news.ycombinator.com).
Anthropic released Mythos / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, Mythos, Then; overlapping topics (anthropic, control, have).
Anthropic released Mythos / Shared entities / Earlier coverage
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, Hacker News, Mythos, Then; earlier Anthropic coverage from 2026-06-10.
Anthropic released Mythos / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, August, Mythos; overlapping topics (anthropic, control).
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, April, Hacker News; overlapping topics (anthropic, found).
Linked by a graph relationship (Anthropic released Mythos); both cover Anthropic, April, Mythos; overlapping topics (anthropic, million).