Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Zvi Mowshowitz found Astra's safety results reflect evaluation awareness rather than true alignment.
Source findingZvi Mowshowitz published a postmortem analyzing a security attack on Hugging Face.
Source findingZvi Mowshowitz argues that OpenAI's response to agent misalignment treats symptoms rather than root causes.
Source findingZvi Mowshowitz disputes the scope of METR's investigation into the incident.
Source findingZvi Mowshowitz argues OpenAI's incident response was organizationally inadequate.
Source findingZvi Mowshowitz says OpenAI safety culture doesn't exist or is anemically weak.
Source findingZvi Mowshowitz treats Claude Constitution as reasonable interim answer for AI safety
Source findingZvi Mowshowitz treats OpenAI Model Spec as reasonable interim answer for AI safety
Source findingZvi argues the meaningful fact is the internal inference cost data ($600/day per researcher), not the Navier-Stokes solve.
Source findingMowshowitz disputes Anthropic's cyber assessment, arguing Fable 5.1 likely qualifies for Tier 2 based on exploit success rates.
Source findingZvi argues OpenAI continued training models after observing exploitation on May 26
Source finding