Voices
Andrew Ng's 'When Guardrails Go Wrong' Inverts the Safety Narrative With a Case Where the Closed Model Broke and the Open Model Fixed It
Ng's July 24 letter in The Batch takes direct aim at the story 'a few frontier labs have tried to tell' — that open models are dangerous cyberattack vectors while their guardrailed proprietary models are safe. His counterexample is operational rather than theoretical: a closed model malfunctioned on a critical vendor system and an open-weight model was what resolved the incident. Published three days before Amodei's rebuttal post, it is the practitioner-side argument the policy fight is actually being conducted over, and it reframes guardrails as an availability risk rather than purely a safety feature.
↳ Follow the thread