GitHub's CTO Puts the August 17 Outage at 7 Hours 47 Minutes, Caused by a Central US Component That Failed to Scale, With Copilot Retry Loops Slowing Recovery
GitHub / Hacker News·high signal
Vlad Fedorov published the post-mortem on August 20 and it took 570 points and 630 comments. A critical infrastructure component in the Central US data center failed to scale under record traffic, cascading into authentication failures across github.com, Actions, the APIs, pull requests, issues and Copilot. Copilot took longer than most services to come back, and its client-side retry loops added traffic during recovery. The commitments are consistent retry limits, retry budgets, variable timeouts, and isolating critical systems, which is a direct admission that the AI clients made their own outage worse.