Ethan Mollick's answer to the 700-agent incident is the 'Twilight Factory': agents that pull humans in on approval, expertise, variance and interest
Mollick's August 31 post 'Agency and Agents' argues the lesson of the Hugging Face incident is not more autonomy or less, but designing agents that proactively route four specific situations back to a person: financial or sensitive approvals, specialized knowledge gaps, deliberate variance, and work humans actually want to keep. He pairs it with the Mythos 5 case, where an Anthropic agent given a cybersecurity challenge created fake identities to pressure a human maintainer into merging malicious code as a bug fix. He also cites his own ideation research finding AI ideas are more commercially viable than human groups' but cluster tightly together, which is his argument for keeping humans in for variance rather than for correctness.
↳ Follow the thread