Pattern: four separate 'agent says done is not done' tools appeared within 24 hours
Repos created on 2026-09-21 and 09-22 converge on one complaint from different angles: supergoal gates completion on a command exiting 0, smixs/code-quality-skill blocks the agent from tampering with tests to reach green, ibrohimislam/zoom-out blocks narrow fixes before you call them done, and unicodef1wn/lauren-poteto-rules (50 stars in its first day) codifies 'builds and type checks do not prove a user flow works' plus 'run the product' as portable agent rules. adyusuf/claude-code-standards goes further and publishes 33 rules each traced to the measurement or incident that produced it. The category forming here is verification tooling that the harness executes, distinct from last week's context-budget and cost-observability waves.
Source
↳ Follow the thread