Launch HN: Bullet (YC S26) Claims 95.8% on SWE-bench Verified at 119s Per Task — Commenters Call the Benchmark Meaningless
Bullet, from YC S26 founders Adi and Alex (ex-AppLovin and Citadel, after six pivots), launched on HN August 13 as a speed-focused coding agent harness rather than a new model, layering model routing, targeted code search instead of whole-repo embedding, context management, and turn batching that they say cut round trips 16% and costs 27% internally. They claim 479/500 (95.8%) on SWE-bench Verified in one attempt at 119s average per task, 35–67% faster than mini-SWE-agent, and it plugs into Claude, Codex, and Grok via API keys or existing subscriptions. The top critique in the 74-comment thread questioned whether a saturated SWE-bench leaderboard means anything without disclosing model selection; the founders pointed to Terminal-Bench and CursorBench for future validation.
Source
↳ Follow the thread