OSS
deep-swe measures coding agents on original long-horizon tasks and has not been pushed in eight days
datacurve-ai/deep-swe is a benchmark for frontier coding agents on original, long-horizon engineering tasks, at 1,578 stars and 105 forks with 72 open issues. It was created 2026-05-15 and last pushed 2026-08-26, so today's Python trending appearance at 21 stars comes from attention rather than activity. A benchmark whose value depends on tasks staying uncontaminated going eight days without a commit is worth watching, because the models it scores ship weekly.
↳ Follow the thread