Polymath Labs Horizon-SWE: First Benchmark for Multi-Task End-to-End Software Engineering Workflows
Polymath Labs·medium signal
Horizon-SWE evaluates AI agents on multi-task, multi-step software engineering workflows (5+ sequential tasks per instance) rather than single issue resolution, exposing where SWE-bench performance fails to transfer to real development cycles. Published February 2, 2026 by Polymath Labs, the benchmark targets the 'multi-to-multi' gap — agents that score well on isolated tasks but collapse when chaining changes across interdependent steps. First empirical framework for measuring end-to-end software engineering agent capability.