Skills
SWE-Serve: end-to-end serving tests reject about a third of agent patches that pass every other test on real SGLang features
arXiv 2609.26777 (22 Sep) builds 53 tasks from recent production SGLang changes, scored with hidden functional, regression and end-to-end serving tests. Across 11 models and 31 model-effort configurations the best reached 75% pass@1. On the 19 tasks with E2E coverage, the pass rate falls from 69.4% to 45.9% once E2E tests count. The practical lesson is that agent patches to infrastructure code need a real end-to-end run in CI, because unit and regression tests miss roughly a third of the failures.
↳ Follow the thread