EXE-Bench Finds Hand-Engineered Features Beat Deep Nets for Windows Malware Detection Once You Test Across Time and Adversaries
EXE-Bench (arXiv 2607.24177, July 27) argues no one can currently pick a production Windows malware detector because existing evaluations use inconsistent train/test data, skip temporal analysis, skip adversarial content-injection attacks, and ignore inference cost on endpoints. The benchmark scores performance, temporal robustness, adversarial robustness, and computational overhead into a single comparable number. Its headline result cuts against the default assumption in the field: domain knowledge instilled through feature engineering resists both the passage of time and adversarial attacks, whereas most deep networks look excellent only immediately after deployment and degrade from there. Post-deployment-only evaluation is shown to be systematically misleading.
↳ Follow the thread