InfoResearchPeer-reviewed
EXE-Bench: Ranking the Tradeoffs of AI-Based Windows Malware Detectors for Real-World Usability
- Published
- Record updated
Summary
Existing evaluations of AI-based Windows malware detectors differ in training and test data, lack temporal analysis, skip adversarial content-injection tests, and ignore deployment compute costs, so they cannot show which detector to deploy. The authors introduce EXE-Bench, which assesses performance, temporal and adversarial robustness, and computational overhead, combining them into one score for direct comparison. Their analysis finds that feature-engineered domain knowledge remains highly useful, resisting both time and adversarial attacks, while most deep networks excel only right after deployment.