InfoResearchPreprintLLM-specific
On the Reliability of LLM-Based Vulnerability Patching Benchmarks
- Published
- Record updated
Summary
This paper examines whether LLM-based vulnerability patching benchmarks give reliable results. The authors curate 112 historical bugs from 84 open-source C/C++, Go, and Rust projects, each with PoC, regression, and developer tests. They find that LLMs reach high PoC passing rates under ideal conditions, but developer-test passing rates stay low and improve only marginally with newer models.