Skip to content
InfoResearchPreprintLLM-specific

On the Reliability of LLM-Based Vulnerability Patching Benchmarks

Published
Record updated
View JSON

Summary

This paper examines whether LLM-based vulnerability patching benchmarks give reliable results. The authors curate 112 historical bugs from 84 open-source C/C++, Go, and Rust projects, each with PoC, regression, and developer tests. They find that LLMs reach high PoC passing rates under ideal conditions, but developer-test passing rates stay low and improve only marginally with newer models.