A benchmark that requires a patch to block a working exploit finds the best LLM repairs only 21.7% of 23 real CVEs, with most failures caused by missed localization and malformed patches.
git-apply Documentation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
A benchmark that requires a patch to block a working exploit finds the best LLM repairs only 21.7% of 23 real CVEs, with most failures caused by missed localization and malformed patches.