Pith. sign in

REVIEW 4 cited by

VulnLLMEval: A Framework for Evaluating Large Language Models in Software Vulnerability Detection and Patching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.10756 v1 pith:NBBTEINK submitted 2024-09-16 cs.SE cs.AI

classification cs.SEcs.AI
keywords codellmstasksevaluatingmodelspatchingvulnerabilitiesdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have shown promise in tasks like code translation, prompting interest in their potential for automating software vulnerability detection (SVD) and patching (SVP). To further research in this area, establishing a benchmark is essential for evaluating the strengths and limitations of LLMs in these tasks. Despite their capabilities, questions remain regarding whether LLMs can accurately analyze complex vulnerabilities and generate appropriate patches. This paper introduces VulnLLMEval, a framework designed to assess the performance of LLMs in identifying and patching vulnerabilities in C code. Our study includes 307 real-world vulnerabilities extracted from the Linux kernel, creating a well-curated dataset that includes both vulnerable and patched code. This dataset, based on real-world code, provides a diverse and representative testbed for evaluating LLM performance in SVD and SVP tasks, offering a robust foundation for rigorous assessment. Our results reveal that LLMs often struggle with distinguishing between vulnerable and patched code. Furthermore, in SVP tasks, these models tend to oversimplify the code, producing solutions that may not be directly usable without further refinement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond

    cs.CR 2025-06 conditional novelty 6.0 of 10

    Mono reports that 31% of MegaVul patches are non-security and about 16.7% of CVEs are 'undecidable', while its added context raises LLM vulnerability detection F1 by up to 15%.

  2. BioPose: Biomechanically-accurate 3D Pose Estimation from Monocular Videos

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A monocular-video pipeline converts a learned 3D body mesh into virtual markers and regresses them through a neural inverse kinematics model to output biomechanically accurate joint angles.

  3. INVARLLM: LLM-assisted Physical Invariant Extraction for Cyber-Physical Systems Anomaly Detection

    cs.CR 2024-11 reject novelty 5.0 of 10

    INVARLLM automates extraction of physical invariants from CPS documentation via LLMs, then uses PCMCI+ scores and K-means to validate them, reporting case-level 100 percent precision on SWaT and WADI despite low raw s...

  4. LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LLM-based vulnerability detection covering 58 papers, with a taxonomy, dataset overview, and gap analysis, but limited by non-transparent selection and unsupported quantitative claims.

Pith tools