Pith. sign in

REVIEW 1 cited by

Software Vulnerability and Functionality Assessment using LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.08429 v1 pith:NCDSNOZE submitted 2024-03-13 cs.SE cs.AI

classification cs.SEcs.AI
keywords codevulnerabilitiesfunctionalityllmsmodelssecuritysoftwaredescriptions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While code review is central to the software development process, it can be tedious and expensive to carry out. In this paper, we investigate whether and how Large Language Models (LLMs) can aid with code reviews. Our investigation focuses on two tasks that we argue are fundamental to good reviews: (i) flagging code with security vulnerabilities and (ii) performing software functionality validation, i.e., ensuring that code meets its intended functionality. To test performance on both tasks, we use zero-shot and chain-of-thought prompting to obtain final ``approve or reject'' recommendations. As data, we employ seminal code generation datasets (HumanEval and MBPP) along with expert-written code snippets with security vulnerabilities from the Common Weakness Enumeration (CWE). Our experiments consider a mixture of three proprietary models from OpenAI and smaller open-source LLMs. We find that the former outperforms the latter by a large margin. Motivated by promising results, we finally ask our models to provide detailed descriptions of security vulnerabilities. Results show that 36.7% of LLM-generated descriptions can be associated with true CWE vulnerabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Effective Complementary Security Analysis using Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    Using Chain-of-Thought and Self-Consistency prompts, some LLMs removed over half of SAST false positives on a benchmark while missing no genuine weaknesses, and ensembling three models removed about 79%.

Pith tools