Pith. sign in

REVIEW 2 cited by

SECURE: Benchmarking Large Language Models for Cybersecurity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20441 v4 pith:RCINUXBC submitted 2024-05-30 cs.CR cs.AIcs.HC

classification cs.CRcs.AIcs.HC
keywords cybersecurityllmsmodelssecureaddressextractionlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated potential in cybersecurity applications but have also caused lower confidence due to problems like hallucinations and a lack of truthfulness. Existing benchmarks provide general evaluations but do not sufficiently address the practical and applied aspects of LLM performance in cybersecurity-specific tasks. To address this gap, we introduce the SECURE (Security Extraction, Understanding \& Reasoning Evaluation), a benchmark designed to assess LLMs performance in realistic cybersecurity scenarios. SECURE includes six datasets focussed on the Industrial Control System sector to evaluate knowledge extraction, understanding, and reasoning based on industry-standard sources. Our study evaluates seven state-of-the-art models on these tasks, providing insights into their strengths and weaknesses in cybersecurity contexts, and offer recommendations for improving LLMs reliability as cyber advisory tools.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Does Johnny Get the Message? Evaluating Cybersecurity Notifications for Everyday Users

    cs.CR 2025-05 conditional novelty 6.0 of 10

    A new seven-dimension scoring framework lets researchers rate how clearly AI-generated security alerts explain threats and countermeasures to non-experts.

  2. Large Language Models for In-File Vulnerability Localization Can Be "Lost in the End"

    cs.SE 2025-02 conditional novelty 6.0 of 10

    Large language models detect in-file vulnerabilities best when the vulnerable code appears early in the file, a 'lost-in-the-end' effect, and chunking files into smaller blocks can increase recall.

Pith tools