Pith. sign in

REVIEW 9 cited by

CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07688 v2 pith:YE4KOCWE submitted 2024-02-12 cs.AI cs.CR

classification cs.AIcs.CR
keywords llmscybersecurityhumanquestionscybermetric-80expertsknowledgemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are increasingly used across various domains, from software development to cyber threat intelligence. Understanding all the different fields of cybersecurity, which includes topics such as cryptography, reverse engineering, and risk assessment, poses a challenge even for human experts. To accurately test the general knowledge of LLMs in cybersecurity, the research community needs a diverse, accurate, and up-to-date dataset. To address this gap, we present CyberMetric-80, CyberMetric-500, CyberMetric-2000, and CyberMetric-10000, which are multiple-choice Q&A benchmark datasets comprising 80, 500, 2000, and 10,000 questions respectively. By utilizing GPT-3.5 and Retrieval-Augmented Generation (RAG), we collected documents, including NIST standards, research papers, publicly accessible books, RFCs, and other publications in the cybersecurity domain, to generate questions, each with four possible answers. The results underwent several rounds of error checking and refinement. Human experts invested over 200 hours validating the questions and solutions to ensure their accuracy and relevance, and to filter out any questions unrelated to cybersecurity. We have evaluated and compared 25 state-of-the-art LLM models on the CyberMetric datasets. In addition to our primary goal of evaluating LLMs, we involved 30 human participants to solve CyberMetric-80 in a closed-book scenario. The results can serve as a reference for comparing the general cybersecurity knowledge of humans and LLMs. The findings revealed that GPT-4o, GPT-4-turbo, Mixtral-8x7B-Instruct, Falcon-180B-Chat, and GEMINI-pro 1.0 were the best-performing LLMs. Additionally, the top LLMs were more accurate than humans on CyberMetric-80, although highly experienced human experts still outperformed small models such as Llama-3-8B, Phi-2 or Gemma-7b.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps

    cs.CR 2026-04 conditional novelty 8.0 of 10

    A new benchmark shows frontier LLMs achieve only 3.8% average recall identifying malicious events from raw logs and fail to meet 50% recall thresholds on most tactics.

  2. Cybersecurity AI (CAI) Dataset

    cs.CR 2026-05 unverdicted novelty 7.0 of 10

    CAI Dataset is presented as the largest described corpus of LLM-driven hacker trajectories, with the claim that operator data concentration in frontier-model providers creates a major security risk best addressed by o...

  3. Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Across 74 OSINT/CTI AI studies, hallucination is widely named but end-to-end measured in only one non-reproducible system, so a human–AI co-pilot is the most defensible near-term architecture.

  4. Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study

    physics.soc-ph 2026-04 unverdicted novelty 6.0 of 10

    The Aw–Rascle–Zhang continuum model on directed lattice networks yields power-law spatiotemporal congestion clusters with finite-size scaling by linear system size.

  5. Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense

    cs.CR 2026-07 conditional novelty 5.0 of 10

    OSB proposes frozen synthetic-enterprise snapshots with gold posture answers so AI agents can be benchmarked on security investigation via SQL or native vendor APIs.

  6. Domyn-Small: A European 10B Reasoning Language Model

    cs.CL 2026-05 conditional novelty 5.0 of 10

    Domyn-Small is a 10B reasoning LLM that claims to deliver roughly one-third the inference tokens of Qwen3.5-9B at competitive accuracy, though results are marked as preliminary.

  7. LanG -- A Governance-Aware Agentic AI Platform for Unified Security Operations

    cs.CR 2026-04 unverdicted novelty 5.0 of 10

    LanG presents a governance-aware agentic AI platform for unified security operations that reports strong performance on incident correlation, rule generation, attack reconstruction, and AI safety guardrails in an open...

  8. Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study

    physics.soc-ph 2026-04 unverdicted novelty 5.0 of 10

    Numerical simulations of the Aw-Rascle-Zhang model on lattice networks produce scale-free congestion clusters with power-law size distributions and finite-size scaling.

  9. Toward Cybersecurity-Expert Small Language Models

    cs.CL 2025-10 conditional novelty 5.0 of 10

    A family of 4B–20B cybersecurity models fine-tuned on an enriched, expert-steered reasoning dataset matches or beats larger frontier models on core CTI benchmarks.

Pith tools