Pith. sign in

Paper Citation Record · LEDGER

CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2402.07688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07688 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:04:52.247233Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:13:26.877796Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 33e472e4-b258-493c-b30a-664313e2f358 · inbound

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models cites this paper.

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:52.247233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:52.247233Z digest=sha256:e781f1975e7804ce55468c329bdcfa6b8e0da20d9aec935afda7a896f8fd61f5

Observation b4d33c2e-9f31-47a8-9b8c-9f4f12074890 · inbound

SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis cites this paper.

SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:16.074470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:16.074470Z digest=sha256:a3144dd4e26499c27c01d0125cf90fc6429e6489ac9f36b219d2467516b4ada8

Observation e5476660-c479-4587-a8a3-f1d1bbf5073d · inbound

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy cites this paper.

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:00.964558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:00.964558Z digest=sha256:145de1d8bc24e6efc1b892d9e93ebf6068de44aa19e7e2392a14e709ca9c36f2

Observation 7a79103d-b8de-4a37-8e5e-96ceec3b28e6 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.975106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.975106Z digest=sha256:a6ae997816244aaa4d623944b52fc26aa3793e2e6849f96cd8c7ab57e326e68c

Observation c3df7809-948b-4c09-870c-f03c3475cdef · inbound

Toward Cybersecurity-Expert Small Language Models cites this paper.

Toward Cybersecurity-Expert Small Language Models CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:27.499437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:27.499437Z digest=sha256:1ae7fe8dba186a4869fb561eb5f6ea851f73be2083ebf1edefad3e8665195153

Observation 22beb45a-bdeb-450a-9897-39e20bb5387f · inbound

Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study cites this paper.

Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T09:27:26.581889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:27:26.581889Z digest=sha256:67801fd7c08f1c803f1cc851ec97d5b3828e113cfb39a3ce77c0a7f4141bb43b

Observation 9543e35c-6174-4138-a8c3-81976a8aecf8 · inbound

LanG -- A Governance-Aware Agentic AI Platform for Unified Security Operations cites this paper.

LanG -- A Governance-Aware Agentic AI Platform for Unified Security Operations CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:51.431008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:53:05.325203Z digest=sha256:eef8de2d8def12f12bf3d2ce098543416b1d00ed1e6c07ea102671add20d69b4

Observation c466733a-76f9-4862-8270-38b85169ffa9 · inbound

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps cites this paper.

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:02.859959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T02:23:56.777888Z digest=sha256:daa4e292904020a1c36b3f80f0df79fa335060a73dc99e96ec0da84d6c16d5ea

Observation 27707ee9-35a7-4e53-8b8d-058ee486ea93 · inbound

Cybersecurity AI (CAI) Dataset cites this paper.

Cybersecurity AI (CAI) Dataset CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:26.879232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T12:07:49.656453Z digest=sha256:51c2e15c089d033a917246069e3210362f77e28fd9ce5dce118c400cbbbfa380

Observation 60c72e94-c0ec-4823-a9c0-766365907f03 · inbound

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions cites this paper.

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T03:57:47.560737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:57:47.560737Z digest=sha256:ba25872f5c095539aaf3923c7a7dc6a735201b1a94afd32aead8a0c95634b2c9

Observation 7d2948a5-2c5f-4e72-9694-a7aa6f1f5348 · inbound

Domyn-Small: A European 10B Reasoning Language Model cites this paper.

Domyn-Small: A European 10B Reasoning Language Model CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T14:07:59.382149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:07:59.382149Z digest=sha256:6985f65893242e2f98242f80907c680e602a76f9ab6fbbd62ea14a048df99a25

Observation a1347f63-ea6a-451e-ab4b-6f7f93fa301d · inbound

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense cites this paper.

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T10:23:21.194170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:23:21.194170Z digest=sha256:58187634734961e5b0952a5bba9ee5f053428a839cc00bb2d2b6d1e8c7f9272c