Pith. sign in

Paper Citation Record · LEDGER

CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2402.07688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07688 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:31.442971Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:13:26.877796Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 33e472e4-b258-493c-b30a-664313e2f358 · inbound

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models cites this paper.

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:52.247233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:52.247233Z digest=sha256:785ea52b6977f3336efe2b43a6d54ed19d854efffc9681a996c6e144839f0ff1

Observation 0d61b9a3-360a-440f-8743-57679b614f7f · inbound

MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks cites this paper.

MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:31.442971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:31.442971Z digest=sha256:2bce5b5a0c08b61a3f477e546d37d55df313d6a71ba5fd68b8da770ab52e3e8a

Observation 1c4f3760-6060-4832-82e2-3b20d32836e0 · inbound

ACSE-Eval: Can LLMs threat model real-world cloud infrastructure? cites this paper.

ACSE-Eval: Can LLMs threat model real-world cloud infrastructure? CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:00.734118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:00.734118Z digest=sha256:0de0d7f3a6eb8a147c7f599a07102062bc16c28eeecf7e1e2e16733bfc55a951

Observation b4d33c2e-9f31-47a8-9b8c-9f4f12074890 · inbound

SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis cites this paper.

SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:16.074470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:16.074470Z digest=sha256:0350ff3da71bfc62d82ad67d2474d4709fc4014507f0e9e476e11c50b87f2a08

Observation e5476660-c479-4587-a8a3-f1d1bbf5073d · inbound

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy cites this paper.

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:00.964558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:00.964558Z digest=sha256:63aa0ae144e03d4d2cd8f66b136ed909ee34541795a48da7104012fe51055154

Observation 7a79103d-b8de-4a37-8e5e-96ceec3b28e6 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.975106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.975106Z digest=sha256:e6b4ff1a582be17bcbadffd72038529e09e78d6bf3c8de48b58acc475604deb0

Observation c3df7809-948b-4c09-870c-f03c3475cdef · inbound

Toward Cybersecurity-Expert Small Language Models cites this paper.

Toward Cybersecurity-Expert Small Language Models CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:27.499437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:27.499437Z digest=sha256:5bfb241c70622d20b351eeadeba6a65081af79ca29504f811ff47596d3dcd5e8

Observation 22beb45a-bdeb-450a-9897-39e20bb5387f · inbound

Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study cites this paper.

Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T09:27:26.581889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:27:26.581889Z digest=sha256:ed419a33c01c2f6527bad37d5fb24ea3a9af8e837dadeb118ebefdc012c5f473

Observation 9543e35c-6174-4138-a8c3-81976a8aecf8 · inbound

LanG -- A Governance-Aware Agentic AI Platform for Unified Security Operations cites this paper.

LanG -- A Governance-Aware Agentic AI Platform for Unified Security Operations CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:51.431008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:53:05.325203Z digest=sha256:18686a3a06f92b71e8bf16bc22d81989c9bc909fab4fe687f21c46c47a576893

Observation c466733a-76f9-4862-8270-38b85169ffa9 · inbound

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps cites this paper.

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:02.859959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:23:56.777888Z digest=sha256:b8ba16daecc9f3257a215e65a3bae68e7c3abe39a957f42ac537dd5afdd8cf29

Observation 27707ee9-35a7-4e53-8b8d-058ee486ea93 · inbound

Cybersecurity AI (CAI) Dataset cites this paper.

Cybersecurity AI (CAI) Dataset CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:26.879232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T12:07:49.656453Z digest=sha256:3da8c406422f6d6967c82f92798e75068acaee88e2f0621373f791e00ee930c8

Observation 60c72e94-c0ec-4823-a9c0-766365907f03 · inbound

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions cites this paper.

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T03:57:47.560737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:57:47.560737Z digest=sha256:c843fcd6ef9f6563cd0379d31fdbdb0099bdf205ca032b93254a31545f99fa1c

Observation 7d2948a5-2c5f-4e72-9694-a7aa6f1f5348 · inbound

Domyn-Small: A European 10B Reasoning Language Model cites this paper.

Domyn-Small: A European 10B Reasoning Language Model CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T14:07:59.382149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:07:59.382149Z digest=sha256:6ea24e8a0297755456cfc39283a6e6e17c480132d7a929ff9f904fdad7bb92ad

Observation a1347f63-ea6a-451e-ab4b-6f7f93fa301d · inbound

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense cites this paper.

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T10:23:21.194170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:23:21.194170Z digest=sha256:c10a3a1b719f7614dbbcb78af168ee32fdc43858e30c759c68d190a7a92c9796