Pith. sign in

Paper Citation Record · LEDGER

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response

As of 23 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2505.19973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19973 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:01.393974Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:21:11.287455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:41:07.580051Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact4
  • verified fuzzy6
  • unresolved18
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e7400f4-1739-431f-b46a-09b65903ab69 · outbound

This paper cites In: Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track (2024).

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track (2024)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:09.459205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:21.853495Z digest=sha256:82aaaf2ef42b1b952dbb305036993c3f62e3c4f9f476a1cdd2e69c2e23328f52

Observation 7c401da4-505c-4185-88f3-09e4dca02f06 · outbound

This paper cites The DeepSpeak Dataset.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response The DeepSpeak Dataset

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:21.909923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:21.909923Z digest=sha256:899f67db517d839ff5225cda7abd36c46ace69b9f58ce7088b95f4e89dbeb76b

Observation ce3740db-48bf-4642-80df-11a6ad02864e · outbound

This paper cites In: Proceedings of the 8th International Con- ference on Information Systems Security and Privacy.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Proceedings of the 8th International Con- ference on Information Systems Security and Privacy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:09.244776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:22.005344Z digest=sha256:6f825cc7a85c085c6dd974f4acf6950df1c51b95a74faa50273a60909b56153d

Observation 42bfa150-170b-405d-be2d-e18e8abd04f9 · outbound

This paper cites an unresolved cited work.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:08.864757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:22.099945Z digest=sha256:8505c77941e01a5ae041cc3c207081955043165b02e597a20f86d7f06ee77423

Observation 56870997-740c-4c89-8fb4-37d16de7c9d1 · outbound

This paper cites an unresolved cited work.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Unresolved cited work

Reference 5

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:09:06.944745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:22.173948Z digest=sha256:e65c22452e38cfd685be8f05bb07fa5ab4bb4b7b515085f12255d3bdbbd5efb0

Observation c8fbc282-17a1-4067-81e6-1c9bb0a5c761 · outbound

This paper cites In: Al-Onaizan, Y., Bansal, M., Chen, Y.N.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Al-Onaizan, Y., Bansal, M., Chen, Y.N

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:22.268933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.268933Z digest=sha256:1734db833d677e1866c1721c6454867cab664c32b8b29f319c9bbb335970e167

Observation 686507d1-4c77-4c22-a672-4981dd42c0e7 · outbound

This paper cites Internet of Things and Cyber-Physical Systems5, 1–46 (2025).https://doi.org/10.1016/j.iotcps.2025.01.001.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Internet of Things and Cyber-Physical Systems5, 1–46 (2025).https://doi.org/10.1016/j.iotcps.2025.01.001

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:22.387683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.387683Z digest=sha256:b9a9f18cce690dad47b0e13738fe1af9ee12ab809c302563558f57e5ba875c90

Observation 522d94ff-2cf8-4a0a-b126-814867508d6f · outbound

This paper cites IEEE Access12, 23733–23750 (2024).https://doi.org/10.1109/ACCESS.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response IEEE Access12, 23733–23750 (2024).https://doi.org/10.1109/ACCESS

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:08:22.483604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.483604Z digest=sha256:9b753c8edd1ad1adde158fe59a8832934d2ffe0ba336e93307347ac6042cb76e

Observation 4857d67b-0746-4bc9-ad44-c737b17a1ac4 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:22.589495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.589495Z digest=sha256:e0511d6efe73b894e9e1d932c93a8e751c22babcae4a5822a429156a1bb8b149

Observation d5079905-c200-4c49-bdcf-a3a00b93d006 · outbound

This paper cites Forensic Science International: Digital Investigation38, 301264 (Sep 2021).https: //doi.org/10.1016/j.fsidi.2021.301264.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Forensic Science International: Digital Investigation38, 301264 (Sep 2021).https: //doi.org/10.1016/j.fsidi.2021.301264

Reference 10

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:09:06.382074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:22.674613Z digest=sha256:98fd9c8ce10f7f74c1b72dbdc474c5e2fc4350c125a0b5923c6571e649f9084c

Observation ac2167b3-5f91-457d-99b2-919a03ed2e68 · outbound

This paper cites Packt Publishing, Birm- ingham, England, 2 edn.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Packt Publishing, Birm- ingham, England, 2 edn

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:08.564416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:22.783874Z digest=sha256:dd5f6251dca3dc30290cfdadd5d6b003c5c7600b8301e5faf70918540c699676

Observation e35c8751-ab69-4505-8fde-fc46e66d96df · outbound

This paper cites an unresolved cited work.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:22.852771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.852771Z digest=sha256:4bea1f42d616ebecd3cb2d77a42151d088d6e1c93ffd0725163a69cb7a05b575

Observation 0c4b10b2-35ba-456a-a8fa-bc92a2cab5d7 · outbound

This paper cites an unresolved cited work.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:08.114762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:22.916291Z digest=sha256:4ec5243571dbac19ec3a8fb16191ab580001af8173fbaa87cf79c912d1bef74f

Observation 9e7c063d-4d7d-4325-bfa6-57f05b040eb7 · outbound

This paper cites IEEE Networking Letters4(3), 162–166 (Sep 2022).https://doi.org/10.1109/ LNET.2022.3185553 14 B.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response IEEE Networking Letters4(3), 162–166 (Sep 2022).https://doi.org/10.1109/ LNET.2022.3185553 14 B

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:09:06.053146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:23.036548Z digest=sha256:01afa2ca9fdc9a36fe12d33edd5e4ec62cdd1b7eabac6ada217b35ae17abe5a7

Observation 72a33af0-4852-4e67-99d9-f8b2a6b444e6 · outbound

This paper cites In: GLOBECOM 2022 - 2022 IEEE Global Communications Conference.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: GLOBECOM 2022 - 2022 IEEE Global Communications Conference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:23.148460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:23.148460Z digest=sha256:78ba07b081ae095a7f1174e180c7d65d92e4209eff41c7f5d194710cb17414b0

Observation 1a0a9d54-458c-4caf-b6ba-91cb97dc44b2 · outbound

This paper cites Computers14(2), 67(Feb2025).https://doi.org/10.3390/computers14020067, number: 2 Publisher: Multidisciplinary Digital Publishing Institute.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Computers14(2), 67(Feb2025).https://doi.org/10.3390/computers14020067, number: 2 Publisher: Multidisciplinary Digital Publishing Institute

Reference 16

Resolution
verified exact
doi, observed 2026-08-07T14:09:02.304749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:23.254903Z digest=sha256:c12f25439ee9748bd0eab5a1b5ed39920f8689f691caf421cb37a68c39356c82

Observation 22996b99-e82f-405c-810c-79745a1871aa · outbound

This paper cites Forensic Science International: Digital Investigation48, 301683 (Mar 2024).

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Forensic Science International: Digital Investigation48, 301683 (Mar 2024)

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:09:05.635934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:59.406401Z digest=sha256:5bd5a3b72732aaaa64f0ae9e1bafbf0e758ff0b1aca86c636d55f534e65981cb

Observation e6d8b77f-7fc2-4c84-bca1-87ff85087c27 · outbound

This paper cites Ad Hoc Networks174, 103840 (Jul 2025).

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Ad Hoc Networks174, 103840 (Jul 2025)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:59.482147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:59.482147Z digest=sha256:aa941253693f59d816eb868a7a47dca8b29c85d2f7cfa0e3f59567a8de5b73d9

Observation 044247eb-f8c9-4e0d-8efa-6cdf17a71cdd · outbound

This paper cites Computer Networks227, 109688 (May 2023).https://doi.org/10.1016/j.comnet.2023.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Computer Networks227, 109688 (May 2023).https://doi.org/10.1016/j.comnet.2023

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T14:09:02.097738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:59.575213Z digest=sha256:eb0072885dfe3ea1b3e00fb9904795ffabaab162fa9d927f8ca4ba37840ee53f

Observation e2e559d2-a13e-442e-b6e0-aaa4a46e923e · outbound

This paper cites In: 2024 5th International Conference in Electronic Engineering, Information Technology & Education (EEITE).

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: 2024 5th International Conference in Electronic Engineering, Information Technology & Education (EEITE)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:59.724751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:59.724751Z digest=sha256:4dffb05b16ff72915f00f285c60a7b59434cd5043b65fe2b8fb319eb63ef203c

Observation 3ac27fb3-1426-47d9-9540-d256df0ae217 · outbound

This paper cites IEEE Software40(3), 4–8 (2023).https: //doi.org/10.1109/MS.2023.3248401.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response IEEE Software40(3), 4–8 (2023).https: //doi.org/10.1109/MS.2023.3248401

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:59.798695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:59.798695Z digest=sha256:fee085b78375454cc8a3d48718feabea441d8bbd923f3ae8e4a5b82cb4cad2cb

Observation d3630059-5892-48b5-a009-7b699e10ae81 · outbound

This paper cites In: Proceedings of the 2016 Conference on Em- pirical Methods in Natural Language Processing.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Proceedings of the 2016 Conference on Em- pirical Methods in Natural Language Processing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:07.830251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:08:59.865085Z digest=sha256:835503897e23dfc3152754ee65c2317a2d4b792716f5bad42d3cecf7679d60a0

Observation a1959e6e-bbdc-4ade-a614-fe6958d5d6b9 · outbound

This paper cites Forensic Science International: Digital Investigation46, 301609 (Oct 2023).https://doi.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Forensic Science International: Digital Investigation46, 301609 (Oct 2023).https://doi

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.014749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.014749Z digest=sha256:11285115d127a3c3cf007e6f994a4ad8b8dce5b43ade921f442c7052e25456c8

Observation 0c74d775-9016-4dce-a428-bfd8ed546358 · outbound

This paper cites Forensic Science International: DFIR-Metric: A Benchmark Dataset for Evaluating LLMs in DFIR 15 Digital Investigation52, 301872 (Mar 2025).https://doi.org/10.1016/j.fsidi.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Forensic Science International: DFIR-Metric: A Benchmark Dataset for Evaluating LLMs in DFIR 15 Digital Investigation52, 301872 (Mar 2025).https://doi.org/10.1016/j.fsidi

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:09:00.264750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.264750Z digest=sha256:7ae433e33a236fb240b43c02f26ad3311b70e3548211185e28172bf70270bdb3

Observation 85a067d4-6cd1-4b97-82c1-89b6252237ea · outbound

This paper cites EURASIP Journal on Information Security 2017(1), 15 (Oct 2017).https://doi.org/10.1186/s13635-017-0067-2.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response EURASIP Journal on Information Security 2017(1), 15 (Oct 2017).https://doi.org/10.1186/s13635-017-0067-2

Reference 27

Resolution
verified exact
doi, observed 2026-08-07T14:09:01.953142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:09:00.354748Z digest=sha256:33c9a11894c74e498c3b621c22de6b9ffdddb35d59abc5cf5fabc72d9747a4f6

Observation f040064d-e91f-473a-a599-44e0a74bf403 · outbound

This paper cites Computers and Electrical Engineering124, 110307 (2025).https://doi.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Computers and Electrical Engineering124, 110307 (2025).https://doi

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:09:04.234750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:09:00.444855Z digest=sha256:7e80e12e35377d24a3de693b519fb1e9cb558514c4f240680d68e1def2a97697

Observation 0247944f-51a0-41c1-bd3e-770c7270a65d · outbound

This paper cites https://doi.org/10.48550/arXiv.2505.03100.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response https://doi.org/10.48550/arXiv.2505.03100

Reference 29

Resolution
verified exact
doi, observed 2026-08-07T14:09:01.800745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:09:00.528611Z digest=sha256:220b8136aca3f30f40dae5e7a473aba392bb3a2016fa6bae294d698fed151820

Observation 7d854e0d-28d3-41c4-81fc-89bda38cd9f1 · outbound

This paper cites In: 2024 IEEE Interna- tional Conference on Big Data (BigData).

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: 2024 IEEE Interna- tional Conference on Big Data (BigData)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.598209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.598209Z digest=sha256:cf8d28d75ac6fb61f95e3fcb009c2bc8b36a3d4758518d5f6a6c4eb1680a6d86

Observation 3ec31bb7-ac50-4844-8102-82982d5a4821 · outbound

This paper cites In: 2024 IEEE International Conference on Cyber Security and Resilience (CSR).

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: 2024 IEEE International Conference on Cyber Security and Resilience (CSR)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.641388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.641388Z digest=sha256:28bb9315933c705fda543b6ab0618a3d33d7603e12bb101ff2fb270657c671fa

Observation 1b4fafc2-d381-4dc3-a0d5-33a84d82a272 · outbound

This paper cites In: Ideas That Cre- ated the Future, pp.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Ideas That Cre- ated the Future, pp

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:07.574818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:09:00.811978Z digest=sha256:52002bb2a3a45fef1c940fbe902bbf8036d9d6279b8827fd537e2247964a3ee1

Observation 35d555da-003d-4796-af14-32b1c69a8536 · outbound

This paper cites In: Proceedings of the 31st International Conference on Neural Information Processing Systems.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Proceedings of the 31st International Conference on Neural Information Processing Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.925258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.925258Z digest=sha256:fa664139ef5d6b81c615dd95b73fc69a04d08ee541f6bd8d40425cc113bf9bd8

Observation 9d499b04-88be-4839-bfbb-c7e9997de976 · outbound

This paper cites In: Linzen, T., Chrupała, G., Alishahi, A.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Linzen, T., Chrupała, G., Alishahi, A

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.994772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.994772Z digest=sha256:0f2459c00c71bc46460c47e0c843a40f34d846a02ac5cbc911218cd7def30089

Observation 19180516-dd0c-4787-8f96-833fe3e02d5a · outbound

This paper cites In: Proceedings of the 31st Inter- national Conference on Computational Linguistics.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Proceedings of the 31st Inter- national Conference on Computational Linguistics

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:07.253729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:09:01.065223Z digest=sha256:69b44a0c422f9894ed29ee6b063319860c7e5a8eff7c8e86fddb0fa5a34c1988

Observation 4c631ead-22b0-454f-bfe9-5ca53e402b26 · outbound

This paper cites In: Proceedings of the Digital Forensics Doctoral Sym- posium.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: Proceedings of the Digital Forensics Doctoral Sym- posium

Reference 36

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:09:03.280184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:09:01.209874Z digest=sha256:5d93ab1ffdc7fa66355f5d4c0ea5695aaa019ccb90ac22d7911d95b47fbb6efb

Observation 1497f6d1-788d-4d2c-90cb-1fd6a276aad2 · outbound

This paper cites In: 2024 12th International Sympo- sium on Digital Forensics and Security (ISDFS).

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response In: 2024 12th International Sympo- sium on Digital Forensics and Security (ISDFS)

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:01.314815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:01.314815Z digest=sha256:ae8fc347677041a03bd3aebb83f59b8e2d55508fd02220f011fd452961212a23

Observation 4eaaa9b9-46bd-4c3d-98e5-f2b8a6a11824 · outbound

This paper cites Digital Forensics in the Age of Large Language Models.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response Digital Forensics in the Age of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:01.393974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:01.393974Z digest=sha256:9bb06d85ebccb0bfdbb353b1862ec235a70158e13b70ebbce50f770020c80528

Pith citing papers

Observation 3ee7520f-eeab-46dd-a785-7edb08ccdea1 · inbound

SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents cites this paper.

SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:07.583631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:21:11.287455Z digest=sha256:d90f71dfd0885f67b6d2463b252de7dd4010b4151248f11d37336ad26712bf26