Pith. sign in

Paper Citation Record · LEDGER

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity

As of 12 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2411.16239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16239 v3

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:27:06.770620Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6223f76-b8bf-47fc-9623-9235a82aa926 · outbound

This paper cites The multiple-choice question should provide four answer options, with non-correct options being similar or related to the correct answer.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity The multiple-choice question should provide four answer options, with non-correct options being similar or related to the correct answer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.079989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.701093Z digest=sha256:ea4f50fab2b0e8f3d598bf0d468b84231ea4eeb325e917f0f541bdf47c2eb3b2

Observation 38462fc9-d789-48ce-8a39-a724681605a7 · outbound

This paper cites Holistic Evaluation of Language Models.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Holistic Evaluation of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.690101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.690101Z digest=sha256:a6e3442f5fe72fabee63e79a7d9878c50068294eaf368d6131d7cdd1f3f6270d

Observation d0bfada8-5a51-4057-b44e-477677586e40 · outbound

This paper cites Offer four potential answers, making sure that the incorrect options are similar to the correct one.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Offer four potential answers, making sure that the incorrect options are similar to the correct one

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.049896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.711343Z digest=sha256:91bc94ee747dbf23072628e540b514c7ff3b3fd1b0f2ab140a0764c449b8ab32

Observation 6df72d61-22f0-4855-ab47-4fc5765188aa · outbound

This paper cites Offer yes or no answers.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Offer yes or no answers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.035024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.716027Z digest=sha256:81cec8ed617f14b6523032f710e6ab136313fd8c853a432b6f296a243565ba33

Observation 8c4f9870-5a72-4fed-bbdf-e1c023fc4010 · outbound

This paper cites Provide four possible answers, including the correct one, and make sure the incorrect choices are similar or related to the right answer.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Provide four possible answers, including the correct one, and make sure the incorrect choices are similar or related to the right answer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.065007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.706171Z digest=sha256:11c5dca0f6235a16353baccb1254858eaa6e3854920840c6b77300ffd357feba

Observation d68ade81-76a4-41a5-99d2-0bfc6aa4c32d · outbound

This paper cites For example, if the original question asks about the result of a specific action, the reversed question could ask about the conditions re- quired for that result to occur.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity For example, if the original question asks about the result of a specific action, the reversed question could ask about the conditions re- quired for that result to occur

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.942721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.747016Z digest=sha256:37fcf40bb1132752eb6586053ef3dbe3f5db9a480a1f92f7b8771433297d651d

Observation 98dceb1b-d886-491e-bdd1-fc5ac4448c79 · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.924730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.751817Z digest=sha256:2ae1a5540f3911233f477427ac7c307431fcc800ac5c632acf9f1088e13462b2

Observation d10afb71-f4bd-43cc-a5be-7651a4b6a3a4 · outbound

This paper cites Ensure the question still tests the same core concept.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Ensure the question still tests the same core concept

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.020332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.721287Z digest=sha256:9cb317a762af0a987f5f1b3cefd7282663fbab95ccb1e1fad5e08691b67ec628

Observation 6f5c699c-3d24-4cb3-8fd2-4efc46a1bb4b · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.003477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.726245Z digest=sha256:deac1cffdee9609af0365ae51715f9e0489796e9f678004f848a89fd90a6e5d9

Observation 4b3bccfe-1558-42fa-8dc6-b1f553948c51 · outbound

This paper cites Ensure the question remains rel- evant to the original cybersecurity concept.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Ensure the question remains rel- evant to the original cybersecurity concept

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.988704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.731701Z digest=sha256:103e213f54e8b4e587eb336affc9f093711f35244acc7dac5b781cb6ed74ada3

Observation 6f6fb752-538b-4d2b-9ff7-889b47deb4b2 · outbound

This paper cites Ensure the question tests the same concept.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Ensure the question tests the same concept

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.973659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.737061Z digest=sha256:954f7cac07899fe46a1ef169e7e1c96e3f282d3701287bf59ee82189651bddaa

Observation cfa993c3-36fb-4edb-9287-05d0d6f95d69 · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.958427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.741964Z digest=sha256:35d1c286394709c40332104adae8ed6147842431c8f588010f8fb02d414c28bc

Observation 323ec978-3e5a-440d-88a4-d35f6f3171bd · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.909168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.756418Z digest=sha256:45c052ba31d8c54a62de145778af46b091efe149ba300d7cfbdc5edd779a00d8

Observation 44b35c47-9be7-4d36-9956-5958f5560fa7 · outbound

This paper cites {Original Question} Table 8: Prompts for Question Reformulation Prompts for Dynamic Question Generation, Part 2.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question} Table 8: Prompts for Question Reformulation Prompts for Dynamic Question Generation, Part 2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.893060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.761073Z digest=sha256:3006c1a8aa370a1d56406ec901f6b8b8a679108a44354ed2179d5965980050d0

Observation 916c879d-c879-41e5-b1dc-62767bbb9eee · outbound

This paper cites The questions are: {questions} Reply to me in the following format: ‘‘‘json [’knowledge point 1’, ’knowledge point 2’, ......] ‘‘‘.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity The questions are: {questions} Reply to me in the following format: ‘‘‘json [’knowledge point 1’, ’knowledge point 2’, ......] ‘‘‘

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.875943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.765625Z digest=sha256:9940884489bcb06689de28d30d0163a8df75f37d8bffca14baa49abcbdac617d

Observation 578995e8-0ca9-487b-8545-69d61c8db5c8 · outbound

This paper cites Then rewrite it as a multiple-choice question by leav- ing one key position blank.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Then rewrite it as a multiple-choice question by leav- ing one key position blank

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.859231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:27:06.770620Z digest=sha256:c40c06b4c6b96f90003d28eb5754fabf6dd3eb00836a03d4151b00c5bbd1f03a

Observation cae6d39c-90b5-48ae-abba-8f30f0f0f551 · outbound

This paper cites A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.683416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.683416Z digest=sha256:05e96473489033ef2cfa74a87c7ada985f7ab343d79c09d9b3d013e93dd6ccf0

Observation c1ad2b14-06d5-4e86-a267-0758ad84bb32 · outbound

This paper cites FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.695774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.695774Z digest=sha256:29a34b32d2ef26570bb1a84c1fba72922b74acf83f5e5b8eda7eb382b7160432

Pith citing papers

No inbound Pith citation observations are available.