Pith. sign in

Paper Citation Record · LEDGER

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.03597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03597 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:21:36.988483Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T07:33:13.390618Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a78e576-5327-4cb2-abe6-488df5659def · inbound

Rethinking the Understanding Ability across LLMs through Mutual Information cites this paper.

Rethinking the Understanding Ability across LLMs through Mutual Information The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:36.988483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:36.988483Z digest=sha256:3e5f5e235098d968bea8de813841d7d813d4226c544d95fcf8d8e2f8fad7e451

Observation 4bd3f3c2-d3b9-4673-bab5-53e4d143ff46 · inbound

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models cites this paper.

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:41:55.997298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:39:02.015913Z digest=sha256:c0396279bcca292452c9ce678bc7275cdbf4c49fc6495582e5aad2173d3ec863

Observation 027116f9-b917-4296-870b-6cba7ef60d16 · inbound

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia cites this paper.

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:26:24.835532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T13:25:17.313704Z digest=sha256:700ba145b8f0dccb72c5d5838f69f665cfe9e1771391462decb036acfc376900

Observation 3680f679-e001-4f0e-bc07-05f693db9aae · inbound

Human-aligned AI Model Cards with Weighted Hierarchy Architecture cites this paper.

Human-aligned AI Model Cards with Weighted Hierarchy Architecture The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:11:09.673896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T09:09:11.154051Z digest=sha256:2b35128847bcf655e42827539115fd6308619b9fdfff975eaeb312b0bc22a975

Observation cd72ad2b-5c21-4586-ba94-6a6d7070fc1a · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.437860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:020a4a5ecf0d04a04ce67269cd5ba682a1cbd7edc78b47f668e3311f9386fc55

Observation 20ac7aa0-9ded-4b97-a1af-337206e4eebe · inbound

Training a General Purpose Automated Red Teaming Model cites this paper.

Training a General Purpose Automated Red Teaming Model The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:08.329246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T11:20:09.224745Z digest=sha256:dc9898200ee7347afaf0fa7febd7d29ce15b95b92ef23c75fbe26a179c42375c

Observation bd84c14f-df89-4b4d-9160-73d8af3b0f08 · inbound

Latent Performance Profiling of Large Language Models cites this paper.

Latent Performance Profiling of Large Language Models The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.392079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T07:31:02.595386Z digest=sha256:f33dcf14156e523987ff30095da0b5554d2934afb3dc73eeb188146380f1e8a7

Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · inbound

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) cites this paper.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.307850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.307850Z digest=sha256:497bbf9839857837034ae334fe076cad9340bf4fb636e346fa344a4656af45e0