Pith. sign in

Paper Citation Record · LEDGER

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.03492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03492 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:08:06.512628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:14:03.236562Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 772c4ac6-2bcc-4f8e-9b10-16db216a2aa6 · inbound

Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations cites this paper.

Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:06.512628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:08:06.512628Z digest=sha256:0a6ae694d6a8be19d3b23b25aac2af98331e196d1d5c713fd106b2d5453dca65

Observation f7e1caf3-44f0-4c07-86a0-d19a3b528224 · inbound

From Queries to Criteria: Understanding How Astronomers Evaluate LLMs cites this paper.

From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:37.798046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:37.798046Z digest=sha256:69d8c5a0003b6c7273e7c6290cebb2f1a55c022ecc853bd9c20b38f6cbf7f9d7

Observation 464e1b2e-033f-4181-b6cd-8445e4645deb · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:44.325768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:44.325768Z digest=sha256:91733ab7a94c72699f52c5a7771a8f6646c44326dec54c216e31a6013cb6556f

Observation 985e0aaf-b029-4c94-806b-29d130e9e035 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:53.503157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:53.503157Z digest=sha256:fa3c23b0c6808f72dac77e1422017e155703e35c7ec8b7e7d723209ae0f01f54

Observation 74b179ef-447e-4276-a3e8-e63b291a653c · inbound

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito cites this paper.

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T08:03:52.276129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:03:52.276129Z digest=sha256:551f2cf6f76d5b1f6a2073741b94e675f920de18a71645339edc8ec7a2e4fadc

Observation 59ea88af-9ed3-44f4-a313-28e2d64dd207 · inbound

Inspectable AI for Science: A Research Object Approach to Generative AI Governance cites this paper.

Inspectable AI for Science: A Research Object Approach to Generative AI Governance Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:20:59.433259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:06:30.197232Z digest=sha256:8db5a4498f1a0029ad5fc9175b045d3bcec80e33aecc886f850bd0e27dcdfd49

Observation 9b3f1e80-46a1-4959-8f7a-6b08e0af7eff · inbound

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing cites this paper.

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.558726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:19:18.667967Z digest=sha256:4a1c28e28a00dfc1390bb44493a61e0159ae4b0667b9a52c3d809e2ecc617b20

Observation 84a8b33b-3879-4803-bac8-d08a11bfb0a4 · inbound

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi cites this paper.

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.608918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:18:08.326304Z digest=sha256:4ae4cced09b50087b091742ab578ac5e4ea00c65be8308dceabfc4b4a6b072a0

Observation 54f3f90f-0e1b-4d45-87d8-e5942a63277f · inbound

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility cites this paper.

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.947735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:55:05.581205Z digest=sha256:8f7d7a1d16cb920e4b4daf50288a3b2bbcb72ef0582139c16c0ecf80ca1f022f

Observation a2561c61-f772-4ee2-bdac-2a9207f60677 · inbound

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility cites this paper.

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:14:03.238763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:13:00.601818Z digest=sha256:fbe9832bb8cb05f2af82374485992b1d4d97d4af66cb6758c3884a85d9871983

Observation d722c6cf-7d54-4b14-88a4-d94c794e860a · inbound

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings cites this paper.

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T15:02:05.907440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T15:02:05.907440Z digest=sha256:1f04bf17459bc8b6810e0c5bd048507e0cf53ae388aef97f15d378f275a8cc2a

Observation 332ff4db-a0c9-4568-97eb-ad4a60e7fae0 · inbound

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings cites this paper.

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:20.114741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:20.114741Z digest=sha256:38fe025881ab2848be12d118e66bfb9686ba68c53d362b28ddab0ba6ad0bf60a