Pith. sign in

Paper Citation Record · LEDGER

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.03492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03492 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:08:06.512628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:14:03.236562Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 772c4ac6-2bcc-4f8e-9b10-16db216a2aa6 · inbound

Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations cites this paper.

Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:06.512628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:08:06.512628Z digest=sha256:e8a5e6cb8e57f1b4fd2377a5f66b7665d15440494327e03614b0004c3cb7b866

Observation f7e1caf3-44f0-4c07-86a0-d19a3b528224 · inbound

From Queries to Criteria: Understanding How Astronomers Evaluate LLMs cites this paper.

From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:37.798046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:37.798046Z digest=sha256:90f6b2c578d9347ccd0421d02d5771fb176915ccd477c22a02306818eeda11fa

Observation 464e1b2e-033f-4181-b6cd-8445e4645deb · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:44.325768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:44.325768Z digest=sha256:b11d79430518a68d39376db4cc901921b105f6add0bfa4c4d54f43485d31a0ad

Observation 985e0aaf-b029-4c94-806b-29d130e9e035 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:53.503157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:53.503157Z digest=sha256:681f06b776e6575892df49b61d41ad513188767a8fc366fb18b90e82edb63f04

Observation 74b179ef-447e-4276-a3e8-e63b291a653c · inbound

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito cites this paper.

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T08:03:52.276129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:03:52.276129Z digest=sha256:634668df69dfe78ded31c76d643b1227e9c22a56f3876d36cd4205af9ded0ae4

Observation 59ea88af-9ed3-44f4-a313-28e2d64dd207 · inbound

Inspectable AI for Science: A Research Object Approach to Generative AI Governance cites this paper.

Inspectable AI for Science: A Research Object Approach to Generative AI Governance Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:20:59.433259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:06:30.197232Z digest=sha256:ae1d33ba49fe6a396b91fd32513a8b7a61ddfa36d820f381eb6cf2a6b0219960

Observation 9b3f1e80-46a1-4959-8f7a-6b08e0af7eff · inbound

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing cites this paper.

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.558726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:19:18.667967Z digest=sha256:24f280ded626131657f6f21b4671fd6edb20ea66139c898410b0bee11673b070

Observation 84a8b33b-3879-4803-bac8-d08a11bfb0a4 · inbound

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi cites this paper.

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.608918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:18:08.326304Z digest=sha256:c055499f31a4f87ea67cdfc877a6bfeac1b940a2a379a848a69162ec31644ea4

Observation 54f3f90f-0e1b-4d45-87d8-e5942a63277f · inbound

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility cites this paper.

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.947735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T06:55:05.581205Z digest=sha256:0ede4de8dd8626f619e037d2244430d56539c08f9c7a67dedfa872b6d17b37de

Observation a2561c61-f772-4ee2-bdac-2a9207f60677 · inbound

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility cites this paper.

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:14:03.238763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:13:00.601818Z digest=sha256:a49416d1fd5ca77f0d3aead3925e4872c4acb05f3432497de3e4be1686e8f90b

Observation d722c6cf-7d54-4b14-88a4-d94c794e860a · inbound

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings cites this paper.

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T15:02:05.907440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T15:02:05.907440Z digest=sha256:1c2b74c75fa1c7489c68ebafdf2a3c0d586cfaedcfc7a40d19815356b36c4bd9

Observation 332ff4db-a0c9-4568-97eb-ad4a60e7fae0 · inbound

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings cites this paper.

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:20.114741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:20.114741Z digest=sha256:e2b26787ee25d528a96e4d8d724e380fc98a18a3fd9279b5a5d362b06dc1d148