Pith. sign in

Paper Citation Record · LEDGER

Quantifying Variance in Evaluation Benchmarks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2406.10229.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10229 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:56.617692Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0b1e145f-0b0c-4cbf-938e-5aa28133c018 · inbound

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws cites this paper.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Quantifying Variance in Evaluation Benchmarks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.051526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:3bc7a091dcb59fd5250f255811ab782e02f80d562279c52c1c0bfca8a504b8f9

Observation b0ea9157-82e9-4f2d-b003-8f3f5de0b70d · inbound

A Study of LLMs' Preferences for Libraries and Programming Languages cites this paper.

A Study of LLMs' Preferences for Libraries and Programming Languages Quantifying Variance in Evaluation Benchmarks

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:55:12.391707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T22:53:16.951417Z digest=sha256:8ede3b83a7669bae0b1760cdc9f1bdc2ce487d7b3d10fa9d865f067b69ca2277

Observation bc38713e-f5ab-4a85-86d9-645a25d093ce · inbound

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks cites this paper.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Quantifying Variance in Evaluation Benchmarks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.617692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.617692Z digest=sha256:8d70e065cf61cde49adc613397f4a071dc197ecfcab38e8788e304ea0c004c23

Observation bef963a3-d7a2-4043-aa23-ec9d4c5e89aa · inbound

Structure-Aware Fill-in-the-Middle Pretraining for Code cites this paper.

Structure-Aware Fill-in-the-Middle Pretraining for Code Quantifying Variance in Evaluation Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:23.000112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:23.000112Z digest=sha256:7167258b46f7314387b8cd581c4197b4e02f9ae3e9a194ad8fecd25e1b2f379b

Observation 5238fcf6-3d49-42d7-95d7-957763d3bf49 · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Quantifying Variance in Evaluation Benchmarks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:10.760594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:10.760594Z digest=sha256:ca5f252ecfdc350479993c0e40924ef0b91d09dc8d6a3ce3ae1e2d72ab04448b

Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.922228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.922228Z digest=sha256:363ce0deab0e1339b13705fd1ba70c71575cac83e796903e9a3186bce34ee3b9

Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.311443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.311443Z digest=sha256:d709d16ed43b062aea7fc9c8fa1d2988b6f9fdb400c2a236dbf7e287794392a8

Observation e2e17ed5-befc-4d79-9cb8-7275906204c8 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking Quantifying Variance in Evaluation Benchmarks

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.439511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.439511Z digest=sha256:2270aea1d318526488fd6036eba722d179f3251c16e20da49a74924407c7bd35

Observation 5c0953d6-f1b9-4059-9f27-cdadff721879 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Quantifying Variance in Evaluation Benchmarks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:47.532043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:47.532043Z digest=sha256:b7cd179bbdac38d120358e6c6ccd21d846c8d850c71075ffccb42fe6629023d1

Observation 38ae454e-4f7b-403d-a66c-cc2cee262063 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs Quantifying Variance in Evaluation Benchmarks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.003804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:e09d980d9f754353e374bb6f872865c399954030bb9c2241b10a46abbfe9784e

Observation 976900a8-7ad3-4a51-b8f9-ca3797e671d2 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Quantifying Variance in Evaluation Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:55.247281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:55.247281Z digest=sha256:56a9265dadfd44e594df307016d3651c0fbb25f05237efe3f5e8d5763af75aca

Observation 9f3ce2f6-22e1-4faa-a8a2-1ac5f18decc0 · inbound

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning cites this paper.

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning Quantifying Variance in Evaluation Benchmarks

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:36:11.793707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:26:01.717280Z digest=sha256:3b14828238069975a020b7cb028f39338d2805800ddc773af300c794d7d48049

Observation 507f4142-380b-435a-aab5-d0e2331614df · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models Quantifying Variance in Evaluation Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.592425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:989ebc47b6edb9df01b8e34ad63b50f896a1cf3749fb63b9c787f2daa55e1ebd

Observation 8bbafc1c-e21b-4b90-93a7-af375f2bddfd · inbound

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness cites this paper.

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.039511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T12:42:53.018183Z digest=sha256:9d6616907d1963d21d239affdffec89d7b7b8fdd6ee46ea3508b1fb9ce270c7f

Observation 94b16516-8599-4888-a6ea-8900097459f0 · inbound

Resolution Diagnostics for Paired LLM Evaluation cites this paper.

Resolution Diagnostics for Paired LLM Evaluation Quantifying Variance in Evaluation Benchmarks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.254548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:15:27.402338Z digest=sha256:3b703e0291f1ae260c20076e3ca57e2ecf974d8b6ceb924cf715c796b87bae49

Observation 30606fe5-3738-4848-a408-cc5b42f6d912 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Quantifying Variance in Evaluation Benchmarks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:36:44.879438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:ec5629cf2ea9a360637e493726ef172be5075947ada16d628d01ba52e3afd076

Observation 93cf316e-d257-42ee-b052-52a7e1bda135 · inbound

Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty cites this paper.

Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty Quantifying Variance in Evaluation Benchmarks

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:09:00.963089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T23:02:58.564465Z digest=sha256:4126df1174dc5ec6b29060d1b7b3b2007e0577c4777bbb37107e18a4070dece2

Observation 6d6830a2-a2fa-452e-af6b-407fe573777f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.638431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:5909cfc36fd3a00bda7f9e29b3e0b16a24f40d18103827abdeaeb0273e377684

Observation cffd2844-7096-4c59-a7f3-1a2b2b322c4f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.993727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:58b598ad20902708b44c5df4f83a1c8068dd18af7732551e1519a48547f5dca7

Observation 7c690aa6-92c3-42d1-886a-5b6d566a38d6 · inbound

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise cites this paper.

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise Quantifying Variance in Evaluation Benchmarks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:58.886380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:29:58.886380Z digest=sha256:8afc9b8d60a4941a3fbfecef7dcf25a566b98616f41922b321f38ccdc10a121c

Observation 96468dd4-d445-4fba-b111-cd7f773baf37 · inbound

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks cites this paper.

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks Quantifying Variance in Evaluation Benchmarks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:07:19.640776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T14:58:38.688149Z digest=sha256:96997c98708573122afe22b581ec369a0aa6751c3ef3f41d324e23f12c3442b4

Observation daf1a889-df9e-4597-bf7c-34b3b488455c · inbound

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally cites this paper.

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally Quantifying Variance in Evaluation Benchmarks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T12:17:16.303388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:17:16.303388Z digest=sha256:a80a5c1e084580d856a88fbc08d8ba89005e6b0dc5c90a51a45190893fceebb5

Observation 8c876188-bedd-4059-9116-3271153f8224 · inbound

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks cites this paper.

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks Quantifying Variance in Evaluation Benchmarks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:00:48.305722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:00:48.305722Z digest=sha256:0ba3fab519089c74f6bd2cba1dfe467f3c39e5579ebdf1515bef236803b89536

Observation c82b8d54-7f25-4e7d-a7c0-596b29880703 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Quantifying Variance in Evaluation Benchmarks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:44.479430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:44.479430Z digest=sha256:8147477a74065b9fbdd58d965577759b390d8e909c607593faf00187f346ba27