Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2406.10229.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:56.617692Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 0b1e145f-0b0c-4cbf-938e-5aa28133c018 · inbound
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Quantifying Variance in Evaluation Benchmarks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0ea9157-82e9-4f2d-b003-8f3f5de0b70d · inbound
A Study of LLMs' Preferences for Libraries and Programming Languages Quantifying Variance in Evaluation Benchmarks
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc38713e-f5ab-4a85-86d9-645a25d093ce · inbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Quantifying Variance in Evaluation Benchmarks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bef963a3-d7a2-4043-aa23-ec9d4c5e89aa · inbound
Structure-Aware Fill-in-the-Middle Pretraining for Code Quantifying Variance in Evaluation Benchmarks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5238fcf6-3d49-42d7-95d7-957763d3bf49 · inbound
Beyond Text Compression: Evaluating Tokenizers Across Scales Quantifying Variance in Evaluation Benchmarks
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · inbound
How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · inbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e17ed5-befc-4d79-9cb8-7275906204c8 · inbound
Fluid Language Model Benchmarking Quantifying Variance in Evaluation Benchmarks
Reference 1983
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c0953d6-f1b9-4059-9f27-cdadff721879 · inbound
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Quantifying Variance in Evaluation Benchmarks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38ae454e-4f7b-403d-a66c-cc2cee262063 · inbound
The Art of Scaling Reinforcement Learning Compute for LLMs Quantifying Variance in Evaluation Benchmarks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 976900a8-7ad3-4a51-b8f9-ca3797e671d2 · inbound
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Quantifying Variance in Evaluation Benchmarks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3ce2f6-22e1-4faa-a8a2-1ac5f18decc0 · inbound
A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning Quantifying Variance in Evaluation Benchmarks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 507f4142-380b-435a-aab5-d0e2331614df · inbound
Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models Quantifying Variance in Evaluation Benchmarks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8bbafc1c-e21b-4b90-93a7-af375f2bddfd · inbound
The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness Quantifying Variance in Evaluation Benchmarks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 94b16516-8599-4888-a6ea-8900097459f0 · inbound
Resolution Diagnostics for Paired LLM Evaluation Quantifying Variance in Evaluation Benchmarks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 30606fe5-3738-4848-a408-cc5b42f6d912 · inbound
Validity Threats for Foundation Model Research Quantifying Variance in Evaluation Benchmarks
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93cf316e-d257-42ee-b052-52a7e1bda135 · inbound
Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty Quantifying Variance in Evaluation Benchmarks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d6830a2-a2fa-452e-af6b-407fe573777f · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks
Reference 197
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cffd2844-7096-4c59-a7f3-1a2b2b322c4f · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks
Reference 197
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7c690aa6-92c3-42d1-886a-5b6d566a38d6 · inbound
How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise Quantifying Variance in Evaluation Benchmarks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96468dd4-d445-4fba-b111-cd7f773baf37 · inbound
DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks Quantifying Variance in Evaluation Benchmarks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation daf1a889-df9e-4597-bf7c-34b3b488455c · inbound
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally Quantifying Variance in Evaluation Benchmarks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c876188-bedd-4059-9116-3271153f8224 · inbound
Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks Quantifying Variance in Evaluation Benchmarks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82b8d54-7f25-4e7d-a7c0-596b29880703 · inbound
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Quantifying Variance in Evaluation Benchmarks
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.