Pith. sign in

Paper Citation Record · LEDGER

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 5 inbound Pith citation observations for arXiv:2501.04234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04234 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:43:06.616805Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:30.024509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:47:27.677181Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ee6cec1-6a14-4895-9c83-9fe2550131a8 · outbound

This paper cites GPT-4 Technical Report.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.499332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.499332Z digest=sha256:9cf5534fd5183689821b4c04a94d08346635e7404a1d550273930e85bc382cae

Observation 6ce775bc-e399-466f-87b3-61d0fddea82c · outbound

This paper cites Bayesian inferences on uncertain ranks and orderings: Application to ranking players and lineups.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Bayesian inferences on uncertain ranks and orderings: Application to ranking players and lineups

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:07.007179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.504544Z digest=sha256:39a57a381f99c4c2d7c725c43d096072c01c8f220ef425e2d89408f31de51aff

Observation 4212ce72-189f-44e7-b5af-716038db8c01 · outbound

This paper cites Time for a change: A tutorial for comparing multiple classifiers through Bayesian analysis.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Time for a change: A tutorial for comparing multiple classifiers through Bayesian analysis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.994702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.508818Z digest=sha256:346d0b82d78a39e28ae89b7e8440c68e65a64b9890b1da98ecf3e4f580a87739

Observation c040c238-9c07-47ee-85a6-7f3826dee3b5 · outbound

This paper cites Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.982741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.513317Z digest=sha256:4ae04c14de6930a53e8882c31635335f90e533721e86495d2065e0d69bb94cbb

Observation b24df635-f720-4d26-88fc-70487398d3e6 · outbound

This paper cites Accounting for variance in machine learning benchmarks.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Accounting for variance in machine learning benchmarks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.970718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.517895Z digest=sha256:cfeb687372a867451dbb9121da63c08e9817e887886554649b13e20ff1e94e94

Observation 5bce9be4-2f6d-4392-92b4-75330a7ec61e · outbound

This paper cites What are the best systems? N ew perspectives on NLP benchmarking.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks What are the best systems? N ew perspectives on NLP benchmarking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.957749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.522408Z digest=sha256:ceebea53162a88c22f8737df7cc4a8d2e3eb54008303c2845d5719001c8ad2bd

Observation 3c0f5b98-4385-4e21-9fbf-bbdac12fa03b · outbound

This paper cites The Benchmark Lottery.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks The Benchmark Lottery

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.526759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.526759Z digest=sha256:8fd83ed490d291221fb73be32a1307c2d8e343ee6015932e5c1dfab2de00d3e4

Observation 86d6398b-416a-41b6-96a4-69e0c5ef7c5f · outbound

This paper cites Statistical comparisons of classifiers over multiple data sets.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Statistical comparisons of classifiers over multiple data sets

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.945859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.530628Z digest=sha256:fe89d688fabe14578a65a8a66f01662700905c7f85657414957cd2847a0543b3

Observation 2ae27b53-2cea-4c6a-a380-6436a2d15947 · outbound

This paper cites Bayesian aggregation of order-based rank data.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Bayesian aggregation of order-based rank data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.933812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.534246Z digest=sha256:2d846514f466ff19b3f657f8f94f9bcbf3edab2279cddfb4fda402acde2e374f

Observation afae105b-2648-43b3-b27e-d32223867a54 · outbound

This paper cites Approximate statistical tests for comparing supervised classification learning algorithms.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Approximate statistical tests for comparing supervised classification learning algorithms

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.920637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.538052Z digest=sha256:1157a6e67bfc671ef945cdc101cd323f548952f122c04d46d450d9ab4bbee435

Observation fa0b86bb-d71f-46a2-9cf1-52fcf4c1d1a3 · outbound

This paper cites Statistical significance testing for natural language processing.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Statistical significance testing for natural language processing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.908376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.541830Z digest=sha256:16502ab70c7f927fb4be6e1ae45f5900926df9cc7e15915a81c99c41952613be

Observation a03c59fb-ff64-451b-b43c-6be47a494543 · outbound

This paper cites An Introduction to the Bootstrap.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks An Introduction to the Bootstrap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.545426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.545426Z digest=sha256:5ca1ec743c8dcd9aeab08a7bd18fba4e3de116ebf4c1af64803868bee6e10127

Observation 6cf60cdd-e389-4cdb-9b7b-6d4f3e37775b · outbound

This paper cites The graphical presentation of a collection of means.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks The graphical presentation of a collection of means

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.889354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.549358Z digest=sha256:ab5bb1f7acedc505d3c323c83adff98fc46f25aeca6cdf1d156bbe9151ce6584

Observation eca38a4f-76de-4952-b108-f88de834458d · outbound

This paper cites League tables and their limitations: S tatistical issues in comparisons of institutional performance.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks League tables and their limitations: S tatistical issues in comparisons of institutional performance

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.878323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.552950Z digest=sha256:802bf54e08a2db60dce1a28217ab0ed6505851fcf342b8e47999bdee14aa4dc4

Observation d7da6371-c1dc-4d90-9473-3a247f0731aa · outbound

This paper cites Randomized significance tests in machine translation.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Randomized significance tests in machine translation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.866528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.556318Z digest=sha256:64a7c09c201367378bf062a1211fbc0dd88bc926b5be926b00d2415bcd0030b9

Observation 651d2b0e-5c5b-4fd7-8a29-f72b5187211c · outbound

This paper cites Modeling the variability of rankings.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Modeling the variability of rankings

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.853267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.559927Z digest=sha256:c4dbf66406d63ebb34f5b084a4508f5454d54bc3ae5f2ffa93b1d4885b3c61ba

Observation 2e5a9dd2-22fb-4e09-b1a6-6779713d3a43 · outbound

This paper cites Statistical comparisons of classifiers by generalized stochastic dominance.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Statistical comparisons of classifiers by generalized stochastic dominance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.841019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.563813Z digest=sha256:0936da3578d0842337f92800a9bde34a379f582b983cdd65534eb19ca0248174

Observation 123283bf-cb14-447c-a32c-e9a9b52625be · outbound

This paper cites Active Bayesian assessment of black-box classifiers.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Active Bayesian assessment of black-box classifiers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.827661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.567510Z digest=sha256:905b98678e4a78885a86bf136e22049f9a62914d300ef8975138cb59b5fa53ac

Observation 0c55c288-6523-4424-abc1-446d5bc5ca62 · outbound

This paper cites Theory of Point Estimation.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Theory of Point Estimation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.814772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.571098Z digest=sha256:ebeb7371bc67b527af16cf430a128844d8353f4bbae6cdef3dc36d08857161a8

Observation acdeee00-44d8-4ef0-b3d4-c74dbc1a7e7b · outbound

This paper cites Bayesian analysis of rank data with covariates and heterogeneous rankers.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Bayesian analysis of rank data with covariates and heterogeneous rankers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.802155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.574833Z digest=sha256:d9938404443f5cdbc7a060e10b3c9f67707a42df0f5c1f19b8acdbce7427c1a5

Observation 000e0b1b-144c-424f-8db9-63accd29a40c · outbound

This paper cites Slice sampling.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Slice sampling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.579923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.579923Z digest=sha256:3c2365fb558d2b9485f2faa15fddaf023978704a5b3592dac1ce50bcc583a5f3

Observation 8dca4664-9298-47f1-83da-5451cd1b1a9e · outbound

This paper cites Uncertainty in Ranking.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Uncertainty in Ranking

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.583624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.583624Z digest=sha256:6bf2ce205b98edbe890d2a55e9354a5bae760eaea365355a3d09939675abc45d

Observation 730be737-3b13-4fff-97fa-a7ed45edcdc4 · outbound

This paper cites On comparing classifiers: Pitfalls to avoid and a recommended approach.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks On comparing classifiers: Pitfalls to avoid and a recommended approach

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.782703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.587733Z digest=sha256:0348c7b2e9508688f548640b0315158ae68dc3944536da6a679307a712a42379

Observation 2790ecde-d15d-4e1c-b6a3-1a9f61e32af1 · outbound

This paper cites an unresolved cited work.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Unresolved cited work

Reference 24

Resolution
verified exact
doi, observed 2026-08-10T21:43:06.651434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.591539Z digest=sha256:b5e8436819af8679cbf52a3c71f04cfd6822ff57f5bfbe0c2f840fb5cb833cfa

Observation a8292bee-30c6-4278-a8ea-7e7cb8bbfc9e · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.595642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.595642Z digest=sha256:3d92358519ddaeec889cb857ce9bcbd6d9bda6b478ef09f75ce882b0eb1322c2

Observation aafa2a5d-bd87-4698-895c-8ba22576d0b3 · outbound

This paper cites Classifier uncertainty: E vidence, potential impact, and probabilistic treatment.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Classifier uncertainty: E vidence, potential impact, and probabilistic treatment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.770303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.600120Z digest=sha256:6737859307a9fc18f4cae239cb55838571d7d4d794840d89f032fb77d25851c8

Observation 56df0d95-102f-4570-a73d-4924f262a84b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.604550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.604550Z digest=sha256:99d10fc1d4c58eb83dd2ef3e79af05e89335f8f7b8f2b4ed835ab3c9df8d64bd

Observation 6edff260-ef13-4ad0-9033-af101c7eba67 · outbound

This paper cites Follow the leader (board) with confidence: E stimating p-values from a single test set with item and response variance.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Follow the leader (board) with confidence: E stimating p-values from a single test set with item and response variance

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.757704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.608568Z digest=sha256:1c6434343e1454960a7c86012c72ad74a3ed3a76f6b2a4fb2f2d41d789595e48

Observation 86801887-c27a-478a-b0c3-9e707c3d3173 · outbound

This paper cites Confidence intervals for population ranks in the presence of ties and near ties.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Confidence intervals for population ranks in the presence of ties and near ties

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.744142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.612790Z digest=sha256:53376763d5a2d950f533da7196e8e167da5bbc06c47246ed13e33c7f6b889d8d

Observation de1b81ab-b83a-4acf-955c-cdf7dd101341 · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.616805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.616805Z digest=sha256:a15ba3c656af98d293b636a7b999d7b7259ae292ee6987daaa1cdd77dd2520e5

Pith citing papers

Observation 665aba2b-438d-435f-a8da-e66539ee39eb · inbound

Rethinking LLM Parametric Knowledge as Post-retrieval Confidence for Dynamic Retrieval and Reranking cites this paper.

Rethinking LLM Parametric Knowledge as Post-retrieval Confidence for Dynamic Retrieval and Reranking Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:30.024509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:30.024509Z digest=sha256:b80d2496f0d2d213c6b082559939e94eb54e5bdbc0a96203f09a91d3457fb5f1

Observation d087617a-4524-4cd8-a690-bc1ca6097709 · inbound

Unstable Rankings in Bayesian Deep Learning Evaluation cites this paper.

Unstable Rankings in Bayesian Deep Learning Evaluation Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:31:15.186122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:34:44.254637Z digest=sha256:8122e6c11697021c37b95e97dd481f87eb8864dbbd666d74593fd9455dad8191

Observation 83b60adc-67a0-46a5-a3fc-5402f697f8ad · inbound

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning cites this paper.

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.788732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:26:01.717280Z digest=sha256:54ba4d1760f12b9fc7e4454f632b3c4dd8861f6624fa551ce5e022d06b8efa74

Observation c3a5bc77-6ef5-4ef8-bb6c-936e0b756488 · inbound

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation cites this paper.

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:47:27.679532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T17:54:22.974336Z digest=sha256:b8d82ec3fc1de03551359ee7df8932d504f4f3b41d1f3f19c0f29d6a951e275f

Observation 3514908c-2504-42d2-9ea3-a515cef37245 · inbound

Quantifying Ranking Uncertainty in LLM Benchmarks cites this paper.

Quantifying Ranking Uncertainty in LLM Benchmarks Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:26.760955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:26.760955Z digest=sha256:5e63c0d94cfdfc512c239031bf782ee0f2e1fe84ef7875acefbd2fb79d80a65a