Pith. sign in

Paper Citation Record · LEDGER

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

As of 5 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2605.10405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10405 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T05:05:35.282725Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:08:20.340476Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact9
  • verified fuzzy24
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35ffab3c-84d1-49c4-b07d-a695f14400c7 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization On the Opportunities and Risks of Foundation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:24.206597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:bbcc7c0d2b502cbd4dd311b4e5c9edac5142b0705ae2216eb0e2787498b099b2

Observation 316257e9-30d7-40a7-98b1-c31a53bb7926 · outbound

This paper cites A survey on large language models for code generation.ACM Transactions on Software Engineering and Methodology, 35 (2):1–72.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization A survey on large language models for code generation.ACM Transactions on Software Engineering and Methodology, 35 (2):1–72

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.748037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:48e968d83c4d3c2abeff192c64eb5c5b35a6fce4b508e98d4d872f0934f18e94

Observation c4d6e241-5531-453e-b276-623b4ec3143d · outbound

This paper cites Emergent Abilities of Large Language Models.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Emergent Abilities of Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:24.214692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:8ae70f5e43e91d8b0f7436c724b5069da1f270a49f549354bbf7d32ee1b84ef3

Observation d5992728-8131-404c-a682-f577f6324e1b · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.735059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:362a38343d56776d3aa971dec05879d987e88c616a32b3b2ffdc7c755ef8a007

Observation aad61b36-5ad1-4a0a-9c34-6d2967b90d8f · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Measuring mathematical problem solving with the math dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.739298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:23f231afc36de1d53ea1f4565e4f4f7577c6d622dd825d51027e81b3dfae7f22

Observation e49e4c1b-370b-4f1f-ab97-a5358aca0ca5 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Gaia: a benchmark for general ai assistants

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.764615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:372f43d0cda295b45b5b497dc17d5e51abc2106649b70d9e3f5df8f47ddaca5f

Observation 8be03e6f-8839-4fc1-b5be-e5d924563abc · outbound

This paper cites Efficient benchmarking of ai agents.arXiv preprint arXiv:2603.23749.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Efficient benchmarking of ai agents.arXiv preprint arXiv:2603.23749

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:24.255501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:251a65a20aabaa90a86a51e69477b36d6fdad09f7dc954fe01802130cf410280

Observation 9bbcb3ba-097f-4cf3-b695-5255d4a8604e · outbound

This paper cites Holistic agent leaderboard: The missing infrastructure for ai agent evaluation.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Holistic agent leaderboard: The missing infrastructure for ai agent evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:24.276169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:7dd1f2aae3741f8f3ef75cd450fb98412d653f7324f05ac0e5121b85d4c17211

Observation acf2c947-c9a8-40dd-a171-1758aac8c417 · outbound

This paper cites Best arm identification in multi-armed bandits.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Best arm identification in multi-armed bandits

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.701844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:210a1d07694f45b809bba4517dcd706f42c10433cb1ded242ac448a2c17e52d3

Observation 303d3a0d-bf7f-4a4f-abae-86a46f7af6a0 · outbound

This paper cites On speeding up language model evaluation.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization On speeding up language model evaluation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.696353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:dd97b19be392e065a85eec83db707a3a4daadf168be689c81c232926f735d924

Observation fa348ac7-0015-44e9-bc62-fa746b701407 · outbound

This paper cites Semiparametric efficiency in multivariate regression models with missing data.Journal of the American Statistical Association, 90(429):122–129.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Semiparametric efficiency in multivariate regression models with missing data.Journal of the American Statistical Association, 90(429):122–129

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.744057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:69f54679a15542f0f6a2c06df10362eb6328a86cf53642328536cd24fb2b9384

Observation 716fd16d-738c-4263-89a1-1aacdd9da477 · outbound

This paper cites Springer.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Springer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.670494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:116be77900b35e1d81bebb70e6ca64fe98e474301c87cf1ee5bfa454d5990a91

Observation abf59e41-4263-498c-8f5d-4904a46a2c1f · outbound

This paper cites Prediction-powered inference.Science, 382(6671):669–674.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Prediction-powered inference.Science, 382(6671):669–674

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.688720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:de5f6f98d3b37a9e6d11361a928836ae9ecb7efc7a0d61850b7bcc08c81d5e9c

Observation 5b4c5ffb-fd23-4c44-a4d3-77edb54a8e04 · outbound

This paper cites PPI++: Efficient Prediction-Powered Inference.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization PPI++: Efficient Prediction-Powered Inference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:22:26.080624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:8e346a8fe9143f98380828de82d76a9f2baee8351dd8a649998d57ef3a573688

Observation 79ed0348-cb2d-41aa-a504-c5dd42e43561 · outbound

This paper cites Confidence intervals for policy evaluation in adaptive experiments.Proceedings of the national academy of sciences, 118(15):e2014602118.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Confidence intervals for policy evaluation in adaptive experiments.Proceedings of the national academy of sciences, 118(15):e2014602118

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.656675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:eaf9485691b6d0de5b782cb0fc9c3d2e9a09d95536c86f7277da8e041fef85ad

Observation 47d474e3-e995-4591-a498-1a4ba04e74aa · outbound

This paper cites Doubly Robust Policy Evaluation and Learning.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Doubly Robust Policy Evaluation and Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:24.266806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:c32c413da4ab90722db4e43fa7509994ca86e692d9080582ffb9815d532e644d

Observation 0fd182fa-279c-46fc-9206-a925070d7efd · outbound

This paper cites Optimal and adaptive off-policy evalua- tion in contextual bandits.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Optimal and adaptive off-policy evalua- tion in contextual bandits

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.666173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:aecadea51124f138c18e9e3c06720cc26970beae4f59eb647cc944b1408095e9

Observation 25fb32f0-da85-4992-85f9-4b99eb513c93 · outbound

This paper cites Online multi-armed bandits with adaptive inference.Advances in Neural Information Processing Systems, 34:1939–1951.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Online multi-armed bandits with adaptive inference.Advances in Neural Information Processing Systems, 34:1939–1951

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.730428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:e54f74b2853c0f53bdc400248d6b0bd77968fbb438db08fe7fa746e2d6540b0a

Observation 837b5998-84e8-4a13-a895-318a55fe5944 · outbound

This paper cites The adaptive doubly robust estimator and a paradox concerning logging policy.Advances in neural information processing systems, 34:1351–1364.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization The adaptive doubly robust estimator and a paradox concerning logging policy.Advances in neural information processing systems, 34:1351–1364

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.769100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:2f57f4fe4dd04036ba4f4f4ff381c666267b2e3fe3717b61447502afa47f72e2

Observation ae715d7a-a9c0-4aaf-8249-adc99d41ee08 · outbound

This paper cites Post-contextual-bandit inference.Advances in neural information processing systems, 34:28548–28559.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Post-contextual-bandit inference.Advances in neural information processing systems, 34:28548–28559

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.682464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:bfbb409c12422652f6d9409a6d7d6893611eb5840b6b7259833bb01aa1da644f

Observation 555dc1b2-47d2-4077-a03d-aa9854fc4409 · outbound

This paper cites Off-policy evaluation via adaptive weighting with data from contextual bandits.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Off-policy evaluation via adaptive weighting with data from contextual bandits

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.718175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:cff50500a3bc670547590ec5900712fc8e758d513102309d729deb2d7f2d18bc

Observation 6febb63f-f64f-4b31-b921-ee252bd6dbaf · outbound

This paper cites Doubly-robust lasso bandit.Advances in Neural Information Processing Systems, 32.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Doubly-robust lasso bandit.Advances in Neural Information Processing Systems, 32

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.722310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:f7212dc491bbc609aef55966fbc55b10e37e0cf44488fcfb25dfeddec3532f8b

Observation 2a0bd780-9b91-4f5e-9142-35f19babaa46 · outbound

This paper cites Doubly robust thompson sampling with linear payoffs.Advances in neural information processing systems, 34:15830–15840.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Doubly robust thompson sampling with linear payoffs.Advances in neural information processing systems, 34:15830–15840

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.754953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:ce2b089b1115c6ebc559a8cd9141e99adf874f176aa1d25751ffd65e90b0a96a

Observation 33bff643-cf89-47ca-87dd-2985cbc21fb8 · outbound

This paper cites Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:24.243690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:6fdf3a81ce806f775ef72ce5f7378c93650f4f19b1f39ba915d74a0da37fdd21

Observation 754caf53-6c2c-41c0-9f28-953f3c3c872c · outbound

This paper cites Best arm identification with llm judges and limited human audits.Available at SSRN 6147806.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Best arm identification with llm judges and limited human audits.Available at SSRN 6147806

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.711506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:e9d59d9ab0ee88f8549ea85863c06751364bb3fef5bce46cf824e9b90aa7f7f3

Observation 3332347f-b3d2-41ef-b8b1-ee9a1ba432db · outbound

This paper cites Efficient Evaluation of LLM Performance with Statistical Guarantees.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Efficient Evaluation of LLM Performance with Statistical Guarantees

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:24.220209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:e53313da378ab6736aae67307f66494b79317db71310042421f49aaa02fa8b24

Observation 8de8400c-9ee4-4b32-9c5d-8a32c63d5e47 · outbound

This paper cites Concentration inequalities for sampling without replacement.Bernoulli.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Concentration inequalities for sampling without replacement.Bernoulli

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.661407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:d13a25ba5507e9592ed281c0e29d2e17adc97c46c2f5ae86c56216bd4e3efa95

Observation 3027fe90-6a4f-4563-b486-37712814b645 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:41:24.233296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:e9fbe413900f1c9a69df27afb367684b08250ea5358a403ea3bd504a6c048b66

Observation 6eaa682e-e3b5-4688-b2c3-222dcb314d2f · outbound

This paper cites an unresolved cited work.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-12T12:21:33.678085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:35330cddecfe6ed5596c0c4eb920698a0ceb7b378d3b85ca2f18586e36ddd3a6

Observation 18cc7a51-c3c0-4424-8280-d83feb37826b · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Instruction-Following Evaluation for Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:24.249491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:a4791f236349e8cfaae98f5b7202ca1b5eed1d7e35ad9256ed106d56dbe86809

Observation 7738c768-1d29-4896-800a-4adc16406be9 · outbound

This paper cites Camel: Communicative agents for" mind" exploration of large language model society.Advances in neural information processing systems, 36:51991–52008.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Camel: Communicative agents for" mind" exploration of large language model society.Advances in neural information processing systems, 36:51991–52008

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.705940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:a121042d4fa449fe006d3d5e1cc4aa21b99818641ebb385da1d7cf0b0043dfa3

Observation 4f1ebc1f-899e-413c-992a-4bb39f93f397 · outbound

This paper cites Musr: Testing the limits of chain-of-thought with multistep soft reasoning.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Musr: Testing the limits of chain-of-thought with multistep soft reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.674483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:e947ed03e4374ba5c55c02cfccbf215f803be57276a44d931bb006a2a9e8df01

Observation 7d61f4a5-640d-4220-a72f-6636721ecb00 · outbound

This paper cites Regularization paths for generalized linear models via coordinate descent.Journal of statistical software, 33:1–22.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Regularization paths for generalized linear models via coordinate descent.Journal of statistical software, 33:1–22

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.759625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:6b0f7d5b97315cfc07adae25c772fa683c365b0b94cccf1f4bcea2e7b53ce8af

Observation 9f85d91c-89bb-4002-8f26-7e73c88f0ae0 · outbound

This paper cites On bernstein-type inequalities for martingales.Stochas- tic processes and their applications, 93(1):109–117.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization On bernstein-type inequalities for martingales.Stochas- tic processes and their applications, 93(1):109–117

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.773705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:18156b56638a9d06e6a829c08e3371077e64a6164abc9b262ff16464234cbc7b

Observation 646e4902-2c99-41bc-978d-84894251507c · outbound

This paper cites Z k i − ¯Z <k i 2 Fk−1 # =E.

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Z k i − ¯Z <k i 2 Fk−1 # =E

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:21:33.726231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:05:35.282725Z digest=sha256:e5506ed0780f9454378978e6b3d07a12e68cbf20d7c4acef891c8543434145e5

Pith citing papers

Observation a08e148e-dd09-4811-8b08-5fe1a3515dd9 · inbound

Efficient Sequential Evaluation of Large Language Models cites this paper.

Efficient Sequential Evaluation of Large Language Models Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:08:20.340476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:08:20.340476Z digest=sha256:bafe296149333c6ba91107098d76f238cd296dba09afe713d748d924862326e2