Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:05:35.282725Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2605.10405.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:05:35.282725Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:08:20.340476Z
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 35ffab3c-84d1-49c4-b07d-a695f14400c7 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization On the Opportunities and Risks of Foundation Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 316257e9-30d7-40a7-98b1-c31a53bb7926 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization A survey on large language models for code generation.ACM Transactions on Software Engineering and Methodology, 35 (2):1–72
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4d6e241-5531-453e-b276-623b4ec3143d · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Emergent Abilities of Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5992728-8131-404c-a682-f577f6324e1b · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aad61b36-5ad1-4a0a-9c34-6d2967b90d8f · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Measuring mathematical problem solving with the math dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e49e4c1b-370b-4f1f-ab97-a5358aca0ca5 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Gaia: a benchmark for general ai assistants
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8be03e6f-8839-4fc1-b5be-e5d924563abc · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Efficient benchmarking of ai agents.arXiv preprint arXiv:2603.23749
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bbcb3ba-097f-4cf3-b695-5255d4a8604e · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Holistic agent leaderboard: The missing infrastructure for ai agent evaluation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation acf2c947-c9a8-40dd-a171-1758aac8c417 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Best arm identification in multi-armed bandits
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 303d3a0d-bf7f-4a4f-abae-86a46f7af6a0 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization On speeding up language model evaluation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa348ac7-0015-44e9-bc62-fa746b701407 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Semiparametric efficiency in multivariate regression models with missing data.Journal of the American Statistical Association, 90(429):122–129
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 716fd16d-738c-4263-89a1-1aacdd9da477 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Springer
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation abf59e41-4263-498c-8f5d-4904a46a2c1f · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Prediction-powered inference.Science, 382(6671):669–674
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b4c5ffb-fd23-4c44-a4d3-77edb54a8e04 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization PPI++: Efficient Prediction-Powered Inference
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 79ed0348-cb2d-41aa-a504-c5dd42e43561 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Confidence intervals for policy evaluation in adaptive experiments.Proceedings of the national academy of sciences, 118(15):e2014602118
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47d474e3-e995-4591-a498-1a4ba04e74aa · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Doubly Robust Policy Evaluation and Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0fd182fa-279c-46fc-9206-a925070d7efd · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Optimal and adaptive off-policy evalua- tion in contextual bandits
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25fb32f0-da85-4992-85f9-4b99eb513c93 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Online multi-armed bandits with adaptive inference.Advances in Neural Information Processing Systems, 34:1939–1951
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 837b5998-84e8-4a13-a895-318a55fe5944 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization The adaptive doubly robust estimator and a paradox concerning logging policy.Advances in neural information processing systems, 34:1351–1364
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae715d7a-a9c0-4aaf-8249-adc99d41ee08 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Post-contextual-bandit inference.Advances in neural information processing systems, 34:28548–28559
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 555dc1b2-47d2-4077-a03d-aa9854fc4409 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Off-policy evaluation via adaptive weighting with data from contextual bandits
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6febb63f-f64f-4b31-b921-ee252bd6dbaf · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Doubly-robust lasso bandit.Advances in Neural Information Processing Systems, 32
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2a0bd780-9b91-4f5e-9142-35f19babaa46 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Doubly robust thompson sampling with linear payoffs.Advances in neural information processing systems, 34:15830–15840
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 33bff643-cf89-47ca-87dd-2985cbc21fb8 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 754caf53-6c2c-41c0-9f28-953f3c3c872c · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Best arm identification with llm judges and limited human audits.Available at SSRN 6147806
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3332347f-b3d2-41ef-b8b1-ee9a1ba432db · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Efficient Evaluation of LLM Performance with Statistical Guarantees
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8de8400c-9ee4-4b32-9c5d-8a32c63d5e47 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Concentration inequalities for sampling without replacement.Bernoulli
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3027fe90-6a4f-4563-b486-37712814b645 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6eaa682e-e3b5-4688-b2c3-222dcb314d2f · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 18cc7a51-c3c0-4424-8280-d83feb37826b · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Instruction-Following Evaluation for Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7738c768-1d29-4896-800a-4adc16406be9 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Camel: Communicative agents for" mind" exploration of large language model society.Advances in neural information processing systems, 36:51991–52008
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4f1ebc1f-899e-413c-992a-4bb39f93f397 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Musr: Testing the limits of chain-of-thought with multistep soft reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d61f4a5-640d-4220-a72f-6636721ecb00 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Regularization paths for generalized linear models via coordinate descent.Journal of statistical software, 33:1–22
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9f85d91c-89bb-4002-8f26-7e73c88f0ae0 · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization On bernstein-type inequalities for martingales.Stochas- tic processes and their applications, 93(1):109–117
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 646e4902-2c99-41bc-978d-84894251507c · outbound
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization Z k i − ¯Z <k i 2 Fk−1 # =E
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a08e148e-dd09-4811-8b08-5fe1a3515dd9 · inbound
Efficient Sequential Evaluation of Large Language Models Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.