Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:27.021795Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 6 inbound Pith citation observations for arXiv:2506.07673.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:27.021795Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:40:23.654993Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:49:46.920258Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 26ce0089-b6d4-4d42-b7b8-b406dacf525f · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Jordan, and Tijana Zrnic
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12e946a6-75f0-4505-a47d-68740937a7ac · outbound
How Benchmark Prediction from Fewer Data Misses the Mark PPI++: Efficient Prediction-Powered Inference
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8998634-c9ec-4efd-a751-da8bce42a030 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb62aaaa-d323-4ab9-b403-cb5c3f267220 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark The fifth PASCAL recognizing textual entailment challenge
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b456a683-73b7-4e2a-9d30-d37b78bf7543 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark AutoEval Done Right: Using Synthetic Data for Model Evaluation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be3af6d-5e00-46bc-b5ee-b7b40d24cfd3 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark A singular value thresholding algorithm for matrix completion
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4900b0b7-bbd8-482c-b3ee-442914c09442 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a51dd97-39e8-48af-ba9d-61a528884761 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad63229-20d8-44ad-9299-46d04ceaf4c6 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Computing the testing error without a testing set
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b98538e5-9647-4134-87d6-3619b1fe2cbf · outbound
How Benchmark Prediction from Fewer Data Misses the Mark The PASCAL recognising textual entailment challenge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d22a3f24-6499-4e4b-86a0-4f7c3537c121 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41110b6f-0e2e-4556-bd50-59291e0a0e13 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Automatically constructing a corpus of sentential paraphrases
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57d1e637-a851-45ce-ac17-aa6b56ab3582 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dce72101-6bdb-446d-9e38-3fdc7196cc92 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Facility location: concepts, models, algo- rithms and case studies
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3eebed2e-4b99-46e3-95ad-6207737980b7 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Open llm leaderboard v2
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1bfc293-0b96-49ef-b95b-6e74709768c6 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Challenges in evaluating AI systems, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fef75381-5e12-4c9b-8e76-4763ed271a4e · outbound
How Benchmark Prediction from Fewer Data Misses the Mark The third PASCAL recognizing textual entailment challenge
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e92bd0f6-4ca0-4717-9731-1ee70e48c1e3 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark An introduction to the augmented inverse propensity weighted estimator
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75e2726-15af-4092-bb23-e4c5e4ba78b6 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Great Models Think Alike and this Undermines AI Oversight
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011eccdc-8773-4472-a298-376579d64cd4 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark A Survey on LLM-as-a-Judge
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4172e1-1add-4cf7-9d9a-ff19aaaa9c42 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e236eb94-6ade-4c5c-8d7c-372701a630ef · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d271dc8-53d5-4f79-a88a-b7ed61a6cf7f · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Test-Time Training on Nearest Neighbors for Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28058209-5f05-40f2-aee6-3bd86c0af9a6 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Measuring Massive Multitask Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3adc9fc-5581-4631-8821-dd64ac57f168 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Measuring Mathematical Problem Solving With the MATH Dataset
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0802203-fea5-4bf8-8d0a-3f8e4ffbd548 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Categorical Reparameterization with Gumbel-Softmax
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a8b5ae-eb96-4c08-9469-491f0d59427c · outbound
How Benchmark Prediction from Fewer Data Misses the Mark What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f0551a2-b906-4595-bfc7-fbf5e87d54f4 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Scaling Laws for Neural Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b45bc4-f92a-41a9-91b7-4d3ab770d763 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Active testing: Sample- efficient model evaluation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0337a2e4-f01d-4d11-bdf5-60f2c848bf25 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7fc985a9-dc69-49a8-8d29-7754f35819ff · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Retrieval- augmented generation for knowledge-intensive nlp tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fea7a5-cf22-4c32-a251-182db853b17e · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Active Evaluation Acquisition for Efficient LLM Benchmarking
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 423e05f4-9dde-4337-9584-8a968046cef3 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Manning, Christopher R’e, Diana Acosta-Navas, Drew A
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16dfc68-05f7-407b-8512-b586c565db80 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Model Similarity Mitigates Test Set Overuse
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef43aa4b-2f20-4461-a61c-8771e5acea8b · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 080a3031-0e61-44d9-a8f4-1d99a863ec85 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark How predictable is language model benchmark performance?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf06c144-35d8-4081-ae7f-9f0e12da39ea · outbound
How Benchmark Prediction from Fewer Data Misses the Mark PredictaBoard: Benchmarking LLM Score Predictability
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ef8940-83f3-4a49-8f21-17c2e82ebbe2 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark LLM Evaluators Recognize and Favor Their Own Generations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d31cfa5-4838-48d6-b3ba-5e84408948d3 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark tinyBenchmarks: evaluating LLMs with fewer examples
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b3f156-3235-4eb2-82b4-51d6389c886e · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Efficient multi-prompt evaluation of LLMs
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2570af8-3b21-4c0a-af04-5be2ba7ee595 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark SQuAD: 100,000+ questions for machine comprehension of text
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed27291c-1724-485f-8517-1abc5e842edb · outbound
How Benchmark Prediction from Fewer Data Misses the Mark GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7c9bc9-ee91-4177-a4b8-dd0edfd42832 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Semiparametric efficiency in multivariate regression models with missing data
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2738fefe-3532-4e29-96a8-666d66c99118 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Lalor, Robin Jia, and Jordan L
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b515bf99-dcf7-4650-b11c-3290558b5a87 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Observational Scaling Laws and the Predictability of Language Model Performance
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60afbc5-1bd7-4340-ad53-90e5ba059210 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Bernstein, Alexander C
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6da416b0-5c7d-4d36-ae7e-bb0990390be9 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Data distillation: A survey
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e72658b-55a9-4ae9-a9f7-4698ca13d93c · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e17da9-1936-4277-9f66-41108ef856ea · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Recursive deep models for semantic compositionality over a sentiment treebank
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ffbcf20-b91c-480a-9863-65da0b63a71c · outbound
How Benchmark Prediction from Fewer Data Misses the Mark MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad17933-bd87-4d4a-910f-86803c968dd7 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Test-time training with self-supervision for generalization under distribution shifts
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1446e7b-b1a7-4633-a1a4-c2523c4af7a0 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Le, Ed H
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a3b525b-d889-45a6-8b7c-05356aa14d5e · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Four lectures on probabilistic methods for data science
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c68714b-37a1-4384-b026-c3e71eb8d677 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Anchor Points: Benchmarking Models with Much Fewer Examples
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ead71a23-d406-4f36-b537-835d2941fa44 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa39132f-b54c-4a1f-9498-b0cc5d5562c2 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b64fdac-d565-4747-8e86-db0fdbf64704 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Self-Preference Bias in LLM-as-a-Judge
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11929c3-9425-4a9c-bf3c-74488c36aad4 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18b48d55-7d95-4d69-8dad-1b3795ef9cab · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff6138b-4a97-4362-b2f6-3655a56c0bba · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Automatic evaluation of attribution by large language models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cbd5450-d806-4bfd-8b28-5ab478823d84 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Instruction-Following Evaluation for Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a454e4-155c-44a6-93ef-63ffc29a4e1c · outbound
How Benchmark Prediction from Fewer Data Misses the Mark On Speeding Up Language Model Evaluation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e104514-0e04-4d17-a20c-fc8609e6259e · outbound
How Benchmark Prediction from Fewer Data Misses the Mark Probabilistic Bilevel Coreset Selection
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e3e299b-9295-4d96-8c3f-399e60260ee5 · outbound
How Benchmark Prediction from Fewer Data Misses the Mark How to select datapoints for efficient human evaluation of nlg models?, 2025
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b27a9196-c3e0-4eb4-9adc-6acd9af2d40c · inbound
Efficient Evaluation of LLM Performance with Statistical Guarantees How Benchmark Prediction from Fewer Data Misses the Mark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6992b4b1-fe2e-435b-85da-cacb5393970f · inbound
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking How Benchmark Prediction from Fewer Data Misses the Mark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 470626fd-38bc-4101-a212-c01eccc4759e · inbound
FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences How Benchmark Prediction from Fewer Data Misses the Mark
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e131494d-8e10-4c20-bf88-bbc35f486f3d · inbound
Validity Threats for Foundation Model Research How Benchmark Prediction from Fewer Data Misses the Mark
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf978fc5-69ca-48b4-a646-c55b8a9574af · inbound
You Don't Need to Run Every Eval How Benchmark Prediction from Fewer Data Misses the Mark
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 497f56ee-452b-4e85-b947-3c7cecea0230 · inbound
Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification How Benchmark Prediction from Fewer Data Misses the Mark
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.