Pith. sign in

Paper Citation Record · LEDGER

How Benchmark Prediction from Fewer Data Misses the Mark

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 6 inbound Pith citation observations for arXiv:2506.07673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07673 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:27.021795Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:40:23.654993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.920258Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved38
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26ce0089-b6d4-4d42-b7b8-b406dacf525f · outbound

This paper cites Jordan, and Tijana Zrnic.

How Benchmark Prediction from Fewer Data Misses the Mark Jordan, and Tijana Zrnic

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.846205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.564352Z digest=sha256:2399d4ed1fef0bb41c5549915294ac25a89ff65ff022cdd98014075eed120fb3

Observation 12e946a6-75f0-4505-a47d-68740937a7ac · outbound

This paper cites PPI++: Efficient Prediction-Powered Inference.

How Benchmark Prediction from Fewer Data Misses the Mark PPI++: Efficient Prediction-Powered Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.571156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.571156Z digest=sha256:45b964e44c3883e4c46423cc564be703a145bb11f3204f67b1107724bf31a8a3

Observation e8998634-c9ec-4efd-a751-da8bce42a030 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

How Benchmark Prediction from Fewer Data Misses the Mark Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.582523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.582523Z digest=sha256:4e3ca0ac1821a612ea586343e4b4806006cfe683f2192863de1b3827e5fa81b5

Observation cb62aaaa-d323-4ab9-b403-cb5c3f267220 · outbound

This paper cites The fifth PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The fifth PASCAL recognizing textual entailment challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.595425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.595425Z digest=sha256:7af1782318e3b281049fa7fe7b1ae2ffaf0320f8bd7c54684d9e3ce8a0d1d860

Observation b456a683-73b7-4e2a-9d30-d37b78bf7543 · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.612857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.612857Z digest=sha256:0433da034477e976be7e663c25a7fb10f48d5e8dd95be7e12dd655d445484154

Observation 4be3af6d-5e00-46bc-b5ee-b7b40d24cfd3 · outbound

This paper cites A singular value thresholding algorithm for matrix completion.

How Benchmark Prediction from Fewer Data Misses the Mark A singular value thresholding algorithm for matrix completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.629844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.629844Z digest=sha256:4b3cec14c40d2e34dbfbc6fed5778855abc710eb068b4e6060fed147ee75a5f1

Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.647558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.647558Z digest=sha256:8a3529fb0b7bae3feb2a06a5e1f357b9ea2e8a126f7bd072ba70ce00b667ff54

Observation 4900b0b7-bbd8-482c-b3ee-442914c09442 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

How Benchmark Prediction from Fewer Data Misses the Mark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.659435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.659435Z digest=sha256:94b196d3d7d36286ca9723c5a38546a7e5754d81de5b85c0f1be77a0474f987b

Observation 4a51dd97-39e8-48af-ba9d-61a528884761 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Benchmark Prediction from Fewer Data Misses the Mark Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.728273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.728273Z digest=sha256:067782e9a95ad06956b52e7c68fb439844790a9a70af4758a959d872a874525b

Observation aad63229-20d8-44ad-9299-46d04ceaf4c6 · outbound

This paper cites Computing the testing error without a testing set.

How Benchmark Prediction from Fewer Data Misses the Mark Computing the testing error without a testing set

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.823011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.804792Z digest=sha256:63816c5fdaaa30e5a50bc91c42f325aa5ae1cb4f00fea227b815954da68de12b

Observation b98538e5-9647-4134-87d6-3619b1fe2cbf · outbound

This paper cites The PASCAL recognising textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The PASCAL recognising textual entailment challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.813136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.845227Z digest=sha256:ff8e0afb8afb5c0e1a63a8a82a4ea918e2f7f86f9f3823148b6ca01ae8cd4b28

Observation d22a3f24-6499-4e4b-86a0-4f7c3537c121 · outbound

This paper cites Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021.

How Benchmark Prediction from Fewer Data Misses the Mark Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.803019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.848483Z digest=sha256:3a65e026a48c90d9f11f458f29d0391349baaf8d44f0f38d1588ee9349355816

Observation 41110b6f-0e2e-4556-bd50-59291e0a0e13 · outbound

This paper cites Automatically constructing a corpus of sentential paraphrases.

How Benchmark Prediction from Fewer Data Misses the Mark Automatically constructing a corpus of sentential paraphrases

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.791643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.851663Z digest=sha256:dd7d0778fd2a48c9d7dbca96f5c7038ecfd0160701789a0154d83a060fe4bd68

Observation 57d1e637-a851-45ce-ac17-aa6b56ab3582 · outbound

This paper cites Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data.

How Benchmark Prediction from Fewer Data Misses the Mark Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.779352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.854801Z digest=sha256:c246ac779e21bdd4575cd21040d614a83f7125071c32b1e4cd901c46b70af83c

Observation dce72101-6bdb-446d-9e38-3fdc7196cc92 · outbound

This paper cites Facility location: concepts, models, algo- rithms and case studies.

How Benchmark Prediction from Fewer Data Misses the Mark Facility location: concepts, models, algo- rithms and case studies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.768782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.858336Z digest=sha256:53f2e9391df03f4b6894fe1cf25ad9d1ddb888cb1f8d3adb31bc5370d489bd9b

Observation 3eebed2e-4b99-46e3-95ad-6207737980b7 · outbound

This paper cites Open llm leaderboard v2.

How Benchmark Prediction from Fewer Data Misses the Mark Open llm leaderboard v2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.758038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.861417Z digest=sha256:e2d687ed667185245ad4d56df1b01109a9d4f63896eb887ce3d1154fec14a040

Observation d1bfc293-0b96-49ef-b95b-6e74709768c6 · outbound

This paper cites Challenges in evaluating AI systems, 2023.

How Benchmark Prediction from Fewer Data Misses the Mark Challenges in evaluating AI systems, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.747484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.864676Z digest=sha256:4ba764aaf653ef0dd8acd20bf6a89a7c24cdc7259cdacaa0ec8ea540f373eb17

Observation fef75381-5e12-4c9b-8e76-4763ed271a4e · outbound

This paper cites The third PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The third PASCAL recognizing textual entailment challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.737314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.868362Z digest=sha256:9b2afe6565b76b102105a194b7cf82311e80cd52d14b50671a75a35a9c7fbbf5

Observation e92bd0f6-4ca0-4717-9731-1ee70e48c1e3 · outbound

This paper cites An introduction to the augmented inverse propensity weighted estimator.

How Benchmark Prediction from Fewer Data Misses the Mark An introduction to the augmented inverse propensity weighted estimator

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.872011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.872011Z digest=sha256:869c00af8beb7dcd3078963c0e394d420259f5c1722c67c25a4e86fd55ef40de

Observation f75e2726-15af-4092-bb23-e4c5e4ba78b6 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

How Benchmark Prediction from Fewer Data Misses the Mark Great Models Think Alike and this Undermines AI Oversight

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.874878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.874878Z digest=sha256:37f5a4773ff80dc5d9bf0e496f68d4ae20cda987cbefbd5dbc1b335053b21b35

Observation 011eccdc-8773-4472-a298-376579d64cd4 · outbound

This paper cites A Survey on LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark A Survey on LLM-as-a-Judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.877994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.877994Z digest=sha256:353eee5b4ddbe72923bfe18b778338a9a9064e10db0d1e3f7900ae3f7e8b3367

Observation 9a4172e1-1add-4cf7-9d9a-ff19aaaa9c42 · outbound

This paper cites Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N.

How Benchmark Prediction from Fewer Data Misses the Mark Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.719690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.880977Z digest=sha256:f916637892244d4cf76969f7619f66458362572b1508df2017bece157504f70b

Observation e236eb94-6ade-4c5c-8d7c-372701a630ef · outbound

This paper cites Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings.

How Benchmark Prediction from Fewer Data Misses the Mark Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.482516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.884061Z digest=sha256:2487985b21c95f2ee7990a322efceb6a5616c0dcc90acae6eda9136e08a9a2dc

Observation 9d271dc8-53d5-4f79-a88a-b7ed61a6cf7f · outbound

This paper cites Test-Time Training on Nearest Neighbors for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Test-Time Training on Nearest Neighbors for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.887369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.887369Z digest=sha256:acafef9fbee8bc20c90e000b223f393c9cbeb069a7c2bf9a38312c3e614c365f

Observation 28058209-5f05-40f2-aee6-3bd86c0af9a6 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Massive Multitask Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.890744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.890744Z digest=sha256:9a055b97f76b18b4a029326a3b2de26e964c1702ad61c91b82db5c1e2bc7c13f

Observation e3adc9fc-5581-4631-8821-dd64ac57f168 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Mathematical Problem Solving With the MATH Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.894434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.894434Z digest=sha256:b268240c29a02f39efc840964280faaeb4c5fa6fdca1290d596df4d71d443f9c

Observation f0802203-fea5-4bf8-8d0a-3f8e4ffbd548 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

How Benchmark Prediction from Fewer Data Misses the Mark Categorical Reparameterization with Gumbel-Softmax

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.897769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.897769Z digest=sha256:44f8eec8240432ca2d0c03233f6eeecb8cf0efd4682772c011ec06217d7e1459

Observation 87a8b5ae-eb96-4c08-9469-491f0d59427c · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

How Benchmark Prediction from Fewer Data Misses the Mark What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.901004Z digest=sha256:e1b90afa262e0fe13abe21d5117ea9129268ead2b4fb3feb69d9799760920c3d

Observation 2f0551a2-b906-4595-bfc7-fbf5e87d54f4 · outbound

This paper cites Scaling Laws for Neural Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling Laws for Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.904300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.904300Z digest=sha256:cceeecd295ba82eca12e491bdfcc4dce6fe9b6d9f0262dfc8a9254e3a43fe69c

Observation 31b45bc4-f92a-41a9-91b7-4d3ab770d763 · outbound

This paper cites Active testing: Sample- efficient model evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active testing: Sample- efficient model evaluation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.709413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.907237Z digest=sha256:afbee049bc306adce10f0f661c05e5040df463e82da919ba301eb76e72e41434

Observation 0337a2e4-f01d-4d11-bdf5-60f2c848bf25 · outbound

This paper cites Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.405042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.909784Z digest=sha256:b01f0a6006fe0eb1ec3295dde7365f7498d790950b5d9894eaab3d382b883a92

Observation 7fc985a9-dc69-49a8-8d29-7754f35819ff · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

How Benchmark Prediction from Fewer Data Misses the Mark Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.912846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.912846Z digest=sha256:80b97df176c3ad13227ecc017e98d7128a2959b5ddb2fe32173bd52709c0bd3b

Observation 75fea7a5-cf22-4c32-a251-182db853b17e · outbound

This paper cites Active Evaluation Acquisition for Efficient LLM Benchmarking.

How Benchmark Prediction from Fewer Data Misses the Mark Active Evaluation Acquisition for Efficient LLM Benchmarking

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.915668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.915668Z digest=sha256:704743b64aed7265eb2d7fda272abfdd23aba22a3d529b06acce608c4d49cbfd

Observation 423e05f4-9dde-4337-9584-8a968046cef3 · outbound

This paper cites Manning, Christopher R’e, Diana Acosta-Navas, Drew A.

How Benchmark Prediction from Fewer Data Misses the Mark Manning, Christopher R’e, Diana Acosta-Navas, Drew A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.692361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.919274Z digest=sha256:30e3b95d700c0f2d12eba93ee3ae4c64fd32b262f95b2dad0bb15b685a1ff568

Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.922228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.922228Z digest=sha256:7c96afe2e2a0883dbbe911fb19ce3b933d28a1d2ae3790ddf9e24506224842c2

Observation c16dfc68-05f7-407b-8512-b586c565db80 · outbound

This paper cites Model Similarity Mitigates Test Set Overuse.

How Benchmark Prediction from Fewer Data Misses the Mark Model Similarity Mitigates Test Set Overuse

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.367042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.925314Z digest=sha256:084514414ca6ee7e6d0ee708e804d73cff1c733b36e1b18fef2b4b23f92ab1b6

Observation ef43aa4b-2f20-4461-a61c-8771e5acea8b · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

How Benchmark Prediction from Fewer Data Misses the Mark Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.682228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.928473Z digest=sha256:83cf1505eeac2157a4dc5972eef984f090f43709def2b280cc24be7b6e394f05

Observation 080a3031-0e61-44d9-a8f4-1d99a863ec85 · outbound

This paper cites How predictable is language model benchmark performance?.

How Benchmark Prediction from Fewer Data Misses the Mark How predictable is language model benchmark performance?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.931350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.931350Z digest=sha256:1f2cca93b89ef6219b76fee18d6cc7bc4df0bf63a0901ab3d92f4c59b186f128

Observation bf06c144-35d8-4081-ae7f-9f0e12da39ea · outbound

This paper cites PredictaBoard: Benchmarking LLM Score Predictability.

How Benchmark Prediction from Fewer Data Misses the Mark PredictaBoard: Benchmarking LLM Score Predictability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.934415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.934415Z digest=sha256:b911f3636523c8ced681da2098dffe8f686f6047056f2981dbead2d370963d8a

Observation 30ef8940-83f3-4a49-8f21-17c2e82ebbe2 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

How Benchmark Prediction from Fewer Data Misses the Mark LLM Evaluators Recognize and Favor Their Own Generations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.937721Z digest=sha256:0ce694fc20b2c2e3e69483a9802d167c18c1286f6d0682c74fb9a3cb9aa1a1c3

Observation 7d31cfa5-4838-48d6-b3ba-5e84408948d3 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

How Benchmark Prediction from Fewer Data Misses the Mark tinyBenchmarks: evaluating LLMs with fewer examples

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.940907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.940907Z digest=sha256:4def5ef88f5ab9c5c5dca5278e39e21b041019a4f4acd483d0eea835029b40a2

Observation 98b3f156-3235-4eb2-82b4-51d6389c886e · outbound

This paper cites Efficient multi-prompt evaluation of LLMs.

How Benchmark Prediction from Fewer Data Misses the Mark Efficient multi-prompt evaluation of LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.944282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.944282Z digest=sha256:ebe803d564a2eaaafd4118a6b965ac9c1c88500236ab52596f2d69fbe83489c9

Observation c2570af8-3b21-4c0a-af04-5be2ba7ee595 · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

How Benchmark Prediction from Fewer Data Misses the Mark SQuAD: 100,000+ questions for machine comprehension of text

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.671738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.947725Z digest=sha256:3875644acd9a8a14cc087f61bfb4120cbb40d3b3fbd2fee067a0357db4d929c3

Observation ed27291c-1724-485f-8517-1abc5e842edb · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.950951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.950951Z digest=sha256:c049ff6ff3238f2b2a94233c6d6b333609d15f9fe7c7e17a86f6ec3ca1d7c335

Observation 2f7c9bc9-ee91-4177-a4b8-dd0edfd42832 · outbound

This paper cites Semiparametric efficiency in multivariate regression models with missing data.

How Benchmark Prediction from Fewer Data Misses the Mark Semiparametric efficiency in multivariate regression models with missing data

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.954437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.954437Z digest=sha256:fde70623e79de6ec2aaac986a10805c24993daa47298014385643d3a7d32708c

Observation 2738fefe-3532-4e29-96a8-666d66c99118 · outbound

This paper cites Lalor, Robin Jia, and Jordan L.

How Benchmark Prediction from Fewer Data Misses the Mark Lalor, Robin Jia, and Jordan L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.655964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.957403Z digest=sha256:31bcfc19d5759b1e72a9d4658bcbd22c2ac38a0009c5d42491cc1c1cca027653

Observation b515bf99-dcf7-4650-b11c-3290558b5a87 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

How Benchmark Prediction from Fewer Data Misses the Mark Observational Scaling Laws and the Predictability of Language Model Performance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.960631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.960631Z digest=sha256:78416ed843ddc0f7d1c1c59bcde87c09afbe2dde561971a039df5f76fafd3f5d

Observation b60afbc5-1bd7-4340-ad53-90e5ba059210 · outbound

This paper cites Bernstein, Alexander C.

How Benchmark Prediction from Fewer Data Misses the Mark Bernstein, Alexander C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.646469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.963910Z digest=sha256:16995834f6b215ed49d72cf1b87bbf1e9f885ee54e441d5c11d48015e9ee965f

Observation 6da416b0-5c7d-4d36-ae7e-bb0990390be9 · outbound

This paper cites Data distillation: A survey.

How Benchmark Prediction from Fewer Data Misses the Mark Data distillation: A survey

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.636832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.966967Z digest=sha256:90f56e3862704f88fade67a68dec30338bd2a66ff5bbdbaedd24d1b688875881

Observation 1e72658b-55a9-4ae9-a9f7-4698ca13d93c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.970380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.970380Z digest=sha256:f03eb1ae8327d625cc03e3fa894abef8ba8b615b04a5e2de4d145cb2f8203ddb

Observation 97e17da9-1936-4277-9f66-41108ef856ea · outbound

This paper cites Recursive deep models for semantic compositionality over a sentiment treebank.

How Benchmark Prediction from Fewer Data Misses the Mark Recursive deep models for semantic compositionality over a sentiment treebank

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.627625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.973877Z digest=sha256:109961a33c51ec317224e8b2a941a450a73af04ef2799ecc4f26a624877aa8a2

Observation 0ffbcf20-b91c-480a-9863-65da0b63a71c · outbound

This paper cites MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning.

How Benchmark Prediction from Fewer Data Misses the Mark MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.976895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.976895Z digest=sha256:af5d9efbaa6f403192b1384f4ee1996297490ae4d7c8fb87bb72cb1ca50c23a5

Observation fad17933-bd87-4d4a-910f-86803c968dd7 · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

How Benchmark Prediction from Fewer Data Misses the Mark Test-time training with self-supervision for generalization under distribution shifts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.980129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.980129Z digest=sha256:fcecd0d06b1b36c550fc055b005782c8684344d42f790f89bc7a4b1b2d4f6b19

Observation a1446e7b-b1a7-4633-a1a4-c2523c4af7a0 · outbound

This paper cites Le, Ed H.

How Benchmark Prediction from Fewer Data Misses the Mark Le, Ed H

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.610789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.983164Z digest=sha256:b457dccacecd0faa6c4b96633f7e87d803829c19efab12e140a3d01fddc5a426

Observation 6a3b525b-d889-45a6-8b7c-05356aa14d5e · outbound

This paper cites Four lectures on probabilistic methods for data science.

How Benchmark Prediction from Fewer Data Misses the Mark Four lectures on probabilistic methods for data science

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.141439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.986668Z digest=sha256:ac8d0654134804b3c3bf7eecc0a0389844375d789c0067387da771ea3e29f269

Observation 5c68714b-37a1-4384-b026-c3e71eb8d677 · outbound

This paper cites Anchor Points: Benchmarking Models with Much Fewer Examples.

How Benchmark Prediction from Fewer Data Misses the Mark Anchor Points: Benchmarking Models with Much Fewer Examples

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.127869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.989961Z digest=sha256:43df36c0421fd1055f4af0bcd6860247208f3063a1630db0f907873481b3d1d6

Observation ead71a23-d406-4f36-b537-835d2941fa44 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.601129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:26.993144Z digest=sha256:03233f7c1caddcaa4da1e6a8075d9a8c1975269a0273bd7082e95b059ec65d63

Observation aa39132f-b54c-4a1f-9498-b0cc5d5562c2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.996332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.996332Z digest=sha256:4bc767d9d7f78381874cfb3b9ef4ad14c3e2f65963ef8a4af10a0d2081b30384

Observation 1b64fdac-d565-4747-8e86-db0fdbf64704 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark Self-Preference Bias in LLM-as-a-Judge

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.999567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.999567Z digest=sha256:b57130e7c14b7d365a10aac18876c6f46ac417ec0abd287c7e8d704018502c41

Observation a11929c3-9425-4a9c-bf3c-74488c36aad4 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.591844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:27.002653Z digest=sha256:f404d46d941a401bacc59eeb985188f40748f200ce8706f77131e1fa8aad974e

Observation 18b48d55-7d95-4d69-8dad-1b3795ef9cab · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

How Benchmark Prediction from Fewer Data Misses the Mark Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.005749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.005749Z digest=sha256:c70ecc8e1e5691350099b78512ab9d0ae543acf8245f8410c5c32f1f82583f66

Observation dff6138b-4a97-4362-b2f6-3655a56c0bba · outbound

This paper cites Automatic evaluation of attribution by large language models.

How Benchmark Prediction from Fewer Data Misses the Mark Automatic evaluation of attribution by large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.582423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:27.008828Z digest=sha256:954e8e3b1183d1ab3dd307b997ed93983702e76eefd777904ae56789c3efdc5d

Observation 9cbd5450-d806-4bfd-8b28-5ab478823d84 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Instruction-Following Evaluation for Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.011892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.011892Z digest=sha256:d1b63c38178d077383aac9efa5e351a36681bda395db003bfb2adf1576b7ef13

Observation e8a454e4-155c-44a6-93ef-63ffc29a4e1c · outbound

This paper cites On Speeding Up Language Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark On Speeding Up Language Model Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.015178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.015178Z digest=sha256:4bc3980b2681947a9158f697a92a3661e4b05220eedd1cc0a252d6479584c778

Observation 8e104514-0e04-4d17-a20c-fc8609e6259e · outbound

This paper cites Probabilistic Bilevel Coreset Selection.

How Benchmark Prediction from Fewer Data Misses the Mark Probabilistic Bilevel Coreset Selection

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.059462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:27.018282Z digest=sha256:e73a13001dd852f40ba2b0ad73f26c1bba4adab470b73bb5d1b67f5c10681fba

Observation 1e3e299b-9295-4d96-8c3f-399e60260ee5 · outbound

This paper cites How to select datapoints for efficient human evaluation of nlg models?, 2025.

How Benchmark Prediction from Fewer Data Misses the Mark How to select datapoints for efficient human evaluation of nlg models?, 2025

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:34:27.572476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:27.021795Z digest=sha256:ac0a66555425d7ade81c4e4d8fbf105946336bf0416ebe980f5c43e450f5145b

Pith citing papers

Observation b27a9196-c3e0-4eb4-9adc-6acd9af2d40c · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees How Benchmark Prediction from Fewer Data Misses the Mark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:0885c3e9a6ded8a0fa38bf8919e87cd09184d65eb01b23a03d8bc14d972b1f20

Observation 6992b4b1-fe2e-435b-85da-cacb5393970f · inbound

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking cites this paper.

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking How Benchmark Prediction from Fewer Data Misses the Mark

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T05:26:41.529866Z digest=sha256:a8a8ffc09c952cfdadc85e62b92928ac5cea5f6456d47894ff3ef15b67fae0a6

Observation 470626fd-38bc-4101-a212-c01eccc4759e · inbound

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences cites this paper.

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences How Benchmark Prediction from Fewer Data Misses the Mark

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:24:01.547119Z digest=sha256:103b94a747e3d1769dccbf6a4e66e4e97291d0556b3c24c6a87920b1c72e0a40

Observation e131494d-8e10-4c20-bf88-bbc35f486f3d · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research How Benchmark Prediction from Fewer Data Misses the Mark

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:797f221e8e79d80f4eccb573852eb39fbc46bd9ac985275940a355c6796cc477

Observation cf978fc5-69ca-48b4-a646-c55b8a9574af · inbound

You Don't Need to Run Every Eval cites this paper.

You Don't Need to Run Every Eval How Benchmark Prediction from Fewer Data Misses the Mark

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:23:12.145516Z digest=sha256:57b5c7e191aa98b7dfd5c8e3e477cf1e672da048e9812b73ae36de0bbd45c5ff

Observation 497f56ee-452b-4e85-b947-3c7cecea0230 · inbound

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification cites this paper.

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification How Benchmark Prediction from Fewer Data Misses the Mark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T06:40:23.654993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:40:23.654993Z digest=sha256:44f4ff77eeea0a920a50bcad447a02eaaf3733c2bdb59a4354740266a55b61cb