Pith. sign in

Paper Citation Record · LEDGER

How Benchmark Prediction from Fewer Data Misses the Mark

As of 20 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 6 inbound Pith citation observations for arXiv:2506.07673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07673 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:27.021795Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:40:23.654993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.920258Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved38
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26ce0089-b6d4-4d42-b7b8-b406dacf525f · outbound

This paper cites Jordan, and Tijana Zrnic.

How Benchmark Prediction from Fewer Data Misses the Mark Jordan, and Tijana Zrnic

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.846205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.564352Z digest=sha256:e6ed9fb63f8a97d0e086b0138658726d5305604d3f5423d42fe11920e9b32094

Observation 12e946a6-75f0-4505-a47d-68740937a7ac · outbound

This paper cites PPI++: Efficient Prediction-Powered Inference.

How Benchmark Prediction from Fewer Data Misses the Mark PPI++: Efficient Prediction-Powered Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.571156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.571156Z digest=sha256:9111851c2ab611f79d411ce05bc98738cf83fde2dc0a68804adc8ef9e1168551

Observation e8998634-c9ec-4efd-a751-da8bce42a030 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

How Benchmark Prediction from Fewer Data Misses the Mark Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.582523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.582523Z digest=sha256:1e2797d486048287bb752759f5a0b19c7459cd7df89a0c6a387c1a98f54c0e85

Observation cb62aaaa-d323-4ab9-b403-cb5c3f267220 · outbound

This paper cites The fifth PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The fifth PASCAL recognizing textual entailment challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.595425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.595425Z digest=sha256:7033e7b81b8970b32b5f6869d5d463005ecddc34e8469c23b36c7427475de93a

Observation b456a683-73b7-4e2a-9d30-d37b78bf7543 · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.612857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.612857Z digest=sha256:968df7a2f262ead4a916a24454e57bb9a7af750c69329374a336d64f4967a2dd

Observation 4be3af6d-5e00-46bc-b5ee-b7b40d24cfd3 · outbound

This paper cites A singular value thresholding algorithm for matrix completion.

How Benchmark Prediction from Fewer Data Misses the Mark A singular value thresholding algorithm for matrix completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.629844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.629844Z digest=sha256:f3fb7629774f4c512a95c5a4d141aff4bdc0db97dd6ad5ec48f16329a92009ca

Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.647558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.647558Z digest=sha256:fe98586f6d620b07165c29843f005e86d221281c2086b0b6cdd3fd1804693cb2

Observation 4900b0b7-bbd8-482c-b3ee-442914c09442 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

How Benchmark Prediction from Fewer Data Misses the Mark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.659435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.659435Z digest=sha256:f9db87777d855728871bf63ce4194a2b9cee744a28fec66287f593c203bdfd71

Observation 4a51dd97-39e8-48af-ba9d-61a528884761 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Benchmark Prediction from Fewer Data Misses the Mark Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.728273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.728273Z digest=sha256:9e5a90770bc3f9b8cb0ca08c082b79d2ab8d70410901a64f88b7af43a18e49a6

Observation aad63229-20d8-44ad-9299-46d04ceaf4c6 · outbound

This paper cites Computing the testing error without a testing set.

How Benchmark Prediction from Fewer Data Misses the Mark Computing the testing error without a testing set

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.823011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.804792Z digest=sha256:93dd40e37ae68b2919886d43451b458d1cbf4f724206d468400d29a5ba18c3b7

Observation b98538e5-9647-4134-87d6-3619b1fe2cbf · outbound

This paper cites The PASCAL recognising textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The PASCAL recognising textual entailment challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.813136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.845227Z digest=sha256:11b8db6a72e0f24122c8ef39671549a0aef2c45a4b2ca1cf7340961d87064130

Observation d22a3f24-6499-4e4b-86a0-4f7c3537c121 · outbound

This paper cites Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021.

How Benchmark Prediction from Fewer Data Misses the Mark Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.803019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.848483Z digest=sha256:c559a7b8e990c95c5dd962b44915f1bb14ef48744e2c004f7d3bf2554ed96e2a

Observation 41110b6f-0e2e-4556-bd50-59291e0a0e13 · outbound

This paper cites Automatically constructing a corpus of sentential paraphrases.

How Benchmark Prediction from Fewer Data Misses the Mark Automatically constructing a corpus of sentential paraphrases

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.791643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.851663Z digest=sha256:b63365ce30db9800a3c4028cdeb92b892c0d6334f400375aa17f59527d8f2503

Observation 57d1e637-a851-45ce-ac17-aa6b56ab3582 · outbound

This paper cites Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data.

How Benchmark Prediction from Fewer Data Misses the Mark Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.779352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.854801Z digest=sha256:fa2767f671e47c32bfeead9dbe22005e0df3e2e633a82460812c9b3bbb8444dc

Observation dce72101-6bdb-446d-9e38-3fdc7196cc92 · outbound

This paper cites Facility location: concepts, models, algo- rithms and case studies.

How Benchmark Prediction from Fewer Data Misses the Mark Facility location: concepts, models, algo- rithms and case studies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.768782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.858336Z digest=sha256:3b454b5fed69610969ace03f4787bf10f56771ee0ca482d5d7fd99c860534cde

Observation 3eebed2e-4b99-46e3-95ad-6207737980b7 · outbound

This paper cites Open llm leaderboard v2.

How Benchmark Prediction from Fewer Data Misses the Mark Open llm leaderboard v2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.758038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.861417Z digest=sha256:43ef6da5edf3f7e76d6c96b6385eab6d0b8288dd29bcb0d8db2f99e563135800

Observation d1bfc293-0b96-49ef-b95b-6e74709768c6 · outbound

This paper cites Challenges in evaluating AI systems, 2023.

How Benchmark Prediction from Fewer Data Misses the Mark Challenges in evaluating AI systems, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.747484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.864676Z digest=sha256:299efcb2281534a1b261d72bf82a8506664969ed6e829978d9535f69e76aabc3

Observation fef75381-5e12-4c9b-8e76-4763ed271a4e · outbound

This paper cites The third PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The third PASCAL recognizing textual entailment challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.737314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.868362Z digest=sha256:c58a8a8a4bb3ff1092aa877b2374bee6229b57390019e77edfea590871474b16

Observation e92bd0f6-4ca0-4717-9731-1ee70e48c1e3 · outbound

This paper cites An introduction to the augmented inverse propensity weighted estimator.

How Benchmark Prediction from Fewer Data Misses the Mark An introduction to the augmented inverse propensity weighted estimator

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.872011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.872011Z digest=sha256:639f950e7a8cfedc37b4c84f5f3406df95504817dfec5f6b0a660c7dc22ac69e

Observation f75e2726-15af-4092-bb23-e4c5e4ba78b6 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

How Benchmark Prediction from Fewer Data Misses the Mark Great Models Think Alike and this Undermines AI Oversight

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.874878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.874878Z digest=sha256:87052d4c0d316f854547269bcfec9e4c542ee66b59fe7784f4b6e6ee197b9418

Observation 011eccdc-8773-4472-a298-376579d64cd4 · outbound

This paper cites A Survey on LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark A Survey on LLM-as-a-Judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.877994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.877994Z digest=sha256:9b1482b3671f0a86026727e027d8d62040ff74345ad68290a1c5311edf6ff6e4

Observation 9a4172e1-1add-4cf7-9d9a-ff19aaaa9c42 · outbound

This paper cites Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N.

How Benchmark Prediction from Fewer Data Misses the Mark Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.719690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.880977Z digest=sha256:0e418b6bc24e4c9fd4c17aaa3593ca29a75e3fb0cc185ef4d09497c365a001e4

Observation e236eb94-6ade-4c5c-8d7c-372701a630ef · outbound

This paper cites Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings.

How Benchmark Prediction from Fewer Data Misses the Mark Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.482516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.884061Z digest=sha256:382472e2ceca8589c6163f43cc5678260329922283dd037628434a5db427cf83

Observation 9d271dc8-53d5-4f79-a88a-b7ed61a6cf7f · outbound

This paper cites Test-Time Training on Nearest Neighbors for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Test-Time Training on Nearest Neighbors for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.887369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.887369Z digest=sha256:1a5c0f005dae21a023196fbdb23f85cb30347f97753293bb64430052125a20e5

Observation 28058209-5f05-40f2-aee6-3bd86c0af9a6 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Massive Multitask Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.890744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.890744Z digest=sha256:0cf1df456b8e4c6c87ec4667b2355a576163425d9bf92c36d1303e9fb66acbf8

Observation e3adc9fc-5581-4631-8821-dd64ac57f168 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Mathematical Problem Solving With the MATH Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.894434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.894434Z digest=sha256:ae7c8c6c9633bef8b7c9dbbcc796c96e1f61ae003d3e65e9669f62dc6fb78cdb

Observation f0802203-fea5-4bf8-8d0a-3f8e4ffbd548 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

How Benchmark Prediction from Fewer Data Misses the Mark Categorical Reparameterization with Gumbel-Softmax

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.897769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.897769Z digest=sha256:d9b54f64f0c9757eaeab29ab6bf53c1cd876163721beab7ef2368add3892dbda

Observation 87a8b5ae-eb96-4c08-9469-491f0d59427c · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

How Benchmark Prediction from Fewer Data Misses the Mark What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.901004Z digest=sha256:095ef405919621ecbf5a93b526af4dd8fb38f83f79798ec8c90e009f96a22f04

Observation 2f0551a2-b906-4595-bfc7-fbf5e87d54f4 · outbound

This paper cites Scaling Laws for Neural Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling Laws for Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.904300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.904300Z digest=sha256:9dca0fda191a27bda89a51fcb9b8a97d0f915b1389d2786ef5a17570531a357c

Observation 31b45bc4-f92a-41a9-91b7-4d3ab770d763 · outbound

This paper cites Active testing: Sample- efficient model evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active testing: Sample- efficient model evaluation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.709413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.907237Z digest=sha256:967fc53351301d9078639be2688f64cb59a7efb7e7640975e8a7a7146216cb75

Observation 0337a2e4-f01d-4d11-bdf5-60f2c848bf25 · outbound

This paper cites Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.405042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.909784Z digest=sha256:a6e90033263ea548139d2002aa632dc15179db840595f45eeb495eff02a88612

Observation 7fc985a9-dc69-49a8-8d29-7754f35819ff · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

How Benchmark Prediction from Fewer Data Misses the Mark Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.912846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.912846Z digest=sha256:0eddbe9548eced44109d79039b6da2cc60e3f606a0562c35c50531c3e1f974c5

Observation 75fea7a5-cf22-4c32-a251-182db853b17e · outbound

This paper cites Active Evaluation Acquisition for Efficient LLM Benchmarking.

How Benchmark Prediction from Fewer Data Misses the Mark Active Evaluation Acquisition for Efficient LLM Benchmarking

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.915668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.915668Z digest=sha256:0e63928a4a7d829e2b3fab1525f5bffabd396ff71c08cc06e33ad3562fa809d7

Observation 423e05f4-9dde-4337-9584-8a968046cef3 · outbound

This paper cites Manning, Christopher R’e, Diana Acosta-Navas, Drew A.

How Benchmark Prediction from Fewer Data Misses the Mark Manning, Christopher R’e, Diana Acosta-Navas, Drew A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.692361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.919274Z digest=sha256:eb7ac823cb6df3a4218a948c4c6aec20faea761e1105bc0a3d32a0636efe4b2b

Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.922228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.922228Z digest=sha256:ad8c536e01ac9f7b950d3f0fd80b77d056567af40b3a8dec2f58e274e14e218a

Observation c16dfc68-05f7-407b-8512-b586c565db80 · outbound

This paper cites Model Similarity Mitigates Test Set Overuse.

How Benchmark Prediction from Fewer Data Misses the Mark Model Similarity Mitigates Test Set Overuse

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.367042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.925314Z digest=sha256:b36070f28276af0c905c8dad22d155ae8fc3cf1c4299bcd47080bd8a2a6c6496

Observation ef43aa4b-2f20-4461-a61c-8771e5acea8b · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

How Benchmark Prediction from Fewer Data Misses the Mark Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.682228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.928473Z digest=sha256:ea7542711a28c297ace9cfbd9f503f4a7cad3e69fccbaca5651e517798515a93

Observation 080a3031-0e61-44d9-a8f4-1d99a863ec85 · outbound

This paper cites How predictable is language model benchmark performance?.

How Benchmark Prediction from Fewer Data Misses the Mark How predictable is language model benchmark performance?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.931350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.931350Z digest=sha256:5c7485f7cac2fcfa9114f86551df09f4bd6b75908c7d6ad4d89dc3e00750eb95

Observation bf06c144-35d8-4081-ae7f-9f0e12da39ea · outbound

This paper cites PredictaBoard: Benchmarking LLM Score Predictability.

How Benchmark Prediction from Fewer Data Misses the Mark PredictaBoard: Benchmarking LLM Score Predictability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.934415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.934415Z digest=sha256:adefb461ebafa121755dd5c2c53429041780762201a4bc540f7f2ef852cb0013

Observation 30ef8940-83f3-4a49-8f21-17c2e82ebbe2 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

How Benchmark Prediction from Fewer Data Misses the Mark LLM Evaluators Recognize and Favor Their Own Generations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.937721Z digest=sha256:2fd060851869558c6432bb5f5ca6884eeac44ff9f14aaf69b9a58708517db452

Observation 7d31cfa5-4838-48d6-b3ba-5e84408948d3 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

How Benchmark Prediction from Fewer Data Misses the Mark tinyBenchmarks: evaluating LLMs with fewer examples

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.940907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.940907Z digest=sha256:3969a91330e212376488b2dd4d897eba4a8fa9714cd5a8685efbb6a65c0b1f07

Observation 98b3f156-3235-4eb2-82b4-51d6389c886e · outbound

This paper cites Efficient multi-prompt evaluation of LLMs.

How Benchmark Prediction from Fewer Data Misses the Mark Efficient multi-prompt evaluation of LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.944282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.944282Z digest=sha256:6e5b502529d61091633a120cb8b08bd4149ce158bc44e8f972147003bb569fa5

Observation c2570af8-3b21-4c0a-af04-5be2ba7ee595 · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

How Benchmark Prediction from Fewer Data Misses the Mark SQuAD: 100,000+ questions for machine comprehension of text

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.671738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.947725Z digest=sha256:f5aefaf1d1acec5c45e3cb5bd1cdbdc726939dcc62df73c21546648b6e6bcfbc

Observation ed27291c-1724-485f-8517-1abc5e842edb · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.950951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.950951Z digest=sha256:3c339f95e1ef3b3f5dfc501cfa69fd86ee82397b533ecf13a6edbf384ca395b8

Observation 2f7c9bc9-ee91-4177-a4b8-dd0edfd42832 · outbound

This paper cites Semiparametric efficiency in multivariate regression models with missing data.

How Benchmark Prediction from Fewer Data Misses the Mark Semiparametric efficiency in multivariate regression models with missing data

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.954437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.954437Z digest=sha256:d0aee10fa5fb82fd6f7b14150e8ada9e20dcfd463c113de537eb904bd60562a0

Observation 2738fefe-3532-4e29-96a8-666d66c99118 · outbound

This paper cites Lalor, Robin Jia, and Jordan L.

How Benchmark Prediction from Fewer Data Misses the Mark Lalor, Robin Jia, and Jordan L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.655964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.957403Z digest=sha256:a482fa23da7346f27a63647a3d5e3c76eb0094255b042fd703086e4f8b7d1f85

Observation b515bf99-dcf7-4650-b11c-3290558b5a87 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

How Benchmark Prediction from Fewer Data Misses the Mark Observational Scaling Laws and the Predictability of Language Model Performance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.960631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.960631Z digest=sha256:1af4d0ce995518b80b89accf4e5f38ea8f2114d083eb612365e681d93e7274ac

Observation b60afbc5-1bd7-4340-ad53-90e5ba059210 · outbound

This paper cites Bernstein, Alexander C.

How Benchmark Prediction from Fewer Data Misses the Mark Bernstein, Alexander C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.646469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.963910Z digest=sha256:99b159964b0845126a5ee393c9db92aa84d26a64b1dd4705f4d38aa0f4d9a0fc

Observation 6da416b0-5c7d-4d36-ae7e-bb0990390be9 · outbound

This paper cites Data distillation: A survey.

How Benchmark Prediction from Fewer Data Misses the Mark Data distillation: A survey

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.636832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.966967Z digest=sha256:4ab48ac2d000fbd3046e27a1b0531256ab7f441f280d63f1258e6b819c4c0b16

Observation 1e72658b-55a9-4ae9-a9f7-4698ca13d93c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.970380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.970380Z digest=sha256:4972bd224dcb905eb318014f9df56dc2e56ca1496fdf01c62d4ec15b04ab2877

Observation 97e17da9-1936-4277-9f66-41108ef856ea · outbound

This paper cites Recursive deep models for semantic compositionality over a sentiment treebank.

How Benchmark Prediction from Fewer Data Misses the Mark Recursive deep models for semantic compositionality over a sentiment treebank

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.627625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.973877Z digest=sha256:cf20acf244f769c8d39f1cad5c2bc8d13d6a2c2953f9efe7d14601b1e2bd86ce

Observation 0ffbcf20-b91c-480a-9863-65da0b63a71c · outbound

This paper cites MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning.

How Benchmark Prediction from Fewer Data Misses the Mark MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.976895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.976895Z digest=sha256:c5fbd2e0fa8d58dc5bfc2d6b056b159dbd8673e297c0acd993266df7254ccb0f

Observation fad17933-bd87-4d4a-910f-86803c968dd7 · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

How Benchmark Prediction from Fewer Data Misses the Mark Test-time training with self-supervision for generalization under distribution shifts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.980129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.980129Z digest=sha256:e08e66a6f26e5ab38058908055ff66d2093d9f006f71e918527ad9a7f5d799ec

Observation a1446e7b-b1a7-4633-a1a4-c2523c4af7a0 · outbound

This paper cites Le, Ed H.

How Benchmark Prediction from Fewer Data Misses the Mark Le, Ed H

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.610789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.983164Z digest=sha256:90845593f36214d6c85785af4aa73263b1ce88f76c141e3e1e34202d6fea36b5

Observation 6a3b525b-d889-45a6-8b7c-05356aa14d5e · outbound

This paper cites Four lectures on probabilistic methods for data science.

How Benchmark Prediction from Fewer Data Misses the Mark Four lectures on probabilistic methods for data science

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.141439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.986668Z digest=sha256:a1d84ae5bf2ba660b05da436aa2e02b3622a78ee9653f0f5e0d70cc0704f1299

Observation 5c68714b-37a1-4384-b026-c3e71eb8d677 · outbound

This paper cites Anchor Points: Benchmarking Models with Much Fewer Examples.

How Benchmark Prediction from Fewer Data Misses the Mark Anchor Points: Benchmarking Models with Much Fewer Examples

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.127869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.989961Z digest=sha256:4cd79c7ec80ebb32ec12ea5aba0ec8da8205d62a1e95cfe79c5d9c2302e5b2ae

Observation ead71a23-d406-4f36-b537-835d2941fa44 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.601129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:26.993144Z digest=sha256:db0f3efef0ad1cd484d982985070241a1ffe918025120ba80f6ef080f0c881af

Observation aa39132f-b54c-4a1f-9498-b0cc5d5562c2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.996332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.996332Z digest=sha256:cb8cf3fde054c71b8a01260248a15eefb427a82b782281d46622a4af10c455b2

Observation 1b64fdac-d565-4747-8e86-db0fdbf64704 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark Self-Preference Bias in LLM-as-a-Judge

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.999567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.999567Z digest=sha256:ae75d68ae6d19763c647f88d86029357278092559d8503a74d9eda6cb8552db8

Observation a11929c3-9425-4a9c-bf3c-74488c36aad4 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.591844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:27.002653Z digest=sha256:a8b19ea1a7b312fe5cb65597b465e6c09aa86d455cb884eec2756bc63caa1221

Observation 18b48d55-7d95-4d69-8dad-1b3795ef9cab · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

How Benchmark Prediction from Fewer Data Misses the Mark Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.005749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.005749Z digest=sha256:6a4f6243849413655e2ac1e837097f5474f5ee427927628bbc994bd58edb9da4

Observation dff6138b-4a97-4362-b2f6-3655a56c0bba · outbound

This paper cites Automatic evaluation of attribution by large language models.

How Benchmark Prediction from Fewer Data Misses the Mark Automatic evaluation of attribution by large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.582423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:27.008828Z digest=sha256:1d7ebc908f9d3483445ec23a6b74d5d8fe5d6478e5e3531aab7b52fc10822acd

Observation 9cbd5450-d806-4bfd-8b28-5ab478823d84 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Instruction-Following Evaluation for Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.011892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.011892Z digest=sha256:6e6a2e6f354a54ad75692e39f6fbdb90805e52dd895ba1ad3c1a68514fc64af0

Observation e8a454e4-155c-44a6-93ef-63ffc29a4e1c · outbound

This paper cites On Speeding Up Language Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark On Speeding Up Language Model Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.015178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.015178Z digest=sha256:5958bebd3bf77950103d1b6ab4d1dc4281c8575ae8918a086d2630198693490b

Observation 8e104514-0e04-4d17-a20c-fc8609e6259e · outbound

This paper cites Probabilistic Bilevel Coreset Selection.

How Benchmark Prediction from Fewer Data Misses the Mark Probabilistic Bilevel Coreset Selection

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.059462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:27.018282Z digest=sha256:8ae3b6c137e57bcb3b91d0b4d4cede5725910086d2505d746e123b9eb32e2255

Observation 1e3e299b-9295-4d96-8c3f-399e60260ee5 · outbound

This paper cites How to select datapoints for efficient human evaluation of nlg models?, 2025.

How Benchmark Prediction from Fewer Data Misses the Mark How to select datapoints for efficient human evaluation of nlg models?, 2025

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:34:27.572476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:34:27.021795Z digest=sha256:7d0ac31c06ff0cbaec0d3a9980ace18340bfb96f5eae26424f7babd2aad472ba

Pith citing papers

Observation b27a9196-c3e0-4eb4-9adc-6acd9af2d40c · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees How Benchmark Prediction from Fewer Data Misses the Mark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:350a0388103d774a8779c92d1628cd627dc65e2380ad3699b13df6d7be7180bc

Observation 6992b4b1-fe2e-435b-85da-cacb5393970f · inbound

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking cites this paper.

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking How Benchmark Prediction from Fewer Data Misses the Mark

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T05:26:41.529866Z digest=sha256:16b51be167aadb7857ecb3aa0ef5b830a1af1e9b21d7ba367ca52997157543cd

Observation 470626fd-38bc-4101-a212-c01eccc4759e · inbound

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences cites this paper.

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences How Benchmark Prediction from Fewer Data Misses the Mark

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T11:24:01.547119Z digest=sha256:66dbf4cfd159e640e5a000fd6cffa533e690203f46248f34d4ad9771e38dbb4d

Observation e131494d-8e10-4c20-bf88-bbc35f486f3d · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research How Benchmark Prediction from Fewer Data Misses the Mark

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:617cfa1502c534498072ca32fcd941196e506e029735975c8d65ab5fecda69d3

Observation cf978fc5-69ca-48b4-a646-c55b8a9574af · inbound

You Don't Need to Run Every Eval cites this paper.

You Don't Need to Run Every Eval How Benchmark Prediction from Fewer Data Misses the Mark

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T08:23:12.145516Z digest=sha256:4fc3d770f8c5550b03d923ed51a93e2229b741008ebdf53ba74cfa0e4a3a6351

Observation 497f56ee-452b-4e85-b947-3c7cecea0230 · inbound

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification cites this paper.

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification How Benchmark Prediction from Fewer Data Misses the Mark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T06:40:23.654993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:40:23.654993Z digest=sha256:eddf815a0a036620e92e16cd14edbbbcd9053ceb23a5b4604b2893c9b7f52082