Pith. sign in

Paper Citation Record · LEDGER

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

As of 11 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2507.17747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17747 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:44:09.471855Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact5
  • verified fuzzy30
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a83d6f7c-cc0d-4c07-9845-7e11451947d1 · outbound

This paper cites Claude 3.5 sonnet.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.909879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.233992Z digest=sha256:48d8bada4b33c62a05d37067a389564dfa5eba79e3e4fe43a06f48bb15cb30bb

Observation 0c351428-9b2f-4c72-bcd6-0b4c6cf93302 · outbound

This paper cites Claude 3.5 haiku.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 haiku

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.901402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.237996Z digest=sha256:b8e34f104c00e6579d84f6166dc3409c0e576b7ed1157d449770faa3fcc74d16

Observation 00efc9cd-aaad-4095-8e28-dba4fd18b6a8 · outbound

This paper cites ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.893059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.241713Z digest=sha256:5ee5b1271ca4b245838b1d0791aae8b3b5a5e9e41c6d8fb19dc96a793c3d3f58

Observation 168020df-faa3-4617-bcb0-3f661ff0428e · outbound

This paper cites Benchmarking Foundation Models with Language-Model-as-an-Examiner.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.245058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.245058Z digest=sha256:e2e2454488702eb71bac2d1429d859ad7b2487938a1846493e5837b1e0527e51

Observation d2899a6c-dd33-44ef-9b7f-4fa173d23ace · outbound

This paper cites Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.884260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.248532Z digest=sha256:1b1f7b2a621aae4ca7af93de19fd2e6a34c3d3f4fce537c77585e4b2eb614949

Observation 8e75f131-7bc5-43fa-b6b7-b20bf3ce0b61 · outbound

This paper cites Adversarial multi-agent evaluation of large language models through iterative debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Adversarial multi-agent evaluation of large language models through iterative debates

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.252377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.252377Z digest=sha256:16921ead90b935e278964c910e6dbcf1b1d22cb2ab49b844775fdd55124b8070

Observation 325cfa56-0309-4fc2-bdd2-eb2b2dc9dbb9 · outbound

This paper cites Flageval.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Flageval

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.875447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.255787Z digest=sha256:12a93fde4106cc48a1fb138b5a82fee867ea238228040336bc5e02a827ece346

Observation 28e252be-bc05-477b-9b7c-6b1eb1892a54 · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.258694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.258694Z digest=sha256:547608f27e832641e1be14dc1fe09f47b06f8b42eeed916f39cea9b0208ebf55

Observation 91128ddb-5cc4-4299-be72-ffd00b990fdd · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.261955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.261955Z digest=sha256:db68ab5210dc29f6ff7931e662020ec647650c16fd0d278566c1b4d9a4d57975

Observation f00b4b03-06ee-4086-9adc-237c2aaeaa5e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.264731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.264731Z digest=sha256:97a6fd21096aad426a0be1235f79bb9512ca529f964dac363a799ee9bf1af4a9

Observation d8b5ddd1-7732-46aa-a3de-03b2b08f4d5c · outbound

This paper cites The Role of Deductive and Inductive Reasoning in Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Role of Deductive and Inductive Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.268106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.268106Z digest=sha256:67cb4d71f4b43a233551ae5e5270059d94ab27e7e6461c78d5f2ab443a6a6524

Observation 9dca74ff-0d91-4e32-ae9c-e9c730ce254e · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Are we on the right way for evaluating large vision-language models? In A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.866662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.271284Z digest=sha256:c4ae696bda2e30493cae70e802c4e7c364039431b65e4400221da29beeaf0395

Observation 2000b724-3372-452c-975c-a70ce8d73640 · outbound

This paper cites Jordan, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Jordan, Joseph E

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.858241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.274046Z digest=sha256:cea4e976603ff828cb7ff91e0609876c44ec5fb162182c12a3b030023f953a06

Observation 8877c7e4-7337-4901-9ce7-33441fd4fad0 · outbound

This paper cites ARC Prize 2024: Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC Prize 2024: Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.277264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.277264Z digest=sha256:b3445128775a0d9794fb26b725cb0adf29fc97c30ecc8a13f9aea117d35ee915

Observation 8e8b6024-00f3-460a-a420-2b2ccca90a02 · outbound

This paper cites Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.849645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.280482Z digest=sha256:603010a2b25da7e1315cb94152cf9bf023472f75d7948200be96b856df6a4a92

Observation 85429f59-fd84-4ab2-a982-cdf409324331 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.283458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.283458Z digest=sha256:af48761ed3b8cdbdfafe9e8e1d71261d8ff5a2b727e5d0b93a5e0f1304a2c186

Observation 48812bb5-00e1-424c-b0e9-4dafced156f7 · outbound

This paper cites DeepSeek-V3 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.286882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.286882Z digest=sha256:9c326615036617387a41f25012052cf34d9579a2e166215ca7415a670e965529

Observation a3ba679b-47be-4e9a-b1ff-285680b552b7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.289691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.289691Z digest=sha256:53508fd5312c26a8eee99efed2c54253e42c00dfac56a88a8191c14be6597ff7

Observation ed6de936-dfbd-410d-b224-bb625d1ae89a · outbound

This paper cites Investigating data contamination in modern benchmarks for large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Investigating data contamination in modern benchmarks for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.292859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.292859Z digest=sha256:ca0539f37f0552daf88a9c8cd1d07f9ca2284d7805afd9654565af36eb1e5272

Observation a740c12f-d3a0-49c5-b8d7-ef683acdcc95 · outbound

This paper cites Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.840252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.295646Z digest=sha256:ddc22e22be731d4273a1840050df0a87ab9cd1caab7c4b168dc078fd1509bd9f

Observation 8f4b08d6-f95c-48ea-9b21-134bb8dd3aba · outbound

This paper cites Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.830837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.298363Z digest=sha256:eb25439b7d06813477e42ac004bdcdc44cdfe605cab3b87a9149d839628c6d01

Observation 6bb92465-f91d-41de-9fff-5c67a3152038 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.301684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.301684Z digest=sha256:667f6000275339f2c24103dd9a8fef963058653a2b1dd2c44b5b8d26a800baaa

Observation a810d3f1-b2a7-4cbb-9524-d02ddea554e3 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 23

Resolution
verified exact
raw_fallback, observed 2026-08-06T14:44:09.824746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.304648Z digest=sha256:d2cca82b75649ea9cb2277391203e943448932eace97084064f2c0ccf32bcc3c

Observation 834de928-4ab6-4b41-a962-c2fad6bd8d13 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.307985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.307985Z digest=sha256:3c62870246d8083a7c3be28b5777d05690dc8992c9168170bf54324e88dc00ed

Observation 0a8ff0aa-536d-4f47-a10f-ed347fddadcb · outbound

This paper cites Time travel in llms: Tracing data contamination in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Time travel in llms: Tracing data contamination in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.821333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.311518Z digest=sha256:cd5403e6df52671d94ac4ab989bfa7377589eacd97d4612b073c700bd49b0bcc

Observation 29f26e3a-ff53-43ea-8d89-28c698a42493 · outbound

This paper cites The Llama 3 Herd of Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.314332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.314332Z digest=sha256:80f131f5c3d9cc97f115595e20666672a4793af9b74475a2571d9d3159596a25

Observation 3f84371c-551c-4272-9c46-4b5923fea23e · outbound

This paper cites A Survey on LLM-as-a-Judge.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Survey on LLM-as-a-Judge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.317974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.317974Z digest=sha256:2db68efddcd3e6cb045e2416e76bef76791e7799b5d4bc1078f39b0bd23dfbb1

Observation 52a345a1-6f49-40e1-810e-a7640175aa05 · outbound

This paper cites Improving Model Evaluation using SMART Filtering of Benchmark Datasets.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Model Evaluation using SMART Filtering of Benchmark Datasets

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.727362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.321034Z digest=sha256:4cff725b305aa02d1d74f745150722330e1378b813386563a3e789fe47210073

Observation e85e20e6-37bf-4433-a5fb-1c42a9f3b091 · outbound

This paper cites Measuring massive multitask language understanding.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Measuring massive multitask language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.812461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.323991Z digest=sha256:645fc5c529ecb779923c115deef39da1517cd694b67198674cffad7b3e2ffa16

Observation 1af5e0d3-7f8a-4b50-b586-318c67eb5b8b · outbound

This paper cites Trueskill : A bayesian skill rating system.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Trueskill : A bayesian skill rating system

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.803349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.327209Z digest=sha256:41082b86da8947574ea2c084e30669697c2a35a79e5c7c9c13ddc0a7ad345c02

Observation 5126eeb1-cf26-4ab6-a34d-14f7327f25db · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Lo RA : Low-rank adaptation of large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.330561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.330561Z digest=sha256:0c1593d77e17a389de2e5ac0fb83390bb3b495cf813a6bcd3b0e9583c31f89c7

Observation 47a8eb31-62ee-40b4-88d8-7030916e14a2 · outbound

This paper cites AI safety via debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks AI safety via debate

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.333252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.333252Z digest=sha256:cf0a1163558770f72d88944060d99175337d6cd008ea5e85b142e01b2d87169b

Observation c9630232-1386-4a80-9503-c007f5e98ef4 · outbound

This paper cites Mistral 7B.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.336865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.336865Z digest=sha256:e767fd1e6e66392b2ba7af7027bc7e8ab0b42e6297ae849cf326ed823aca76a5

Observation b86f7a43-fbad-48c8-ab59-5f59d018b4c0 · outbound

This paper cites Mixtral of Experts.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.340177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.340177Z digest=sha256:bee176d1b4234984115b3eb86a2dc42e5e532da27d943389bda42d713ba79a54

Observation 87e694a9-d61b-4e1a-be24-bfeeeeac2a80 · outbound

This paper cites Bowman, Tim Rockt \"a schel, and Ethan Perez.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Bowman, Tim Rockt \"a schel, and Ethan Perez

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.787423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.343125Z digest=sha256:6d86e52565490c264bfa5b4ad361fa804dd2d286975f69ec0b4c2097856a9d11

Observation 325bd2a9-eb0f-4a67-a6c4-a13267cefb10 · outbound

This paper cites Debate Helps Weak-to-Strong Generalization.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Debate Helps Weak-to-Strong Generalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.346312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.346312Z digest=sha256:70d7757c57b5852e75cec5a8018d4a73f02b0e25f34e2083fd28903fb8767511

Observation dcdbfae7-ea09-477a-93c7-932389ee3b53 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.349366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.349366Z digest=sha256:813e50e4b4522f36a3db9576ea63a7b8dcb3a840399e1838fa791bbb28a8ba36

Observation 88498f70-5335-422b-be70-d39ac2d85e91 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CMMLU: Measuring massive multitask language understanding in Chinese

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.352447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.352447Z digest=sha256:6f723fcd7b2181b9e3a0f7fa15fea100349573cb205374da2ca70b0022761924

Observation 87452799-2260-4345-bd6e-15bd193c9117 · outbound

This paper cites A Debate-Driven Experiment on LLM Hallucinations and Accuracy.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Debate-Driven Experiment on LLM Hallucinations and Accuracy

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.659401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.355365Z digest=sha256:afec1c52a510124d7e554b0be085c8629051c02e07997a8cf9804516b9e8ded1

Observation ed2720b2-bae9-47a1-b47c-1b33256a0864 · outbound

This paper cites Manning, Christopher R \' e , Diana Acosta - Navas, Drew A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Manning, Christopher R \' e , Diana Acosta - Navas, Drew A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.777855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.358457Z digest=sha256:5fdee097847e5f01573f63f2c4d576237bc77e211939a83585e61837827dfbf6

Observation b16022a5-2bc7-40ea-b7db-0dbdc8a78cd6 · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Encouraging divergent thinking in large language models through multi-agent debate

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.754017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.361156Z digest=sha256:f86b1ede117348a02bfcfe807325e87b8e049bcb014d7dff465272335ff24e64

Observation bcd0f28c-7357-451f-a019-43dd1ffd8c48 · outbound

This paper cites An empirical analysis on large language models in debate evaluation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks An empirical analysis on large language models in debate evaluation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.699063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.364110Z digest=sha256:5e70c5c30ecca1f5f7b6b1d51afaaadf2dc2179d1f468101c97ac41b06fceaa8

Observation c70ea09d-581d-4baf-a8ff-a4403bcf8c35 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.638130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.367183Z digest=sha256:ade9dfeef4acb75ea4aeaa8e32e5c3fcb58e264029bf355e53b360936cf82137

Observation 658269b3-bea3-41b0-9ee6-c5226b858a26 · outbound

This paper cites Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.550141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.370593Z digest=sha256:bc45b68cbe4f0cebf5d6818c250c0a27bc6ae0f82b1e25cce34cd6821165ae6a

Observation 963a0a2e-6217-4657-988a-4b2724e01846 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.373495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.373495Z digest=sha256:ddcb2bd00cfffb1bb000114b6c94c76b02033f77b7f33f277f9696ca6786675b

Observation 471f0037-bffb-4f85-927f-10aaadc82ad1 · outbound

This paper cites Cheaper, better, faster, stronger.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Cheaper, better, faster, stronger

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.454555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.376438Z digest=sha256:804813c71442697695969110683112025705e28c06f66e270174177bd9b34df4

Observation 48907f32-9cc1-47f7-a2bb-4fc3af06a2d2 · outbound

This paper cites Mistral large.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral large

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.325353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.379133Z digest=sha256:84b55c412ebecaa9a5b2a5f8ea9a2e4bcd82e88cd581c8186ec641771708a914

Observation e966dac2-cd50-4d15-baaa-94f3454d0b18 · outbound

This paper cites Evaluating the Performance of Large Language Models via Debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Evaluating the Performance of Large Language Models via Debates

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.381842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.381842Z digest=sha256:664df003695bd4b434ccdf687f6b502c1d756fbbfd05e681f0268cb6f58c08dd

Observation 0933da57-b03d-4223-9983-e35edb8fc8f0 · outbound

This paper cites GPT-4 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.384917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.384917Z digest=sha256:9ad03709384d9bd4943ce6f3a0a5a4b8bf8f027a965ca2facee77057d58718ad

Observation ec0c7b9f-1c02-404d-b2f5-9f6bb2b5342d · outbound

This paper cites GPT-4o System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o System Card

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.387932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.387932Z digest=sha256:242acd73052aeaef243f816d6e5d7f35c9d146c60c87420bfbcaab4f92955ae2

Observation f817549b-8f37-4ac2-800d-c183768200b8 · outbound

This paper cites GPT-4o mini: advancing cost-efficient intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o mini: advancing cost-efficient intelligence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.063257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.391394Z digest=sha256:e70bb292a26bc2571f0b71d46cef25e6cedbda4f75d5c17fd5bd773f839f4950

Observation 8c4d6f4f-37f0-4b11-a6fc-4272bd79a562 · outbound

This paper cites OpenAI o1 System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks OpenAI o1 System Card

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.394084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.394084Z digest=sha256:1b469ef75fdcf558da18c3681b27365fcbd8417925fddf59e741a4b81c30826b

Observation 81b8e3f6-eb82-4a21-94f3-71c9f31dfe3d · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.396998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.396998Z digest=sha256:5c3b62b5586e692421518f75bdc61811b22bbb33663fb0b06548d01a4b2bb868

Observation 067938fe-4300-44b8-b2e4-4beba9569ad5 · outbound

This paper cites Mapping global dynamics of benchmark creation and saturation in artificial intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mapping global dynamics of benchmark creation and saturation in artificial intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.399824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.399824Z digest=sha256:c99058699651091b9debce60d9ea70ffa2089c3fa21bede069301f88f478ff2a

Observation 7053c87d-d62c-44aa-b800-fd743facd11e · outbound

This paper cites Humanity's Last Exam.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Humanity's Last Exam

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.402625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.402625Z digest=sha256:258d999aa382dde8b8e5be36bbbde70dab8b072f569e046402aba6e6d15f8df7

Observation f12d1973-fa38-4f70-9902-f1b6e486e20c · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Introducing gemini 2.0: our new ai model for the agentic era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.854039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.405513Z digest=sha256:2ac3dd28285d8ca461e51c4852c31ade90d115c4ca39f4852c3ea7cf3780500e

Observation 36bb7fd3-ed87-4e09-b768-ccd3d1c8dc91 · outbound

This paper cites Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.471219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.408457Z digest=sha256:2ced19e921749b4ad684858258877666ef0231db202be7beaf3f629f22108cb8

Observation 07618cbf-e2a8-490c-bec3-fbc58bbe8bf3 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.411020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.411020Z digest=sha256:48398f828f7ddcb35c39b987f646196fa6c53eff474ed501b26266272c25b7b5

Observation 8bb4c05e-68cc-4951-855a-2f612619ae98 · outbound

This paper cites Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.413880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.413880Z digest=sha256:d69eff57638b6b0c04d7863565c2fdac8bed0bf78682be674955b4013dc42dbd

Observation 3cf9e6df-393b-4b02-adde-6800a7912adf · outbound

This paper cites Pretraining on the Test Set Is All You Need.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Pretraining on the Test Set Is All You Need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.416573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.416573Z digest=sha256:d3b3b8e9c1bc2d975f65b7731f21cbd54a9a3ebdcf0a80d2643d784c9a634c01

Observation d5675478-26e6-415f-b752-bc3c18d2885d · outbound

This paper cites Detecting pretraining data from large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Detecting pretraining data from large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.319102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.419654Z digest=sha256:41e7702367893efac1b205f1b95850a2e3dc23f496426b6f07033b3eb859a360

Observation db9419df-7567-4932-9c0b-f3dd29fabb58 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.271308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.422236Z digest=sha256:1d4c88d79c94fbd5dfcbb692b08b7a1d3d08412bc37c619bde2cafdf09b2a746

Observation 9717d2a1-315a-4ac1-b51d-f7940e0fb2e5 · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.424893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.424893Z digest=sha256:7f1c6e9a4aa6749302a8f13b15cbbf3522def64f8d6c38f9c9aa6cc8985c0849

Observation 1647dcfc-81fe-4705-a412-fd81c0ab1680 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.208592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.427865Z digest=sha256:55a1b554b63b65881a98a7f335b705f6abf3b1f2984881bab4b20f05795c14e6

Observation e7e1d3ac-fd53-430f-8ac1-ee4cd9cb21f5 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.148317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.430445Z digest=sha256:cd60b397b1bb5a672f2674bd0d8f2ae93f3d67ef64cba255567eb7887eea8508

Observation 29e7c9eb-c1a0-4b40-a774-be1a12be7040 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.433045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.433045Z digest=sha256:52c89b0f2df7575e604222984c60bb335acd6a96a7a9d641298ae578887d42bb

Observation afab2f8b-6539-46ef-83d0-73bada5393a3 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.124690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.436267Z digest=sha256:5caacb37f35ee58abccc15f545971396e090248693819bb54d4b0438d37d8f02

Observation ff3d821e-b5f0-47d2-aae8-11173c64161c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chain-of-thought prompting elicits reasoning in large language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.115093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.439020Z digest=sha256:688cb962365aa5aa3a5bee4659594932168b274b3ad9d118bd543ce40e1d7c6e

Observation 6c16eb36-7f0b-4507-a872-929d4fc131ee · outbound

This paper cites Livebench: A challenging, contamination-free LLM benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Livebench: A challenging, contamination-free LLM benchmark

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.105651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.441678Z digest=sha256:34f7c6ecb60f90d0a8fd47678ca68e346f77597ef794efd68711694718c33b44

Observation ff34c8d0-bf09-4d70-9d71-6d5258bfea53 · outbound

This paper cites QUD eval: The evaluation of questions under discussion discourse parsing.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks QUD eval: The evaluation of questions under discussion discourse parsing

Reference 70

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.505113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.444655Z digest=sha256:b9e019675d98c0397a7c03067bfa69b32dcb2b77e2926d72f635b1d50831619e

Observation f51edf2e-e3ea-4b73-b8f5-05d222f22c20 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmark Data Contamination of Large Language Models: A Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.448188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.448188Z digest=sha256:9b92221628d937e88f6ec49a3e003a011fb231bf144f90a1cff0dda20d239e25

Observation faaed509-9aa0-41be-8f71-1d8c24b38fac · outbound

This paper cites KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.451179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.451179Z digest=sha256:d87750f378d66968f20acf5b9e4826c581fe55e6c2648e5acd39d7a9e4ea726c

Observation d1754a25-0780-4bf6-8e85-c1d0fa4eb958 · outbound

This paper cites Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.454188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.454188Z digest=sha256:f087445cf1b6dcd528c4266c67d0365a9d8580d8ad38de7746c3c3c6e1fb1abe

Observation 28e1a61d-5759-49a8-a35d-c52eb1b52a35 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Xing, Hao Zhang, Joseph E

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.096336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.457014Z digest=sha256:509b9271eb27f072f7416efbb448a7a0366495a0c48974014222277bf0b4b22b

Observation aa7be3b1-c45d-452b-a424-1fc0ded6b452 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.086077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.459891Z digest=sha256:08dd9d27024200b180a9dcec770a4dc69e300ba14235b3e781a16fe401544696

Observation 7d0e033a-3e88-47c9-bea6-db99e9f6c28f · outbound

This paper cites write newline.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks write newline

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.462617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.462617Z digest=sha256:167ae9e750eb381eb593e0a46b38e0ce9159c7dc1d673ba9fa1fc7fdde620d6a

Observation 6edf6c22-fd91-48f8-8d29-35f64be39ffd · outbound

This paper cites @esa (Ref.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks @esa (Ref

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.465929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.465929Z digest=sha256:e3ae775620c8edc8af1e4e18b21a961ed50d01726ddd3eb14449df2f6e588a2a

Observation 4339ae32-1247-4087-955b-68eaa972ae45 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.468994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.468994Z digest=sha256:5e81f68499e70cf494a79418f4494a8762157440c02ad2d04f5892b81e1d343d

Observation 5fde41a7-2a87-47c6-9831-2c3a0ca7d3a2 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.471855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.471855Z digest=sha256:e9ec9405d7865024de7a0a492e126d3bf701429433324269b04d70a459f493dd

Pith citing papers

No inbound Pith citation observations are available.