Pith. sign in

Paper Citation Record · LEDGER

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

As of 7 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2507.17747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17747 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:44:09.471855Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact5
  • verified fuzzy30
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a83d6f7c-cc0d-4c07-9845-7e11451947d1 · outbound

This paper cites Claude 3.5 sonnet.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.909879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.233992Z digest=sha256:c48dc02933c4a1261e6463a662d91c44f3b905ed2737f27b672baf3d09cb1243

Observation 0c351428-9b2f-4c72-bcd6-0b4c6cf93302 · outbound

This paper cites Claude 3.5 haiku.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 haiku

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.901402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.237996Z digest=sha256:4cbb35737f2153ea8f0fa38f2ff8f720deda6065fc8d39bc1898ac0b1c2ea8e9

Observation 00efc9cd-aaad-4095-8e28-dba4fd18b6a8 · outbound

This paper cites ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.893059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.241713Z digest=sha256:40782715bde0ec146a5f9b879803716d7c9569d29eefa26c05546b1484c11a56

Observation 168020df-faa3-4617-bcb0-3f661ff0428e · outbound

This paper cites Benchmarking Foundation Models with Language-Model-as-an-Examiner.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.245058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.245058Z digest=sha256:f92c8acbee74c2fe926b9fccbad573a2707d2a8cba94d134195a89530fd66637

Observation d2899a6c-dd33-44ef-9b7f-4fa173d23ace · outbound

This paper cites Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.884260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.248532Z digest=sha256:e1cf3cd16291d027374e9fb069d7b3762e4c8bc6ee7cffccfc1b0d7f5d3f351c

Observation 8e75f131-7bc5-43fa-b6b7-b20bf3ce0b61 · outbound

This paper cites Adversarial multi-agent evaluation of large language models through iterative debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Adversarial multi-agent evaluation of large language models through iterative debates

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.252377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.252377Z digest=sha256:6be0c2a24186e98cff8af327f05b57efd1e438f80c465a0b06076d72746a8925

Observation 325cfa56-0309-4fc2-bdd2-eb2b2dc9dbb9 · outbound

This paper cites Flageval.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Flageval

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.875447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.255787Z digest=sha256:c6c2ca554c9cc06206bae72d11dace6fe38849e3d0934234e89c2f410edcc471

Observation 28e252be-bc05-477b-9b7c-6b1eb1892a54 · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.258694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.258694Z digest=sha256:226cb92c75d490275e9ca9aff8e1f868b1d52fdaf0efb296bc15e2c9e82c6a29

Observation 91128ddb-5cc4-4299-be72-ffd00b990fdd · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.261955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.261955Z digest=sha256:602872e2552bef8d66455aab6dab893380ec69b721c912f4c5b5a0bad5298b35

Observation f00b4b03-06ee-4086-9adc-237c2aaeaa5e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.264731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.264731Z digest=sha256:40258123152564922cf2ddab460bb5eece5e3446e9bbee899d6d81b36c48bc24

Observation d8b5ddd1-7732-46aa-a3de-03b2b08f4d5c · outbound

This paper cites The Role of Deductive and Inductive Reasoning in Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Role of Deductive and Inductive Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.268106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.268106Z digest=sha256:b8f0c3309f11eb9fd0f9ba2e3979b4db9c35f43efd253fb02af5f46c65f11371

Observation 9dca74ff-0d91-4e32-ae9c-e9c730ce254e · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Are we on the right way for evaluating large vision-language models? In A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.866662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.271284Z digest=sha256:9a801dbbc54988ec79a0317525e1195506d6166880a5aea3277cff9d14e11a46

Observation 2000b724-3372-452c-975c-a70ce8d73640 · outbound

This paper cites Jordan, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Jordan, Joseph E

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.858241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.274046Z digest=sha256:86cc9de05b832ff37dbae2fc098bdda404226f274031bc8aaf5517b67e008428

Observation 8877c7e4-7337-4901-9ce7-33441fd4fad0 · outbound

This paper cites ARC Prize 2024: Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC Prize 2024: Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.277264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.277264Z digest=sha256:666c359d6cc4c45c2c64b59f33f69eed670b052dfb720e8db92069ba51e4c271

Observation 8e8b6024-00f3-460a-a420-2b2ccca90a02 · outbound

This paper cites Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.849645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.280482Z digest=sha256:50c48dbccea2c9760f9603ddf65828b6df5aa4353cd810815f477b86d0a06fdd

Observation 85429f59-fd84-4ab2-a982-cdf409324331 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.283458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.283458Z digest=sha256:129153a7bd520f46deda3644987caff9c0795ea3e3c340acd6c345371c0edcb7

Observation 48812bb5-00e1-424c-b0e9-4dafced156f7 · outbound

This paper cites DeepSeek-V3 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.286882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.286882Z digest=sha256:ea78a7c9b53435a816f790625fa6ab8f4f0cbabe909243f63164edc487605dbf

Observation a3ba679b-47be-4e9a-b1ff-285680b552b7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.289691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.289691Z digest=sha256:58c26d85ba3189d14c4024fee686f8a236598b78e555c9241861b6b5980ea808

Observation ed6de936-dfbd-410d-b224-bb625d1ae89a · outbound

This paper cites Investigating data contamination in modern benchmarks for large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Investigating data contamination in modern benchmarks for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.292859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.292859Z digest=sha256:f1a90db1e6dc9052e96aff010012de0478456d515348b21ce7da81671354ed0d

Observation a740c12f-d3a0-49c5-b8d7-ef683acdcc95 · outbound

This paper cites Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.840252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.295646Z digest=sha256:67f1aafab91d92f47776983aff2a4a36768ceecadf56a3e4d7c3bc3bf0b74f8b

Observation 8f4b08d6-f95c-48ea-9b21-134bb8dd3aba · outbound

This paper cites Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.830837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.298363Z digest=sha256:eb85e55804a94877a946ffea8e46fbc9ff7fbf0404ab4865fecdebbc443d95c5

Observation 6bb92465-f91d-41de-9fff-5c67a3152038 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.301684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.301684Z digest=sha256:65733ca45f00bb32caf6c074db9d5c07f0129c8cfa52ed1974322ebbbfce68c1

Observation a810d3f1-b2a7-4cbb-9524-d02ddea554e3 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 23

Resolution
verified exact
raw_fallback, observed 2026-08-06T14:44:09.824746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.304648Z digest=sha256:c1d37e7f030237f1779ac1b6c29dd72b69a7af3fbb9bfc9f21279beb62a1893b

Observation 834de928-4ab6-4b41-a962-c2fad6bd8d13 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.307985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.307985Z digest=sha256:60694dfd4b10977a5a3412cd1a5b0bd2999c2f349db2410903f62828a1bbfebe

Observation 0a8ff0aa-536d-4f47-a10f-ed347fddadcb · outbound

This paper cites Time travel in llms: Tracing data contamination in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Time travel in llms: Tracing data contamination in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.821333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.311518Z digest=sha256:8b0aa0ca2d3e8d75a5787000c1e637d9a322ce8de4a9df0e6316c223e372bec8

Observation 29f26e3a-ff53-43ea-8d89-28c698a42493 · outbound

This paper cites The Llama 3 Herd of Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.314332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.314332Z digest=sha256:51415b57150fae2d621e3d3be9761647673d06d11bcf037961ae8d634d7b81fd

Observation 3f84371c-551c-4272-9c46-4b5923fea23e · outbound

This paper cites A Survey on LLM-as-a-Judge.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Survey on LLM-as-a-Judge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.317974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.317974Z digest=sha256:5622fecf4957bd666b766d311ed794d357ef35423c21ac6b3d8ea9987bdc85f2

Observation 52a345a1-6f49-40e1-810e-a7640175aa05 · outbound

This paper cites Improving Model Evaluation using SMART Filtering of Benchmark Datasets.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Model Evaluation using SMART Filtering of Benchmark Datasets

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.727362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.321034Z digest=sha256:7f62721b856970c4cb8f84159e6dfdc3548474056cc2a6c5505e3a6c0d7a360c

Observation e85e20e6-37bf-4433-a5fb-1c42a9f3b091 · outbound

This paper cites Measuring massive multitask language understanding.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Measuring massive multitask language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.812461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.323991Z digest=sha256:d17728a697cdb175149f9d4f9502b9511317aac55052214d70710d3dc964ba94

Observation 1af5e0d3-7f8a-4b50-b586-318c67eb5b8b · outbound

This paper cites Trueskill : A bayesian skill rating system.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Trueskill : A bayesian skill rating system

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.803349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.327209Z digest=sha256:d17c509794a567d7d73635c3a54e5cf5f2ce796a5a71b8f4e1acae70f0354215

Observation 5126eeb1-cf26-4ab6-a34d-14f7327f25db · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Lo RA : Low-rank adaptation of large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.330561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.330561Z digest=sha256:0b2846505e3519e2e66b9cfb378d126f42e70e759217f7005b53be1908c8ef65

Observation 47a8eb31-62ee-40b4-88d8-7030916e14a2 · outbound

This paper cites AI safety via debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks AI safety via debate

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.333252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.333252Z digest=sha256:fe9da43f45ff9a2a3c7f3a9f0b4eb889113a28ac19d414bc26b287b01ecf8d32

Observation c9630232-1386-4a80-9503-c007f5e98ef4 · outbound

This paper cites Mistral 7B.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.336865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.336865Z digest=sha256:85fff108a55b1b17596397fbdf01b7d540609a7ab08f6e97ea6a9ac9d3993da0

Observation b86f7a43-fbad-48c8-ab59-5f59d018b4c0 · outbound

This paper cites Mixtral of Experts.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.340177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.340177Z digest=sha256:e5098a4efbc98a6b4a69c4aa62846306695cbbb0c42050777e41290d893cf53c

Observation 87e694a9-d61b-4e1a-be24-bfeeeeac2a80 · outbound

This paper cites Bowman, Tim Rockt \"a schel, and Ethan Perez.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Bowman, Tim Rockt \"a schel, and Ethan Perez

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.787423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.343125Z digest=sha256:ebe59f8da9beba44153bfcb979043c2bb95d26fcf91b502f562d9b56fcbee77c

Observation 325bd2a9-eb0f-4a67-a6c4-a13267cefb10 · outbound

This paper cites Debate Helps Weak-to-Strong Generalization.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Debate Helps Weak-to-Strong Generalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.346312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.346312Z digest=sha256:fd731c392b7ebc89c777a2289841050d645de0d141479ffa0c0cf3ddfa100876

Observation dcdbfae7-ea09-477a-93c7-932389ee3b53 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.349366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.349366Z digest=sha256:3c8e7e14222a008d015a891b37746e37228d0bf675ddafd7b3ffc8ee3da5db09

Observation 88498f70-5335-422b-be70-d39ac2d85e91 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CMMLU: Measuring massive multitask language understanding in Chinese

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.352447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.352447Z digest=sha256:d7fd99df9b63b20d0c8d218c595e6b2b1ddeb2acb92bb8c7e7647cf96923ffe4

Observation 87452799-2260-4345-bd6e-15bd193c9117 · outbound

This paper cites A Debate-Driven Experiment on LLM Hallucinations and Accuracy.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Debate-Driven Experiment on LLM Hallucinations and Accuracy

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.659401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.355365Z digest=sha256:55a804f02be17651edd8c3c448dadcf54c6dcf1bbb8876208cc47a24fa9b8f1d

Observation ed2720b2-bae9-47a1-b47c-1b33256a0864 · outbound

This paper cites Manning, Christopher R \' e , Diana Acosta - Navas, Drew A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Manning, Christopher R \' e , Diana Acosta - Navas, Drew A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.777855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.358457Z digest=sha256:70b0375b0adac5f43b3038990f6d9ec33046e3cf54ce3ff3f3a8f2777bf11187

Observation b16022a5-2bc7-40ea-b7db-0dbdc8a78cd6 · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Encouraging divergent thinking in large language models through multi-agent debate

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.754017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.361156Z digest=sha256:60d0dfb91a171400027a8f85d399356c48c9f85e1c6798edd51b42becdb09a3d

Observation bcd0f28c-7357-451f-a019-43dd1ffd8c48 · outbound

This paper cites An empirical analysis on large language models in debate evaluation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks An empirical analysis on large language models in debate evaluation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.699063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.364110Z digest=sha256:e8658c4feb3ad8b52f14b63ccdcd961c90926e1910aeedf84ca42a55ed986cd3

Observation c70ea09d-581d-4baf-a8ff-a4403bcf8c35 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.638130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.367183Z digest=sha256:c44d5792baae52ae813689629f02ace8ca09974782d7abcfcde53f1e279567c4

Observation 658269b3-bea3-41b0-9ee6-c5226b858a26 · outbound

This paper cites Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.550141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.370593Z digest=sha256:1b8e3df0e83d3c5ff1479f8c0d69ed28ce971b7d1cec250cb381f1021cca802b

Observation 963a0a2e-6217-4657-988a-4b2724e01846 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.373495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.373495Z digest=sha256:a3765ddc9528c1ea18ae40fb5b3ac68aea3c4d5738df6c0673fab3323023d8cf

Observation 471f0037-bffb-4f85-927f-10aaadc82ad1 · outbound

This paper cites Cheaper, better, faster, stronger.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Cheaper, better, faster, stronger

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.454555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.376438Z digest=sha256:4d55622b98afcf74378b18fe27e32e748a3ad9702927b98dab809a19050c86c5

Observation 48907f32-9cc1-47f7-a2bb-4fc3af06a2d2 · outbound

This paper cites Mistral large.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral large

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.325353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.379133Z digest=sha256:d5e0d45bda5bc137464c2bd64d871f6186a012d2168fb4a22462ecbceb190c21

Observation e966dac2-cd50-4d15-baaa-94f3454d0b18 · outbound

This paper cites Evaluating the Performance of Large Language Models via Debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Evaluating the Performance of Large Language Models via Debates

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.381842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.381842Z digest=sha256:ca71cb96ff4d8eb0554f392c62f460d3546df2b9a87b2e4998c8c80f1ba1d4e8

Observation 0933da57-b03d-4223-9983-e35edb8fc8f0 · outbound

This paper cites GPT-4 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.384917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.384917Z digest=sha256:c21c85ac6de63bcf21ad83b2e2c432964fc74674bd69e228c620597340ee81ad

Observation ec0c7b9f-1c02-404d-b2f5-9f6bb2b5342d · outbound

This paper cites GPT-4o System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o System Card

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.387932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.387932Z digest=sha256:f524eb737c2fd1d6a42927dcdf8e4b60b9258c2c9fcf105cc715f9b57caa40a3

Observation f817549b-8f37-4ac2-800d-c183768200b8 · outbound

This paper cites GPT-4o mini: advancing cost-efficient intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o mini: advancing cost-efficient intelligence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.063257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.391394Z digest=sha256:ed905deda057bab2c4e4d458ce5244f47c477c95d3cb262ce17ced314a0f0bec

Observation 8c4d6f4f-37f0-4b11-a6fc-4272bd79a562 · outbound

This paper cites OpenAI o1 System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks OpenAI o1 System Card

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.394084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.394084Z digest=sha256:6a16398d0b18ef5ae3cf55bc5413bc44e33860e5ba300e384c94d91e541d8059

Observation 81b8e3f6-eb82-4a21-94f3-71c9f31dfe3d · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.396998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.396998Z digest=sha256:c182b400750066e54fb2b0fd24ca74e840b21699542fed8b3a2290407de0dcf3

Observation 067938fe-4300-44b8-b2e4-4beba9569ad5 · outbound

This paper cites Mapping global dynamics of benchmark creation and saturation in artificial intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mapping global dynamics of benchmark creation and saturation in artificial intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.399824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.399824Z digest=sha256:3b101b2ff0dd12a6eb0e13dc2c9c8fb82d6d51a6a4dc971040b1341cb630bab1

Observation 7053c87d-d62c-44aa-b800-fd743facd11e · outbound

This paper cites Humanity's Last Exam.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Humanity's Last Exam

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.402625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.402625Z digest=sha256:df572e60eeacd7d95977ec1c89273eac4f2942bee0fc02d556028480ff3390df

Observation f12d1973-fa38-4f70-9902-f1b6e486e20c · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Introducing gemini 2.0: our new ai model for the agentic era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.854039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.405513Z digest=sha256:015599964f4bcfc4a89b91e3bab0df8811611718833a3a9f2a46238a5db43ac8

Observation 36bb7fd3-ed87-4e09-b768-ccd3d1c8dc91 · outbound

This paper cites Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.471219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.408457Z digest=sha256:2ff30fa193e4d6ad7bbcc10b335b5b3e17f7e782c95e0feda3c03e71d6cad99e

Observation 07618cbf-e2a8-490c-bec3-fbc58bbe8bf3 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.411020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.411020Z digest=sha256:afca59f902fbbb8f3c83f7fffa60a976d36a2a0ab82f5b2203dba400af578a1e

Observation 8bb4c05e-68cc-4951-855a-2f612619ae98 · outbound

This paper cites Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.413880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.413880Z digest=sha256:ba1b5cbbd3919b973fe6f35ab74244f33d4536339d29c1db3b0bbd10c9fbef11

Observation 3cf9e6df-393b-4b02-adde-6800a7912adf · outbound

This paper cites Pretraining on the Test Set Is All You Need.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Pretraining on the Test Set Is All You Need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.416573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.416573Z digest=sha256:1714e8a642400bd2c489d60c3a343e5160148be4a597ebf24cc803b04b98b2d4

Observation d5675478-26e6-415f-b752-bc3c18d2885d · outbound

This paper cites Detecting pretraining data from large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Detecting pretraining data from large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.319102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.419654Z digest=sha256:6f189bcc8d22d8a896ff2967797ab241cd29b58b2a674dfa81e5145824b57af6

Observation db9419df-7567-4932-9c0b-f3dd29fabb58 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.271308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.422236Z digest=sha256:222c3013c2d856465a5431edc0119d45cc7a04827834c2f04c7deaeaa1b9c07b

Observation 9717d2a1-315a-4ac1-b51d-f7940e0fb2e5 · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.424893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.424893Z digest=sha256:40c546fb2bbbc24da08d5d262096626ba70e84068112f2abe497c82ff36c3591

Observation 1647dcfc-81fe-4705-a412-fd81c0ab1680 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.208592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.427865Z digest=sha256:a89b2e20b36aec1532ebc6e6a255c7dbfd3e0e33c83c483aa1b7fa31d707435d

Observation e7e1d3ac-fd53-430f-8ac1-ee4cd9cb21f5 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.148317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.430445Z digest=sha256:504fdbbddb318f7bf430bffc7c508e69112108a95970b45f506cc8fb3648a7a6

Observation 29e7c9eb-c1a0-4b40-a774-be1a12be7040 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.433045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.433045Z digest=sha256:e03bbe0e08ff411e453b686ee7fc70d71247b690cf6edd7b29d244f7459cc36d

Observation afab2f8b-6539-46ef-83d0-73bada5393a3 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.124690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.436267Z digest=sha256:e420213c49344aa3f6426794573b88b187f7073f70aea92a5126739d4be81223

Observation ff3d821e-b5f0-47d2-aae8-11173c64161c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chain-of-thought prompting elicits reasoning in large language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.115093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.439020Z digest=sha256:2dbcedd02b563bb289a80f63e18944d5c032a3ed2be53a199d77403051350441

Observation 6c16eb36-7f0b-4507-a872-929d4fc131ee · outbound

This paper cites Livebench: A challenging, contamination-free LLM benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Livebench: A challenging, contamination-free LLM benchmark

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.105651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.441678Z digest=sha256:5b86c5457a477731b451fb77452a9d089e0a427bef55de5bf2877aadcb61ca20

Observation ff34c8d0-bf09-4d70-9d71-6d5258bfea53 · outbound

This paper cites QUD eval: The evaluation of questions under discussion discourse parsing.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks QUD eval: The evaluation of questions under discussion discourse parsing

Reference 70

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.505113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.444655Z digest=sha256:ad9be9296bccb6a773aa395b8a05faedd137c77d3658ef47da8f761d54d5cb95

Observation f51edf2e-e3ea-4b73-b8f5-05d222f22c20 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmark Data Contamination of Large Language Models: A Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.448188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.448188Z digest=sha256:50fc8974af6d551c3a077372b1bd39eb3a25305a50139f30ea940b8e82293dac

Observation faaed509-9aa0-41be-8f71-1d8c24b38fac · outbound

This paper cites KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.451179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.451179Z digest=sha256:deaadbe599ec94ee6ccc9b27d8a565d2fa48c2510c6fb1fa94322aa70510188d

Observation d1754a25-0780-4bf6-8e85-c1d0fa4eb958 · outbound

This paper cites Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.454188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.454188Z digest=sha256:995ae815fb9a8b7a33bbc8874921606636669c84c4c4d3e83e76c935b1b58f02

Observation 28e1a61d-5759-49a8-a35d-c52eb1b52a35 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Xing, Hao Zhang, Joseph E

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.096336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.457014Z digest=sha256:eea333635bf3309f5acad4ca47720a169465a14f899a93b62f3d6f60f9639de7

Observation aa7be3b1-c45d-452b-a424-1fc0ded6b452 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.086077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.459891Z digest=sha256:11e357adf9d46a7cc088e47c5a9237ab482797f6845d1e0de793c8c09aac532b

Observation 7d0e033a-3e88-47c9-bea6-db99e9f6c28f · outbound

This paper cites write newline.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks write newline

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.462617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.462617Z digest=sha256:ed3640440ee1b48d3d3addab2f2a99373939944935ea750079e7b559eb229629

Observation 6edf6c22-fd91-48f8-8d29-35f64be39ffd · outbound

This paper cites @esa (Ref.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks @esa (Ref

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.465929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.465929Z digest=sha256:b784b944adef4003e7760643b112d6cd437b421b3d81bdaafcca970a09c2fd55

Observation 4339ae32-1247-4087-955b-68eaa972ae45 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.468994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.468994Z digest=sha256:1c3aac01433218b5e6c60f9205a384a78a7b698253e3a4b051d08915d97bc81a

Observation 5fde41a7-2a87-47c6-9831-2c3a0ca7d3a2 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.471855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.471855Z digest=sha256:ddc5345093e68ae4837c141227e3f666da3ea17cc30c3c655455e01342b8520d

Pith citing papers

No inbound Pith citation observations are available.