Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:44:09.471855Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2507.17747.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:44:09.471855Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a83d6f7c-cc0d-4c07-9845-7e11451947d1 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 sonnet
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c351428-9b2f-4c72-bcd6-0b4c6cf93302 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 haiku
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00efc9cd-aaad-4095-8e28-dba4fd18b6a8 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 168020df-faa3-4617-bcb0-3f661ff0428e · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmarking Foundation Models with Language-Model-as-an-Examiner
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2899a6c-dd33-44ef-9b7f-4fa173d23ace · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e75f131-7bc5-43fa-b6b7-b20bf3ce0b61 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Adversarial multi-agent evaluation of large language models through iterative debates
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 325cfa56-0309-4fc2-bdd2-eb2b2dc9dbb9 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Flageval
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28e252be-bc05-477b-9b7c-6b1eb1892a54 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91128ddb-5cc4-4299-be72-ffd00b990fdd · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00b4b03-06ee-4086-9adc-237c2aaeaa5e · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b5ddd1-7732-46aa-a3de-03b2b08f4d5c · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Role of Deductive and Inductive Reasoning in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dca74ff-0d91-4e32-ae9c-e9c730ce254e · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Are we on the right way for evaluating large vision-language models? In A
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2000b724-3372-452c-975c-a70ce8d73640 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Jordan, Joseph E
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8877c7e4-7337-4901-9ce7-33441fd4fad0 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC Prize 2024: Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8b6024-00f3-460a-a420-2b2ccca90a02 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85429f59-fd84-4ab2-a982-cdf409324331 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Training Verifiers to Solve Math Word Problems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48812bb5-00e1-424c-b0e9-4dafced156f7 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ba679b-47be-4e9a-b1ff-285680b552b7 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6de936-dfbd-410d-b224-bb625d1ae89a · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Investigating data contamination in modern benchmarks for large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a740c12f-d3a0-49c5-b8d7-ef683acdcc95 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f4b08d6-f95c-48ea-9b21-134bb8dd3aba · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bb92465-f91d-41de-9fff-5c67a3152038 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Factuality and Reasoning in Language Models through Multiagent Debate
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a810d3f1-b2a7-4cbb-9524-d02ddea554e3 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 834de928-4ab6-4b41-a962-c2fad6bd8d13 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a8ff0aa-536d-4f47-a10f-ed347fddadcb · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Time travel in llms: Tracing data contamination in large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29f26e3a-ff53-43ea-8d89-28c698a42493 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 3 Herd of Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f84371c-551c-4272-9c46-4b5923fea23e · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Survey on LLM-as-a-Judge
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a345a1-6f49-40e1-810e-a7640175aa05 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Model Evaluation using SMART Filtering of Benchmark Datasets
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e85e20e6-37bf-4433-a5fb-1c42a9f3b091 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Measuring massive multitask language understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1af5e0d3-7f8a-4b50-b586-318c67eb5b8b · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Trueskill : A bayesian skill rating system
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5126eeb1-cf26-4ab6-a34d-14f7327f25db · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Lo RA : Low-rank adaptation of large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a8eb31-62ee-40b4-88d8-7030916e14a2 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks AI safety via debate
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9630232-1386-4a80-9503-c007f5e98ef4 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral 7B
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86f7a43-fbad-48c8-ab59-5f59d018b4c0 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mixtral of Experts
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e694a9-d61b-4e1a-be24-bfeeeeac2a80 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Bowman, Tim Rockt \"a schel, and Ethan Perez
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 325bd2a9-eb0f-4a67-a6c4-a13267cefb10 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Debate Helps Weak-to-Strong Generalization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdbfae7-ea09-477a-93c7-932389ee3b53 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88498f70-5335-422b-be70-d39ac2d85e91 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CMMLU: Measuring massive multitask language understanding in Chinese
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87452799-2260-4345-bd6e-15bd193c9117 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Debate-Driven Experiment on LLM Hallucinations and Accuracy
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed2720b2-bae9-47a1-b47c-1b33256a0864 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Manning, Christopher R \' e , Diana Acosta - Navas, Drew A
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b16022a5-2bc7-40ea-b7db-0dbdc8a78cd6 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Encouraging divergent thinking in large language models through multi-agent debate
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcd0f28c-7357-451f-a019-43dd1ffd8c48 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks An empirical analysis on large language models in debate evaluation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c70ea09d-581d-4baf-a8ff-a4403bcf8c35 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 658269b3-bea3-41b0-9ee6-c5226b858a26 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 963a0a2e-6217-4657-988a-4b2724e01846 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 471f0037-bffb-4f85-927f-10aaadc82ad1 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Cheaper, better, faster, stronger
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48907f32-9cc1-47f7-a2bb-4fc3af06a2d2 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral large
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e966dac2-cd50-4d15-baaa-94f3454d0b18 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Evaluating the Performance of Large Language Models via Debates
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0933da57-b03d-4223-9983-e35edb8fc8f0 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec0c7b9f-1c02-404d-b2f5-9f6bb2b5342d · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o System Card
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f817549b-8f37-4ac2-800d-c183768200b8 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o mini: advancing cost-efficient intelligence
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c4d6f4f-37f0-4b11-a6fc-4272bd79a562 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks OpenAI o1 System Card
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b8e3f6-eb82-4a21-94f3-71c9f31dfe3d · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chatterji, Faisal Ladhak, and Tatsunori Hashimoto
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 067938fe-4300-44b8-b2e4-4beba9569ad5 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7053c87d-d62c-44aa-b800-fd743facd11e · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Humanity's Last Exam
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12d1973-fa38-4f70-9902-f1b6e486e20c · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Introducing gemini 2.0: our new ai model for the agentic era
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36bb7fd3-ed87-4e09-b768-ccd3d1c8dc91 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07618cbf-e2a8-490c-bec3-fbc58bbe8bf3 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb4c05e-68cc-4951-855a-2f612619ae98 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cf9e6df-393b-4b02-adde-6800a7912adf · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Pretraining on the Test Set Is All You Need
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5675478-26e6-415f-b752-bc3c18d2885d · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Detecting pretraining data from large language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db9419df-7567-4932-9c0b-f3dd29fabb58 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9717d2a1-315a-4ac1-b51d-f7940e0fb2e5 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1647dcfc-81fe-4705-a412-fd81c0ab1680 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7e1d3ac-fd53-430f-8ac1-ee4cd9cb21f5 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29e7c9eb-c1a0-4b40-a774-be1a12be7040 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afab2f8b-6539-46ef-83d0-73bada5393a3 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff3d821e-b5f0-47d2-aae8-11173c64161c · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chain-of-thought prompting elicits reasoning in large language models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c16eb36-7f0b-4507-a872-929d4fc131ee · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Livebench: A challenging, contamination-free LLM benchmark
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff34c8d0-bf09-4d70-9d71-6d5258bfea53 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks QUD eval: The evaluation of questions under discussion discourse parsing
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f51edf2e-e3ea-4b73-b8f5-05d222f22c20 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmark Data Contamination of Large Language Models: A Survey
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faaed509-9aa0-41be-8f71-1d8c24b38fac · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1754a25-0780-4bf6-8e85-c1d0fa4eb958 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e1a61d-5759-49a8-a35d-c52eb1b52a35 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Xing, Hao Zhang, Joseph E
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa7be3b1-c45d-452b-a424-1fc0ded6b452 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Dyval: Dynamic evaluation of large language models for reasoning tasks
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d0e033a-3e88-47c9-bea6-db99e9f6c28f · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks write newline
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edf6c22-fd91-48f8-8d29-35f64be39ffd · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks @esa (Ref
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4339ae32-1247-4087-955b-68eaa972ae45 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fde41a7-2a87-47c6-9831-2c3a0ca7d3a2 · outbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.