Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.504623Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2502.06655.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.504623Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b4409f8-0ba5-488f-a87f-87c585bf5257 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f31d78a-fdb8-4696-9c3f-572500fe735a · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3aa460-0158-40b5-a0e4-024c112f9eaa · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4c9d7f-1845-4048-833c-52fbfdf98479 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47abcf14-5598-461a-91cf-003ed3205f83 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d7e468-59d7-4412-9762-3a346131bc1c · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92115fe6-1869-4852-b587-e4e35eba4996 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Mistral 7B
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf07b56c-9562-4a43-a90a-cfadc57ea01f · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1618b515-b5b2-43a2-bc67-6ac3d2c5ac4d · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Deduplicating Training Data Makes Language Models Better
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d291a9b5-6e1c-49d9-b7c1-70ce3558911b · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f5b0c50-5873-4c8b-ae0f-29c48757c09d · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective An Open Source Data Contamination Report for Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaeada13-0cc4-4762-9209-69f34cdc99f6 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Holistic Evaluation of Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ecd899-cc2d-42fc-91b0-27259ba69af6 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a417e7-a81b-403a-b639-4071f017c38e · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Training on the Benchmark Is Not All You Need
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce1ddb3-57b9-482a-9d4c-84c632a00216 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38930be-bc42-42ce-8362-ea6b0e0573df · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315e78dc-d08d-4f21-9280-599eececdb96 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938b5a87-2aa1-4590-8fc9-d387d09541cf · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Gemini: A Family of Highly Capable Multimodal Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93570609-b863-4b35-ba27-d7e9024919df · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e462bad4-5da1-4923-93a4-47d4ec74d808 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6af2ff-2f14-4613-9844-ee7acd779fe9 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4909b1b-cc7f-473c-8b42-83f020171eef · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45fb792-6052-4f45-a804-274711e4840e · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d3f199f-97ef-4ba1-aebd-2a112b542616 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcff33de-f9e2-46cd-9075-7d3edb614343 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Training Verifiers to Solve Math Word Problems
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb2c226-877e-4409-9356-16042d1529cf · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Measuring short-form factuality in large language models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c775374b-a7bc-4ac4-85ee-856e99153b1f · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Language (Technology) is Power: A Critical Survey of "Bias" in NLP
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b45f01e-ca1f-4342-a913-6ad12f7562f0 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Let's Verify Step by Step
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced1362a-0f76-47d9-91cd-532dc3219114 · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective M., Gebru, T., McMillan-Major, A., and Shmitchell, S
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f4d370ee-0a30-47f5-87d1-08a5c0976daf · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Qwen Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246859ee-fbaf-4caa-ba72-24d790e7d46a · outbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.