Pith. sign in

Paper Citation Record · LEDGER

Unbiased Evaluation of Large Language Models from a Causal Perspective

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2502.06655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06655 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.504623Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b4409f8-0ba5-488f-a87f-87c585bf5257 · outbound

This paper cites GPT-4 Technical Report.

Unbiased Evaluation of Large Language Models from a Causal Perspective GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.341060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.341060Z digest=sha256:c4c0b6403ceb7b8ff7fea5161e5ab7776eb5c28e1a90f26336addec58e9f50ac

Observation 6f31d78a-fdb8-4696-9c3f-572500fe735a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Unbiased Evaluation of Large Language Models from a Causal Perspective Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.366197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.366197Z digest=sha256:51fc577dcbb78f8b8fe74ef664fe4ab8c3452672c12476e737bdbc33d249c2e6

Observation 9a3aa460-0158-40b5-a0e4-024c112f9eaa · outbound

This paper cites Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era.

Unbiased Evaluation of Large Language Models from a Causal Perspective Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.378498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.378498Z digest=sha256:88c128793d422b76ab7db9a3ffff2f04187674c562da0359482bd90e8a028338

Observation fc4c9d7f-1845-4048-833c-52fbfdf98479 · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

Unbiased Evaluation of Large Language Models from a Causal Perspective NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.383999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.383999Z digest=sha256:66ff165af3ef1dfe935eb9b3fa3a467460a40e9b1361ffb6bc89bc9cbba5a034

Observation 47abcf14-5598-461a-91cf-003ed3205f83 · outbound

This paper cites Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.389935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.389935Z digest=sha256:f11a6a828c81852f203b21e93bf785b1eba4d43a7a23b61f6c97f22e197c5c63

Observation 98d7e468-59d7-4412-9762-3a346131bc1c · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Unbiased Evaluation of Large Language Models from a Causal Perspective OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.395521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.395521Z digest=sha256:378185783a193c8c1e9786086b59e78ca38824bfe31e018480a0375ce50f30eb

Observation 92115fe6-1869-4852-b587-e4e35eba4996 · outbound

This paper cites Mistral 7B.

Unbiased Evaluation of Large Language Models from a Causal Perspective Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.400935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.400935Z digest=sha256:0031771f7b598f9620c25a4b12adc8cfa0ff27a7b6e97f1faf6e32514d4a0f9b

Observation cf07b56c-9562-4a43-a90a-cfadc57ea01f · outbound

This paper cites Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:51:16.014831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:51:15.405782Z digest=sha256:4b335b3790822a7ee1a7f999245a2e66b131267449c22734c56893d921c0176c

Observation 1618b515-b5b2-43a2-bc67-6ac3d2c5ac4d · outbound

This paper cites Deduplicating Training Data Makes Language Models Better.

Unbiased Evaluation of Large Language Models from a Causal Perspective Deduplicating Training Data Makes Language Models Better

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.410849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.410849Z digest=sha256:d78bbf8afb79df6babba9d52a88b8145740f4fc1f7b27a3b3e7b303a8168295b

Observation d291a9b5-6e1c-49d9-b7c1-70ce3558911b · outbound

This paper cites S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models.

Unbiased Evaluation of Large Language Models from a Causal Perspective S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.415940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.415940Z digest=sha256:11c1f266470ddaf55404cbe6a0db9d4e104bb29d8923732fecd30ef64fc75fa3

Observation 2f5b0c50-5873-4c8b-ae0f-29c48757c09d · outbound

This paper cites An Open Source Data Contamination Report for Large Language Models.

Unbiased Evaluation of Large Language Models from a Causal Perspective An Open Source Data Contamination Report for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.421100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.421100Z digest=sha256:371609e325326e9ef6c385e7a7b0ca1d7a8918a4fa45a66277dc9ccfa8b8a3b4

Observation eaeada13-0cc4-4762-9209-69f34cdc99f6 · outbound

This paper cites Holistic Evaluation of Language Models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Holistic Evaluation of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.426227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.426227Z digest=sha256:d98d405c6134daa501873422fbf9fd234cc4a910910bdb7b9bbe0d44184771c6

Observation 46ecd899-cc2d-42fc-91b0-27259ba69af6 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Unbiased Evaluation of Large Language Models from a Causal Perspective Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.436864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.436864Z digest=sha256:a569353a8c1b505de8892e555aaca48588c05dd9202d4b2d1fab56e2c202c60e

Observation 11a417e7-a81b-403a-b639-4071f017c38e · outbound

This paper cites Training on the Benchmark Is Not All You Need.

Unbiased Evaluation of Large Language Models from a Causal Perspective Training on the Benchmark Is Not All You Need

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.442357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.442357Z digest=sha256:d4abc983757e690d561a7f0ac1e6010179db904b18c8357104074b3711ea335c

Observation 9ce1ddb3-57b9-482a-9d4c-84c632a00216 · outbound

This paper cites Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.447817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.447817Z digest=sha256:921631e1813198aae613aa0f10b399a42efeb3b4f29dae91b7f37dc8827cf99f

Observation d38930be-bc42-42ce-8362-ea6b0e0573df · outbound

This paper cites NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark.

Unbiased Evaluation of Large Language Models from a Causal Perspective NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.453440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.453440Z digest=sha256:932ebe10619c5a3c7422716365d8b80c3fa734dc3d909a4ea6d4d29b916abef0

Observation 315e78dc-d08d-4f21-9280-599eececdb96 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.458616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.458616Z digest=sha256:7ce42202983a9b08f810ab2da6931c68820e947110a05eedc81ec12d2738724e

Observation 938b5a87-2aa1-4590-8fc9-d387d09541cf · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Gemini: A Family of Highly Capable Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.463959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.463959Z digest=sha256:cd40ec9bc6a9ab10ef0f555af067b6ea3549f43ecfed6ae251154337c88f4217

Observation 93570609-b863-4b35-ba27-d7e9024919df · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.469266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.469266Z digest=sha256:3eca3913a1be898d9afce5ccfd47461fdbe2c5248adccf2c61f1f2d980d60d0e

Observation e462bad4-5da1-4923-93a4-47d4ec74d808 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Unbiased Evaluation of Large Language Models from a Causal Perspective GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.474266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.474266Z digest=sha256:125c9e22886ce2c2f491ccbab5ccec82893a68323a82a26a221171df815542ca

Observation 2d6af2ff-2f14-4613-9844-ee7acd779fe9 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Unbiased Evaluation of Large Language Models from a Causal Perspective LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.484517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.484517Z digest=sha256:18fe58dbe57135422f3e3fe952f97dac8dac3343184deb92233654f4f4e964ef

Observation d4909b1b-cc7f-473c-8b42-83f020171eef · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Unbiased Evaluation of Large Language Models from a Causal Perspective Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.494390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.494390Z digest=sha256:a27b75cb2172975f6de94331f7c2684502a67b4c95cbc9c898e4f585b6811780

Observation a45fb792-6052-4f45-a804-274711e4840e · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Unbiased Evaluation of Large Language Models from a Causal Perspective Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.499571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.499571Z digest=sha256:b087412d225ef5a5cf98ac5c3314e0c65700297d563888b8038e0c9003192826

Observation 4d3f199f-97ef-4ba1-aebd-2a112b542616 · outbound

This paper cites DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks.

Unbiased Evaluation of Large Language Models from a Causal Perspective DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.504623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.504623Z digest=sha256:e01e4a567fa924df1de4b6395255c97173d2f96c464472dadd9230f4b1ce6376

Observation fcff33de-f9e2-46cd-9075-7d3edb614343 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Unbiased Evaluation of Large Language Models from a Causal Perspective Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.372544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.372544Z digest=sha256:a684a60df4eab65f02eb0533a8bc03ced490870003c6ccb96caf67fb007e7878

Observation 1bb2c226-877e-4409-9356-16042d1529cf · outbound

This paper cites Measuring short-form factuality in large language models.

Unbiased Evaluation of Large Language Models from a Causal Perspective Measuring short-form factuality in large language models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.479245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.479245Z digest=sha256:a77924a9785975a250a7db6ef9e1ce6a2cf6e6939e25c288d884aee20c5ca5b2

Observation c775374b-a7bc-4ac4-85ee-856e99153b1f · outbound

This paper cites Language (Technology) is Power: A Critical Survey of "Bias" in NLP.

Unbiased Evaluation of Large Language Models from a Causal Perspective Language (Technology) is Power: A Critical Survey of "Bias" in NLP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.358571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.358571Z digest=sha256:e2443cb467a8a5a3653836464bd7eae80b70ea6e2a15af844eda11250f44c297

Observation 2b45f01e-ca1f-4342-a913-6ad12f7562f0 · outbound

This paper cites Let's Verify Step by Step.

Unbiased Evaluation of Large Language Models from a Causal Perspective Let's Verify Step by Step

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.431556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.431556Z digest=sha256:551b8a5883a05a91bd76347e8b6b294eb3b84f0684f40a19e2db85b0a2aa5bca

Observation ced1362a-0f76-47d9-91cd-532dc3219114 · outbound

This paper cites M., Gebru, T., McMillan-Major, A., and Shmitchell, S.

Unbiased Evaluation of Large Language Models from a Causal Perspective M., Gebru, T., McMillan-Major, A., and Shmitchell, S

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:51:16.032953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:51:15.352793Z digest=sha256:0e37f1b020a96addc185cc484b137e68ac89f278b944d0f88455f01a402e1f88

Observation f4d370ee-0a30-47f5-87d1-08a5c0976daf · outbound

This paper cites Qwen Technical Report.

Unbiased Evaluation of Large Language Models from a Causal Perspective Qwen Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.346879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.346879Z digest=sha256:69b2ba8fadc32209d325a1d3bd7d12ad0171a1c726b43986b00e504c52d38775

Observation 246859ee-fbaf-4caa-ba72-24d790e7d46a · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Unbiased Evaluation of Large Language Models from a Causal Perspective Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.489455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.489455Z digest=sha256:c033cdeece96aa5d7458054f04fd802826b0ce7255677ac71d8611b7bc1d1a19

Pith citing papers

No inbound Pith citation observations are available.