Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Cognitive Biases in Large Language Models as Evaluators

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2309.17012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17012 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:34:00.009661Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:53:50.367532Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f899883b-0a27-470b-8d38-f46471bf40bf · inbound

LLM Evaluators Recognize and Favor Their Own Generations cites this paper.

LLM Evaluators Recognize and Favor Their Own Generations Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.816877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:6cddc76aaeb1ba017f49c952b4347fd85301d7d77305f621b436935bf472e2b2

Observation ab7b2c8e-08ff-4d0d-838d-9aed89e9baa7 · inbound

Better & Faster Large Language Models via Multi-token Prediction cites this paper.

Better & Faster Large Language Models via Multi-token Prediction Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:26:09.810803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T12:26:09.731664Z digest=sha256:5d6818841b5f524c2ba712a28570f0fdb53f347c6d51334b821f3853245df615

Observation fc4e8733-be44-49d0-9fec-7d700a4c89f8 · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.988267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:29bc929a3523b140b406b566790b346baa1b046cc4a2757ee52d0e0305b18fc7

Observation 5a3abb5e-cfdd-4880-860c-93fc57e331bd · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.032235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:2da86e2a4f7f38edcfe838602b368b8f26575f470877a4f9a3759fd8001f9110

Observation 9e50b7f7-7e23-4f03-9e13-a8302fd0db5b · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.018788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:0139598d332227f9464d7d4eabc9de6ae8f533d0399d724761ee41b4ed04dfd9

Observation 6b7b52ec-d510-4b2a-ae3e-41f8d6522a8a · inbound

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models cites this paper.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.722030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.722030Z digest=sha256:3770f5e5a450ea45f7b47be58f07bf64bdce8bbfe970424c693e44109451cd36

Observation 3bb4f7c7-5569-427d-a7e3-9a8fec855943 · inbound

Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer cites this paper.

Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:00.009661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:34:00.009661Z digest=sha256:da5f7ef134370a33abf1bcf92d42bdb996f2d72fc0b6e13d4de978aa09515db9

Observation 5648ec8d-639b-4df0-bcf0-3c04f82cfe5a · inbound

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward cites this paper.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.832116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.832116Z digest=sha256:47f9e19577f426348356f125777504f897af6a0de5f6ba5b73fd76081a2df9a7

Observation d3d9a890-367e-4753-b7db-98dd3c0c463c · inbound

The Leaderboard Illusion cites this paper.

The Leaderboard Illusion Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:22:55.548240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:22:55.548240Z digest=sha256:369de3ef74f90f03277d98455a16afd8030cd766f186f331d7b1444bd00abce0

Observation 505f286c-531a-419f-b90c-58386211ef30 · inbound

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP cites this paper.

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:51.836687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:51.836687Z digest=sha256:3ac07518fc23b4bed4a0b9a3738b2cd700a8bf28f49c0b0b7720b7feaebb039f

Observation 2f4d8d5a-931a-4bb6-81bc-60471a111347 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.842968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.842968Z digest=sha256:97ae2dbbd752e0fd6452011ede4462b135d167eb567bdfdc1ba90b3f2b32be71

Observation 6c134385-792b-4c51-9ea6-155b1800f4e3 · inbound

AbsenceBench: Language Models Can't Tell What's Missing cites this paper.

AbsenceBench: Language Models Can't Tell What's Missing Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:13:17.718512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:13:17.718512Z digest=sha256:ec6b4847bc9a645d063885c61f72395857adfeb0c44695deb725ea5b70c614b2

Observation 4568b5f1-b150-4c37-96a1-ad7f735d3388 · inbound

CRISP: Complex Reasoning with Interpretable Step-based Plans cites this paper.

CRISP: Complex Reasoning with Interpretable Step-based Plans Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:01:31.053765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:01:31.053765Z digest=sha256:761aac2d6d0b7b0f29d7f0f5a48bd72635743a2e99665f2f1bd2fdb4fc2c7410

Observation 602d4ffe-c342-4bef-9e26-f28960e02120 · inbound

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming cites this paper.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.064778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.064778Z digest=sha256:1f2a1e531be53e7e818028aa396aedd3eb04eb9a2a1e5762b23cc6c69ce37e1e

Observation 44abe90e-85c8-43a4-9e08-15a1b0802d92 · inbound

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge cites this paper.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.206258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.206258Z digest=sha256:f3b07deed57e6cf1f6c0d540e0c7b2133463a523cbca1d8091ebb49a7992a560

Observation a1b445b0-1406-4459-b511-3d2c5ce62956 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:58.213182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:58.213182Z digest=sha256:ed8bc78a5273676abf90ac44c50377fe6401e6c55bb443bc166196aaa35d2ce8

Observation 4e1e2dec-7704-4b8c-935e-200e8f2c817e · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.732735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.732735Z digest=sha256:230a63c0c124ec2d40e45e811c2f83ef82154db345cd039d1eccfcbba62df874

Observation e4a5ec78-d9dc-4d66-ba76-3f0446d240da · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.892235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.892235Z digest=sha256:9ed6d6a0f25c032112957827e46deb5b79debaffd13cedae0b19d83534f56d64

Observation 038f4fa5-de01-4194-bad5-b298963f42eb · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.481371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:fcfc032937f3f3758f9db5b30a520655c2c9464919f41919794a020179fe12dc

Observation b59a23e2-396e-4e0c-a714-32830efcc736 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.426512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.426512Z digest=sha256:7c96f322979ebfa443b4c5d1cd81281f9f234ad4b45ddcd746dcbb90ce0c6900

Observation 8ed94691-671b-45a0-9090-854487c94667 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.332913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.332913Z digest=sha256:fd60d6a799242aad3edd2df5dda67c309751c5293012e54fdbbc0c430f895043

Observation 38c552b8-0b23-44f9-86b9-13366418ef4a · inbound

Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations cites this paper.

Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:44:37.607656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T10:41:21.406201Z digest=sha256:2b174e90b0cfcca78f00aae958e45e746b2b6379e1820517db537d48d0f0e96a

Observation 2ffbd726-6c6a-476d-9f0e-873e74becc45 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:41:13.976357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:4ad39c06725c43ec67d404077e40c71fc6902b9a78f8e49beab7885b5b4fc9e6

Observation e1a4976e-a1e9-4a7e-94f5-0d076211b651 · inbound

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning cites this paper.

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.077576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:19:55.849451Z digest=sha256:90425bc2bc0ddc71e56636a8cd50c4f937feefb39679c6e05997d601cee364ac

Observation 66475279-6144-48ed-9d7e-c63011334597 · inbound

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps cites this paper.

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:33:16.779705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T12:32:22.536994Z digest=sha256:b97dd7b33318d2621afe4ef7581b8ec65aefacc293a2dc19fb126531e3ed3983

Observation 8bbb5c07-bfda-4855-9660-e75ed527d963 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.715349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:8754fb832d544dee1328070f283d45c0a8474ae469b2b1dc927c622dbf7c7078

Observation fb9f1fd1-16a0-48c7-81b5-28a198b0db93 · inbound

Show, Don't TELL: Explainable AI-Generated Text Detection cites this paper.

Show, Don't TELL: Explainable AI-Generated Text Detection Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.555612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T12:56:04.003337Z digest=sha256:4e5c7d3fbacb2873eaf82a31246be2f4a89482649281e77d53e8bff94e6363e5

Observation 50aad0d0-743e-43ca-9fad-4a0f00ca67f3 · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.556059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:39fe4bdddd062bc1c48847e0e61b594556a3b38df3a5ab58e3336803e4027c46

Observation 1892a6d9-f0ed-461b-86c8-9c57c92a65c1 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T12:53:50.369284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:6143623d235828c72945004365f6270e7dab0a6df14eab9a70a999eb706e9d13

Observation c9b73bbf-03de-4ba3-b178-d538abc065d4 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:06b0c810642def3db1132a54ef3757db313da81ab5d5f57f64d4875e9839aa44

Observation cb7ec13a-8c12-44aa-9f13-a015ba53d632 · inbound

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation cites this paper.

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T06:40:17.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:40:17.865408Z digest=sha256:63c2fe45e0bb3c6cf9e9e774896c485bfdd7facf6e38dfb564ccac45796f796a

Observation 2cb7b65c-3a5a-4d79-9294-95f4a01daa4d · inbound

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias cites this paper.

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T02:33:34.084111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T02:33:34.084111Z digest=sha256:b2e671f306b73a809c5881788d6b2f90b9289b79ab5b1d6c10dbcca7c77a583c

Observation 48c6a736-6094-48ba-be94-2cf8e6e86dbc · inbound

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces cites this paper.

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T16:42:05.597565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T16:42:05.597565Z digest=sha256:8bdb7455e3ba3379f594352051c22760a4c87ad5939f09198ca27fe6f259f71f