Pith. sign in

Paper Citation Record · LEDGER

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2508.06709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06709 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:36:43.365289Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b56aec3f-1b02-45fa-8577-3816428ecdc9 · inbound

RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Lifestyle Social Media cites this paper.

RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Lifestyle Social Media Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:51:26.124111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T13:46:48.348859Z digest=sha256:afbd60b25541b83dfa7dcdc5dadba4fd74839f43e2b071ed65e6b50bb9e71989

Observation 1af5245e-b6bd-4899-9cf6-12a53103a301 · inbound

Extreme Self-Preference in Language Models cites this paper.

Extreme Self-Preference in Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.082645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T20:57:37.199128Z digest=sha256:06d10cd9b8290668b57124caee46b4b0534e5c7f98e071c751b68f1a6fdb606a

Observation 6022807f-8a57-4574-bd20-c2abb734a229 · inbound

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning cites this paper.

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:51:08.533941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T08:50:53.486588Z digest=sha256:ef5b52569caf9d9470605fd254cc1d47166078ae19a4b02dc66969b9b19ddeb6

Observation 40c16883-2fe6-48d7-a8f9-8420bfe1dbba · inbound

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations cites this paper.

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T06:35:37.070203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:35:37.070203Z digest=sha256:ff9305bd7ec6c556c6dd2914d36daa50d33538276a3e033da8c69dc4139e7e70

Observation 6bcbd550-bf76-455c-8374-2092d0880441 · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:28:48.785362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:dab700ef764a2bacec3505dff21788b59386119c6fbbe3e730ab6e89a1625ea7

Observation 4d2d3d9e-46ad-4bc4-8ec8-5f727af6279c · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:51.666192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:7315696be4b9fe4a9ef4cf3fc36866c0f5974ff8db0ce82f23bb157e0a39f862

Observation 2cfb26b7-9e07-417b-906a-14e481c77c59 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.434134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:d6658aeaf30eb5708598c4aa31b4c37c42240de0d1489c84c72281fe18076174

Observation 44be8642-c2da-490b-b916-08fbd24dc89a · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.448677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.448677Z digest=sha256:418766793dc913f7f5586f772ba0dd9f2a625cbbacdb7072bb53991f4f922c20

Observation e0c2e89e-0807-40cc-aad1-04f92cd4a65d · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.365289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.365289Z digest=sha256:e2bd58567e07abe86363ed358017a2602e1da10aa1a86c93f6dbae3927ed1457

Observation 3e25a0e9-63c4-4f43-a7de-2c632048062e · inbound

Why Do Safety Guardrails Degrade Across Languages? cites this paper.

Why Do Safety Guardrails Degrade Across Languages? Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:18:21.034567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T14:17:41.347193Z digest=sha256:0a2504759285bd3c6f3a5fd616959316437857e5aff21a2cd5e4910614616a55

Observation 328827ab-7ee7-488b-8ec1-322163c9e6d9 · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.462423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:56eaae7faadca500ca6e46feb20b51b8d831f5372d221304ee79bd9030acdf3d

Observation daf7bb2d-8931-4544-87ec-51046d0acc16 · inbound

Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill cites this paper.

Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:47:04.549846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T00:08:06.720586Z digest=sha256:eb7bee98002121f3eb46ee524bda64af6fc5f177cdbf7fe93a9502bdfaa34a13

Observation 2bfa219f-0831-4e2d-8252-5e3a3f8ba9d5 · inbound

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG cites this paper.

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T10:22:55.558755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:22:55.558755Z digest=sha256:a29fb61c318078282a24440c619391ee2cd5541ccdcb1b6480ed8ef92e4d24aa

Observation fa650952-5d5a-4684-b682-9a9d4e91de11 · inbound

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding cites this paper.

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T11:54:20.710057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:54:20.710057Z digest=sha256:a7c94565a2843dc34874e9e822cf5111ec1f96d42f95213c9e6bb3e8ff9bdc5d