Pith. sign in

Paper Citation Record · LEDGER

Beyond the Surface: Measuring Self-Preference in LLM Judgments

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2506.02592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02592 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:07.442164Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:16:42.324571Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T21:00:39.084405Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1aab9281-65d6-486d-8ba9-3327a25e6110 · outbound

This paper cites online" 'onlinestring :=.

Beyond the Surface: Measuring Self-Preference in LLM Judgments online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.399959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.399959Z digest=sha256:e40b5d167bea9200dee30742364dea9ab238ac5d3953165fd38fe6a9257f7632

Observation 916ecf34-c6b3-4acd-b948-635d662e7df4 · outbound

This paper cites write newline.

Beyond the Surface: Measuring Self-Preference in LLM Judgments write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.465840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.465840Z digest=sha256:f2a552e6d8b6ce88ad761ff89c0947111b83b27a1091aad9c4e4070696d1039e

Observation 3024f0e2-9035-4cba-ad29-7d3a29d9bcb5 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:09.309397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.557460Z digest=sha256:188ad51b5ee43e89c9f0677b15e7fd03e9769fd1b32b3f70bb2ee9471ada0767

Observation fcc45e8a-b65a-47b3-949a-abdeb19eecd6 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.632516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.632516Z digest=sha256:c21ae19b452d6e89fd44c29596713a438750c9a8c91a69ed67fb475ffe5e9474

Observation e93e67a4-bb96-4958-b89f-6ac226594aba · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.729009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.729009Z digest=sha256:da5ce4508355a35fe26c191b6c6e6347366731eed6197dd304c3f7b5f558e1b5

Observation a6ef2e73-9a5d-487e-ac02-bcb0f2e3ec14 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.796164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.796164Z digest=sha256:257cce26cac6d760d8202429c42f06be7e27e30404b6b387629aeaa93ad2bd83

Observation 2a704e88-0a57-490d-aa55-a99138c00bb1 · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.891765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.891765Z digest=sha256:281b4bc1dcd0167f82d0a394d0774648f16874a51b2a23780e14b44c7b96ace1

Observation 44e04843-4344-4b95-9591-919fa420439a · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.992851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.992851Z digest=sha256:d358010d9d28d09e1047a37a10fe1e7e2ac33ffc5f537cfba29dfde33cb4db41

Observation bfd26f89-e201-463e-94c4-0b7f99aec333 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:09.130400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.089341Z digest=sha256:184ea005c264789e426c9a1dcef6add2ef871f6ae32c98aeedc2654e12e33614

Observation df786909-59ae-4eb3-a216-77e507584322 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Beyond the Surface: Measuring Self-Preference in LLM Judgments UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.214831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.214831Z digest=sha256:f4fc2606748e61b993edf7f3f8104c9cbd2611f02a183a26c9486e40e53fd0df

Observation 511d3411-70d5-4f68-85cd-17437e3cfd96 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.375895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.375895Z digest=sha256:e334a63f95cea057474a67d86274346ded7ddbd03ae575fe4a00128c2cf14eab

Observation 91f45ef9-147b-48e1-9b2f-0a96f1876e88 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.502570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.502570Z digest=sha256:d717d22a2c214612371d90b6647f3cc187179dfaff7b0e08ad019804b586ff1e

Observation 9031e36a-c68a-45b5-8448-46253566c8a1 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:08.966253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.638673Z digest=sha256:6c96964bbdae2ffa611bfe2d376e0dfbbcc7f0028fe52b3c55974a4f043a6f85

Observation 8d0dee2f-f657-4628-8a3d-c5f7d77129dd · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Beyond the Surface: Measuring Self-Preference in LLM Judgments ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.787662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.787662Z digest=sha256:3ecd568b5ef60d35404b77697c77060d1056403f9ddf88f1dfb87700759f5190

Observation 6611b524-41ca-4621-94f6-bd82bdd9031d · outbound

This paper cites The Llama 3 Herd of Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.942801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.942801Z digest=sha256:314b1fcb2484396af3cdb32838517a033cdb88b11e7e60dbca5a2c7d465497e3

Observation 2849b9fb-d3d2-4036-9aa9-29a6a1d57b8f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond the Surface: Measuring Self-Preference in LLM Judgments DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.056652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.056652Z digest=sha256:3e5d30e46c9597b51fbe7f7b5e249b6bb868d9aa4d699e26f8a49c86208c7785

Observation 99aba7be-d3b8-49e9-a7fd-c8ef588971ae · outbound

This paper cites Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.173700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.173700Z digest=sha256:f62d7c76fad09cff3a4434d1e605287b791b5f2e0ac93e08e2aa72cec751922c

Observation d8c09241-ecda-4029-b8cc-3e301a904ac1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.351926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.351926Z digest=sha256:9e8da4ef32acee9ded4b44a8087894813c5ad170fbbabdcbe6d998b9d5f4baba

Observation 6fb0e45e-169b-42a5-9970-6a79fd6a0456 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:08.754298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:26:03.467505Z digest=sha256:52c041f87a9f7a65ec06db0ea4802398b757d182a6f352feb2c4babbde65a67a

Observation 568c9826-03d8-4c44-94b8-b56c6042850c · outbound

This paper cites GPT-4o System Card.

Beyond the Surface: Measuring Self-Preference in LLM Judgments GPT-4o System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.625941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.625941Z digest=sha256:313e16f0b86b2db3d2c40f3c7643a4bb8cf0cd1c2df9f9b08d2a2525e0781026

Observation fa632138-9ef8-49ee-a22b-2b698d4f6bf8 · outbound

This paper cites OpenAI o1 System Card.

Beyond the Surface: Measuring Self-Preference in LLM Judgments OpenAI o1 System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.732146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.732146Z digest=sha256:bcc2c7ac0d91b0466c85e6457023daa7775409d8e3716c24f85af39c643b83f9

Observation 2f4d8d5a-931a-4bb6-81bc-60471a111347 · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.842968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.842968Z digest=sha256:39b6e88437435631ff905678fa8f60fb3b96bd511230358dfc7cf0287d2bfbe2

Observation e36f4cae-64be-45db-9f0f-1e7b3cfff12a · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Beyond the Surface: Measuring Self-Preference in LLM Judgments RewardBench: Evaluating Reward Models for Language Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.953001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.953001Z digest=sha256:156cc3ab061bec3f6fcd540e0479ab48ab8b27cdfe9ba58bca12289d0360cfa5

Observation a2a151bb-abec-4435-91ae-e2868aaa889b · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Beyond the Surface: Measuring Self-Preference in LLM Judgments RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.092617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.092617Z digest=sha256:cc4a0a61a6897b06ee65870526c2ea85e9efa0f01ec82fe0c77d1c4618b6f266

Observation 9193983d-d82f-418e-b27d-2ebefab0ee75 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.199628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.199628Z digest=sha256:e631182ddbc7b814943361da58cefe0a54ed70800587c5b0bd0ec91156ea3466

Observation fe2a4650-f233-4da7-a0b6-7a32f078144c · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.279674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.279674Z digest=sha256:4533de94655f748620d5502b4a32d06d04b867a6078a030aa5c1077a922cb195

Observation b64f8235-083e-4ee9-b0c7-f7b01fb26ca2 · outbound

This paper cites Hashimoto.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Hashimoto

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.363339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.363339Z digest=sha256:ebab523306f97f38cdcc183b27b19fed7bf659b56ff98034d3cbb85400f1ad47

Observation 3469f3d8-c65e-4e37-bb2e-5e0afe358758 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

Beyond the Surface: Measuring Self-Preference in LLM Judgments The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.503845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.503845Z digest=sha256:ec5efe359a3b966f017bd9c43436d60fa21d05ed94238871faf29342e3105b3d

Observation 59248575-72ef-42be-9662-9c1db09f8e83 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Beyond the Surface: Measuring Self-Preference in LLM Judgments TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.592298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.592298Z digest=sha256:728673494cf2586a8c860d0ac49a06719edb6706685ed2d2f3b609a8e924039e

Observation 42b79dbf-0220-4ead-866c-e13a1984be0e · outbound

This paper cites DeepSeek-V3 Technical Report.

Beyond the Surface: Measuring Self-Preference in LLM Judgments DeepSeek-V3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.730166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.730166Z digest=sha256:cfe9595b8f65df08513cceccd4cf6afca7994c6d0292b198a62ddee7ed8dad4f

Observation df62893e-d5f8-4916-a8f8-3c71b13c46a2 · outbound

This paper cites AlignBench: Benchmarking Chinese Alignment of Large Language Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.837491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.837491Z digest=sha256:cac1ea1290d82fe4f580bfcd8e00e6d74642b74e606d25216f57c898a037d3e6

Observation c5a216dd-79c7-4ed1-89f9-9c84f203cf38 · outbound

This paper cites LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores.

Beyond the Surface: Measuring Self-Preference in LLM Judgments LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.946179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.946179Z digest=sha256:236e4a7159e45578a7eebec9e2bcaf3e69386d17033deb40925d771f6f6739e6

Observation 2652d21b-ea36-4892-a7e1-769f3ea96b18 · outbound

This paper cites Evaluating Style Transfer for Text.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Evaluating Style Transfer for Text

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:07.976111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:26:05.025643Z digest=sha256:bb2465c3bf49af0583594574c4f39132fce405958418e2e169568b1b0c1277d2

Observation 250c8b4a-aeea-47ce-bf7a-78d8d52a83fd · outbound

This paper cites Text Style Transfer Evaluation Using Large Language Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Text Style Transfer Evaluation Using Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:07.833592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:26:05.130234Z digest=sha256:b388430ef4364734d6523ec9ae911b3000d74784d6c0b32ab4b6884c270761ac

Observation 846299db-4fcb-47a2-9645-7711fc940f2d · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:08.558455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:26:05.237736Z digest=sha256:7054da0be790933b1e95335b1236d9f0c73617786652596e122233c3dd3d6ac1

Observation be9bacc0-b406-498a-a6ab-cd665dc04a99 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.347061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.347061Z digest=sha256:7467282181709f95e8b3519858fb9665ced9465cbb52d8ae9ab47adae71887cf

Observation e2bc21e0-b692-4fb5-8259-61f7dd50d2a5 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Beyond the Surface: Measuring Self-Preference in LLM Judgments ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.458126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.458126Z digest=sha256:39526953b4129a3a24a6aaa0743cf7c322544eb1ab895819cd29486ee6754da6

Observation 6d918236-3f33-4060-8d64-fac47a29f405 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.547451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.547451Z digest=sha256:3abf0ad4eb25acbc5dfaca204c045b083b58d36028ea6b39b143bcacafe0f384

Observation 7b650efc-f0e8-4153-9842-6765752cf9c5 · outbound

This paper cites SALMON: Self-Alignment with Instructable Reward Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments SALMON: Self-Alignment with Instructable Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.656399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.656399Z digest=sha256:136259426451db42944ff5bb5e30b662657c536a74a79756864b972584657150

Observation f7d88fb0-a0cb-4aa9-b9e8-82264ccdc5cf · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.795023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.795023Z digest=sha256:d15afdd6ab662b85b3100dc2c279384e27888f920523ce22bcda269214a35474

Observation d7a6a6c4-1e8f-4520-9ee9-fd9191237295 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Gemma 2: Improving Open Language Models at a Practical Size

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.913103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.913103Z digest=sha256:f27a2ac6e3f91b19f87b8830933f948bdf7ae9fec8cdfbddaa93937785faf23d

Observation 3f66f203-92c3-4d48-bd0f-961683027ee7 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.015237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.015237Z digest=sha256:e298ca4065e4945ab42abdadf1ffd1dbee31781225d0819d5dc4c0d8bfd1bda7

Observation e79cbff3-326f-42d4-b14f-0f10c3b4e1bd · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.158196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.158196Z digest=sha256:ee8ae467cbe3da5ec0252eccbe73fe6969632db0d91f1097f2f72a60d0c0aae3

Observation 962f9a8b-8e50-429b-93dd-db8482425c0d · outbound

This paper cites Large Language Models are not Fair Evaluators.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Large Language Models are not Fair Evaluators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.235975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.235975Z digest=sha256:f7b48bc638dff1b1b841123a48bb92d1ae1df4752725257280d0238bb4e70740

Observation ce379c9d-1747-46e9-9e0e-bbdd090ed7be · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

Beyond the Surface: Measuring Self-Preference in LLM Judgments PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.346963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.346963Z digest=sha256:fc98dc7d25f4d90ba72c5b8590ee1ce607d0ab9a173987e1dca2f10a1434f4b5

Observation 2d3a899b-48da-4516-b36b-31962dfbab41 · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.467953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.467953Z digest=sha256:9f4bd2f76d4be06206a5e8bd7e994a909f29356c3b65a0df8f11642a3ba63775

Observation ed677fea-c2ee-4a76-8c26-83b2f27c212e · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Self-Preference Bias in LLM-as-a-Judge

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.576445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.576445Z digest=sha256:279cdd3d6ca4a39357c87392b859866d2a71eb59a54790144bd9ad4f000c013d

Observation e66676d5-37b4-4900-9a0c-5c695b23be66 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.652545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.652545Z digest=sha256:e9c9e36a0f67c79e5c8a84eea371f4f4dec175374fd992291caae97cb54f0562

Observation ac304e2e-cc51-4890-80b2-814a8191dc1f · outbound

This paper cites Evaluating Mathematical Reasoning Beyond Accuracy.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Evaluating Mathematical Reasoning Beyond Accuracy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.792641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.792641Z digest=sha256:f261b6c04f2bff5ba9f5082032a199e98e089eb85978fe0df5948ee1fed05e1a

Observation 7987081c-db41-4372-a1f6-16cf1e2027c9 · outbound

This paper cites Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.902686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.902686Z digest=sha256:52e444236008d9076f9c14f2c45dd905cf6568578535372bfe0d4ddd949f2745

Observation 8e1b10f2-16c8-40f6-ab06-84fa8b35661b · outbound

This paper cites Qwen2.5 Technical Report.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Qwen2.5 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.980884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.980884Z digest=sha256:efe758d6f312014ee01e93e5f66e75158787d8b7919582c193fc8bd82fe3a980

Observation 63155177-c41f-45d0-9a61-73881ae447db · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.054479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.054479Z digest=sha256:5423d6474b9fe6c1c654a66271070a015ecbccf005639e0653137e30e838e94b

Observation 873dcaaf-0b14-48e5-94f7-3d2ec62cc584 · outbound

This paper cites mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval.

Beyond the Surface: Measuring Self-Preference in LLM Judgments mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.174272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.174272Z digest=sha256:ee2d88c651410e813661bf88b92c4677c1a9cb3a4cdefac5e9aa28c7cce8f638

Observation ffef1924-14c8-4f42-ac8b-2302eeb5382b · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.266045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.266045Z digest=sha256:add356b6c208b88b8d3dad8fc2ef2936700be3410de69cb3914a4dc950bbc01d

Observation ec679e26-bc0e-4ae3-b46e-2637423e0d94 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Beyond the Surface: Measuring Self-Preference in LLM Judgments JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.442164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.442164Z digest=sha256:0d987aee5e4648990c9e13f448eb94092018df3db9353870a0a8f12489397b7e

Pith citing papers

Observation 01dc12c8-47bc-4c49-b235-192caa7839ca · inbound

Extreme Self-Preference in Language Models cites this paper.

Extreme Self-Preference in Language Models Beyond the Surface: Measuring Self-Preference in LLM Judgments

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.086277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T20:57:37.199128Z digest=sha256:13e279f9185ca2bcafddd89804c2715f8276a6135cb4ddcacc727594f8b7f376

Observation 6ad59c7e-f8f6-4c67-bea4-dd910d61e0f8 · inbound

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning cites this paper.

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning Beyond the Surface: Measuring Self-Preference in LLM Judgments

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:51:08.567479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T08:50:53.486588Z digest=sha256:fdc6b1eb65b98d1e5e9867f24f10c163d7504bcd98013896b4c81eabd807c8d6

Observation a4352c5a-0e5b-440b-af5c-2a8cb2ba9922 · inbound

Memory Reward Inflation in Self-Improving LLM Agents cites this paper.

Memory Reward Inflation in Self-Improving LLM Agents Beyond the Surface: Measuring Self-Preference in LLM Judgments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T02:16:42.324571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:16:42.324571Z digest=sha256:9e0426c39e04eded15a1eeabb6a107710d978821eea1e27277cfc7c3896b93de