Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:07.442164Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2506.02592.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:07.442164Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T02:16:42.324571Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T21:00:39.084405Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1aab9281-65d6-486d-8ba9-3327a25e6110 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916ecf34-c6b3-4acd-b948-635d662e7df4 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3024f0e2-9035-4cba-ad29-7d3a29d9bcb5 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fcc45e8a-b65a-47b3-949a-abdeb19eecd6 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e93e67a4-bb96-4958-b89f-6ac226594aba · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ef2e73-9a5d-487e-ac02-bcb0f2e3ec14 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a704e88-0a57-490d-aa55-a99138c00bb1 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e04843-4344-4b95-9591-919fa420439a · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfd26f89-e201-463e-94c4-0b7f99aec333 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df786909-59ae-4eb3-a216-77e507584322 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511d3411-70d5-4f68-85cd-17437e3cfd96 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f45ef9-147b-48e1-9b2f-0a96f1876e88 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9031e36a-c68a-45b5-8448-46253566c8a1 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d0dee2f-f657-4628-8a3d-c5f7d77129dd · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6611b524-41ca-4621-94f6-bd82bdd9031d · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2849b9fb-d3d2-4036-9aa9-29a6a1d57b8f · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99aba7be-d3b8-49e9-a7fd-c8ef588971ae · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c09241-ecda-4029-b8cc-3e301a904ac1 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Measuring Mathematical Problem Solving With the MATH Dataset
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb0e45e-169b-42a5-9970-6a79fd6a0456 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 568c9826-03d8-4c44-94b8-b56c6042850c · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments GPT-4o System Card
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa632138-9ef8-49ee-a22b-2b698d4f6bf8 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments OpenAI o1 System Card
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f4d8d5a-931a-4bb6-81bc-60471a111347 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Benchmarking Cognitive Biases in Large Language Models as Evaluators
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36f4cae-64be-45db-9f0f-1e7b3cfff12a · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments RewardBench: Evaluating Reward Models for Language Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a151bb-abec-4435-91ae-e2868aaa889b · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9193983d-d82f-418e-b27d-2ebefab0ee75 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe2a4650-f233-4da7-a0b6-7a32f078144c · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64f8235-083e-4ee9-b0c7-f7b01fb26ca2 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Hashimoto
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3469f3d8-c65e-4e37-bb2e-5e0afe358758 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59248575-72ef-42be-9662-9c1db09f8e83 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b79dbf-0220-4ead-866c-e13a1984be0e · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments DeepSeek-V3 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df62893e-d5f8-4916-a8f8-3c71b13c46a2 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments AlignBench: Benchmarking Chinese Alignment of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a216dd-79c7-4ed1-89f9-9c84f203cf38 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2652d21b-ea36-4892-a7e1-769f3ea96b18 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Evaluating Style Transfer for Text
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 250c8b4a-aeea-47ce-bf7a-78d8d52a83fd · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Text Style Transfer Evaluation Using Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 846299db-4fcb-47a2-9645-7711fc940f2d · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be9bacc0-b406-498a-a6ab-cd665dc04a99 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2bc21e0-b692-4fb5-8259-61f7dd50d2a5 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d918236-3f33-4060-8d64-fac47a29f405 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b650efc-f0e8-4153-9842-6765752cf9c5 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments SALMON: Self-Alignment with Instructable Reward Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d88fb0-a0cb-4aa9-b9e8-82264ccdc5cf · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a6a6c4-1e8f-4520-9ee9-fd9191237295 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Gemma 2: Improving Open Language Models at a Practical Size
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f66f203-92c3-4d48-bd0f-961683027ee7 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e79cbff3-326f-42d4-b14f-0f10c3b4e1bd · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962f9a8b-8e50-429b-93dd-db8482425c0d · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Large Language Models are not Fair Evaluators
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce379c9d-1747-46e9-9e0e-bbdd090ed7be · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3a899b-48da-4516-b36b-31962dfbab41 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed677fea-c2ee-4a76-8c26-83b2f27c212e · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Self-Preference Bias in LLM-as-a-Judge
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e66676d5-37b4-4900-9a0c-5c695b23be66 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac304e2e-cc51-4890-80b2-814a8191dc1f · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Evaluating Mathematical Reasoning Beyond Accuracy
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7987081c-db41-4372-a1f6-16cf1e2027c9 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1b10f2-16c8-40f6-ab06-84fa8b35661b · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Qwen2.5 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63155177-c41f-45d0-9a61-73881ae447db · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873dcaaf-0b14-48e5-94f7-3d2ec62cc584 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffef1924-14c8-4f42-ac8b-2302eeb5382b · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec679e26-bc0e-4ae3-b46e-2637423e0d94 · outbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01dc12c8-47bc-4c49-b235-192caa7839ca · inbound
Extreme Self-Preference in Language Models Beyond the Surface: Measuring Self-Preference in LLM Judgments
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ad59c7e-f8f6-4c67-bea4-dd910d61e0f8 · inbound
When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning Beyond the Surface: Measuring Self-Preference in LLM Judgments
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4352c5a-0e5b-440b-af5c-2a8cb2ba9922 · inbound
Memory Reward Inflation in Self-Improving LLM Agents Beyond the Surface: Measuring Self-Preference in LLM Judgments
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.