Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:37.128884Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 10 inbound Pith citation observations for arXiv:2506.13639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:37.128884Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:57:42.489401Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
28 of 28 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
Observation 4ac6696d-1805-40bd-9dcd-954d938735fd · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 8301–8327
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3342a1ff-0da4-4c63-886e-f002b45b3315 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability DeepSeek-V3 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f636348e-0aa3-4358-9ac0-529ed93c7a90 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e604453-680f-4eaa-a9fb-b77ef8f92ea3 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e29f9c5b-0188-49b5-bae4-ae976161266a · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc8685d-4bd6-4dff-b760-dfcb40673684 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability A Survey on LLM-as-a-Judge
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2d45e3-c5f7-4aaa-af0a-d45d682d4127 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Mixtral of Experts
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ccd1d5c-745c-4c72-a8c4-a1c8588a3896 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 840b7fc9-1e0e-4dcc-99f5-49bfd2af4791 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e41943-29f7-41ca-b7c5-9d96d54aeec3 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377d08fb-1e9b-42ab-91c7-0b051cccca8d · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a5c216-4747-45e8-91ec-72ca80aed3cd · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8ff90b-378a-49fc-a3c1-5e875cc9bbc9 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 60835722-30b9-4029-95ec-185286065f5e · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 631e4c20-e8c3-48bd-9989-3bcb3f7350d9 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Qwen2.5 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 231646de-e959-454f-8438-dadbef107fb7 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378b0edc-7941-49ea-9496-d35e4823bf24 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 95a9b814-a5ff-43fc-8d30-8d1b280ba898 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability BERTScore: Evaluating Text Generation with BERT
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6713a3b-c396-46e6-8795-96fe2f5bc088 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2a6f2e-a85c-4faa-95fe-8e40091b9e68 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c125dec4-740c-4de0-b143-a7f6f052dd4f · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c62f108-f991-4eef-a9d4-91eaa0926d9f · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Goal-Oriented Prompt Attack and Safety Evaluation for LLMs
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65da1aa6-5c74-41db-969d-19ad9127f48d · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92140abe-0aa6-4d6b-88e3-032f50a74885 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability CIDEr: Consensus-based Image Description Evaluation
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4b84ba2-f15a-4f92-8c20-7ce6e8b4046a · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability COMET: A Neural Framework for MT Evaluation
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a031a9b-11da-440f-91ab-11ddad9dbc81 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Sys- tems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16,
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e54480e0-3d9a-48ea-bd97-8f6f87743c9c · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbccbcef-a0e0-4a25-a3d0-e0ab3cfabc98 · outbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4abf00ce-8520-4104-89e0-460594eeca8c · inbound
Evaluating LLM Agent Collusion in Double Auctions An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5181625e-f7a5-4a69-9903-c0fb3420f29b · inbound
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c8a979-7c5b-4517-b045-2a493b8cc496 · inbound
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9afa4fa9-d2fd-44cf-87c1-d72498563e7b · inbound
Multi-Turn Neural Transparency: Surfacing Neural Activations Improves User Calibration to LLM Behavioral Drift An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 52999629-05d3-480f-a99f-0b9e298fcd31 · inbound
Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6db26b5f-812e-42eb-a53b-e31c082d7b5c · inbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f82759b6-9b63-4e90-9e73-a73c515e77f3 · inbound
Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1b36da1e-73a1-4f21-8e99-62204d41cadd · inbound
ComplexConstraints and Beyond: Expert Rubrics for RLVR An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e7d8b4d4-a882-4740-bdaa-f2bdb818f1cb · inbound
Are LLMs Bad at Moral Reasoning? An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2037b41f-1d8a-41fa-9a14-187003592721 · inbound
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.