Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

As of 23 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 10 inbound Pith citation observations for arXiv:2506.13639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13639 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:37.128884Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:57:42.489401Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4ac6696d-1805-40bd-9dcd-954d938735fd · outbound

This paper cites InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 8301–8327.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 8301–8327

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:38.221562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:31:34.852140Z digest=sha256:1e499004ebca964c0a29f1d4645a50cea99df53331d10532c50c503657347f36

Observation 3342a1ff-0da4-4c63-886e-f002b45b3315 · outbound

This paper cites DeepSeek-V3 Technical Report.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability DeepSeek-V3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:34.925180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:34.925180Z digest=sha256:700155bef540c5efefe1e698bfd77dd6ac73e3388c3c3e78f1c506591dacc2c4

Observation f636348e-0aa3-4358-9ac0-529ed93c7a90 · outbound

This paper cites Finding Blind Spots in Evaluator LLMs with Interpretable Checklists.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.022146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.022146Z digest=sha256:5b84ac5cc17fc3c62f424cabe9ca0d6f21941cc5df9c76044b321babff4e34f8

Observation 8e604453-680f-4eaa-a9fb-b77ef8f92ea3 · outbound

This paper cites The Llama 3 Herd of Models.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.146627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.146627Z digest=sha256:e86dc29fd3cf92c03916a21e39157f0e2bd4170e5873a4631ff8489c2024261e

Observation e29f9c5b-0188-49b5-bae4-ae976161266a · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.214458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.214458Z digest=sha256:0ab122f39b236e06136bbcc51f142a89c6ec89ca367e4e58d9edeaaa3b545d59

Observation dbc8685d-4bd6-4dff-b760-dfcb40673684 · outbound

This paper cites A Survey on LLM-as-a-Judge.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability A Survey on LLM-as-a-Judge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.273814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.273814Z digest=sha256:872938f49ae5d5186785d726ec528a6d8e7375a3c410a9ccb7f3cbbd7662f17e

Observation 3d2d45e3-c5f7-4aaa-af0a-d45d682d4127 · outbound

This paper cites Mixtral of Experts.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Mixtral of Experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.408153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.408153Z digest=sha256:fc7dac37a76058b4b2ad140ce4bc08a8c3e4e26518f15cd0afc024191d017d99

Observation 4ccd1d5c-745c-4c72-a8c4-a1c8588a3896 · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.477675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.477675Z digest=sha256:84ae714d587fbb05097e47d9cc5a1bd6765715bf8aebdda03fd2d3f336d0bbb4

Observation 840b7fc9-1e0e-4dcc-99f5-49bfd2af4791 · outbound

This paper cites The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.564117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.564117Z digest=sha256:ed241358ceb62c809f7849839b835bac0cbbd83cfe18dbb149e2388de2ffa67d

Observation c6e41943-29f7-41ca-b7c5-9d96d54aeec3 · outbound

This paper cites GPT-4 Technical Report.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.904893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.904893Z digest=sha256:36dc02cdbe6082976f1cc85bcf02dbd5e16b179532c234d964084585e79d8c12

Observation 377d08fb-1e9b-42ab-91c7-0b051cccca8d · outbound

This paper cites Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.166184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.166184Z digest=sha256:3679508bb98b760435bf73c9fb9e1e8b5a296cb0a17c1fefabac685b706cccc6

Observation e2a5c216-4747-45e8-91ec-72ca80aed3cd · outbound

This paper cites Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.251369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.251369Z digest=sha256:19016eafa928306cd91cbff33ab188196363d8aae78b9e3f551358cc6052024f

Observation 8e8ff90b-378a-49fc-a3c1-5e875cc9bbc9 · outbound

This paper cites InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5,.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:37.916376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:31:36.563627Z digest=sha256:5bd72c0c32781ce3e913a4e915272215fef301419a49bea0948d89509f068a8f

Observation 60835722-30b9-4029-95ec-185286065f5e · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.601561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.601561Z digest=sha256:27947c842a9a28456bc71bf28abca8784735102e2ccbae7b15be669040b88fd5

Observation 631e4c20-e8c3-48bd-9989-3bcb3f7350d9 · outbound

This paper cites Qwen2.5 Technical Report.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.692401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.692401Z digest=sha256:6ba7ea0dad3f32013de445057dafdf19541f7366ebf0f3390e306996dce13d85

Observation 231646de-e959-454f-8438-dadbef107fb7 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.795367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.795367Z digest=sha256:04a66db233d16a3858a12b4c63140f60a20ac639b4e53725563ae09351aca306

Observation 378b0edc-7941-49ea-9496-d35e4823bf24 · outbound

This paper cites InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:37.729676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:31:36.882454Z digest=sha256:1ab92a20b7933e6986d325771367f36b52583d6585a951a640ed6b9bfe89d2b1

Observation 95a9b814-a5ff-43fc-8d30-8d1b280ba898 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability BERTScore: Evaluating Text Generation with BERT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.976720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.976720Z digest=sha256:ac86e0e81c6622ec7db61856c8d9e7033af5a0754d7d852a7d1b3ea00b0808ff

Observation d6713a3b-c396-46e6-8795-96fe2f5bc088 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:37.043951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:37.043951Z digest=sha256:330ca0944ee212dc60019dfedf2b6d80517f65cb3d67346ba9ad2d303da687df

Observation 9a2a6f2e-a85c-4faa-95fe-8e40091b9e68 · outbound

This paper cites an unresolved cited work.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:37.558801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:31:37.128884Z digest=sha256:58ee06911a1b90a4697348486041dce065b716df0a22ba35e44253128c8cd8b2

Observation c125dec4-740c-4de0-b143-a7f6f052dd4f · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.029108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.029108Z digest=sha256:28a2c0acc23d6bf00e23507a44e281201f9bfa0df3ee94ff07b6169cdd164195

Observation 0c62f108-f991-4eef-a9d4-91eaa0926d9f · outbound

This paper cites Goal-Oriented Prompt Attack and Safety Evaluation for LLMs.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Goal-Oriented Prompt Attack and Safety Evaluation for LLMs

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.732858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.732858Z digest=sha256:de019a87ca8ee28b5e140e84c257a33c2aa034c02accff93d9ff7ccbb85d4f84

Observation 65da1aa6-5c74-41db-969d-19ad9127f48d · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.647892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.647892Z digest=sha256:c9202dfbb519bbf52cd5079cafcd139f6359ef7b563ae0a8746e7bd7dfe550db

Observation 92140abe-0aa6-4d6b-88e3-032f50a74885 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability CIDEr: Consensus-based Image Description Evaluation

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.490899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.490899Z digest=sha256:f1a54216a3fcd15929d6cd4f23ec318a464e3ab84dfea612251ae6744a34b265

Observation f4b84ba2-f15a-4f92-8c20-7ce6e8b4046a · outbound

This paper cites COMET: A Neural Framework for MT Evaluation.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability COMET: A Neural Framework for MT Evaluation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.112514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.112514Z digest=sha256:ef1407f4da2fb397ea86118447c0f4c4398a468f46a526475595244b0003a6a7

Observation 6a031a9b-11da-440f-91ab-11ddad9dbc81 · outbound

This paper cites InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Sys- tems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16,.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Sys- tems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16,

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:38.083315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:31:35.832802Z digest=sha256:2a60fdcd02516ef3dc083697bc4d0adf34c39f3b3b4a04c16036c4fba167b142

Observation e54480e0-3d9a-48ea-bd97-8f6f87743c9c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:34.708896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:34.708896Z digest=sha256:99bc95bcb7d8256fab4d55fb5f20faa6e6409a5c3c04f77501361126635cb113

Observation dbccbcef-a0e0-4a25-a3d0-e0ab3cfabc98 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.353198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.353198Z digest=sha256:f2b721ee2c6173605a270233fd9e27c5e59c26c489665f659e808d7fdbf56b33

Pith citing papers

Observation 4abf00ce-8520-4104-89e0-460594eeca8c · inbound

Evaluating LLM Agent Collusion in Double Auctions cites this paper.

Evaluating LLM Agent Collusion in Double Auctions An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:42.489401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:57:42.489401Z digest=sha256:bb4c47ed3adea74af183ed1edc93ec386aa1dfcca4e870fb321ba062f9efb202

Observation 5181625e-f7a5-4a69-9903-c0fb3420f29b · inbound

Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence cites this paper.

Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:50.983413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:50.983413Z digest=sha256:71005f2644f7a4fe00bbb368c49b7ff5c87e97328a68b3cc5f7b3fb2a0d64946

Observation 94c8a979-7c5b-4517-b045-2a493b8cc496 · inbound

A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability cites this paper.

A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:19.529541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T03:03:05.715652Z digest=sha256:c89657ebb399aec18b1f828a5a3310f458f2c1c0be6861adb9170b3ce0fe7011

Observation 9afa4fa9-d2fd-44cf-87c1-d72498563e7b · inbound

Multi-Turn Neural Transparency: Surfacing Neural Activations Improves User Calibration to LLM Behavioral Drift cites this paper.

Multi-Turn Neural Transparency: Surfacing Neural Activations Improves User Calibration to LLM Behavioral Drift An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:42:37.636491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T14:37:45.304949Z digest=sha256:4bd49213c621895cad654f90745292e151ccff78f4b2974c29fa522dd64b8793

Observation 52999629-05d3-480f-a99f-0b9e298fcd31 · inbound

Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making cites this paper.

Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:14:45.956409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T14:04:51.856877Z digest=sha256:6a33c69ad2a3dd8471df2c7a242780fadfb5717d72f0ed4f1d9771770036a1b8

Observation 6db26b5f-812e-42eb-a53b-e31c082d7b5c · inbound

Learning from Mistakes: Can LLM Self-Recover after Misalignment? cites this paper.

Learning from Mistakes: Can LLM Self-Recover after Misalignment? An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T18:51:10.298187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:51:10.298187Z digest=sha256:71ba555e77ddfa83625ffab62a1e34698c8077fd243fb5b2e755b47616949b99

Observation f82759b6-9b63-4e90-9e73-a73c515e77f3 · inbound

Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework cites this paper.

Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:29.477423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T18:21:57.096578Z digest=sha256:95949308d3cc7530792f8c7c3ab6e0df00f111bfc7e74a67b7070f3bd5f264e3

Observation 1b36da1e-73a1-4f21-8e99-62204d41cadd · inbound

ComplexConstraints and Beyond: Expert Rubrics for RLVR cites this paper.

ComplexConstraints and Beyond: Expert Rubrics for RLVR An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:31.387512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:37:11.141846Z digest=sha256:4b3b277fe43db50749875b97eb2697352c3a98dc0a385877e1d9205cd9370718

Observation e7d8b4d4-a882-4740-bdaa-f2bdb818f1cb · inbound

Are LLMs Bad at Moral Reasoning? cites this paper.

Are LLMs Bad at Moral Reasoning? An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.618678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T08:20:24.251540Z digest=sha256:19bdc5ec7b178cdb09fc1b802f5ecacd25704de1a0d584b6cec63ad2dbe6466c

Observation 2037b41f-1d8a-41fa-9a14-187003592721 · inbound

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability cites this paper.

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:56:50.477026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T05:47:32.670216Z digest=sha256:1e14b32513897cf736dde1098f02340cc1d6bfc34a21babaa230eb51e4882d82