Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:04:26.439792Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 28 inbound Pith citation observations for arXiv:2412.12509.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:04:26.439792Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:13.015050Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
76 of 76 outbound references displayed
External citation measurements
5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 3f65fe95-a1e8-4c60-8f6d-1fc02984aa06 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d3aa71-6a72-48bc-b351-afa0ce234c9e · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d19240-2d42-4122-849d-e5f8cdc2786e · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4bef4719-b426-4054-937f-bdbb38fcd748 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41426fd2-e2b3-4197-b50d-90e2d36d5ff8 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e6156c2-f345-41ae-b409-d593b0b58985 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6362d61-155c-4f5e-9f29-290a423d64a7 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f02719c-6832-491d-a57d-29855bc180eb · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Non-Determinism of "Deterministic" LLM Settings
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1adfb48-cc4c-4e71-80ad-747c2372d03a · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fde6601-e8d4-4fc3-bcec-7c9938f941ef · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4861a1e4-1e44-48e4-ab92-ba860b165a52 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 80ce023a-ef11-4ca8-8e30-87d9360f1eb9 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c916ac0d-21f9-4e97-8897-5f766b526d9b · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Assessing the Usability of GutGPT: A Simulation Study of an AI Clinical Decision Support System for Gastrointestinal Bleeding Risk
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2054bfaf-0f8b-4756-894b-d4bd25d69240 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb96bfe0-9aa6-4acb-bb98-aa98ae1bacc1 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Gonzalez, Ion Stoica, and Eric P
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 94e3c324-8b82-45bd-a795-e27249e5d1f9 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ec95df-747f-4fcc-b3e7-1723081ceda7 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8b96f72-bade-49f0-a319-8d6cb5e25295 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ab7738a-0efc-44f9-9493-af9423cedcf6 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e63ec38-fe52-4dd3-a39a-f29c82bacce5 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6673ea66-7798-4655-bfc0-407cf3d8bd11 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 32ef8beb-ea30-491e-99d8-69db21ba1f55 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Can LLM be a Personalized Judge?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d907131-1e26-4774-9a01-20fc7cba4637 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge The Llama 3 Herd of Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a4802b6-40ca-4f5d-b723-e812421a0c75 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3ae35a0a-9323-4ea5-939d-15a4463e99c2 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87bf65d0-6130-4589-9d76-bef9543a5f93 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5c2fe765-0a6b-4ddb-a64d-ffc1ed84a54c · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4a90643b-6192-42ea-907d-561a7dc7dc56 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15610268-167c-47da-82a9-a06be1f82aa4 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge An Automated Survey of Generative Artificial Intelligence: Large Language Models, Architectures, Protocols, and Applications
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671e9d4c-0a45-4a95-9e0e-be46453e2b4b · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge A Survey on LLM-as-a-Judge
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c730290-85a6-4d6d-8377-38ebd4d3cef9 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b307189-afcc-4d1b-ac38-54c038d2c24c · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da6c8035-5a83-4618-9483-00e210adf2b3 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ae659450-e93f-4197-b151-e955b7fa210a · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 22f1974d-5748-4616-ac23-cca8accc6f90 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7760dc1e-5d28-40ab-aa8d-f3cd65ee9f3c · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a82f31fe-cce2-40bf-83c8-56b117a486a6 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 75a55387-d9f1-4a56-bac7-87f3cdf0b900 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Methods to Estimate Large Language Model Confidence
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ec6c66c8-414b-4f79-ade6-dd1f91edc147 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a6396b-fbfc-4f06-b44d-e5a465b8ec51 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dbe0a34-d9b6-4d16-9108-b80186e3b1d3 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aec1faed-113e-4995-8e1a-edf8d0830b27 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge MDDial: A Multi-turn Differential Diagnosis Dialogue Dataset with Reliability Evaluation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781bfa25-0dd7-4dfe-985f-6c7838c5138a · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 940d2f36-2541-4bc5-9e0f-6ea756371786 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5a74f683-6d3a-4d8a-898b-c067183023d4 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98a39769-0f69-4cf9-880f-10983920de19 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c506b427-5351-411f-8261-9cfecc8d43f6 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge An Empirical Study of the Non-determinism of ChatGPT in Code Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4362e961-78ee-4b5b-b255-ab6a5980d808 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8a3632-cf28-4f6d-9697-2b3ff880143e · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Is Temperature the Creativity Parameter of Large Language Models?
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60797a6-bcfb-4b7f-b150-b3c021014c23 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Semantic Consistency for Assuring Reliability of Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3ee178-71ad-4913-b670-d03415d4f7aa · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e280c4-1b20-4423-8ec6-11e80aa6d3c3 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge o rg Rahnenf \
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c74ba101-823b-43c2-b737-d23881e9d5a4 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Reliability of Topic Modeling
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d291bcef-2cca-4f2b-9e64-db3d5928605b · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c68cc77f-6828-40e6-90be-145bd210fa1c · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 139d63d6-5156-4b45-8700-4a9b20864a75 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Calibration and Correctness of Language Models for Code
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb82ca0-95f9-4a0a-9e58-3d020e6017f4 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f03fe839-5613-432f-8a38-d672afe20cbb · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66531181-c97b-45de-982d-4bf7dd1faef3 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63fe2f30-9075-4174-899b-f79055ea0f04 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22561902-6e91-4aa9-8bfc-90cf28d2f792 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22003b49-d8ac-462a-b0a2-2cf8fd2e3865 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Gemma: Open Models Based on Gemini Research and Technology
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d77c40a-0fe7-4716-b6ef-1813d5f90c6b · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Gemma: Open Models Based on Gemini Research and Technology
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf62d123-8a6a-45f8-bc30-bad822d73584 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2a4333-b396-4e20-a8d9-b810a0a200de · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b18f4ecd-bf24-47b8-9574-3313ab95869c · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cf5d044e-83cb-4661-a71e-71100984f6ca · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 717fec84-1dc1-4046-a27b-d65922246322 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6387e8d1-da56-4929-aa61-ee1e91069c33 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 03db526d-ceeb-49ff-8a7e-66ad750c6e8c · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8c3565-afbc-41c7-b394-1cc139df6780 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c8704a-dea7-4808-8826-9404f9201b22 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cdf915e-f171-4b0c-be54-2b0e2523af36 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9820e8db-02d7-47a9-9c05-d3fa32b2a478 · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dbe090d6-eaaf-4828-8b46-f8c6891c72be · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3410c44b-ba9f-4869-8e01-1d563eaff84e · outbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7aa2c22f-8847-4e96-a123-cca44ace6ce4 · inbound
Reasoning LLMs are Wandering Solution Explorers Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377d08fb-1e9b-42ab-91c7-0b051cccca8d · inbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89508a75-c5c4-4b99-a2e0-b798d2b9cdeb · inbound
RecGPT Technical Report Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be00877a-4412-41e5-8544-11daeb8df5d2 · inbound
Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f0165a7-90cd-4660-b6ab-46638cbf535c · inbound
Guidelines for Empirical Studies in Software Engineering involving Large Language Models Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1a1fe4f8-e899-43c2-91dd-807fc59c67b1 · inbound
Guidelines for Empirical Studies in Software Engineering involving Large Language Models Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5f0f1c80-3e02-483f-889d-98728815c6f1 · inbound
IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 210cc5ad-3c76-4b9c-acb3-9ff08272e63f · inbound
SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf681782-f8ee-4fa3-84e2-41dc5df3ab22 · inbound
Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9a234b33-9b17-4179-b58a-432ce1af1df6 · inbound
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b2143566-bc8f-46b0-babc-aadb24da9869 · inbound
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9e4c2215-202a-4f8d-9dc3-e3495f1de088 · inbound
Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a506623b-98f4-4159-9a21-1e3775c73ae6 · inbound
Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9bc9f2bf-31e1-444d-9fd4-8a6c4a14de61 · inbound
Mixed response geometry and critical crossover in the Ising model Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbbfbae8-c536-4762-84c9-9b5bbcfecc41 · inbound
Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e9fd0f6-791c-4bac-80d2-3532eb091302 · inbound
A Systematic Investigation of RL-Jailbreaking in LLMs Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f95f96aa-bc93-4936-935e-bf0f03a6df6e · inbound
A Systematic Investigation of RL-Jailbreaking in LLMs Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 073fcf87-a936-4e32-8649-d782e3a5e264 · inbound
A Systematic Investigation of RL-Jailbreaking in LLMs Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa0f76a-a8f6-4f10-95b4-4040b442bef9 · inbound
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c054c141-fd1e-42c2-b791-d2cc8bdb2c31 · inbound
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2a5a7f47-d2ca-40f3-8387-c5a8392c2013 · inbound
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d6cb879-eb2b-49f8-b934-f36a4d892ec7 · inbound
A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation df36f15d-1dab-42b4-98e9-0144a15dea3e · inbound
ComplexConstraints and Beyond: Expert Rubrics for RLVR Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fc3ef303-5c24-48e1-a078-e054506492d3 · inbound
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 58858e04-0d09-4ed6-8eb9-fe3cb0179867 · inbound
Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7de55c08-2825-4751-9ff4-6e4fcd5c413a · inbound
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c7724a-5157-4c4d-a965-2f801e888550 · inbound
Codifying the Judge: Scalable Evaluation via Program Distillation Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff66f522-10ee-44a2-83a0-0e9740e01355 · inbound
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.