Pith. sign in

Paper Citation Record · LEDGER

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

As of 15 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 28 inbound Pith citation observations for arXiv:2412.12509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12509 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:04:26.439792Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:13.015050Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact15
  • verified fuzzy1
  • unresolved57
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3f65fe95-a1e8-4c60-8f6d-1fc02984aa06 · outbound

This paper cites online" 'onlinestring :=.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.070645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.070645Z digest=sha256:5f5ed27866dd589030417d55ef311ead840714da5dc85a59158e4a620d19bd51

Observation b1d3aa71-6a72-48bc-b351-afa0ce234c9e · outbound

This paper cites write newline.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.076166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.076166Z digest=sha256:946f5ee05979c6e8ce11fbe06609f7f35f0dfe2cb95e1f2eaddb5c3a6f53fb42

Observation 29d19240-2d42-4122-849d-e5f8cdc2786e · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:28.228940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.081514Z digest=sha256:47ce8d06ebc603d8b62b7b1d04f5c3d7e83bf21ced033f2d524a38820c9bcc20

Observation 4bef4719-b426-4054-937f-bdbb38fcd748 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.086321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.086321Z digest=sha256:6cac3d4f1000ae9b0dc902c7879912fad81ab942461490178bd4f0afadd12708

Observation 41426fd2-e2b3-4197-b50d-90e2d36d5ff8 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 5

Resolution
verified exact
doi, observed 2026-08-11T14:04:27.502966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.090900Z digest=sha256:d36939efacee17ed6e32fefceeeb3f8171d712d93fc0b16939b54c2ad7bcedaa

Observation 3e6156c2-f345-41ae-b409-d593b0b58985 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.095747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.095747Z digest=sha256:3f13da5d71276988a0660e18c34d8986e848f04bf63d3e47d81c30075d061f32

Observation c6362d61-155c-4f5e-9f29-290a423d64a7 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 7

Resolution
verified exact
doi, observed 2026-08-11T14:04:27.487527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.100071Z digest=sha256:f3e381e58b7deaab1766807bbbc10be3c4836201db2f279bbae56a7d8beda279

Observation 1f02719c-6832-491d-a57d-29855bc180eb · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Non-Determinism of "Deterministic" LLM Settings

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.104221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.104221Z digest=sha256:256c4a3478244f83f6b27676e8d453596a7f31c77f4748a94d4dc38430ece437

Observation d1adfb48-cc4c-4e71-80ad-747c2372d03a · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.109376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.109376Z digest=sha256:48649e85994910359c2f5bc8b13b3dc4ec93483f2e1354e0718e3caa594ec5de

Observation 9fde6601-e8d4-4fc3-bcec-7c9938f941ef · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.471331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.113779Z digest=sha256:563bd1adf885c46622140ac75d02614c1fd0ca95aedb48ad34d38130573d260d

Observation 4861a1e4-1e44-48e4-ab92-ba860b165a52 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 11

Resolution
verified exact
doi, observed 2026-08-11T14:04:27.455835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.118237Z digest=sha256:2e3f520897cd6476e140e31ad0d604b2bb95e6e9df4b3a6de56fd3f77e24f33e

Observation 80ce023a-ef11-4ca8-8e30-87d9360f1eb9 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.122473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.122473Z digest=sha256:fcdc9c335d7452f8c185464570eb0b2b282e3ff74ad2f5ee245ce89eb0b5d310

Observation c916ac0d-21f9-4e97-8897-5f766b526d9b · outbound

This paper cites Assessing the Usability of GutGPT: A Simulation Study of an AI Clinical Decision Support System for Gastrointestinal Bleeding Risk.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Assessing the Usability of GutGPT: A Simulation Study of an AI Clinical Decision Support System for Gastrointestinal Bleeding Risk

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.127287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.127287Z digest=sha256:fa9807280797beff0b107eb9fc443038b5a1b9ee49965d85a644d3e2aad8d135

Observation 2054bfaf-0f8b-4756-894b-d4bd25d69240 · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.132660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.132660Z digest=sha256:2526473e1a039399ee3e3567354b345cd9aa3053e37f3f2f75f4409ca860e2f1

Observation cb96bfe0-9aa6-4acb-bb98-aa98ae1bacc1 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Gonzalez, Ion Stoica, and Eric P

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:28.450953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.138138Z digest=sha256:84f75d21752a328a1c95f8df31d2117d8915350d8370ab3d9f0f206f7b08ca4d

Observation 94e3c324-8b82-45bd-a795-e27249e5d1f9 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.142910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.142910Z digest=sha256:e7f3fabeb100385a81a890fda0acf740544652db4b1b9a64d720f6a73b7e6127

Observation 99ec95df-747f-4fcc-b3e7-1723081ceda7 · outbound

This paper cites Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.148041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.148041Z digest=sha256:0d96e554f366f0092369cab6881dd66238af7c501e85b6f7cf71123840c25292

Observation b8b96f72-bade-49f0-a319-8d6cb5e25295 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.432558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.153434Z digest=sha256:5730e3a48ec6eb5a7de78168fd2bc1b6c2b757e3fb9330f9c3acd7c7f4c7d6e5

Observation 1ab7738a-0efc-44f9-9493-af9423cedcf6 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.158305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.158305Z digest=sha256:8f3860cc62cb3effd83af7e58afea2b8add5fb6a9b8fd2654b9ac996b8982f73

Observation 1e63ec38-fe52-4dd3-a39a-f29c82bacce5 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:04:26.163353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.163353Z digest=sha256:b37f3b9a08e08824c09d511e20b8697b08201fb1de0fb188ed175179531f59de

Observation 6673ea66-7798-4655-bfc0-407cf3d8bd11 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 21

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T14:04:28.025265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.168312Z digest=sha256:0b9a5859af1efc2f602c596c78afb616abad516294c8ef39460d1193732e75ff

Observation 32ef8beb-ea30-491e-99d8-69db21ba1f55 · outbound

This paper cites Can LLM be a Personalized Judge?.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Can LLM be a Personalized Judge?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.172870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.172870Z digest=sha256:6f6f39e78b091a5b3aa8ce8b953320495c5d6394370ed735fef5758e5e78de37

Observation 6d907131-1e26-4774-9a01-20fc7cba4637 · outbound

This paper cites The Llama 3 Herd of Models.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.178088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.178088Z digest=sha256:c226ea4e2d0191644c8e0f5b755f5bda0da4f346ec777081f45575839a43a9a8

Observation 1a4802b6-40ca-4f5d-b723-e812421a0c75 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.414973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.182994Z digest=sha256:1c536541dc12f98ca73281a21513550c60775e3be3d0ca222a4495db9911e779

Observation 3ae35a0a-9323-4ea5-939d-15a4463e99c2 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.187610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.187610Z digest=sha256:adbef02071b053cf8d40663f136568ad19ae10d22ad87c99460d17ccedde08c7

Observation 87bf65d0-6130-4589-9d76-bef9543a5f93 · outbound

This paper cites Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:04:27.283261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.191704Z digest=sha256:6ebfde77594c5ea3aa5982131afb067cb5f620cbe0413761a58f8c2ff0ef9298

Observation 5c2fe765-0a6b-4ddb-a64d-ffc1ed84a54c · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.398876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.196645Z digest=sha256:a3b75ed62fc3a8bd9b89878a08762d0193b0f1f50cdc313dfafd8c348fad5dd5

Observation 4a90643b-6192-42ea-907d-561a7dc7dc56 · outbound

This paper cites COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.201098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.201098Z digest=sha256:bf115eb0d6de36d973bb08e9a84737121efd2ba803d32723cc11293fd48758ae

Observation 15610268-167c-47da-82a9-a06be1f82aa4 · outbound

This paper cites An Automated Survey of Generative Artificial Intelligence: Large Language Models, Architectures, Protocols, and Applications.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge An Automated Survey of Generative Artificial Intelligence: Large Language Models, Architectures, Protocols, and Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.205923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.205923Z digest=sha256:2af75317c5eed290a840daa11997aa88580d9f56f5012b3f3afbce2f8851d364

Observation 671e9d4c-0a45-4a95-9e0e-be46453e2b4b · outbound

This paper cites A Survey on LLM-as-a-Judge.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge A Survey on LLM-as-a-Judge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.210297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.210297Z digest=sha256:1e65a7be25898daf70cafa17c0bfb3845fc0f92973712282f20c793e4d61d646

Observation 3c730290-85a6-4d6d-8377-38ebd4d3cef9 · outbound

This paper cites LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.215060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.215060Z digest=sha256:a39931924f0f9704fc51c319f553e7e9b228d00fd2013c13eedb09fdd9ca8e30

Observation 1b307189-afcc-4d1b-ac38-54c038d2c24c · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.219545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.219545Z digest=sha256:af49e1adc663a8fec780b0b891a7299bc8b7ad1fbff8fb51d7638863c2968402

Observation da6c8035-5a83-4618-9483-00e210adf2b3 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.382703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.224208Z digest=sha256:fb9a21bee8ddc80b558ea6ad331fd20f680de626291a81e66f52572a4a06bb22

Observation ae659450-e93f-4197-b151-e955b7fa210a · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 34

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:27.904959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.228463Z digest=sha256:7c29f0a210e2429f6b4667a98d8f9fc3c4ffaf2eb68e286ee4cbc48b228da0c2

Observation 22f1974d-5748-4616-ac23-cca8accc6f90 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.232737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.232737Z digest=sha256:dc78067882367b3a2627895595c3f4920ff393d92feb4030ee8a4724e41e26af

Observation 7760dc1e-5d28-40ab-aa8d-f3cd65ee9f3c · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.236817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.236817Z digest=sha256:014376d926025b7cfce45afa4f717a3c69254de18c458b81ee1d28206e9655bf

Observation a82f31fe-cce2-40bf-83c8-56b117a486a6 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 37

Resolution
verified exact
doi, observed 2026-08-11T14:04:27.154764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.241768Z digest=sha256:937ff56ecf929b8982613fa21cb8743347c5e22fa5b9dbb77233379ba245c345

Observation 75a55387-d9f1-4a56-bac7-87f3cdf0b900 · outbound

This paper cites Methods to Estimate Large Language Model Confidence.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Methods to Estimate Large Language Model Confidence

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:04:27.140094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.247067Z digest=sha256:872bb75879b9ad66c8e8aba0c1a9651a37d3ab0346db620a62ae87c844c0450d

Observation ec6c66c8-414b-4f79-ade6-dd1f91edc147 · outbound

This paper cites An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.252447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.252447Z digest=sha256:04ff8258201af95a266fed94415467df84eabf642b58d119d372c8454e4600f7

Observation d6a6396b-fbfc-4f06-b44d-e5a465b8ec51 · outbound

This paper cites Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.257770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.257770Z digest=sha256:4c1594f17ec8a29d1f1dbefe90bfdd77a1e4d5112f3ecf2a4a072f670899d731

Observation 3dbe0a34-d9b6-4d16-9108-b80186e3b1d3 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 41

Resolution
verified exact
doi, observed 2026-08-11T14:04:27.080584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.263043Z digest=sha256:0debccebd31eb45d3c91480d7359883a32b8fdfd3bce01453121be7f531bb2e3

Observation aec1faed-113e-4995-8e1a-edf8d0830b27 · outbound

This paper cites MDDial: A Multi-turn Differential Diagnosis Dialogue Dataset with Reliability Evaluation.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge MDDial: A Multi-turn Differential Diagnosis Dialogue Dataset with Reliability Evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.268094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.268094Z digest=sha256:2ad1df1407e473615f971d68bbf4968c426e171a9feadc4f8583ce865e2d051a

Observation 781bfa25-0dd7-4dfe-985f-6c7838c5138a · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 43

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T14:04:27.795134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.273806Z digest=sha256:5b5db4b201d9a0d82287eecc31a02e08b07d7c4a8f8af8976d4081c832e8674a

Observation 940d2f36-2541-4bc5-9e0f-6ea756371786 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.353351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.279031Z digest=sha256:df6f8a49dbf20452266fb2acaa9ba62a8961383ed02b86511bd47ef8ba20ed88

Observation 5a74f683-6d3a-4d8a-898b-c067183023d4 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 45

Resolution
verified exact
doi, observed 2026-08-11T14:04:27.049265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.284638Z digest=sha256:a4f7a722900c961dd510899261a3b3e06d006080b6535b38fbfaf51598cb853c

Observation 98a39769-0f69-4cf9-880f-10983920de19 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.290507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.290507Z digest=sha256:5385079699e427c7f837aa60f4277002c73943a901ebf3fdefef03021f885801

Observation c506b427-5351-411f-8261-9cfecc8d43f6 · outbound

This paper cites An Empirical Study of the Non-determinism of ChatGPT in Code Generation.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge An Empirical Study of the Non-determinism of ChatGPT in Code Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.296299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.296299Z digest=sha256:7099529beb3f6bc84d9ad036c20569f4de859d11d4839aa0feac191b1da07146

Observation 4362e961-78ee-4b5b-b255-ab6a5980d808 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.301949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.301949Z digest=sha256:5ee17a53cfc1590659d03eae90715753682808504f8394b2c21701e4d42e5446

Observation de8a3632-cf28-4f6d-9697-2b3ff880143e · outbound

This paper cites Is Temperature the Creativity Parameter of Large Language Models?.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Is Temperature the Creativity Parameter of Large Language Models?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.306934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.306934Z digest=sha256:9196829ca4b2754d6a8161afa700a2b78d0b7c2977bb7811dc0e461dc88ede71

Observation e60797a6-bcfb-4b7f-b150-b3c021014c23 · outbound

This paper cites Semantic Consistency for Assuring Reliability of Large Language Models.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Semantic Consistency for Assuring Reliability of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.313243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.313243Z digest=sha256:f9c93626ecefa2dba34107d290294eba2c9c6acd516ece0a5f8e2e251f6de774

Observation 3e3ee178-71ad-4913-b670-d03415d4f7aa · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.318583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.318583Z digest=sha256:9fda062728ce2f0db1d001acffdbac3d5f5cf036f69ca44e20d129e5a7b9481a

Observation a8e280c4-1b20-4423-8ec6-11e80aa6d3c3 · outbound

This paper cites o rg Rahnenf \.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge o rg Rahnenf \

Reference 52

Resolution
verified exact
doi, observed 2026-08-11T14:04:26.713676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.323443Z digest=sha256:5f74299db875997b7559ccdb2758f5836eef457121e644ff840300c92884c7b3

Observation c74ba101-823b-43c2-b737-d23881e9d5a4 · outbound

This paper cites Reliability of Topic Modeling.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Reliability of Topic Modeling

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:04:26.697854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.328647Z digest=sha256:eed9cde817a0cdb41e648e71580e83aae74708aa52f2e26e03c13db7b7033c5a

Observation d291bcef-2cca-4f2b-9e64-db3d5928605b · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-11T14:04:26.676373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.334731Z digest=sha256:4a99b8ea13bcfcbd386cdd56eb2cf55d8ace09650522c65ed14d23db2fc646a4

Observation c68cc77f-6828-40e6-90be-145bd210fa1c · outbound

This paper cites LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.339002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.339002Z digest=sha256:6fff2e4b3ab635ae32972fdae659b2c816e5b7d9dbd3c3f590ac1e0c5552ba9b

Observation 139d63d6-5156-4b45-8700-4a9b20864a75 · outbound

This paper cites Calibration and Correctness of Language Models for Code.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Calibration and Correctness of Language Models for Code

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.343589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.343589Z digest=sha256:bb6565fabdbac0e217170d8f20e43a01bac9ca8246947da5e3f4a355bc5082e7

Observation cdb82ca0-95f9-4a0a-9e58-3d020e6017f4 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 57

Resolution
verified exact
doi, observed 2026-08-11T14:04:26.643918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.348039Z digest=sha256:bc320b54df510b74993547bfaf459158817e6713447c9f9c31cd0d31c731b2e8

Observation f03fe839-5613-432f-8a38-d672afe20cbb · outbound

This paper cites Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.352155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.352155Z digest=sha256:d207a90a96812bd47a4fac0f6d5b852a0ca06f9874e1db736ec4b1c47d585895

Observation 66531181-c97b-45de-982d-4bf7dd1faef3 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.357383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.357383Z digest=sha256:8efb118965bfa53d19e8ae2505befd84a19e734208264cf910b7e5437e45f0d3

Observation 63fe2f30-9075-4174-899b-f79055ea0f04 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.362213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.362213Z digest=sha256:0e7a8fbd6b295561f362054989c3b68f8acd3c2008a43acd9e221845b7f133e3

Observation 22561902-6e91-4aa9-8bfc-90cf28d2f792 · outbound

This paper cites Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.367830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.367830Z digest=sha256:8d86ad262b477a43e90722bbd14789ee0264e3738a100e153d7d1ca1b4063178

Observation 22003b49-d8ac-462a-b0a2-2cf8fd2e3865 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Gemma: Open Models Based on Gemini Research and Technology

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.372895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.372895Z digest=sha256:babf95b68eaf15d4141f5afedd54126d6d3c47595efc1f9b2dfd4a48aa81ae3a

Observation 9d77c40a-0fe7-4716-b6ef-1813d5f90c6b · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Gemma: Open Models Based on Gemini Research and Technology

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.377902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.377902Z digest=sha256:3a1b91cd24d0ffc12c93b6edf4e4679eb54c1f0467883500890d3e69aaaa4e2f

Observation cf62d123-8a6a-45f8-bc30-bad822d73584 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.382442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.382442Z digest=sha256:a2c728e53db11df0d5d25701e58b01eae46b4ace3d90f8862141603e489b58a9

Observation fc2a4333-b396-4e20-a8d9-b810a0a200de · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.336522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.387119Z digest=sha256:26585869f03eeba098bb3b3aecc1a68717ca633bef5afb32eef0a993d071de36

Observation b18f4ecd-bf24-47b8-9574-3313ab95869c · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 66

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:27.585397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.391672Z digest=sha256:dc8e6af3cd5919aa459ef51df72833a8848224d80099ae86815f4238da1fa11c

Observation cf5d044e-83cb-4661-a71e-71100984f6ca · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.396580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.396580Z digest=sha256:6c70363349a44fc7d8433a7183f015e522cc44636e1ee3e163e1a30a675f4ace

Observation 717fec84-1dc1-4046-a27b-d65922246322 · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.401714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.401714Z digest=sha256:5500e6b517988ea570486e2cca1f824809b61439184588d34516b25385d5eb1c

Observation 6387e8d1-da56-4929-aa61-ee1e91069c33 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.319148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.407205Z digest=sha256:234805502c6f0e1e1a43e3d8ba473ae8f3a3b7fca6c8ef0e05e734ff65e6ff1b

Observation 03db526d-ceeb-49ff-8a7e-66ad750c6e8c · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.412482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.412482Z digest=sha256:e2043a45de173b4a274dfee8c7bdece0e7d47cbaea446bfb3be9f2791f6c1c3e

Observation 6b8c3565-afbc-41c7-b394-1cc139df6780 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.417029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.417029Z digest=sha256:b69d1f18e40d25c9c52ebb145a900b69692bf57c3bbae8745e190b25c6da141a

Observation 48c8704a-dea7-4808-8826-9404f9201b22 · outbound

This paper cites Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.421602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.421602Z digest=sha256:39175611f7fb1a02b66e737bbabe5a59f13436a80bc787176479fdd23cc8cf62

Observation 0cdf915e-f171-4b0c-be54-2b0e2523af36 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.300488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.426067Z digest=sha256:ba7b689e26363d3e29d1853acb5f08674450aa3ed8452a8b566b4eac2aea3e87

Observation 9820e8db-02d7-47a9-9c05-d3fa32b2a478 · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.281852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.430767Z digest=sha256:6447163d35b58fc03f39ea96fae15314ac0272e8dc63da9e87dda05fba6e13ed

Observation dbe090d6-eaaf-4828-8b46-f8c6891c72be · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.264027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.435520Z digest=sha256:7d1adaffe8170c33dd502a0948b3a268c64509b97cef65b23d5b10c9d5d8e7e0

Observation 3410c44b-ba9f-4869-8e01-1d563eaff84e · outbound

This paper cites an unresolved cited work.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:28.246278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:04:26.439792Z digest=sha256:780b14dcb0e1b60df9f185278302370c446a4cf774cf60c45986a3c99bf7f993

Pith citing papers

Observation 7aa2c22f-8847-4e96-a123-cca44ace6ce4 · inbound

Reasoning LLMs are Wandering Solution Explorers cites this paper.

Reasoning LLMs are Wandering Solution Explorers Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:13.015050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:13.015050Z digest=sha256:cf666c72f0a21f4ae49ee2b38e48b88b91991b792890d0b5ff74f0103459fb0b

Observation 377d08fb-1e9b-42ab-91c7-0b051cccca8d · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.166184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.166184Z digest=sha256:4533db7b51733f92cfaeb9391f8b4b6de088386bf3f12a62f431cffb050fba15

Observation 89508a75-c5c4-4b99-a2e0-b798d2b9cdeb · inbound

RecGPT Technical Report cites this paper.

RecGPT Technical Report Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:04.910071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:04.910071Z digest=sha256:44347bd4fd5a53064fc230801b2f0a58ca215584651698b6a22ffc7fc80b2e1a

Observation be00877a-4412-41e5-8544-11daeb8df5d2 · inbound

Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale cites this paper.

Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:09:11.546963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:09:11.546963Z digest=sha256:769539f19e1b606d1913e19348544f6664a85a273cf4217e749acd43e7686b8e

Observation 9f0165a7-90cd-4660-b6ab-46638cbf535c · inbound

Guidelines for Empirical Studies in Software Engineering involving Large Language Models cites this paper.

Guidelines for Empirical Studies in Software Engineering involving Large Language Models Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:02:52.228552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T22:02:36.307598Z digest=sha256:5d6efc850dc18c25811d56192a000f0b3b0929dff11a5220c1424c880f6cdc0e

Observation 1a1fe4f8-e899-43c2-91dd-807fc59c67b1 · inbound

Guidelines for Empirical Studies in Software Engineering involving Large Language Models cites this paper.

Guidelines for Empirical Studies in Software Engineering involving Large Language Models Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:20:31.427886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T08:18:18.448122Z digest=sha256:a12cbd933f379e19a71e7ec8bc6c7dd1a1a34197462c7698becec8521f83257d

Observation 5f0f1c80-3e02-483f-889d-98728815c6f1 · inbound

IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations cites this paper.

IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T11:26:27.495885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:26:27.495885Z digest=sha256:22e1eafdcbdc84720d9352eab05b6d7d367e8ad790c73de091737d0a38abce2c

Observation 210cc5ad-3c76-4b9c-acb3-9ff08272e63f · inbound

SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling cites this paper.

SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:30.591819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:30.591819Z digest=sha256:c61af3d313385267389fdc28f9f4c34408cd0ac694f4451875e84624d1570c1d

Observation cf681782-f8ee-4fa3-84e2-41dc5df3ab22 · inbound

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users cites this paper.

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:23:39.892183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T23:22:42.431997Z digest=sha256:9c1b026246fde9e8f7e244b6c58cc318b7472d82ceebd072324680947a321a38

Observation 9a234b33-9b17-4179-b58a-432ce1af1df6 · inbound

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection cites this paper.

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:48.058694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:37:19.360378Z digest=sha256:da1a0474f647f46ddc434db02c486ae2519a4783f79673a604e9e5259a4a212b

Observation b2143566-bc8f-46b0-babc-aadb24da9869 · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:51.745869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:81baf8a4f0bcad974ac1f0e325966759e3e4f832327842d32cd2616efdcb5666

Observation 9e4c2215-202a-4f8d-9dc3-e3495f1de088 · inbound

Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems cites this paper.

Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:00:10.111008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T15:59:41.632090Z digest=sha256:65ab6748cba6ad7226a445551b70eca31b2b0be91eacf2d42bfd8ae8c69fdc18

Observation a506623b-98f4-4159-9a21-1e3775c73ae6 · inbound

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents cites this paper.

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.999729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:13cb6884c6aac7ef548d816638983ffb46266526a8cf39eee6b52045d70ccae0

Observation 9bc9f2bf-31e1-444d-9fd4-8a6c4a14de61 · inbound

Mixed response geometry and critical crossover in the Ising model cites this paper.

Mixed response geometry and critical crossover in the Ising model Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T19:18:49.672005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:18:49.672005Z digest=sha256:285c4a43146af027d6e947d5a54e8aac0f65858b327a44de1da76fd19b5425f6

Observation fbbfbae8-c536-4762-84c9-9b5bbcfecc41 · inbound

Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure cites this paper.

Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:06:02.544064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T23:40:00.572491Z digest=sha256:3f8bccd94b71376bd5c28fada31a7a4502ce564f8be5490e6b9cc7f88b78567c

Observation 1e9fd0f6-791c-4bac-80d2-3532eb091302 · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:52.619473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:32:42.151644Z digest=sha256:5793737a19d9dd60da987c949c7f6f2e215e67d26a2891ddc954afeb167468b2

Observation f95f96aa-bc93-4936-935e-bf0f03a6df6e · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.505086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T22:59:07.861941Z digest=sha256:64c30cd219fe74e67ec9b3e1f3435f1e5e9fe54d8099db9fda0f2a45d1e8f0ac

Observation 073fcf87-a936-4e32-8649-d782e3a5e264 · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T14:41:06.754830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:41:06.754830Z digest=sha256:25805a5dfb48311ed10979aa1ad2b16b24df794f6da2ef71587603016b473882

Observation 5fa0f76a-a8f6-4f10-95b4-4040b442bef9 · inbound

A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability cites this paper.

A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:06:19.501659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-12T03:03:05.715652Z digest=sha256:275ec52343d4f6253344c14fe8496817d971b69b91ba07610eb0540454ab1c31

Observation c054c141-fd1e-42c2-b791-d2cc8bdb2c31 · inbound

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On cites this paper.

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.402693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T10:31:16.368065Z digest=sha256:8b9bead2c00ee12c9e97313ba69e39c3b54d1a906970ecaf8cafad9610b1d76d

Observation 2a5a7f47-d2ca-40f3-8387-c5a8392c2013 · inbound

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions cites this paper.

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.458667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T15:17:37.904831Z digest=sha256:80da2a94dba6c876e5996715778ecd3484d1e55c3cf031d4f8c91d8fc88441af

Observation 0d6cb879-eb2b-49f8-b934-f36a4d892ec7 · inbound

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback cites this paper.

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.998548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:27:26.911873Z digest=sha256:bc88938ea570a76698362bc71381bee4e44859b3bacf6fdb0195127c2c4767ba

Observation df36f15d-1dab-42b4-98e9-0144a15dea3e · inbound

ComplexConstraints and Beyond: Expert Rubrics for RLVR cites this paper.

ComplexConstraints and Beyond: Expert Rubrics for RLVR Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:31.332133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:37:11.141846Z digest=sha256:761f7a0f7d51f6e3f9ae2a52debce260e979d60ba560a4d8b2a87c182b0b7520

Observation fc3ef303-5c24-48e1-a078-e054506492d3 · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:19:59.995041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:ba08a03aef5d41c438a7b7b62f0c109d61b0f66e80f23ef9337bd22194bfd263

Observation 58858e04-0d09-4ed6-8eb9-fe3cb0179867 · inbound

Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments cites this paper.

Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:44.282114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-01T01:18:53.736430Z digest=sha256:0477b7ebf80296869c4067e78f35f7694894b3d510f27d9df3715ff5daf484d3

Observation 7de55c08-2825-4751-9ff4-6e4fcd5c413a · inbound

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers cites this paper.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.119758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.119758Z digest=sha256:839f26a0894e4199cebd0648c3a6d529560f6600390a70d3cc1428b3b45dc6d6

Observation 98c7724a-5157-4c4d-a965-2f801e888550 · inbound

Codifying the Judge: Scalable Evaluation via Program Distillation cites this paper.

Codifying the Judge: Scalable Evaluation via Program Distillation Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T12:49:10.833195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:49:10.833195Z digest=sha256:da12e7049a6833146d829065c7167dceda80f559e29b21c4fbc31b598b34ccda

Observation ff66f522-10ee-44a2-83a0-0e9740e01355 · inbound

Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias cites this paper.

Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:00:08.279354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:00:08.279354Z digest=sha256:9d5aa88c795724eb06d60cca9ed3a20f7097c2ccbe86aa7d5c16843b7d62d950