Pith. sign in

Paper Citation Record · LEDGER

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2406.12624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12624 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:34:54.754562Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 96b355a5-e339-4e58-97a2-688587ece62d · inbound

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback cites this paper.

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:38:32.526940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-23T22:37:43.230753Z digest=sha256:037615bf47216e5e01b0e07b6e2486f84483e7c5c61a53dbe00f17aa192c4cae

Observation 961e4675-97a3-4ad4-8d3b-a6b2d11d2e1a · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.910382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:3a2c52a080ae226404bb5b5cb186346f13dbdcac17d2c5ae8dbfa029091bc0f5

Observation 5a1dd21d-8d36-4be1-bf11-05e3f4b196c8 · inbound

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? cites this paper.

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:12:47.273316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:12:47.273316Z digest=sha256:eeea3a10217904f8c99c20b1f978b0ab55f462eacb5c3c059a6054981b5fc995

Observation 296e454e-1347-48f9-9d30-ecab492a6167 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T17:35:44.000270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:c6cafb605ed1f9caf205f701910394f4f2cc941754c54b9a2177b90be903d256

Observation 196eee81-2147-4992-9d96-4f60defcddda · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.400788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.400788Z digest=sha256:97d225d32629d3b44a5f50e87f2c338a270e713c025aaa3c8c80a5c88224db66

Observation 5bfeac36-f10b-4f39-9211-329636dd577a · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.444375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:aa0b76617be89a31916abee0b5c84d9f913724cfb79047779865a2d3ef8d22ab

Observation 744cc02c-5b2a-44cf-82e8-4a72f78dd04e · inbound

JuStRank: Benchmarking LLM Judges for System Ranking cites this paper.

JuStRank: Benchmarking LLM Judges for System Ranking Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:49.609141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:49.609141Z digest=sha256:0c5dc30d8ed5dd5f61338f6e32691a317391b0d28a55fa5e7076749a8ec2b916

Observation bc0962ce-3cb7-4fa8-a347-e4bf7c1cc902 · inbound

Let your LLM generate a few tokens and you will reduce the need for retrieval cites this paper.

Let your LLM generate a few tokens and you will reduce the need for retrieval Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:54:57.110679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:54:57.110679Z digest=sha256:a2454cc937888a418ae1df07a25f6b613c87786d75d99691bd7e919c9c6c2220

Observation f3454929-62f1-4f2c-931e-17e0c5fd9b9d · inbound

KARRIEREWEGE: A Large Scale Career Path Prediction Dataset cites this paper.

KARRIEREWEGE: A Large Scale Career Path Prediction Dataset Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:08:18.611372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:08:18.611372Z digest=sha256:bc36b3365aff4590b193ea34ee9d55c4de66034f673527c91527fe4ded3314b5

Observation 4ad801e4-71e0-44e9-bd03-b0ea41956fb3 · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.784558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.784558Z digest=sha256:598c45086e20859d2f19ed78cbbad26f999652335fbd564eb19db34bf33e73cc

Observation 6d91b716-4c8c-4cc3-bd55-2c92f31e323e · inbound

Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses cites this paper.

Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T11:15:13.004759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:15:13.004759Z digest=sha256:491c88406d9559c8e879f1bd5366b4d985204d47374d21a0c4a1c236f6b2ebe5

Observation 73519e5b-d9f2-428b-acd4-993eecbb3a7d · inbound

WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models cites this paper.

WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:16.580429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:16.580429Z digest=sha256:ad8eea565d88316afb5e4f17c5cc3535e1f687c4eb56eb8d35a59fda46a746fe

Observation 6338cc35-3f69-4baf-81b6-61d9ce791fa7 · inbound

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient cites this paper.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.679147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.679147Z digest=sha256:7fa6d200c9f149681ce1cb1ee8c0942c215d4e2987a4ad75f5fec74ab5866d11

Observation 880d571c-3ab7-4bc6-8fb6-19a2c5ab27e3 · inbound

A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs cites this paper.

A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T10:44:41.088451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:44:41.088451Z digest=sha256:e8ca1d9f69a7f196352a23d9889008aa7ba651ba8d257d1f5740fcd3dd09d130

Observation 574feeb1-6552-48a2-b94f-952ed453f52c · inbound

Combining Large Language Models with Static Analyzers for Code Review Generation cites this paper.

Combining Large Language Models with Static Analyzers for Code Review Generation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:53:28.371495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:53:28.371495Z digest=sha256:f0c8635e02ddfee736c533eef70c3cf1dd56d2c1a08a0847819c379712f1ef39

Observation 8115c949-cfaa-4942-9038-22dba9639f4b · inbound

Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment cites this paper.

Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T10:29:05.223746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:29:05.223746Z digest=sha256:2a3cccd204e6ff70ef35b72165f8120fea64e29193446b551c299d27b42022c0

Observation 4bec3727-f646-495a-9b8f-975b8dd0f09d · inbound

NUTSHELL: A Dataset for Abstract Generation from Scientific Talks cites this paper.

NUTSHELL: A Dataset for Abstract Generation from Scientific Talks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:52:26.527289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T02:50:10.307310Z digest=sha256:432a2bcca5272570dfcecec13020174bbf41e83339fa4669af5356a0111673d5

Observation 47b452ea-8934-455f-80ad-a3fb6e8f80a8 · inbound

Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy cites this paper.

Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:31:55.906972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T18:27:19.450554Z digest=sha256:d9b5a9e41fdf306f53fbfbdff8dee23eaca57a6ec2eaca6ffefd19822c87a0f7

Observation 47b3bed6-0941-4f76-8bd6-94c08b34f2d0 · inbound

Cost-Optimal Active AI Model Evaluation cites this paper.

Cost-Optimal Active AI Model Evaluation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:05.279668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:05.279668Z digest=sha256:11de9b80885f691f07a905ce47f681fedde2599e15f3b330ccc69ea2ce1dd51e

Observation 730e3869-d1e8-479d-bd45-66c421fac6da · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.310755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.310755Z digest=sha256:bb12176bb1eaae4051dbe4e218ae842134474eed25cd10c0b43f224fc8752cef

Observation bbcd6eb4-4ad1-4407-b1ff-d7cd2d376c3a · inbound

Revealing Political Bias in LLMs through Structured Multi-Agent Debate cites this paper.

Revealing Political Bias in LLMs through Structured Multi-Agent Debate Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:07.823626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:07:07.823626Z digest=sha256:35481b9c8273b8526865c752eb27d6d1b4fde71767782666711d3ccafa26f440

Observation dbccbcef-a0e0-4a25-a3d0-e0ab3cfabc98 · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.353198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.353198Z digest=sha256:47a33b396bb0230a984f20252e0c593a4f7475883e193d0e0d4e18a4d8618170

Observation 6aeeb5cc-7968-4a71-ad09-0a8f7e651acb · inbound

Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs cites this paper.

Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:14.239923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:14.239923Z digest=sha256:1015ede88ee2514dbc5cd064d7409b1ac480d6ac007d7ab143780fcb7d5e0c4d

Observation 2e9102f2-ea1c-4a32-b45b-7953a01a4fe7 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:35.063138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:35.063138Z digest=sha256:13b3f17a36982bb3635e51cefe220abc654d7d19f5a59ab7350cd6a11ece26b0

Observation 65b8d7fa-94da-4a15-bcf1-17f2e4e41e66 · inbound

Evaluating LLM Agent Collusion in Double Auctions cites this paper.

Evaluating LLM Agent Collusion in Double Auctions Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:42.258308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:57:42.258308Z digest=sha256:2854870c3effbac9d3aa6d2e99128dc02ff5cb40260e044eef73e20e234e7590

Observation 81e637f6-ca8b-4188-bcde-5563b1c381dd · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.485583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.485583Z digest=sha256:37ed48105ea3b02bc9f487a92b270b1bf4ac30d36c058d7479897b778d56339e

Observation be9185e6-f33f-4126-a79d-fa0c512e0962 · inbound

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation cites this paper.

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T12:32:26.094061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:32:26.094061Z digest=sha256:86bb38f329ff1e027bc7792ae1bc2e21dc9c184480697e0d94bae48245335df4

Observation e6808b0f-c3d1-4656-8ec4-40373c23c594 · inbound

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference cites this paper.

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:21.829738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:21:21.829738Z digest=sha256:0fbb7e82932a0135ed01b2b06498da899e5ddbb1a99770335546fe5ddc2bebf7

Observation b6917440-946e-479c-b0d9-ecf3803b79ec · inbound

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers cites this paper.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.682665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:ff7f9314b9366d2e63b32139db98052be60b38e072147d5b7caed6755f57c6df

Observation c79f583a-488a-4755-ab11-770e48249687 · inbound

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation cites this paper.

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:07:43.327951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:04:44.379128Z digest=sha256:927131010a90b89ddc6d2d957a239f6576a30f1cfd68f66a01b8fcf0b10ec820

Observation 155f65ea-abb3-4eec-bf17-4bce3d1a5f27 · inbound

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks cites this paper.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:07a4940f9d479ce2be59fc03e451079eee02d7b0a1021c2ab28cf77d9c4ea325

Observation 256cd814-938b-4ac8-9c5c-756b9e922b4a · inbound

Code Review Agent Benchmark cites this paper.

Code Review Agent Benchmark Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:23:22.027618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T00:21:36.481731Z digest=sha256:d9ff08fa887cbfd502aba04628e3aec4956d4bf9e3c74a9c266b3eb33c7317da

Observation 98a09ffe-cc34-4299-8b87-cb3e37423673 · inbound

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents cites this paper.

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.009796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:78cb2be1706855940692aa5b523b45b109377d9f7dd861a0ba9853ab023f4338

Observation 63347391-bfa9-4d3d-8b1c-439f5c3be01d · inbound

Mixed response geometry and critical crossover in the Ising model cites this paper.

Mixed response geometry and critical crossover in the Ising model Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T19:18:49.672005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:18:49.672005Z digest=sha256:c2af90cfe44ed41f71f7cf2240861aba6dd29826edd39511c7875e549c82640d

Observation b5f7733b-4452-4167-be79-a2af72571ae4 · inbound

Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop cites this paper.

Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:13.502497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T03:32:38.940942Z digest=sha256:53b7ffead75f0d6bf66561b95d4f3f69ac3de59432374129c993c9e3feb435de

Observation 60e1efee-b275-4e11-977a-a93f959dd882 · inbound

Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls cites this paper.

Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:50:40.902897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T18:04:50.972507Z digest=sha256:90219a6f82c09d04f06b0ab9d0acfda858be193447ecf54df2e98daed7d45636

Observation 071c21ff-3874-4b72-a93d-5e865503c0b0 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.923870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:ade5b06f4cdd3125001e4a5602bc6e6aec71cd4359fd41c7abb145f9983acc2b

Observation b5221e9f-b502-49db-967d-6016eb52bf13 · inbound

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents cites this paper.

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:53:26.356204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T12:52:20.911788Z digest=sha256:888c096ed7ff5fb5184f9f8cda673c6c592d88cfa05eb795b49c018bb72efbc6

Observation 7233044c-0e63-43ce-a851-2f150da5f7d3 · inbound

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems cites this paper.

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:19.960925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T14:44:21.487169Z digest=sha256:9fb630af2319030ce4317085fc350e7aa3ba8f5c27fe1e3874455251cb7fb2c5

Observation 47b519b7-7810-4432-81d2-688802ec80e4 · inbound

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing cites this paper.

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:40.098418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T15:48:26.303462Z digest=sha256:c31a8a8f55069409d4e2a1532a16d10bfec5428b1d8f22a8ecd84a154ae50e77

Observation de7874be-7ec5-4114-904d-b80ee5b56c05 · inbound

Leveraging Resolved Incident History for LLM-Assisted Software Bug Diagnosis cites this paper.

Leveraging Resolved Incident History for LLM-Assisted Software Bug Diagnosis Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:34:54.754562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:34:54.754562Z digest=sha256:dfd9f2a35c6667aded1616d2bb41509b138f9718b3f1317cd35ca0a68a4c574c