Pith. sign in

Paper Citation Record · LEDGER

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2406.12624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12624 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:30:05.279668Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 96b355a5-e339-4e58-97a2-688587ece62d · inbound

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback cites this paper.

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:38:32.526940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T22:37:43.230753Z digest=sha256:c305a96c4b1b9e9eed8f7c73ca4061d96742c2b133602385681e4a5c77ba2e33

Observation 961e4675-97a3-4ad4-8d3b-a6b2d11d2e1a · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.910382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:ce709087e8b16fe36d455d71792a00960846089782704bad129ef589c659f881

Observation 296e454e-1347-48f9-9d30-ecab492a6167 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T17:35:44.000270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:c562235358d42fbaf466ab07038a27b452177a9bd4985f699f982285a1e805d1

Observation 5bfeac36-f10b-4f39-9211-329636dd577a · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.444375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:ce273221c1473238c7be9ab563ff41290be6e150a9fc137a35dbd3dde8809dcc

Observation 4bec3727-f646-495a-9b8f-975b8dd0f09d · inbound

NUTSHELL: A Dataset for Abstract Generation from Scientific Talks cites this paper.

NUTSHELL: A Dataset for Abstract Generation from Scientific Talks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:52:26.527289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T02:50:10.307310Z digest=sha256:74deae854decbce2a3acd546c6c969860ec8788a4b969b78244bccc799a986d6

Observation 47b452ea-8934-455f-80ad-a3fb6e8f80a8 · inbound

Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy cites this paper.

Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:31:55.906972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T18:27:19.450554Z digest=sha256:bc9cff558e2a638163c11a0efffe2dca01a10faa46a345072c28cd89801f655d

Observation 47b3bed6-0941-4f76-8bd6-94c08b34f2d0 · inbound

Cost-Optimal Active AI Model Evaluation cites this paper.

Cost-Optimal Active AI Model Evaluation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:05.279668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:05.279668Z digest=sha256:bac37088f5660d240ab7dee169d2fa007a2bbcc34d2ed941586ac0e0bed1b227

Observation 730e3869-d1e8-479d-bd45-66c421fac6da · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.310755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.310755Z digest=sha256:7d77ce850cf0aa69ee25d66f860fb6be539a253d1623e8df8a4b34982426716a

Observation bbcd6eb4-4ad1-4407-b1ff-d7cd2d376c3a · inbound

Revealing Political Bias in LLMs through Structured Multi-Agent Debate cites this paper.

Revealing Political Bias in LLMs through Structured Multi-Agent Debate Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:07.823626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:07:07.823626Z digest=sha256:9f31d84d7e23b674b179aff7c0565c5023f90dfccf7ceccfeeacd11f9dc4b3da

Observation dbccbcef-a0e0-4a25-a3d0-e0ab3cfabc98 · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.353198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.353198Z digest=sha256:e817ef3fd9e254052c07b884d564ea57ffe4c669db814a1fbd891f15da1dad00

Observation 6aeeb5cc-7968-4a71-ad09-0a8f7e651acb · inbound

Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs cites this paper.

Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:14.239923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:14.239923Z digest=sha256:2ed544dfa73ee87ad9ac02de8fd54476e1f05039edf8854ab74c625f223ce7e4

Observation 2e9102f2-ea1c-4a32-b45b-7953a01a4fe7 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:35.063138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:35.063138Z digest=sha256:d04e752bbb0b134fb0d96cc89c341627b3578dadfddef75ab0d15847d7eed3fa

Observation 65b8d7fa-94da-4a15-bcf1-17f2e4e41e66 · inbound

Evaluating LLM Agent Collusion in Double Auctions cites this paper.

Evaluating LLM Agent Collusion in Double Auctions Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:42.258308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:57:42.258308Z digest=sha256:aab5d52065d9af33fc40adedef8fefe420ae040543bcd86761ab89ee37a05c19

Observation 81e637f6-ca8b-4188-bcde-5563b1c381dd · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.485583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.485583Z digest=sha256:d5fd165147a9a1f78fcca133cd27fd4709d1ae57c81938137fd76441cc4d7d8d

Observation be9185e6-f33f-4126-a79d-fa0c512e0962 · inbound

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation cites this paper.

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T12:32:26.094061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:32:26.094061Z digest=sha256:4e4c91cd555e388cc2b64a3d035db9c00ff2f3d8b26e75c2e91a40761791a923

Observation e6808b0f-c3d1-4656-8ec4-40373c23c594 · inbound

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference cites this paper.

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:21.829738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:21:21.829738Z digest=sha256:23d08b9de5fcea992446cdcaec54872267e7a41b1598511b616a4a0cf9f2e9ec

Observation b6917440-946e-479c-b0d9-ecf3803b79ec · inbound

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers cites this paper.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.682665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:1e4dc49e64debf5e1927efbfbfaae01de3b0a465d6c3afe9935800bddf034748

Observation c79f583a-488a-4755-ab11-770e48249687 · inbound

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation cites this paper.

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:07:43.327951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:04:44.379128Z digest=sha256:f794274f7f19feccd4139865362de954da12e7e67f017e56e1406189bee81fba

Observation 155f65ea-abb3-4eec-bf17-4bce3d1a5f27 · inbound

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks cites this paper.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:ee732db3f6309b1ce72c0406ae9849caa8db4532e598587f6d80455943c77ad8

Observation 256cd814-938b-4ac8-9c5c-756b9e922b4a · inbound

Code Review Agent Benchmark cites this paper.

Code Review Agent Benchmark Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:23:22.027618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:21:36.481731Z digest=sha256:beff09d7fe28bc423152b8f0b8ba5d8865266d2633dabafb9091503cffadc5f3

Observation 98a09ffe-cc34-4299-8b87-cb3e37423673 · inbound

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench cites this paper.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.009796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:b63a8a2d84312e55e860d9e344a0c12532053498c30754100b3c6c24c1a6b3de

Observation 63347391-bfa9-4d3d-8b1c-439f5c3be01d · inbound

Mixed response geometry and critical crossover in the Ising model cites this paper.

Mixed response geometry and critical crossover in the Ising model Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T19:18:49.672005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:18:49.672005Z digest=sha256:cc29fab7b6c0bf172b4815b428de562b9b462aca7cc255f640a15f565d8f9b66

Observation b5f7733b-4452-4167-be79-a2af72571ae4 · inbound

Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop cites this paper.

Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:13.502497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:32:38.940942Z digest=sha256:085d63e91040c2854a257a4309ba907d1a9ba9b56695e96464d94f0614059479

Observation 60e1efee-b275-4e11-977a-a93f959dd882 · inbound

Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls cites this paper.

Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:50:40.902897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T18:04:50.972507Z digest=sha256:1cc06d9a04e1fe31a9bc5be582429ef462d4101c94a294d99d5b6d3407e6e4f6

Observation 071c21ff-3874-4b72-a93d-5e865503c0b0 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.923870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:75cce503904e5d1f6f0985da54c2e5522e14960f2e43a7474397f95f20b5107c

Observation b5221e9f-b502-49db-967d-6016eb52bf13 · inbound

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents cites this paper.

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:53:26.356204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:52:20.911788Z digest=sha256:dabe1be2b1533de34418c94078de42af0a326bd9c99bae5c8c4cf47f4795fd60

Observation 7233044c-0e63-43ce-a851-2f150da5f7d3 · inbound

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems cites this paper.

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:19.960925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:44:21.487169Z digest=sha256:4133f6da557df8127ee5b4959990c42a2220ca4e2efc36ad6709140a122b8236

Observation 47b519b7-7810-4432-81d2-688802ec80e4 · inbound

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing cites this paper.

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:40.098418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T15:48:26.303462Z digest=sha256:9a02e836814215cc3d4df5de327fe279e41bf9794c840009dd0cfe6ae14b4df4