Pith. sign in

Paper Citation Record · LEDGER

LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2506.09443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09443 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:09:11.328304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T14:17:10.144588Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 37ec9b1c-1c4c-444f-9ef0-5ce82bf4230c · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:720e399d6cc65e8716d2cf5c82e1e45ab84a4936891bbd39e535a4d128616d7b

Observation 1a800c6c-9f97-4cb9-bd4c-7c1e5c164ddb · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:a5edfecac398340476f553faff0a06ec4b54f32564c1383d02f29335e71f6cd7

Observation ec19ef6e-2772-44da-aae4-b49188f278c2 · inbound

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs cites this paper.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:d269a4370f25b32687f247be1cf6b9a14a1fcf8f0c4970fac35cb6eb1c610c75

Observation 5758604b-1fca-4aaf-a88d-965dea1cfe97 · inbound

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following cites this paper.

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T16:24:33.375282Z digest=sha256:93ca7f1b270dd5690ddaa4597f02fd1c5722c5c28ff3cb438cb4bbfcd5cfedb4

Observation 0a58087e-4ad6-4a36-ad0a-4bcf37cd46a6 · inbound

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges cites this paper.

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T10:23:58.198464Z digest=sha256:664e938bed2a958e3c6ce517e7b3d6bf20ae415d4ffd917460357fb8d360c6ce

Observation cd2a0700-e44f-4903-9643-9dede2d6fcab · inbound

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator cites this paper.

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T08:50:44.523468Z digest=sha256:10e106c1a2b4c7f8e24a2352983756eb2d0ab27c2e3b88593754b21678c5b7ef

Observation 42e23d4e-3b66-4f5d-b430-9de3132b0757 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 138

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:827e48d400062d597c4548e2363c92b12dd222a263a990c9e7bbe5d5ec3a3902

Observation f252d99f-fc69-4c7f-b2f0-429b1dfa1656 · inbound

No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions cites this paper.

No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T06:37:45.175776Z digest=sha256:970e08cacae484ed269eb72c91d096ac403c18fe5ae61c397f04fcae68a3651e

Observation 04277a2a-a34a-40df-ac8d-b9258216665a · inbound

Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems cites this paper.

Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-10T14:09:11.328304Z digest=sha256:d9e4558b85005f4bdb29ba8d963876a6efd2a1569e4eaa7e382caada5fc0aa57