Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2506.09443.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-10T14:09:11.328304Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T14:17:10.144588Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 37ec9b1c-1c4c-444f-9ef0-5ce82bf4230c · inbound
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a800c6c-9f97-4cb9-bd4c-7c1e5c164ddb · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ec19ef6e-2772-44da-aae4-b49188f278c2 · inbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5758604b-1fca-4aaf-a88d-965dea1cfe97 · inbound
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0a58087e-4ad6-4a36-ad0a-4bcf37cd46a6 · inbound
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd2a0700-e44f-4903-9643-9dede2d6fcab · inbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 42e23d4e-3b66-4f5d-b430-9de3132b0757 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f252d99f-fc69-4c7f-b2f0-429b1dfa1656 · inbound
No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 04277a2a-a34a-40df-ac8d-b9258216665a · inbound
Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.