Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2406.12624.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:30:05.279668Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 96b355a5-e339-4e58-97a2-688587ece62d · inbound
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 961e4675-97a3-4ad4-8d3b-a6b2d11d2e1a · inbound
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 296e454e-1347-48f9-9d30-ecab492a6167 · inbound
A Survey on LLM-as-a-Judge Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bfeac36-f10b-4f39-9211-329636dd577a · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 223
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bec3727-f646-495a-9b8f-975b8dd0f09d · inbound
NUTSHELL: A Dataset for Abstract Generation from Scientific Talks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47b452ea-8934-455f-80ad-a3fb6e8f80a8 · inbound
Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47b3bed6-0941-4f76-8bd6-94c08b34f2d0 · inbound
Cost-Optimal Active AI Model Evaluation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730e3869-d1e8-479d-bd45-66c421fac6da · inbound
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcd6eb4-4ad1-4407-b1ff-d7cd2d376c3a · inbound
Revealing Political Bias in LLMs through Structured Multi-Agent Debate Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbccbcef-a0e0-4a25-a3d0-e0ab3cfabc98 · inbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aeeb5cc-7968-4a71-ad09-0a8f7e651acb · inbound
Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9102f2-ea1c-4a32-b45b-7953a01a4fe7 · inbound
Enterprise Large Language Model Evaluation Benchmark Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b8d7fa-94da-4a15-bcf1-17f2e4e41e66 · inbound
Evaluating LLM Agent Collusion in Double Auctions Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e637f6-ca8b-4188-bcde-5563b1c381dd · inbound
Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be9185e6-f33f-4126-a79d-fa0c512e0962 · inbound
SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6808b0f-c3d1-4656-8ec4-40373c23c594 · inbound
Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6917440-946e-479c-b0d9-ecf3803b79ec · inbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c79f583a-488a-4755-ab11-770e48249687 · inbound
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 155f65ea-abb3-4eec-bf17-4bce3d1a5f27 · inbound
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 256cd814-938b-4ac8-9c5c-756b9e922b4a · inbound
Code Review Agent Benchmark Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98a09ffe-cc34-4299-8b87-cb3e37423673 · inbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63347391-bfa9-4d3d-8b1c-439f5c3be01d · inbound
Mixed response geometry and critical crossover in the Ising model Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f7733b-4452-4167-be79-a2af72571ae4 · inbound
Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60e1efee-b275-4e11-977a-a93f959dd882 · inbound
Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 071c21ff-3874-4b72-a93d-5e865503c0b0 · inbound
Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5221e9f-b502-49db-967d-6016eb52bf13 · inbound
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7233044c-0e63-43ce-a851-2f150da5f7d3 · inbound
POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47b519b7-7810-4432-81d2-688802ec80e4 · inbound
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.