Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T01:24:02.052803Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 39 inbound Pith citation observations for arXiv:2410.12784.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T01:24:02.052803Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:40:03.461298Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T17:09:58.526716Z
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 22e41b80-1d6d-44ee-834b-0e09ebfc8657 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges 10/10”, “Neither A nor B
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c203191a-bd16-4eb5-b2a9-8dee1cb61036 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges knowledge
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a66ec54-bfd4-4d96-bc51-1dbb4b058d16 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Each side’s lateral pterygoid has a different function during this movement
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dadbf0f8-afce-4ed4-b5d6-126c97d5ffb5 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4406d922-8fec-4c77-8c71-8752311c2bda · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Output (a)
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49b9ea51-6e31-40e1-94c3-0c07e9563525 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fe14601-6ef7-4601-bff3-83033f4981c2 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c72140b1-2a06-4ac1-953b-f1c43386206c · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d1b2595-9978-48ac-b65e-3b00f315782b · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80a5f0f9-fdb1-45e0-840e-81cb9ebae46c · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges My final verdict is tie: [[A=B]]
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19303786-670b-4a39-9c58-f6e3fbef159d · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9b1432f-66fc-4bc3-a9fe-90c56b394e53 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges You should refer to the score rubric
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e43529e3-e0bb-4cd2-9852-7707b7a9caef · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges (write a feedback for criteria) [RESULT] (A or B)
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a16053e1-5fbe-4422-b118-183795266389 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af812f28-b742-4c25-96a3-9916772928c7 · outbound
JudgeBench: A Benchmark for Evaluating LLM-based Judges So, the final decision is Response 1 / Response 2 / Tie
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 877b95d4-104b-4d9a-b190-241e22a2d80d · inbound
A Survey on LLM-as-a-Judge JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a44a7667-9683-416d-a720-4c3a3a793636 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1127d743-065d-4dd2-a4b3-fe8c51514e52 · inbound
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac8d15c4-c40b-430a-a911-560fb805d8a8 · inbound
Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68938448-2a4b-4ea9-98d3-3cad8e1b3909 · inbound
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee9cd37-ef97-44e2-a5ff-b6d9c35525db · inbound
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbdbe3dd-7eae-40c2-b34e-81e813040f4a · inbound
NVIDIA Nemotron 3: Efficient and Open Intelligence JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd097b4a-db0f-4555-af2d-6ea21860b988 · inbound
AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b50371a8-365e-43e5-8085-354baf51de63 · inbound
MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12cebc39-8826-4396-97af-2baf8d07d4d1 · inbound
Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ad28c37-11bd-4c01-8427-06b41a8221bf · inbound
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 612a4f1e-701e-41a8-85f5-43f1c6a57e61 · inbound
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e24a3111-f640-4a69-a58e-70922bd160f1 · inbound
Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92b89e58-b68e-4f01-b646-8839c3ab19e7 · inbound
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c430462-4d8b-4fca-9ac5-7798a5fc2c2c · inbound
Green Shielding: A User-Centric Approach Towards Trustworthy AI JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 182870d9-e3c5-4940-8fe9-9f5de1b3dbe1 · inbound
Training Computer Use Agents to Assess the Usability of Graphical User Interfaces JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ca0f88c-7962-4e8e-a25a-3b8221eac6e6 · inbound
LATTICE: Evaluating Decision Support Utility of Crypto Agents JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 125ea4ac-d648-4461-8879-cc69a105ebb8 · inbound
LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f24a1c1b-8d2a-4bfd-9a06-9682f53b81e8 · inbound
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d612516c-8963-4972-aeea-dadb6b54358e · inbound
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9588355e-9fba-4ddb-9d11-f94ee8d924c6 · inbound
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 828fc486-253a-4beb-9349-2a51aa29cb17 · inbound
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6939b2af-23ce-4106-b771-cbac7f8f19a2 · inbound
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8d17776-c9e9-4a32-821e-3c80a1799a81 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 453009ee-8fc6-42b3-956d-c4809fabfdfb · inbound
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef892dda-8de7-4890-a369-2d2ce812d574 · inbound
CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa4d5695-7367-4495-9499-ae4c0f70d1b7 · inbound
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2074f077-bf95-4e0c-8087-3f56d7c4542b · inbound
Are LLMs Bad at Moral Reasoning? JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39937189-1cc9-47a8-b7b5-d222bdd21fd4 · inbound
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa124f31-66f9-4cc6-ac3d-d3e7e9e6d4b2 · inbound
Mind Companion: An Embodied Conversational Agent for Process-Based Psychotherapy JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e90c1f6-5d62-4bc6-a341-dd3dc89dab2f · inbound
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04a7c527-38c5-4e58-a62d-1ad662986a76 · inbound
Counsel: A Meta-Evaluation Dataset for Agentic Tasks JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fcbb5378-2ff7-433e-9377-a60236cc6345 · inbound
Towards Spec Learning: Inference-Time Alignment from Preference Pairs JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26100d5d-96b7-417b-8445-cdc38e93f0b2 · inbound
Towards Spec Learning: Inference-Time Alignment from Preference Pairs JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a43f1ec4-f362-47d9-a625-26365f72cbd2 · inbound
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ccb3788-f29f-45a7-9e4d-5b475ef4c534 · inbound
Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ac25137-7d17-4883-aef9-1070d6762439 · inbound
Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676058d8-0a36-4e19-bad9-c42b2862976d · inbound
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1abe627d-4e0d-4c7c-854b-c78201c607fd · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.