Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:13.658964Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 3 inbound Pith citation observations for arXiv:2506.10403.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:13.658964Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-30T15:14:05.529138Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
35 of 35 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation ef315790-9b74-4f37-80c1-6a150431efaa · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6acd8fde-133d-4d07-9f32-eedf388fa975 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.; 5 Gonzalez, J
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 113822d6-e3d9-4b00-a9c9-05185fb377f3 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Advances in Neural Information Processing Systems 2023, 36, 46595–46623
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a029ec1d-50a9-4bef-9736-cd09ca5b1c94 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Agent-as-a-Judge: Evaluate Agents with Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290c994f-6f21-4700-9db7-cf29e9336f1d · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation QuRating: Selecting High-Quality Data for Training Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b0b47b-a7e8-4a48-85b5-8d71ab66010a · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a74c5cd5-db0d-4001-9307-d494038856ad · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f980c253-bbe6-46ce-83bc-1e9c8ad4a062 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5285b6-e5ed-4824-bb79-f08bef70f2f6 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19dac899-d440-48c8-8b57-7f5e6991fc71 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Discovering Bias in Latent Space: An Unsupervised Debiasing Approach
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d4530c-185d-4291-8809-faa66fc8ed31 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation H.; Chen, S.; Liu, Z.; Jiang, F.; Wang, B
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f607aac-142f-4e00-b349-0551a68d1b97 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07bf243-e9b2-49b9-9f36-fdb29e17b4f4 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation J.; De Sa, C
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e8c7f540-ff3c-4302-b88c-0909217293e3 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation H.; Ehrenberg, H.; Fries, J.; Wu, S.; Ré, C
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ca798256-8249-4e82-8760-09a04738b7a4 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Training complex models with multi-task weak supervision
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation affb8cf4-53ab-49ed-bdea-b30ec53d5f5a · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Fast and three-rious: Speeding up weak supervision with triplet methods
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d07bcf6c-d883-4dbd-83f3-8cbd8345e7f8 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation RewardBench: Evaluating Reward Models for Language Modeling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e7e9cc8-1810-4351-af7f-ae67533c8332 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation The Twelfth International Conference on Learning Representations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08063b97-5a06-43dc-b46d-d612811d8fcc · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation GPT-4o System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7849f027-4f10-4708-9357-5ee5d7c19757 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 292117a4-594a-45ad-826b-945007f1d7f5 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation P.; Fishburne Jr, R
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89546b02-b98c-41dd-a339-c03f10bba31b · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8562dfc4-bc17-4ef9-bcd6-9bcc0311d637 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bdd328f-b03a-4208-8ac1-24a3eca19819 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Universalizing weak supervision
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5711fb54-8843-4ca0-8982-a659fa51253a · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Weak-to-Strong Generalization Through the Data-Centric Lens
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6833448-268e-4a5a-8a56-f5b92c2d2589 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb382b29-30e9-41dc-8491-3df9c32420b7 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation GPT-4 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7646b5-cff3-46c1-b0d7-07c9a51d592e · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Gemma 2: Improving Open Language Models at a Practical Size
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 733cb78b-219d-4772-a6d9-9a4b30d992f6 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e8dac7-4bd3-4ca3-bfa5-b16cfe8053b8 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2ad085-1fa2-444a-9242-9cd64b26d309 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data Annotators
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abc0dae9-433c-4b1f-96ed-5fa5cfe702e3 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation ScriptoriumWS: A Code Generation Assistant for Weak Supervision
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e2df8e1-3ce8-4888-8d33-589c4ba98afa · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Autows-bench-101: Benchmarking automated weak supervision with 100 labels
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc286cef-bb27-4d85-80b2-35448f585fb9 · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997eca7d-7afe-4568-bec9-8d04b8b2df7c · outbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation ""Calculate readability metrics for response
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f017e17-0894-4340-9915-1d2ef361703d · inbound
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 286f4207-5d0a-40c4-b6ab-05c718a80828 · inbound
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b7d92186-e247-4ba5-ae91-d64872b60041 · inbound
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.