Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2408.13006.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:25.799258Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation c736d4fb-3f19-43d8-bef9-b167f1e36ba6 · inbound
Towards an AI co-scientist Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57ba14f9-bb84-44fb-b442-9d6f7946c53a · inbound
Large Language Models for Predictive Analysis: How Far Are They? Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b472d78c-92ea-4f16-8e06-1c1a56c278ee · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9c9731-28a4-45d0-914a-91332507933e · inbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef315790-9b74-4f37-80c1-6a150431efaa · inbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60835722-30b9-4029-95ec-185286065f5e · inbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da944267-1197-4876-a579-307bef239571 · inbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fef8a8-1a7f-49d7-b3b8-a8bf0d604315 · inbound
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dabb476-796a-40e7-a1a1-95629d87f9c2 · inbound
Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0bb244-766b-40a7-bb26-8341a1390907 · inbound
AI Propaganda factories with language models Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c465ca9-2c19-4f63-a87d-aee427466f19 · inbound
IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb117fe-1c82-4dba-8867-5cee665d83b0 · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49d3437-8437-407b-8c8d-2b048b58bf33 · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d85efb-00b7-4abd-8b08-b003eaafd557 · inbound
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfcb7436-1bf7-4185-b89a-df84914418af · inbound
Iterative Finetuning is Mostly Idempotent Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2209723-09be-42a7-9623-9f00e582529d · inbound
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7133626b-ec70-40b2-b733-4667d82b46c1 · inbound
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ff3a2b6-4a2a-453b-8bed-405c8d29267d · inbound
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.