Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2403.02839.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:06:32.364901Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 440556cc-00a4-4867-a40e-587797f21bc3 · inbound
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 619b6d6e-f015-48cd-ab29-8af910c92656 · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fba2e7e-e1bc-4ebc-b016-358942c9def7 · inbound
Benchmark Data Contamination of Large Language Models: A Survey An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ca8af0d-7ab7-4a8b-97ac-91e9263a8cb8 · inbound
ShieldGemma: Generative AI Content Moderation Based on Gemma An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b368de2-740e-4e4d-86c4-5ba9d9cdd820 · inbound
A Survey on LLM-as-a-Judge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33615456-3847-4488-b199-59963e60b2e0 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 338cffb2-af94-4e6e-9237-ea82b87c11ee · inbound
VLM@school -- Evaluation of AI image understanding on German middle school knowledge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35245c9b-6ac1-4d25-abec-e0eb5ed53118 · inbound
A Conceptual Framework for AI Capability Evaluations An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6630e635-cbe4-40e3-8524-e2f4199fe69d · inbound
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e7aece-e3a1-4f62-a2cb-de4e76c4482a · inbound
Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094d9722-3e48-41b7-b780-5842449971b8 · inbound
Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2333b4db-20bd-4583-a8f7-2d4feab28d01 · inbound
U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2d22ba6-a5c3-4f42-a572-30e2337e13eb · inbound
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d67d5d2-beea-4717-858a-1893de695dc3 · inbound
Section-Weighted Hybrid Approach for Legal Case Retrieval An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d426666-40ad-4b7c-9c99-16f0bd40c3e4 · inbound
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd79bc58-d3bd-4eb0-8bed-8afeccf7dde5 · inbound
LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919a65f8-29a8-430a-a1ec-2d0a3744c8f5 · inbound
(Towards) Scalable Reliable Automated Evaluation with Large Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27885a75-59e5-46f3-9d69-832c82e6d4cd · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 272
Source-reported events for the cited work
Unavailable: canonical work link unavailable.