Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2403.02839.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:20:18.631186Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 440556cc-00a4-4867-a40e-587797f21bc3 · inbound
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 619b6d6e-f015-48cd-ab29-8af910c92656 · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8fba2e7e-e1bc-4ebc-b016-358942c9def7 · inbound
Benchmark Data Contamination of Large Language Models: A Survey An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6ca8af0d-7ab7-4a8b-97ac-91e9263a8cb8 · inbound
ShieldGemma: Generative AI Content Moderation Based on Gemma An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d7674992-009d-440c-8e01-33007f920f1b · inbound
Engagement-Driven Content Generation with Large Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3c33c6-80cc-46e2-967b-4e9f011bbede · inbound
AI Tailoring: Evaluating Influence of Image Features on Fashion Product Popularity An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b368de2-740e-4e4d-86c4-5ba9d9cdd820 · inbound
A Survey on LLM-as-a-Judge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5761062c-4fdb-49db-8c14-245720ef9903 · inbound
LLM Augmentations to support Analytical Reasoning over Multiple Documents An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33615456-3847-4488-b199-59963e60b2e0 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 90e7b958-6a00-4b54-b523-3f09bac08a1c · inbound
Let your LLM generate a few tokens and you will reduce the need for retrieval An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 420787e1-dded-45f6-a46b-35829b5a81fd · inbound
An Exploratory Study of ML Sketches and Visual Code Assistants An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f8fbf1b-aae6-4303-a4a9-1e1a3577189c · inbound
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fddcf01c-8429-4d67-9ebe-15c82ae464ae · inbound
Efficient Multi-Agent Collaboration with Tool Use for Online Planning in Complex Table Question Answering An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee596d2-8933-468b-855d-7af07fcd1ab4 · inbound
Towards a scalable AI-driven framework for data-independent Cyber Threat Intelligence Information Extraction An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5f9fb0-9d97-4ce2-98f3-c9f04f6c252e · inbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddaebaf-ba4a-4199-a5ec-917255b161ed · inbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b26c0e-edf8-4f6f-aab4-72e2cdac15a4 · inbound
Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b67cbc0-95b8-43ef-86fd-5d34564de8ec · inbound
Combining Large Language Models with Static Analyzers for Code Review Generation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf976ba1-edc4-46d4-bc82-8340d2bdc577 · inbound
FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eddbdcaf-db0e-4888-8247-92ef29f2589b · inbound
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338cffb2-af94-4e6e-9237-ea82b87c11ee · inbound
VLM@school -- Evaluation of AI image understanding on German middle school knowledge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35245c9b-6ac1-4d25-abec-e0eb5ed53118 · inbound
A Conceptual Framework for AI Capability Evaluations An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6630e635-cbe4-40e3-8524-e2f4199fe69d · inbound
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e7aece-e3a1-4f62-a2cb-de4e76c4482a · inbound
Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094d9722-3e48-41b7-b780-5842449971b8 · inbound
Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2333b4db-20bd-4583-a8f7-2d4feab28d01 · inbound
U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d2d22ba6-a5c3-4f42-a572-30e2337e13eb · inbound
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0d67d5d2-beea-4717-858a-1893de695dc3 · inbound
Section-Weighted Hybrid Approach for Legal Case Retrieval An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9d426666-40ad-4b7c-9c99-16f0bd40c3e4 · inbound
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bd79bc58-d3bd-4eb0-8bed-8afeccf7dde5 · inbound
LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919a65f8-29a8-430a-a1ec-2d0a3744c8f5 · inbound
(Towards) Scalable Reliable Automated Evaluation with Large Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27885a75-59e5-46f3-9d69-832c82e6d4cd · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 272
Source-reported events for the cited work
Unavailable: canonical work link unavailable.