Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2403.02839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.02839 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:06:32.364901Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 440556cc-00a4-4867-a40e-587797f21bc3 · inbound

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models cites this paper.

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:32:17.998336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T23:32:17.279836Z digest=sha256:d297363ed32777b91452b29b2a6978f4b5361c433848a6706333d6b129a4b4a9

Observation 619b6d6e-f015-48cd-ab29-8af910c92656 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:44:49.795483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:a4523d6bd7f96a05b01be1318b34db0e8e3c9b3cef7499d7694af30a8f9de1f6

Observation 8fba2e7e-e1bc-4ebc-b016-358942c9def7 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.974779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:083e23012294418fa392ab5ee12e55271dbaac5b235805331cc30b2c3797ac4d

Observation 6ca8af0d-7ab7-4a8b-97ac-91e9263a8cb8 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.501674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:3b0cf87adc8521570bd138eda581ec586c74e63d02a5b0cba69b325e063f8418

Observation 6b368de2-740e-4e4d-86c4-5ba9d9cdd820 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:43.845001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:6446c8a4593dd4522de53559d2116574d09128440c325fe09277566a338f80ac

Observation 33615456-3847-4488-b199-59963e60b2e0 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.802863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:36dc8383fcfce7715793aded2b500316b20db3a0510b48c32cf2110abaa10cd7

Observation 338cffb2-af94-4e6e-9237-ea82b87c11ee · inbound

VLM@school -- Evaluation of AI image understanding on German middle school knowledge cites this paper.

VLM@school -- Evaluation of AI image understanding on German middle school knowledge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:06:32.364901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:06:32.364901Z digest=sha256:f4715f5018ae254ee021b28c9575e7769339d2deddb8c1fe8965805ff6e4504a

Observation 35245c9b-6ac1-4d25-abec-e0eb5ed53118 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.829032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.829032Z digest=sha256:3400ebd4343d75529d7abb93ac8b1dce333ea3bed6df0def992790f0eb02737a

Observation 6630e635-cbe4-40e3-8524-e2f4199fe69d · inbound

On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization cites this paper.

On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:10:23.222553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:10:23.222553Z digest=sha256:967bc6abdfef58045eb2d1c1e3ff2f293757e3230c8c667ee1f4198df982b7c8

Observation f9e7aece-e3a1-4f62-a2cb-de4e76c4482a · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:56.912415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:56.912415Z digest=sha256:bafd0064c9c74bdfa2d5e2cf6d2d1d744004e50bcfce41608a831190b26c83de

Observation 094d9722-3e48-41b7-b780-5842449971b8 · inbound

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation cites this paper.

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:20.133943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:20.133943Z digest=sha256:779abbabe34f1b4ebc094f4e543126cc24ee860a0c2f01c783a558996db556a5

Observation 2333b4db-20bd-4583-a8f7-2d4feab28d01 · inbound

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning cites this paper.

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.005139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:19:55.849451Z digest=sha256:6cfefa6e0e192373b030d9919b1fb2fee00ca63fa19477868c994653de5326e0

Observation d2d22ba6-a5c3-4f42-a572-30e2337e13eb · inbound

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability cites this paper.

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:17:02.696947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T01:12:44.476894Z digest=sha256:0cba46b5bc3bbae721d7c43100d9db69cc2dbcb39d820c31ebe549c7f191c307

Observation 0d67d5d2-beea-4717-858a-1893de695dc3 · inbound

Section-Weighted Hybrid Approach for Legal Case Retrieval cites this paper.

Section-Weighted Hybrid Approach for Legal Case Retrieval An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:56:39.300020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T08:39:28.510255Z digest=sha256:da894e2f10baaaedee64db8b105347a7341cf47c86448d458e7d1cb18fc654f1

Observation 9d426666-40ad-4b7c-9c99-16f0bd40c3e4 · inbound

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators cites this paper.

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-06-27T21:41:18.571899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T21:39:26.268337Z digest=sha256:0b9b691b24d1301dc1d4739c68fd9a5a67d82781476e4467b064fe1f77ad347d

Observation bd79bc58-d3bd-4eb0-8bed-8afeccf7dde5 · inbound

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition cites this paper.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.492775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.492775Z digest=sha256:affa724bd326a2e1b9fe3540edce5d7a877e30f53db5e8433859053420267b3b

Observation 919a65f8-29a8-430a-a1ec-2d0a3744c8f5 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.090862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.090862Z digest=sha256:6a6960fcdb630d29edea6e7469c6604998364beab986ae79c1da50ced02e79fc

Observation 27885a75-59e5-46f3-9d69-832c82e6d4cd · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 272

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:40.748915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:40.748915Z digest=sha256:5edd4d0c718e7da8d0661ebeb03d3283f126efe7f219d4878317f118387dd14f