Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2403.02839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.02839 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:20:18.631186Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 440556cc-00a4-4867-a40e-587797f21bc3 · inbound

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models cites this paper.

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:32:17.998336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-15T23:32:17.279836Z digest=sha256:5e051b46c2d4f9422b9a937ad677bad49e310b16b4523be5b7a22e3efe9d1f44

Observation 619b6d6e-f015-48cd-ab29-8af910c92656 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:44:49.795483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:1862d739eb3fded66af32e061584bcfbcfb163d2d762917125380f36cf2be2ce

Observation 8fba2e7e-e1bc-4ebc-b016-358942c9def7 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.974779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:646d75ea25b2851ba12525b279e85b12459af1f0b5b9a6cbca1df619a7cce0bd

Observation 6ca8af0d-7ab7-4a8b-97ac-91e9263a8cb8 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.501674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:4474fb5d47f55b3c9bcd8de2e477cb1b2742b80b6f0efc8e01834197ebecfa75

Observation d7674992-009d-440c-8e01-33007f920f1b · inbound

Engagement-Driven Content Generation with Large Language Models cites this paper.

Engagement-Driven Content Generation with Large Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:48:50.691200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:48:50.691200Z digest=sha256:745e471822b1e6dc60ab01728c479967e59cc43e7d1627ef7eacb1ef1ee987f5

Observation 9f3c33c6-80cc-46e2-967b-4e9f011bbede · inbound

AI Tailoring: Evaluating Influence of Image Features on Fashion Product Popularity cites this paper.

AI Tailoring: Evaluating Influence of Image Features on Fashion Product Popularity An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:43.493145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:43.493145Z digest=sha256:7686e20cb3dedb0c6f274756140a2a101a2bc3168085bb0e1f35d349442d5b39

Observation 6b368de2-740e-4e4d-86c4-5ba9d9cdd820 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:43.845001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:cbf4125eac20cb0c93796d8040e826ca31970698b856d324ffe03bc6ddf5e21a

Observation 5761062c-4fdb-49db-8c14-245720ef9903 · inbound

LLM Augmentations to support Analytical Reasoning over Multiple Documents cites this paper.

LLM Augmentations to support Analytical Reasoning over Multiple Documents An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:48.948723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:48.948723Z digest=sha256:3886ee215e3a79b9106914dd2ea89ac8da31c80aca26377a38b9b95ff7ed48df

Observation 33615456-3847-4488-b199-59963e60b2e0 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.802863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:b03f0a9474feef0da09846067f3d1bb034ce5f23c46501b558924f3e5ba104fb

Observation 90e7b958-6a00-4b54-b523-3f09bac08a1c · inbound

Let your LLM generate a few tokens and you will reduce the need for retrieval cites this paper.

Let your LLM generate a few tokens and you will reduce the need for retrieval An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:54:56.998877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:54:56.998877Z digest=sha256:77203c1361e8691aa9e556d97b70516b56bfe980d96011ee70c8f2450309170f

Observation 420787e1-dded-45f6-a46b-35829b5a81fd · inbound

An Exploratory Study of ML Sketches and Visual Code Assistants cites this paper.

An Exploratory Study of ML Sketches and Visual Code Assistants An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:16:19.408399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:16:19.408399Z digest=sha256:7e26024fee3323a8b43b2d6d5409ac79758eccd701778b7b6d594a801fbbe1c1

Observation 5f8fbf1b-aae6-4303-a4a9-1e1a3577189c · inbound

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems cites this paper.

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:22:56.827813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:22:56.827813Z digest=sha256:8e611f42b6a6b03c930f2b69dcd966838c084ab1d5b4b83709d7982ca8fd25ea

Observation fddcf01c-8429-4d67-9ebe-15c82ae464ae · inbound

Efficient Multi-Agent Collaboration with Tool Use for Online Planning in Complex Table Question Answering cites this paper.

Efficient Multi-Agent Collaboration with Tool Use for Online Planning in Complex Table Question Answering An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T23:33:57.625654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:33:57.625654Z digest=sha256:736184d6784a885c2f36cd3c3d12973101ace4f3e46639544818cda7facbca89

Observation 5ee596d2-8933-468b-855d-7af07fcd1ab4 · inbound

Towards a scalable AI-driven framework for data-independent Cyber Threat Intelligence Information Extraction cites this paper.

Towards a scalable AI-driven framework for data-independent Cyber Threat Intelligence Information Extraction An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:36:45.186550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:36:45.186550Z digest=sha256:d3ce84ce77ec788770969c5f5de06babee7a367b37faa5be497696d91b11438f

Observation eb5f9fb0-9d97-4ce2-98f3-c9f04f6c252e · inbound

IC-Cache: Efficient Large Language Model Serving via In-context Caching cites this paper.

IC-Cache: Efficient Large Language Model Serving via In-context Caching An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.302934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.302934Z digest=sha256:1f30f8db874b62a1e58ec7ae91b248c1fe3b83771e659603a41b4755e7c75a0e

Observation 6ddaebaf-ba4a-4199-a5ec-917255b161ed · inbound

Tuning LLM Judge Design Decisions for 1/1000 of the Cost cites this paper.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.196827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.196827Z digest=sha256:936b4e53e9d7dc1c68dfe0e953cb3e12583b6f9f1456daf72351831a3eac0f88

Observation 50b26c0e-edf8-4f6f-aab4-72e2cdac15a4 · inbound

Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation cites this paper.

Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T05:31:14.311143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:31:14.311143Z digest=sha256:84f13169484e207a87d7e1818208e42273f232e77bfb178447de05c341225747

Observation 9b67cbc0-95b8-43ef-86fd-5d34564de8ec · inbound

Combining Large Language Models with Static Analyzers for Code Review Generation cites this paper.

Combining Large Language Models with Static Analyzers for Code Review Generation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:53:28.643250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:53:28.643250Z digest=sha256:6e430c3730816bcedaabdae58f69f366e5f18cc88af5336c22032cb4fcdb1114

Observation cf976ba1-edc4-46d4-bc82-8340d2bdc577 · inbound

FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation cites this paper.

FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T11:20:18.631186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:20:18.631186Z digest=sha256:177a1f5fd6100d13d43c0406540e2a9ef48ed067d6baf08487b845621d6cb459

Observation eddbdcaf-db0e-4888-8247-92ef29f2589b · inbound

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression cites this paper.

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.118892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.118892Z digest=sha256:87ff5455be215f5c2135a81472b4eba9aeb282e0b4d6b1da3d6ceeda2ab4698a

Observation 338cffb2-af94-4e6e-9237-ea82b87c11ee · inbound

VLM@school -- Evaluation of AI image understanding on German middle school knowledge cites this paper.

VLM@school -- Evaluation of AI image understanding on German middle school knowledge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:06:32.364901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:06:32.364901Z digest=sha256:51dd1d0ed8ef483875ca62984d837b0a09fba5fb0b0dfc1f1db5861dd5008f94

Observation 35245c9b-6ac1-4d25-abec-e0eb5ed53118 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.829032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.829032Z digest=sha256:940cfab6da053cc84ee5da4bd089f373f3650f920f1cb73268964e7b418cb7e6

Observation 6630e635-cbe4-40e3-8524-e2f4199fe69d · inbound

On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization cites this paper.

On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:10:23.222553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:10:23.222553Z digest=sha256:3b508a0eb4230e1e0f356a13c1b1668246c82aeff33ba80d4624f3924142d7e3

Observation f9e7aece-e3a1-4f62-a2cb-de4e76c4482a · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:56.912415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:56.912415Z digest=sha256:edd9728fc3bc05dd969b230ae22bf02eef03695e024ebe741631c8eb873d32cf

Observation 094d9722-3e48-41b7-b780-5842449971b8 · inbound

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation cites this paper.

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:20.133943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:20.133943Z digest=sha256:06d204ec637d2baea23c6264eaa580c512c41ac7716b8059b9d6067c4c9ec421

Observation 2333b4db-20bd-4583-a8f7-2d4feab28d01 · inbound

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning cites this paper.

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.005139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:19:55.849451Z digest=sha256:3c6923b95fbc95351f43149c43c97b2bc8d8f9cbc820a5eee5c82a0cae821402

Observation d2d22ba6-a5c3-4f42-a572-30e2337e13eb · inbound

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability cites this paper.

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:17:02.696947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T01:12:44.476894Z digest=sha256:4f94835b14e0cd256f1be50d554526210376470750c8f2a33b9530a19d6f8cd5

Observation 0d67d5d2-beea-4717-858a-1893de695dc3 · inbound

Section-Weighted Hybrid Approach for Legal Case Retrieval cites this paper.

Section-Weighted Hybrid Approach for Legal Case Retrieval An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:56:39.300020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T08:39:28.510255Z digest=sha256:ab49940a57c12f2961b4ad5e2074179613495fe531491048350216a7966feb10

Observation 9d426666-40ad-4b7c-9c99-16f0bd40c3e4 · inbound

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators cites this paper.

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-06-27T21:41:18.571899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T21:39:26.268337Z digest=sha256:e142d33151cac1b02afd912a3218eb8cf6a01dc963ffa2787f35b866aea21647

Observation bd79bc58-d3bd-4eb0-8bed-8afeccf7dde5 · inbound

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition cites this paper.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.492775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.492775Z digest=sha256:aa30b1989d91fe0ae3ea100e100540917a154af92ffafa4af223fba962c59368

Observation 919a65f8-29a8-430a-a1ec-2d0a3744c8f5 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.090862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.090862Z digest=sha256:b94dbc0ad0a15114b7cb18baf2c935201d8dfa083d59932562bfa1a6b4695c54

Observation 27885a75-59e5-46f3-9d69-832c82e6d4cd · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 272

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:40.748915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:40.748915Z digest=sha256:a1fc827d868a833e7fe7f1f4579e54568d46f416638169cc9cd78f70e6d7067c