Pith. sign in

Paper Citation Record · LEDGER

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 8 inbound Pith citation observations for arXiv:2504.17087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17087 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:53:25.065089Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:20:29.355664Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.227079Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0dd83741-44ae-4dca-ad01-b80f8ac6a110 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.707459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.707459Z digest=sha256:3ad94c8a95f4ffad3779e72812396834a7535d0c234b5a89f97271b669de2e53

Observation 8abf0c6f-6050-487a-8758-2119a025dd38 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.840206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.840206Z digest=sha256:e06d4fa5c4dc2d61e8532d4b741f6633b0e3b7d71e966f45deaf784262d77223

Observation 8c1e6529-6777-4d0a-aee8-b7d7d03944d9 · outbound

This paper cites MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.848112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.848112Z digest=sha256:1769d15b185ca83313498936a76141ef123a4f68ceb85e7efd78bd7515f546cd

Observation 08bb1446-c4a7-4589-a742-fe0c0c382d0f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.860155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.860155Z digest=sha256:fcbdf2b02d10210b0e1d943f1d9c9e7ee70f99d30a4843a7d5c352e97c2c12b9

Observation d96debcb-a0ab-4ee9-ac07-e64c686188cc · outbound

This paper cites JudgeBench: A Benchmark for Evaluating LLM-based Judges.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.867753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.867753Z digest=sha256:12f9a278fbea88fba8be171c62e6f165b3adbdd6c8f85ea1ec18f6d7402d361b

Observation 164e792e-d0f6-48b5-9097-56408af56e71 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.871674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.871674Z digest=sha256:d31f9d619480590ecb63725e1009253df2671a370ef47a1f24f9285ab1d719bd

Observation 7609c747-f80d-423a-9148-12fccf3a474b · outbound

This paper cites Self-rationalization improves LLM as a fine-grained judge.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Self-rationalization improves LLM as a fine-grained judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.875340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.875340Z digest=sha256:fb538fcade75df2d0b4e62dedd8bf721a96b3e7e3e048f23972a04630079f10b

Observation 9f34a9d7-520f-44d4-9e48-78738a59461b · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.879314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.879314Z digest=sha256:298ca098ee7287870defe9183ee25efa2e006430439cf1feca52048e66e59e7b

Observation 47dca15f-0f4b-40ed-8137-76226752cf07 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Aligning Large Language Models with Human: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.919360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.919360Z digest=sha256:7687a9b9fd3418d4a40a23dbc65f54e11459c084f06ea0540f0b20e4b4fdcfd4

Observation 0e07aea5-8e62-4d70-b3e6-3f623bb50998 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.990467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.990467Z digest=sha256:0e45bb920b51b4aee93a124616bc2f3b1a5b7a6dd450e1466b8b4b455162f959

Observation 246fa49d-6c82-4101-a111-b0abb1c1e14f · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.046352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.046352Z digest=sha256:db2a246d5f54a49bb90c4f3810202d47a12dace33b313843d1b4c94402ee915a

Observation 852a8c22-4244-4b4e-b17f-86327ee6750c · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.049609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.049609Z digest=sha256:05caa2c5a4d5e8e3580a093f8b5e3d23ad366fc3f96a1929b4da034f68684ec3

Observation 181b3b98-32bd-46b0-b91e-2b6f4e16afd7 · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Do Large Language Models Know What They Don't Know?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.053262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.053262Z digest=sha256:3e7824e26fb07c254a34a67026d882a84636133b8dc815e07c6d757e1f41f621

Observation 2e0b0f99-3c9a-46e2-b141-122e026017c2 · outbound

This paper cites MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:53:25.139148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:53:25.056988Z digest=sha256:b8235690e9c0b0baca24ace104e42c5ef2a51be3bb0683288b47291fc857532c

Observation 06e17a48-7e15-4e51-9542-c596ecc4391e · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments BERTScore: Evaluating Text Generation with BERT

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.060888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.060888Z digest=sha256:bb808a33f6a254dfef7d0733e5da545c3e66f54dd4a91dadedca9057fb5d5284

Observation 277ec298-e4d6-44ca-be0b-cf43c770d623 · outbound

This paper cites Detail of different rubrics In the single-agent precision comparison section, we analyzed the impact of four different rubric configurations.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Detail of different rubrics In the single-agent precision comparison section, we analyzed the impact of four different rubric configurations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:53:25.370128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:53:25.065089Z digest=sha256:08853a16d3d4e4f4982283acc1674f94ed169ec02dff5a1bfbf84d4e85065cad

Observation b2e1119a-76ad-4be0-9649-d4222655c614 · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.855417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.855417Z digest=sha256:4a52f33d196ea77cc5bb00f3addb6defb403d285c7d21c92aa22c35a89087534

Observation bd9b1b95-b754-403b-b922-e5608079069c · outbound

This paper cites MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in Summarization.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in Summarization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.852244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.852244Z digest=sha256:14d2b62dacb28f13b33f614cb1d53bd8b083b7ff3f73c77c0770b479738f0b9a

Observation 475da4c9-edf1-407a-b872-92fe5788823d · outbound

This paper cites Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Generation.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.863952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.863952Z digest=sha256:d7e11ba4c9cca317a2fb80d0af14c8baabe95d60ba7c9b5e6f56ce805552ce44

Observation 320b609c-1579-4ee3-afd3-ad55c33d9230 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Evaluating Large Language Models: A Comprehensive Survey

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.776501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.776501Z digest=sha256:0814a6cd1392a490fc450dc163c842b6791f1f8f1df8f594b973028faae3d055

Observation 8ce2fbba-1898-4282-9b13-4a7dff31c2ac · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.844122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.844122Z digest=sha256:3b308b967c7e4b61ffad068bdbf48fe50c7540ca0349dc444fb719d24fa236ea

Pith citing papers

Observation b82cfc1d-e4e6-4d7d-b402-abcc60a7d50d · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.327951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:907ca4a672ac86f6a1c362b9af9e2d553835ab713fcf0a474f8928ac9b31e448

Observation 06775abf-bd4c-4d41-9862-16567315a85c · inbound

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges cites this paper.

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:25.789277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:27:21.060953Z digest=sha256:4eb8599302216ab8faa280c2988a027b71fec3749faa7964f7566cea898d0e27

Observation 5ad57e2c-14b7-4923-bd6c-9e68778c3c2f · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.440905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:09f70afa909a571c3495a68fdc7cea072b2f5c3ab0cf0b762e024545addf2180

Observation 09773501-43c3-4a0a-a3b3-399f1e2445f3 · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:49:38.229229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:968a9b7a9a011d29daeac632ed67f636efb44b634ca4b62d9819f07b0d97b06e

Observation e28d5d64-fb4d-436f-a613-e16f9573ee0c · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:8059f303fa212dd24af6a36eb1b8febc18e93924c52d50a3ad06e663ff338d6d

Observation 2c691ba5-62e3-4848-bc41-35bc1552a65d · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:28.286096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:28.286096Z digest=sha256:b9323b24a0b4d2a65e87970b5766f79a6a617508a8b074f9d897db83c26c5616

Observation 829ec610-9d70-49c4-ad2b-ffa70b67f5d5 · inbound

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges cites this paper.

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T00:32:29.254980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:32:29.254980Z digest=sha256:8c9c31a841581957884235d49feca992d2a723989233af437c125770ab9127c7

Observation 79bcbf52-710e-4541-b49b-13b4aeca0da7 · inbound

HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems cites this paper.

HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:20:29.355664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:20:29.355664Z digest=sha256:51a5498cd2d679c2521ccf53e4d038a6c325b78e1ea26132d47566512ef8824c