Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 6 inbound Pith citation observations for arXiv:2507.21028.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21028 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:07:16.662207Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:18:36.496381Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.866103Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 716e44de-a45b-425e-81ec-33135f596d20 · outbound

This paper cites Children - Emily Thompson Demographic Info: A 4-year-old girl from Seattle, living with her parents, attending preschool and developing strong language skills.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Children - Emily Thompson Demographic Info: A 4-year-old girl from Seattle, living with her parents, attending preschool and developing strong language skills

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:18.756411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.474260Z digest=sha256:08238f51fb8bc46fa33dcc3e0aad70d0fd070f15046d8c164ef479c289ead19b

Observation cafdf5d4-6789-4ec7-b9a2-d7c293089ccc · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.589727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.485997Z digest=sha256:c3ba08da07c57c7ee2d3981863c7fe08b0503d8c21bf8fae0de90ba97abdf4a0

Observation 1e236875-3995-4f26-8206-c7be88fe7143 · outbound

This paper cites [The Start of Assistant 1’s [evaluation content]] Clinician - Dr.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation [The Start of Assistant 1’s [evaluation content]] Clinician - Dr

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T13:07:18.491719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.496469Z digest=sha256:70d79cfefe451d0172594fd6ccdaa114187346de0814b67c02006cf83c8f4f67

Observation 3bc51568-58a8-41c4-90e6-d29a5b35d111 · outbound

This paper cites stakeholder name.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation stakeholder name

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:18.067418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.538905Z digest=sha256:fc28c1f1bdb9abd7578d2ebd4549334f19d27f0a6a5cf576de204d3336bdaa15

Observation fedf7732-ddfc-4f6d-b1c8-f1916c3b6c93 · outbound

This paper cites RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.449854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.449854Z digest=sha256:edcdef89a63e00e0ae4e56e24a94bbbdcd0334f20823cb9400c29dfe30ccc158

Observation 3098da55-bbef-4d16-8d7d-bb40b5170521 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:07:16.461555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.461555Z digest=sha256:35a4e68767acd35b441ae1750e14d24860402e9c13c74c91b12f9dbf583aa4ea

Observation dea287df-0d39-4a14-ae96-f922a9307eed · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.397249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.506924Z digest=sha256:75deae588b7b2667243e7fc0ce73fd667347cac15ae75f97fa4ec9e8ae1ca19a

Observation 9805c255-0191-43a6-9321-c19360f8d55c · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.262717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.517819Z digest=sha256:5a5e29bd8f8e5df4ec1082b194c6e3036266919196f311abc0f5154850899aae

Observation a7a8cda1-9b5b-4618-bef5-bcc3a7b28e95 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.169335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.528319Z digest=sha256:260a9ee87fba1ad742c2454db7f3faa12332deb5f260b34d15f2e037bd817775

Observation b9f1750c-aa7e-41d7-8bd0-e405bddc597f · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.936064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.550550Z digest=sha256:2a2b72ced43738ecfec12199487bc12f9a7484a690f50b53bef7d7d464f99abe

Observation cf8ca2dd-abba-4100-a234-6543f5b72cd3 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.828232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.566947Z digest=sha256:fcadc19ccfe8e9cf4af706618b148301c34587b52786401737f25f1522657711

Observation 8a89f0eb-e735-4355-b849-3ac8c3bd77ce · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.709490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.582300Z digest=sha256:1d5054731ac187f53ea41e2d7f2f8ad5545bae7424ded51cc4ee0820dd826afe

Observation 9987aec2-8cf4-4b6b-ab7d-9a6cedcf1744 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.612913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.599811Z digest=sha256:9ca41856ac6cbe04a28ea4325fbcf1cad3059eccb41bd5dd8d2213b872cbc7c9

Observation 85450635-3253-4409-b20f-ecfe71643ea6 · outbound

This paper cites Stakeholder Name.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Stakeholder Name

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.514777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.615269Z digest=sha256:001c1e0c72138230a601a20c2ced31a2ac6fa0c20f86264f9d0a071b9b1ae130

Observation ce1f668d-c750-4690-b07f-2f95274f4cab · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.419480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.629445Z digest=sha256:f9b3b42a325b49a9fa38b0bd267a733721c9772112116f9535d0cb2a624a01c3

Observation 9c348cef-30ef-4152-9936-fc958c9ac100 · outbound

This paper cites NO MORE COMMENTS.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation NO MORE COMMENTS

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.291123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.643778Z digest=sha256:0247c58493161270294c6d833a7c3cef5af616f09963565e0027d5309d5e005b

Observation 32637be4-9490-4266-95fa-00e89f0511fd · outbound

This paper cites Here are the initial evaluations from all stakeholders: {phase 1 evaluations } Your task is to evaluate these initial assessments based on your perspective and/or specialty.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Here are the initial evaluations from all stakeholders: {phase 1 evaluations } Your task is to evaluate these initial assessments based on your perspective and/or specialty

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.197620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.651451Z digest=sha256:6cc1bd9bd33b1e0871a731a20de4f2a5dbea714bab7fe82911eb0ff83e53b4a5

Observation 680101d2-51c8-41b8-86bf-ceb04b3fcae2 · outbound

This paper cites NO MORE COMMENTS.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation NO MORE COMMENTS

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.034612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.662207Z digest=sha256:86145425d520608b90f713f954e821445d5c5ed108814ad0fb22dff88beb4c59

Observation 4ae65409-9274-4001-8058-71cc0a4b299b · outbound

This paper cites Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:07:16.926691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.413317Z digest=sha256:b915bea834fee57fae7a2992217e07e9f01cd25e030d1f2f8e37d959381a2b67

Observation cad370de-dd9e-4e7e-9ffd-0ff29d511c6d · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.404481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.404481Z digest=sha256:458bbc0186a150c155a532325458349ab857c5611ebb6795b71634be4ae8ded5

Observation 31f1bc9f-2122-4687-894d-ac7021079360 · outbound

This paper cites LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.435775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.435775Z digest=sha256:a5ec1d81d2752568f23ca231bfbfaa754e6b71d16369d48ae7605bd324c27302

Observation eaa8a27d-a441-4481-bcb0-058c0b67b85c · outbound

This paper cites In Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, New York, NY , USA.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation In Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, New York, NY , USA

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.422913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.422913Z digest=sha256:0a3b154c3b2f068bdfaaa59c5232c3a224997b28563f92bc74abb7fe2e649be9

Pith citing papers

Observation 9a198458-32f2-4503-8984-9026ed1cded6 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.685017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:caadeab7a14d9f81ff197e42babb05ff88db73d1744fdc970ca58baac74ede2b

Observation d7e3ba71-1df1-4d01-bf4d-b237be20aaa5 · inbound

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation cites this paper.

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:10.316870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T08:21:55.648930Z digest=sha256:da3437b9675ed69fdb48bb7f3b8a7d3e4843db03de785acd81dd3c105568e662

Observation cf1850b0-462a-4e6e-8579-da1067759e4c · inbound

Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework cites this paper.

Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:06:34.046531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T03:06:05.168589Z digest=sha256:4e7033681bd915c89f53f38c9a7fa4e2baead9498dfc104ef0ecd88e0cfd3a14

Observation e0d301f2-7ff7-4fe7-8bdc-b8f57a1eefc6 · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.750906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:363566d0b8cc8c13cce670ab5285cb50a1789e3a4daf12202054aa38c6af8314

Observation cfb32351-f75a-4e78-8076-1e30c3cd3f09 · inbound

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments cites this paper.

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.867675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:25:51.788544Z digest=sha256:3663bae1e42654964413920684e327b34e445e370d7bf22b725c0ce830ffcffc

Observation aa634f94-6298-4de8-8bb9-44b4a3b196de · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:36.496381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:36.496381Z digest=sha256:d09a2faf2ae0de4e38592750fec06cc91dc33074d4e417efd6b23ba9a2c756f2