Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 6 inbound Pith citation observations for arXiv:2507.21028.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21028 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:07:16.662207Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:18:36.496381Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.866103Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 716e44de-a45b-425e-81ec-33135f596d20 · outbound

This paper cites Children - Emily Thompson Demographic Info: A 4-year-old girl from Seattle, living with her parents, attending preschool and developing strong language skills.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Children - Emily Thompson Demographic Info: A 4-year-old girl from Seattle, living with her parents, attending preschool and developing strong language skills

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:18.756411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.474260Z digest=sha256:f6a1084bc37b21f0afbdbffc1bd1c15bd538777ff20e362833cfc7ecc2146285

Observation cafdf5d4-6789-4ec7-b9a2-d7c293089ccc · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.589727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.485997Z digest=sha256:1a650625681f6c1389c29cda270fba7024e280d12dac3438adc3e73082e4940d

Observation 1e236875-3995-4f26-8206-c7be88fe7143 · outbound

This paper cites [The Start of Assistant 1’s [evaluation content]] Clinician - Dr.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation [The Start of Assistant 1’s [evaluation content]] Clinician - Dr

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T13:07:18.491719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.496469Z digest=sha256:99228a94019899ebb0f58f9350da2638ae5e4aff1f98d7330c837f3fa0d72da3

Observation 3bc51568-58a8-41c4-90e6-d29a5b35d111 · outbound

This paper cites stakeholder name.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation stakeholder name

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:18.067418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.538905Z digest=sha256:3b04ccbddaa6cad6c43a1986a9cc87c11e453ba1a308c02178bea683315d4ab2

Observation fedf7732-ddfc-4f6d-b1c8-f1916c3b6c93 · outbound

This paper cites RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.449854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.449854Z digest=sha256:0a21275561bdb4b387a24cac9c631d8cda7372b8352f232d617e30996323fb33

Observation 3098da55-bbef-4d16-8d7d-bb40b5170521 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:07:16.461555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.461555Z digest=sha256:8b1f36902b25c48130ba9d9abd3cea134cadbd20c8f8ef79003244cddc79c404

Observation dea287df-0d39-4a14-ae96-f922a9307eed · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.397249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.506924Z digest=sha256:867d55c0bfa7c6e66e1ddf8bd9f720f40c44fc3261a67370dc9c36ddb9e4cc65

Observation 9805c255-0191-43a6-9321-c19360f8d55c · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.262717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.517819Z digest=sha256:67075dd00f567430ca2f7d0f29f0772335ae38312db10f7d55ca1c3a647299ad

Observation a7a8cda1-9b5b-4618-bef5-bcc3a7b28e95 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.169335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.528319Z digest=sha256:f0ad2543b9fe370a97a4f6883d4643ac47900f1bdeb956a46665ecc82df48229

Observation b9f1750c-aa7e-41d7-8bd0-e405bddc597f · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.936064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.550550Z digest=sha256:b8b4d52859e0b98c7174d283ee8e032372c0d5b78812247ad8ed6949fac86c24

Observation cf8ca2dd-abba-4100-a234-6543f5b72cd3 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.828232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.566947Z digest=sha256:e429c76b490505230d3586713c2aefb75923881ef10323513c945413139059d4

Observation 8a89f0eb-e735-4355-b849-3ac8c3bd77ce · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.709490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.582300Z digest=sha256:19e40b71531e3f0a20371e4ca569560f35a823df2790b8851e0cf163c96e0a5e

Observation 9987aec2-8cf4-4b6b-ab7d-9a6cedcf1744 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.612913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.599811Z digest=sha256:7529a4fb02629df64dfd8c50ee9ffff9f4ea622b09da7d97ddfb6e018a384345

Observation 85450635-3253-4409-b20f-ecfe71643ea6 · outbound

This paper cites Stakeholder Name.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Stakeholder Name

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.514777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.615269Z digest=sha256:5f7e7b2e8937f7bfa5610aeb34c6e04b821126acbc7052b5cdd91df235c1fb9c

Observation ce1f668d-c750-4690-b07f-2f95274f4cab · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.419480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.629445Z digest=sha256:7efd40b500b5ca8247c9c2c51e03fd15ff4c2af7aada50e8d395dcc160dddce7

Observation 9c348cef-30ef-4152-9936-fc958c9ac100 · outbound

This paper cites NO MORE COMMENTS.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation NO MORE COMMENTS

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.291123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.643778Z digest=sha256:d863f653ac913a25893e9c63029267f896fc871afb87776c4accb7c57d53d9a9

Observation 32637be4-9490-4266-95fa-00e89f0511fd · outbound

This paper cites Here are the initial evaluations from all stakeholders: {phase 1 evaluations } Your task is to evaluate these initial assessments based on your perspective and/or specialty.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Here are the initial evaluations from all stakeholders: {phase 1 evaluations } Your task is to evaluate these initial assessments based on your perspective and/or specialty

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.197620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.651451Z digest=sha256:dcca1f99d6bc5ada1c531c9ba2d536bf5f67e6f38904dbc9593afcc530ed5344

Observation 680101d2-51c8-41b8-86bf-ceb04b3fcae2 · outbound

This paper cites NO MORE COMMENTS.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation NO MORE COMMENTS

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.034612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.662207Z digest=sha256:c67f8cf0b8fff2e79156361336f7a03056cda3f5175862814f2868dc26c5d30f

Observation 4ae65409-9274-4001-8058-71cc0a4b299b · outbound

This paper cites Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:07:16.926691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:07:16.413317Z digest=sha256:59bb3cdefe9faf7ebd07561718bb97abf284c98af83c14fbe3dc791c0d783376

Observation cad370de-dd9e-4e7e-9ffd-0ff29d511c6d · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.404481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.404481Z digest=sha256:8e031098e019a2101aec9d1627ec2ae313e20d25d8e661085fcbf62ebfa0115a

Observation 31f1bc9f-2122-4687-894d-ac7021079360 · outbound

This paper cites LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.435775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.435775Z digest=sha256:76fb1f1d9c67d2be7de2c4d6d864187c282f001296af7f6def0834df7d216c6a

Observation eaa8a27d-a441-4481-bcb0-058c0b67b85c · outbound

This paper cites In Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, New York, NY , USA.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation In Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, New York, NY , USA

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.422913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.422913Z digest=sha256:eb3c6ccebce77c50fce43c0d6639e38489dd384e377eb5ece5a82a2a2d284070

Pith citing papers

Observation 9a198458-32f2-4503-8984-9026ed1cded6 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.685017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:d96217e93d55d19bbbae1fc585b1877c7a3262544305eed04fbb5fbe8ddb2399

Observation d7e3ba71-1df1-4d01-bf4d-b237be20aaa5 · inbound

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation cites this paper.

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:10.316870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T08:21:55.648930Z digest=sha256:5f394b9befc46515d4b2dd75c61743cbcd90bd1efa617efbac77692fe21622fe

Observation cf1850b0-462a-4e6e-8579-da1067759e4c · inbound

Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework cites this paper.

Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:06:34.046531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T03:06:05.168589Z digest=sha256:335e0d69612d82eb51635b44421724c963384559612864e524b00abbb9970d22

Observation e0d301f2-7ff7-4fe7-8bdc-b8f57a1eefc6 · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.750906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:7d7dca80c0eee72d28980263c3f66af7f21db7cbdccaaf3549154455f6cea8fd

Observation cfb32351-f75a-4e78-8076-1e30c3cd3f09 · inbound

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments cites this paper.

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.867675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:25:51.788544Z digest=sha256:a6f02ac055ca63ed4087eecd4b39619a0917c583b40a33f615f79b290332388b

Observation aa634f94-6298-4de8-8bb9-44b4a3b196de · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:36.496381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:36.496381Z digest=sha256:4c1d42e5a3c7e6ec397c73f2969664b49252f48d00ee9c4da23c2bba5c3cf4d3