Pith. sign in

Paper Citation Record · LEDGER

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

As of 14 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 11 inbound Pith citation observations for arXiv:2501.10970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10970 v4

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:52:06.177196Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:04:04.898211Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 13f32c02-28a1-49fc-ade2-e7e6933f5785 · outbound

This paper cites an unresolved cited work.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:52:06.519040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.123413Z digest=sha256:3a4f92cf9a4ba7a316be82a6cb0983ba61ec5ef016939e66563767f6c27650b2

Observation 90eae042-3e38-400e-a4c5-ccc256358b7c · outbound

This paper cites See https://vicuna.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs See https://vicuna

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:52:06.592482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.030738Z digest=sha256:3ba04838c3dd31a8acf9b102e76b31fad315760c2df9d3775dad791eb20bb80d

Observation b12c757d-94c0-4ee4-a313-be85b1e2c78d · outbound

This paper cites an unresolved cited work.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:52:06.474938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.149247Z digest=sha256:7cb095b4806d9b1b84191a1cc5bfd68f27f706f233983f3b2f7174c2f21ea3aa

Observation 08de9c5b-7a2f-4a60-9a21-c4b46433b7db · outbound

This paper cites ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:06.049232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:06.049232Z digest=sha256:7eb7d926d4b25ed3462dbc4ab1087a885d64f7330f9f2bcaae0a6ff9c94bec4e

Observation a1ea44f5-ba6d-4e30-8419-d7e9981f9060 · outbound

This paper cites an unresolved cited work.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:52:06.436005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.169709Z digest=sha256:17f99e9b4d13b3b35ed1b95b2d5bcbced9816659c0b712b820f3996683188c11

Observation ca41a017-c094-4fdd-8620-5be56b0a9fe6 · outbound

This paper cites Cutting Through the Clutter: The Potential of LLMs for Efficient Filtration in Systematic Literature Reviews.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Cutting Through the Clutter: The Potential of LLMs for Efficient Filtration in Systematic Literature Reviews

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:06.066702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:06.066702Z digest=sha256:80d08ccda5ad452fd0a0a8f27407821e21e7579ef5a46f83a626e6c3b37a9569

Observation f7fbe811-d9ac-4fa0-b3eb-363f61e305f8 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs RewardBench: Evaluating Reward Models for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:06.075421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:06.075421Z digest=sha256:c720c61b66c59ab67407d138e3af176cc327cf144f70475dbc32501e33f20991

Observation 5d338e07-2382-4d6e-895e-b3229c4c8e43 · outbound

This paper cites Evaluating the Performance of Large Language Models via Debates.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Evaluating the Performance of Large Language Models via Debates

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:06.085036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:06.085036Z digest=sha256:e2f873546ef722d5f8bd7d769700e0277815116ee4f670365f0395092b01132a

Observation 7ee777f6-f7bc-4444-b576-0c1340fb7850 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:06.114924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:06.114924Z digest=sha256:7a8dabea0412e0c82b2d501bc91357a8f98d4ca618b1c2160343fa5cda7c99ea

Observation 90334cac-4812-4cba-8f86-a8b044eff7e5 · outbound

This paper cites an unresolved cited work.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:52:06.492527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.132724Z digest=sha256:460d4e8d2d39c0333b01008a3337f6dfefee31a39ba71fe2063a5e57bf8433bc

Observation 0594e65e-0025-4dff-96e8-0b0d2503fad3 · outbound

This paper cites This is essential to ensure that the annotators are suf- ficiently reliable and the ε value is appropriate.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs This is essential to ensure that the annotators are suf- ficiently reliable and the ε value is appropriate

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:52:06.454676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.159393Z digest=sha256:60183f11e8d07a73074565e8357b6e1f50078ce584e30be98c59d29384ff3c31

Observation 567d8ad9-d58d-4ef1-b6ad-d695186b219f · outbound

This paper cites combined annotator.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs combined annotator

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:52:06.416222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.177196Z digest=sha256:ba4fe5d07fd65639abe9a2d560c093eb5b13fcbb058997c276ceff7cf9756c6f

Observation 5f33ddf9-04aa-46ed-b9f8-5acbd5b6d39c · outbound

This paper cites In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 582– 601, Abu Dhabi, United Arab Emirates.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 582– 601, Abu Dhabi, United Arab Emirates

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:52:06.567942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.059367Z digest=sha256:076798395141266e8a7745f2efa067931752a0f611b9b39f5d2c5345f37c712e

Observation c6c8a067-0351-4cc4-9187-f22ece3bf4bc · outbound

This paper cites Can LLMs Replace Manual Annotation of Software Engineering Artifacts?.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Can LLMs Replace Manual Annotation of Software Engineering Artifacts?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:06.019729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:06.019729Z digest=sha256:091d70dd1c6d8f08fbcb4e2fad606b6b1c411a819f51f7cd3312abd650da143e

Observation 2f6198b9-16d7-44dd-812a-6d2be74286b1 · outbound

This paper cites Can LLM be a Personalized Judge?.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Can LLM be a Personalized Judge?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:06.039752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:06.039752Z digest=sha256:6067bf328022fe452fe1cb4b30e3113e3543c773179b59ac2f081fa17e0a591a

Observation ece62fb5-22e2-41d9-94c4-c41905e10c3c · outbound

This paper cites Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al

Reference 4635

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:52:06.546772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:52:06.096610Z digest=sha256:912b9191f5f5601c74c524ecd6420a0e9dd6053eeb8f040bed82e1b5863d8c1b

Pith citing papers

Observation ba4a2884-781c-47f3-be0c-4ac0a44f9252 · inbound

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective cites this paper.

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:15:21.289225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-23T03:14:00.526112Z digest=sha256:336f66a614eb30233e722ca014d035df64e2f971b8cba266697405ef2b1b90ba

Observation 23413ca2-6ffe-41c4-ab1f-82c3100367c4 · inbound

Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs cites this paper.

Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:04:04.898211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:04:04.898211Z digest=sha256:30f21b8bf943204acb151e99c402ebf39725ce33fdcefe2d57bd8d976647055b

Observation f76b4449-d78a-4ec0-baf0-65a16d2ba282 · inbound

Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods cites this paper.

Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.541169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.541169Z digest=sha256:6d453f4cfed71d36c7286c8e30f454fa49c479b923e9d36c4147080c1b4ccda6

Observation 88be4fb8-bd12-4cb5-88ef-726c42c7928a · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:06.332229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:06.332229Z digest=sha256:67fe6c3daafd4ace5b0cef08503dc3b0f187895f5167cc0cc69c0aee7762a38b

Observation bb2a8f9b-2df4-4727-9cb3-89df77f03b3c · inbound

ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering cites this paper.

ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.542384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.542384Z digest=sha256:8a3c341a7633d900a7a0340e4844b4d24afec74684361f46bd1345a0132717cb

Observation b958903d-046f-4147-8229-74b87bd1188b · inbound

EduCoder: An Open-Source Annotation System for Education Transcript Data cites this paper.

EduCoder: An Open-Source Annotation System for Education Transcript Data The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:37:05.540440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T05:35:52.219876Z digest=sha256:37f5035e9aee4e587465b49500cc52b086cb6616d24517e380b63585e87a32ae

Observation cbd87949-a508-4ff1-9728-fbf3ffe6c530 · inbound

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications cites this paper.

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:40.826384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:49:40.826384Z digest=sha256:420be9b315507feb0bb501f52bf331750f0d8b9afd812a8a623266fead5c852e

Observation 0aeb6484-a2ca-438d-bea8-95f5b98646f7 · inbound

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? cites this paper.

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:21.683550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:21.683550Z digest=sha256:b94bcae5e5fce0be2816981eb3f62974d7b974c7b994d97bcd04480045e77554

Observation 05fbd995-8ae2-4e05-9f4d-6350a5a8bb39 · inbound

Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models cites this paper.

Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:33:07.548115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T15:33:00.984344Z digest=sha256:c24d339d472917b26a5b4918b3bcd8f500b1cbf4c06f4a1b6bbadb2d5dfd8f93

Observation f17c4560-d09f-4ffa-9ef6-05541150e15c · inbound

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations cites this paper.

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.726078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T18:27:50.159580Z digest=sha256:6095a38bc302624468c01ca7137052bb323f815d0d09f28392874a4efb9598d4

Observation 04c52b5d-503c-4852-b7e2-411c3c0f2c35 · inbound

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe cites this paper.

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:58.149738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:58.149738Z digest=sha256:8212acb150da2a13415353ffab298b05c99e6832d2ac74c2c1967fb449417858