Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:52:06.177196Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 11 inbound Pith citation observations for arXiv:2501.10970.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:52:06.177196Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:04:04.898211Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
16 of 16 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 13f32c02-28a1-49fc-ade2-e7e6933f5785 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 90eae042-3e38-400e-a4c5-ccc256358b7c · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs See https://vicuna
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b12c757d-94c0-4ee4-a313-be85b1e2c78d · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 08de9c5b-7a2f-4a60-9a21-c4b46433b7db · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ea44f5-ba6d-4e30-8419-d7e9981f9060 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ca41a017-c094-4fdd-8620-5be56b0a9fe6 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Cutting Through the Clutter: The Potential of LLMs for Efficient Filtration in Systematic Literature Reviews
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7fbe811-d9ac-4fa0-b3eb-363f61e305f8 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs RewardBench: Evaluating Reward Models for Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d338e07-2382-4d6e-895e-b3229c4c8e43 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Evaluating the Performance of Large Language Models via Debates
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee777f6-f7bc-4444-b576-0c1340fb7850 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90334cac-4812-4cba-8f86-a8b044eff7e5 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0594e65e-0025-4dff-96e8-0b0d2503fad3 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs This is essential to ensure that the annotators are suf- ficiently reliable and the ε value is appropriate
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 567d8ad9-d58d-4ef1-b6ad-d695186b219f · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs combined annotator
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5f33ddf9-04aa-46ed-b9f8-5acbd5b6d39c · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 582– 601, Abu Dhabi, United Arab Emirates
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c6c8a067-0351-4cc4-9187-f22ece3bf4bc · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Can LLMs Replace Manual Annotation of Software Engineering Artifacts?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6198b9-16d7-44dd-812a-6d2be74286b1 · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Can LLM be a Personalized Judge?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece62fb5-22e2-41d9-94c4-c41905e10c3c · outbound
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al
Reference 4635
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ba4a2884-781c-47f3-be0c-4ac0a44f9252 · inbound
How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 23413ca2-6ffe-41c4-ab1f-82c3100367c4 · inbound
Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76b4449-d78a-4ec0-baf0-65a16d2ba282 · inbound
Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88be4fb8-bd12-4cb5-88ef-726c42c7928a · inbound
Multi-Domain Explainability of Preferences The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2a8f9b-2df4-4727-9cb3-89df77f03b3c · inbound
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b958903d-046f-4147-8229-74b87bd1188b · inbound
EduCoder: An Open-Source Annotation System for Education Transcript Data The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cbd87949-a508-4ff1-9728-fbf3ffe6c530 · inbound
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aeb6484-a2ca-438d-bea8-95f5b98646f7 · inbound
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05fbd995-8ae2-4e05-9f4d-6350a5a8bb39 · inbound
Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f17c4560-d09f-4ffa-9ef6-05541150e15c · inbound
Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 04c52b5d-503c-4852-b7e2-411c3c0f2c35 · inbound
Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.