Pith. sign in

Paper Citation Record · LEDGER

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2501.10970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10970 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:04:04.898211Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ba4a2884-781c-47f3-be0c-4ac0a44f9252 · inbound

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective cites this paper.

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:15:21.289225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T03:14:00.526112Z digest=sha256:832ae99a6abdb379ae311459570265d40cf65b02fae6cbe9b8ea0b11c676d7b5

Observation 23413ca2-6ffe-41c4-ab1f-82c3100367c4 · inbound

Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs cites this paper.

Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:04:04.898211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:04:04.898211Z digest=sha256:941c372ff3ebba0f6d414aad33ccaa80cd9dd5e46ee223a0be6942278eba5964

Observation f76b4449-d78a-4ec0-baf0-65a16d2ba282 · inbound

Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods cites this paper.

Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.541169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.541169Z digest=sha256:5255acdc2bfb5493359bae61672bd4f65edc3c4b82c507462eb50fb2b9160630

Observation 88be4fb8-bd12-4cb5-88ef-726c42c7928a · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:06.332229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:06.332229Z digest=sha256:86d8122cc4072f42c044674908cbffaa644d953134975bd8ac97c77006634b96

Observation bb2a8f9b-2df4-4727-9cb3-89df77f03b3c · inbound

ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering cites this paper.

ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.542384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.542384Z digest=sha256:f526da01a84ef6011b56a019d672ce92ee6b2e720c6e98c3221d48a5b55f8afb

Observation b958903d-046f-4147-8229-74b87bd1188b · inbound

EduCoder: An Open-Source Annotation System for Education Transcript Data cites this paper.

EduCoder: An Open-Source Annotation System for Education Transcript Data The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:37:05.540440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:35:52.219876Z digest=sha256:44fc5bdfb8cca72ed2a66cca3cbfb9c959c3010dddb7f156fb9955e232376903

Observation cbd87949-a508-4ff1-9728-fbf3ffe6c530 · inbound

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications cites this paper.

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:40.826384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:49:40.826384Z digest=sha256:dad34bdd3e51c0fad6eb06655000936d917d64ea393da93d2a6fce6955a04f33

Observation 0aeb6484-a2ca-438d-bea8-95f5b98646f7 · inbound

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? cites this paper.

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:21.683550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:21.683550Z digest=sha256:e8b94a3a8cbe981b0248a30280cf9e3691b175310d8e0b8e8ccb0dfd4397c3ac

Observation 05fbd995-8ae2-4e05-9f4d-6350a5a8bb39 · inbound

Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models cites this paper.

Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:33:07.548115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T15:33:00.984344Z digest=sha256:b0424526b5f3042ca91aea12757386ae564696c43c1a25ed34fd3e183a7d6cd7

Observation f17c4560-d09f-4ffa-9ef6-05541150e15c · inbound

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations cites this paper.

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.726078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T18:27:50.159580Z digest=sha256:3163a94d5419e215970d7e54501410938e18533549bb5b5b8fe3e06556bcce49

Observation 04c52b5d-503c-4852-b7e2-411c3c0f2c35 · inbound

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe cites this paper.

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:58.149738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:58.149738Z digest=sha256:db0c0d8e7e4761a03b3540fa25f2de00a7599fa514eed27fe87cf67d6c7d8df9