REVIEW 3 major objections 5 minor 12 references
Data Annotation as Measurement
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Data annotation should be treated as measurement, with reliability and validity checked alongside agreement, because labeling problems that look alike often demand different fixes.
desk verdict Useful reframing of annotation as measurement, but the five-source taxonomy rests on prompted interviews and overlapping definitions, so treat it as a heuristic, not a validated result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is measurement theory's four-stage process—systematization (explicitly defining the concept), operationalization (building the annotation instrument), application (producing instance-level annotations), and interrogation (assessing reliability and validity)—paired with a five-source taxonomy of annotation issues and a table of decision points running from task design through adjudication. The taxonomy does the causal work: it connects each observed failure mode to a different stage in the process and to a different class of remedy.
What would settle it
Compare two correction protocols on the same low-quality annotation batches: one that first classifies each issue by the five-source taxonomy and then applies the table's corresponding fix, and one that retrains all annotators and filters low-agreement labels as usual; if the taxonomy-driven protocol does not produce higher validity on held-out gold data, the central claim that source determines the right intervention is contradicted.
Extended reading notes
Core claim
The paper's central claim is that annotation is a form of measurement, and that annotation teams therefore should interrogate reliability and validity rather than stop at agreement. On its account, every annotation project passes through task design, annotator management and collection, quality assessment, and quality improvement or adjudication, and each stage hides decision points that shape outcomes. It proposes a taxonomy of five issue sources: error (annotator did not follow instructions), ambiguity (instructions underdetermine the annotation), impossibility (data underdetermine the annotation), subjectivity (annotation depends on annotators' values or beliefs), and identity (the annotator's identity is misaligned with the desired population). Because these sources can produce identical-looking mislabels, the paper argues, correcting them requires process-level intervention matched to the source, and measurement theory supplies the vocabulary—test-retest reliability, face, content, convergent, predictive, and discriminant validity—for doing so.
Load-bearing premise
The load-bearing premise is that the five-source taxonomy and the decision-point map, built from a literature review plus ten interviews with North American researchers, capture the full range of annotation problems and decision points that actually occur across domains, annotator populations, and geographies.
Editorial extensions
If this is right
- Agreement statistics alone are insufficient for quality reporting; annotation teams should also report reliability checks such as test-retest consistency and validity checks such as face, content, convergent, predictive, and discriminant validity.
- Once a label problem is classified by source, the fix differs: retrain or filter annotators for error, improve instructions for ambiguity, re-process or filter data for impossibility, refine the schema for subjectivity, and re-select the annotator population for identity misalignment.
- Undocumented defaults in the annotation process—unit of annotation, preprocessing choices, annotator population, aggregation rule—become audit targets because they are where quality problems originate.
- The measurement framing gives a shared vocabulary for existing calls to move beyond agreement and for future dataset audits, so quality problems can be compared across teams and domains.
Reading between the lines
- This reader infers that the five-source taxonomy could be tested directly by coding failure reports from a large multi-domain dataset repository; a recurring category outside the five would signal the taxonomy needs revision.
- The paper's emphasis on test-retest reliability implies that a concrete diagnostic protocol (for example, how long to wait between annotations, and how to combine intra- and inter-annotator agreement) would accelerate adoption, though the paper does not specify one.
- As AI annotators replace human crowds, the subjectivity and identity sources become harder to detect because machine annotators produce little natural variation; teams may need to retain some human disagreement specifically to run validity checks.
- A labor-practice corollary the paper gestures at: if low agreement is often a sign of ambiguity or impossibility rather than error, then rejecting and withholding pay from annotators who disagree is both unjust and ineffective; source diagnosis could reduce unjustified rejection.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that data annotation should be understood as a measurement process, with systematic concept definition (systematization), operationalization, application, and interrogation of reliability and validity. Drawing on a literature review (N=132) and semi-structured interviews with ten annotation team members, it maps key decision points across four stages (task design, annotator management and collection, quality assessment, quality improvement/adjudication), proposes five distinct sources of annotation issues (error, ambiguity, impossibility, subjectivity, annotator identity), and translates measurement-theoretic validity concepts into practical recommendations such as test-retest reliability, face validity, content, convergent, predictive, and discriminant validity.
Significance. If accepted, the framework provides a useful conceptual vocabulary for moving beyond agreement as the default quality signal in annotation and offers annotation teams a structured set of decision points to document and interrogate. The paper is refreshingly transparent about its methods: it includes the full interview guide, describes the literature search strategy, states the authors' positionality, and acknowledges sample limitations. The main strength is the integration of measurement theory with practitioner accounts. However, the empirical grounding of the five-source taxonomy is the central weakness, and the practical guidance depends on the taxonomy being both distinct and reliably attributable. The paper is likely to be a valuable reference for researchers studying annotation quality, but only if the taxonomy's status and diagnostic procedures are clarified.
major comments (3)
- [§3.2 / Appendix A] The interview guide pre-specifies the taxonomy it claims to identify. Participants are asked about ambiguity, impossibility, subjectivity, and identity directly, with definitions supplied in the questions (e.g., "By ambiguous, I mean that there could be multiple correct ways..."), and about "errors" whose definition is also given. The open-ended questions ("what challenges did you face?") could have surfaced other sources, but the subsequent prompts are leading, so the interviews cannot independently validate the exhaustiveness or distinctness of the five sources. The Abstract states "we identify five distinct sources" as a contribution; given the protocol, this is at best a framework imported from the literature and operationalized in the interviews. I recommend either (a) showing from the literature review that the taxonomy was derived before the interviews, (b) presenting the taxonomy as an analytic framework rather than an emergent empirical finding, or (c) reporting whether any participant spontaneously named these sources before being prompted. Without one of these, the empirical support for the taxonomy is overstated.
- [Table 3 / §4.3] The definitions of the five sources are not mutually exclusive, which undermines the claim that they are "distinct sources" and that outcome-similar problems require different interventions. "Error" presupposes that instructions were clear enough to be followed; "ambiguity" arises when instructions are under-specified. A single misannotation can therefore be both an error (the annotator did not follow what was specified) and a consequence of ambiguity (what was specified was incomplete). Similarly, "impossibility" and "ambiguity" both refer to insufficient information, and the distinction (data versus instructions) is often not observable after the fact. The paper needs operational decision rules or a discussion of borderline cases; for example, specify that a problem is ambiguity if revising instructions changes annotations, impossibility if no instruction revision resolves the deficit, subjectivity if informed annotators disagree after clarification, and error only if the instruction was determinative and the annotator deviated. As it stands, the prescriptive claim that "annotation problems that appear similar at the level of outcomes often require different process-level interventions based on their sources" (Abstract, §4.3) cannot be applied reliably.
- [§5(a) / §5, Reliability] The paper recommends "determine the source of an annotation error before taking action" and suggests measuring test-retest reliability to distinguish ambiguity from subjectivity, and adding a "cannot annotate" option to distinguish ambiguity from impossibility. These diagnostics are not validated. For test-retest reliability, the paper acknowledges there are no agreed-upon best practices (e.g., optimal time interval), and it does not explain how to interpret low or high test-retest values in terms of the five sources. For the "cannot annotate" option, reliance on annotator self-report may itself be biased and may conflate ambiguity with impossibility. The paper should provide a concrete decision procedure with explicit interpretations, or clearly state that source attribution is an open problem and that the framework is a conceptual scaffold rather than a ready-to-use diagnostic. This issue is load-bearing because the central claim involves source-specific interventions.
minor comments (5)
- [§3.1] The search-term list ("annotat* | label * | coding") contains a space in "label *" and uses vertical-bar notation inconsistently; please clarify the exact boolean query, the wildcard syntax, and the date ranges covered by each database search.
- [Table 1] The checkmark matrix is difficult to read, especially for the row totals; consider a more compact representation such as counts per participant or a separate demographic table.
- [§3.2, Thematic Analysis] The paper reports that the first author performed the coding but does not report a second coder or a reliability check; a sentence explaining why this is appropriate in reflexive thematic analysis, or describing any author cross-checks, would help readers assess the analysis.
- [§4.3] The phrase "we introduce identity as a source of annotation issues" overstates novelty, since the interview guide already asks about identity; consider "we include" or "we highlight" instead.
- [§5, Face Validity] The paper states that "nearly every interview participant" conducted manual review and cites P1–P5, P8–P10; that is nine of ten participants, so consider writing "nine of ten" or "all but two" for precision.
Circularity Check
Five-source taxonomy is pre-specified in the interview guide, so its empirical derivation is partly circular; the measurement-theory framing remains independent.
-
self definitional
[Section 4.3 and Appendix A (Semi-Structured Interview Guide)]
"Based on our literature review and interview study, we propose a taxonomy of five distinct sources of annotation issues (Table 3). ... Did you experience situations where the correct annotation seemed ambiguous? By ambiguous, I mean that there could be multiple correct ways of annotating a particular piece of data, and the annotation instructions were under-specified. ... Did you experience situations where the correct annotation seemed subjective? By subjective, I mean that the correct annotation depends on personal values, feelings, tastes, or opinions. ..."
Section 4.3 presents the five-source taxonomy as an empirical finding based on the interview study, but Appendix A's semi-structured guide explicitly asks about error, ambiguity, impossibility, subjectivity, and identity, supplying definitions that are nearly identical to Table 3. Participants were asked to map their experiences onto these pre-given categories, so a thematic analysis can only recover categories the protocol inserted. The interviews therefore cannot establish that the five categories are distinct, exhaustive, or the natural sources of annotation issues. The 'finding' is, by construction, a restatement of the instrument's prior definitions.
full rationale
The paper's main contribution—treating annotation as measurement—imports a substantive external framework (Adcock & Collier 2001; Jacobs & Wallach 2021) and is not circular: it applies measurement theory to annotation, and the reliability/validity vocabulary comes from that external literature rather than from the paper's own outputs. The self-citations (Harvey et al. 2025; Gosciak et al. 2026; Sloane et al. 2026; Thomas et al. 2026) are contextual and not load-bearing; no uniqueness theorem or fitted parameter is invoked. The one genuine circular step is the empirical derivation of the five-source taxonomy, which is pre-specified in the interview protocol. Because the taxonomy is a central contribution and is presented as a finding from the interview study, while the protocol's definitions essentially equal the reported definitions, this is a real but partial circularity, so the score is 4 rather than 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Measurement theory (Adcock and Collier 2001) provides a valid framework for understanding data annotation.
- domain assumption The sample of 10 North American researchers is sufficient to identify the key decision points and sources of annotation issues.
- domain assumption The first author's manual coding of transcripts reliably captures the themes.
Cite this review
Pith. "Pith review of Data Annotation as Measurement." pith.science (2026). https://pith.science/paper/UDB5B4Z5
@misc{pith2026260807297,
author = {Pith},
title = {Pith review of: Data Annotation as Measurement},
year = {2026},
howpublished = {\url{https://pith.science/paper/UDB5B4Z5}},
note = {Machine review of arXiv:2608.07297}
}
read the original abstract
Modern AI systems depend on annotated data, but annotation is rarely treated as the act of measurement that it is. Instead, annotation quality is commonly reduced to agreement: if multiple annotators assign the same annotation to a data instance, the annotations are taken to be high-quality. Yet agreement does not establish whether annotations validly capture the underlying concept they are meant to represent. In this paper, we argue that data annotation should be understood as a measurement problem. Like other forms of measurement, annotation requires defining a concept, operationalizing it through an instrument, applying that instrument, and evaluating the reliability and validity of the resulting measurements. Drawing on a literature review of annotation quality research (N=132) and semi-structured interviews with annotation team members (N=10), we develop a framework for diagnosing and correcting annotation issues. First, we map key decision points across annotation processes - including task design, annotator management, quality assessment, quality improvement, and adjudication - that shape annotation outcomes. Second, we identify five distinct sources of annotation issues: error, ambiguity, impossibility, subjectivity, and annotator identity. Annotation problems that appear similar at the level of outcomes often require different process-level interventions based on their sources. Finally, we translate measurement theory into practical guidance for annotation teams, showing how assessments of reliability and validity can move beyond agreement alone. By reframing annotation as measurement, we offer a conceptual foundation for improving the quality of annotated data used in AI research and practice.
Figures
Reference graph
Works this paper leans on
-
[6]
InThirty-seventh Conference on Neural Information Processing Systems
Automated Classification of Model Errors on Ima- geNet. InThirty-seventh Conference on Neural Information Processing Systems. Plank, B. 2022. The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. InPro- ceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10671–10682. Abu Dhabi, United Ar...
work page 2022
-
[324]
Seattle, United States: Association for Computational Linguistics. A Semi-Structured Interview Guides Prior to scheduling the interview, participants were invited to review study details and provide their written informed consent to participant. Once the interview started, but prior to the start of recording, participants were given a brief overview of th...
-
[1291]
Beyond Agreement: Rethinking Ground Truth in Educational AI Annotation
Chicago IL USA: ACM. ISBN 979-8-4007-0192-4. Thomas, D. R.; Borchers, C.; and Koedinger, K. R. 2025. Beyond Agreement: Rethinking Ground Truth in Educa- tional AI Annotation. ArXiv:2508.00143 [cs]. Thomas, D. R.; Borchers, C.; Vanacore, K. P.; Koedinger, K. R.; and Kizilcec, R. F. 2026. Modernizing Ground Truth: Four Shifts Toward Improving Reliability an...
work page Pith review arXiv 2025
-
[1998]
Codebook Development for Team-Based Qualitative Analysis.CAM Journal, 10(2): 31–36. Marshall, C. C.; Goguladinne, P. S.; Maheshwari, M.; Sathe, A.; and Shipman, F. M. 2023. Who Broke Amazon Mechan- ical Turk?: An Analysis of Crowdsourcing Data Quality over Time. InProceedings of the 15th ACM Web Science Confer- ence 2023, WebSci, 335–345. Austin TX USA: A...
work page 2023
-
[2013]
Learning Whom to Trust with MACE. In Vander- wende, L.; Daum´e III, H.; and Kirchhoff, K., eds.,Proceed- ings of the 2013 Conference of the North American Chap- ter of the Association for Computational Linguistics: Hu- man Language Technologies, NAACL, 1120–1130. Atlanta, Georgia: Association for Computational Linguistics. Hube, C.; Fetahu, B.; and Gadira...
work page 2013
-
[2014]
Everyone wants to do the model work, not the data work
Corpus Annotation through Crowdsourcing: Towards Best Practice Guidelines. InProceedings of the Ninth In- ternational Conference on Language Resources and Eval- uation (LREC’14). Reykjavik, Iceland: European Language Resources Association (ELRA). Sambasivan, N.; Kapania, S.; Highfill, H.; Akrong, D.; Par- itosh, P.; and Aroyo, L. M. 2021. “Everyone wants ...
work page 2021
-
[2017]
InProceedings of the 2017 ACM International Conference on Management of Data, 1711–1716
Crowdsourced Data Management: Overview and Challenges. InProceedings of the 2017 ACM International Conference on Management of Data, 1711–1716. Chicago Illinois USA: ACM. ISBN 978-1-4503-4197-4. Li, T.; Manns, C. J.; North, C.; and Luther, K. 2019. Drop- ping the Baton?: Understanding Errors and Bottlenecks in a Crowdsourced Sensemaking Pipeline.Proceedin...
work page 2017
-
[2018]
Resolvable vs. Irresolvable Disagreement: A Study on Worker Deliberation in Crowd Work.Proceedings of the ACM on Human-Computer Interaction, 2(CSCW): 1–19. Scheuerman, M. K.; Hanna, A.; and Denton, E. 2021. Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development.Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2): ...
work page 2021
Show all 12 references
-
[2022]
In Carpuat, M.; de Marneffe, M.-C.; and Meza Ruiz, I
Two Contrasting Data Annotation Paradigms for Sub- jective NLP Tasks. In Carpuat, M.; de Marneffe, M.-C.; and Meza Ruiz, I. V ., eds.,Proceedings of the 2022 Confer- ence of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...
2022
-
[2023]
Ho, C.-J.; Slivkins, A.; Suri, S.; and Vaughan, J
How Crowd Worker Factors Influence Subjective An- notations: A Study of Tagging Misogynistic Hate Speech in Tweets.Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 11(1): 38–50. Ho, C.-J.; Slivkins, A.; Suri, S.; and Vaughan, J. W. 2015. Incentivizing...
2015
-
[2024]
Sheng, V
Ambiguous Annotations: When is a Pedestrian not a Pedestrian? ArXiv:2405.08794 [cs]. Sheng, V . S.; Provost, F.; and Ipeirotis, P. G. 2008. Get an- other label? improving data quality and data mining using multiple, noisy labelers. InProceedings of the 14th ACM SIGKDD internat...
2008 arXiv
-
[2026]
Alonso, O
Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior.Proceedings of the AAAI Conference on Artificial Intelligence, 40(44): 37222–37231. Alonso, O. 2015a. Challenges with Label Quality for Super- vi...
2021 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.