Pith. sign in

REVIEW 1 cited by

Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09503 v1 pith:YWE32PQ4 submitted 2022-12-15 cs.CL cs.LG

classification cs.CLcs.LG
keywords tasksacrossannotationcomplexmeasureslabelingagreementalpha
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When annotators label data, a key metric for quality assurance is inter-annotator agreement (IAA): the extent to which annotators agree on their labels. Though many IAA measures exist for simple categorical and ordinal labeling tasks, relatively little work has considered more complex labeling tasks, such as structured, multi-object, and free-text annotations. Krippendorff's alpha, best known for use with simpler labeling tasks, does have a distance-based formulation with broader applicability, but little work has studied its efficacy and consistency across complex annotation tasks. We investigate the design and evaluation of IAA measures for complex annotation tasks, with evaluation spanning seven diverse tasks: image bounding boxes, image keypoints, text sequence tagging, ranked lists, free text translations, numeric vectors, and syntax trees. We identify the difficulty of interpretability and the complexity of choosing a distance function as key obstacles in applying Krippendorff's alpha generally across these tasks. We propose two novel, more interpretable measures, showing they yield more consistent IAA measures across tasks and annotation distance functions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MObyGaze: a film dataset of multimodal objectification densely annotated by experts

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MObyGaze provides 6072 expert-annotated segments across 43 hours of film, with a multimodal thesaurus, and shows that current vision and text models can detect objectification above chance, while audio models cannot.

Pith tools