Pith. sign in

REVIEW 1 cited by

GPM: A Generic Probabilistic Model to Recover Annotator's Behavior and Ground Truth Labeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.00475 v1 pith:6MAN75KM submitted 2020-03-01 cs.AI

classification cs.AI
keywords datamodelannotatorgroundlabelingtruthannotatorsbehavior
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the big data era, data labeling can be obtained through crowdsourcing. Nevertheless, the obtained labels are generally noisy, unreliable or even adversarial. In this paper, we propose a probabilistic graphical annotation model to infer the underlying ground truth and annotator's behavior. To accommodate both discrete and continuous application scenarios (e.g., classifying scenes vs. rating videos on a Likert scale), the underlying ground truth is considered following a distribution rather than a single value. In this way, the reliable but potentially divergent opinions from "good" annotators can be recovered. The proposed model is able to identify whether an annotator has worked diligently towards the task during the labeling procedure, which could be used for further selection of qualified annotators. Our model has been tested on both simulated data and real-world data, where it always shows superior performance than the other state-of-the-art models in terms of accuracy and robustness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Granular feedback merits sophisticated aggregation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    As feedback becomes more granular, sophisticated supervised aggregation outperforms regularized averaging, needing about 44% fewer raters at 5-point scales.

Pith tools