Pith. sign in

REVIEW 8 cited by

Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09743 v2 pith:MGMSSNL2 submitted 2023-11-16 cs.CL

Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks

classification cs.CL
keywords annotatorssubjectivetasksclassificationannotationsannotatorbiasedcapturing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated by employing methods like majority voting to determine a single ground truth label. In subjective tasks, aggregating labels will result in biased labeling and, consequently, biased models that can overlook minority opinions. Previous studies have shed light on the pitfalls of label aggregation and have introduced a handful of practical approaches to tackle this issue. Recently proposed multi-annotator models, which predict labels individually per annotator, are vulnerable to under-determination for annotators with few samples. This problem is exacerbated in crowdsourced datasets. In this work, we propose \textbf{Annotator Aware Representations for Texts (AART)} for subjective classification tasks. Our approach involves learning representations of annotators, allowing for exploration of annotation behaviors. We show the improvement of our method on metrics that assess the performance on capturing individual annotators' perspectives. Additionally, we demonstrate fairness metrics to evaluate our model's equability of performance for marginalized annotators compared to others.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  2. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0

    On the same 720 replies, scoring exposure versus manifestation shifts the auditor-judge gap by ~0.2 AUROC and can reverse their ranking, so single detection AUROCs are under-specified.

  3. Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision

    cs.LG 2026-04 unverdicted novelty 6.0

    By adopting a negative ontology where the true target does not objectively exist, the paper defines Democratic Supervision and derives the EL-MIATTs framework for ML evaluation and learning with Multiple Inaccurate Tr...

  4. Empathy Applicability Modeling for General Health Queries

    cs.CL 2026-01 conditional novelty 6.0

    General health queries can be labeled in advance for whether they call for emotional reactions or interpretive empathy, and classifiers trained on these labels beat simple baselines.

  5. Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision

    cs.LG 2026-04 unverdicted novelty 5.0

    The paper posits that the true target does not exist and introduces the EL-MIATTs framework for evaluation and learning under Democratic Supervision in machine learning.

  6. Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision

    cs.LG 2026-04 unverdicted novelty 4.0

    The true target does not objectively exist in ML, so models should use multiple inaccurate true targets under democratic supervision via the EL-MIATTs framework for evaluation and learning.

  7. Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision

    cs.LG 2026-04 unverdicted novelty 3.0

    Proposes the EL-MIATTs framework for ML predictive modeling by assuming the true target does not exist and defining democratic supervision via multiple inaccurate true targets.

  8. Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision

    cs.LG 2026-04 reject novelty 3.0

    By assuming the true target does not exist, this paper proposes Democratic Supervision with Multiple Inaccurate True Targets (MIATTs), packaging prior author work into the EL-MIATTs framework.