Pith. sign in

REVIEW 1 cited by

Ground-Truth, Whose Truth? -- Examining the Challenges with Annotating Toxic Text Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.03529 v1 pith:G3GYNY45 submitted 2021-12-07 cs.CL

classification cs.CL
keywords datasetstexttoxicapproachmodelsannotatingannotatorschallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The use of machine learning (ML)-based language models (LMs) to monitor content online is on the rise. For toxic text identification, task-specific fine-tuning of these models are performed using datasets labeled by annotators who provide ground-truth labels in an effort to distinguish between offensive and normal content. These projects have led to the development, improvement, and expansion of large datasets over time, and have contributed immensely to research on natural language. Despite the achievements, existing evidence suggests that ML models built on these datasets do not always result in desirable outcomes. Therefore, using a design science research (DSR) approach, this study examines selected toxic text datasets with the goal of shedding light on some of the inherent issues and contributing to discussions on navigating these challenges for existing and future projects. To achieve the goal of the study, we re-annotate samples from three toxic text datasets and find that a multi-label approach to annotating toxic text samples can help to improve dataset quality. While this approach may not improve the traditional metric of inter-annotator agreement, it may better capture dependence on context and diversity in annotators. We discuss the implications of these results for both theory and practice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Data Annotation as Measurement

    cs.CY 2026-08 conditional novelty 4.0 of 10

    Data annotation quality should be assessed with reliability and validity concepts from measurement theory, and annotation problems should be diagnosed by their source (error, ambiguity, impossibility, subjectivity, id...

Pith tools