Pith. sign in

REVIEW

Combating high variance in Data-Scarce Implicit Hate Speech Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.13595 v1 pith:UXH6KARX submitted 2022-08-29 cs.CL cs.LG

classification cs.CLcs.LG
keywords hatespeechclassificationimplicitdatahighlanguageproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hate speech classification has been a long-standing problem in natural language processing. However, even though there are numerous hate speech detection methods, they usually overlook a lot of hateful statements due to them being implicit in nature. Developing datasets to aid in the task of implicit hate speech classification comes with its own challenges; difficulties are nuances in language, varying definitions of what constitutes hate speech, and the labor-intensive process of annotating such data. This had led to a scarcity of data available to train and test such systems, which gives rise to high variance problems when parameter-heavy transformer-based models are used to address the problem. In this paper, we explore various optimization and regularization techniques and develop a novel RoBERTa-based model that achieves state-of-the-art performance.

Discussion (0). Sign in to comment.

Pith tools