Pith. sign in

REVIEW 1 cited by

OSACT4 Shared Task on Offensive Language Detection: Intensive Preprocessing-Based Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.07297 v1 pith:IXTIU63H submitted 2020-05-14 cs.CL

classification cs.CL
keywords detectionarabicclassificationlanguagehateoffensivepreprocessingspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The preprocessing phase is one of the key phases within the text classification pipeline. This study aims at investigating the impact of the preprocessing phase on text classification, specifically on offensive language and hate speech classification for Arabic text. The Arabic language used in social media is informal and written using Arabic dialects, which makes the text classification task very complex. Preprocessing helps in dimensionality reduction and removing useless content. We apply intensive preprocessing techniques to the dataset before processing it further and feeding it into the classification model. An intensive preprocessing-based approach demonstrates its significant impact on offensive language detection and hate speech detection shared tasks of the fourth workshop on Open-Source Arabic Corpora and Corpora Processing Tools (OSACT). Our team wins the third place (3rd) in the Sub-Task A Offensive Language Detection division and wins the first place (1st) in the Sub-Task B Hate Speech Detection division, with an F1 score of 89% and 95%, respectively, by providing the state-of-the-art performance in terms of F1, accuracy, recall, and precision for Arabic hate speech detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-task Learning with Active Learning for Arabic Offensive Speech Detection

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A multi-task Arabic offensive speech detector with entropy-based active learning and weighted emoji tokens reports 85.42% macro F1 on OSACT2022 using roughly 3,300 training samples.

Pith tools