Pith. sign in

REVIEW 1 cited by

Multi-Dialect Arabic BERT for Country-Level Dialect Identification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.05612 v1 pith:43QHSP3P submitted 2020-07-10 cs.CL cs.LG

Multi-Dialect Arabic BERT for Country-Level Dialect Identification

classification cs.CL cs.LG
keywords dialectidentificationarabicmodelsolutionsubtaskwinningbert
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Arabic dialect identification is a complex problem for a number of inherent properties of the language itself. In this paper, we present the experiments conducted, and the models developed by our competing team, Mawdoo3 AI, along the way to achieving our winning solution to subtask 1 of the Nuanced Arabic Dialect Identification (NADI) shared task. The dialect identification subtask provides 21,000 country-level labeled tweets covering all 21 Arab countries. An unlabeled corpus of 10M tweets from the same domain is also presented by the competition organizers for optional use. Our winning solution itself came in the form of an ensemble of different training iterations of our pre-trained BERT model, which achieved a micro-averaged F1-score of 26.78% on the subtask at hand. We publicly release the pre-trained language model component of our winning solution under the name of Multi-dialect-Arabic-BERT model, for any interested researcher out there.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Spam and Sentiment Detection in Arabic Tweets Using MARBERT Model

    cs.CL 2026-06 unverdicted novelty 2.0

    MARBERT is fine-tuned on 24,513 Arabic tweets for sentiment analysis, with the claim that the resulting scheme shows promising accuracy versus prior techniques.