Pith. sign in

REVIEW 2 cited by

AraELECTRA: Pre-Training Text Discriminators for Arabic Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.15516 v2 pith:7VYQHLIO submitted 2020-12-31 cs.CL

classification cs.CL
keywords arabiclanguagemodelrepresentationaraelectratokenscurrentmasked
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advances in English language representation enabled a more sample-efficient pre-training task by Efficiently Learning an Encoder that Classifies Token Replacements Accurately (ELECTRA). Which, instead of training a model to recover masked tokens, it trains a discriminator model to distinguish true input tokens from corrupted tokens that were replaced by a generator network. On the other hand, current Arabic language representation approaches rely only on pretraining via masked language modeling. In this paper, we develop an Arabic language representation model, which we name AraELECTRA. Our model is pretrained using the replaced token detection objective on large Arabic text corpora. We evaluate our model on multiple Arabic NLP tasks, including reading comprehension, sentiment analysis, and named-entity recognition and we show that AraELECTRA outperforms current state-of-the-art Arabic language representation models, given the same pretraining data and with even a smaller model size.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhanced Arabic Text Retrieval with Attentive Relevance Scoring

    cs.CL 2025-07 conditional novelty 4.0 of 10

    An Arabic dense retriever using a trainable attentive scoring module instead of dot-product similarity reports improved top-k passage retrieval on ArabicaQA.

  2. QU-NLP at CheckThat! 2025: Multilingual Subjectivity in News Articles Detection using Feature-Augmented Transformer Models with Sequential Cross-Lingual Fine-Tuning

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A feature-augmented transformer with TF-IDF gating and sequential cross-lingual fine-tuning reaches first place on English and Romanian but mixed scores elsewhere in CheckThat! 2025 subjectivity detection.

Pith tools