Pith. sign in

REVIEW 4 cited by

Pre-Training BERT on Arabic Tweets: Practical Considerations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.10684 v1 pith:HE4KKV7U submitted 2021-02-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsarabicbertdatadownstreamhighlighttaskstraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pretraining Bidirectional Encoder Representations from Transformers (BERT) for downstream NLP tasks is a non-trival task. We pretrained 5 BERT models that differ in the size of their training sets, mixture of formal and informal Arabic, and linguistic preprocessing. All are intended to support Arabic dialects and social media. The experiments highlight the centrality of data diversity and the efficacy of linguistically aware segmentation. They also highlight that more data or more training step do not necessitate better models. Our new models achieve new state-of-the-art results on several downstream tasks. The resulting models are released to the community under the name QARiB.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 85 citations worldwide. Full citation record

  1. Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components

    cs.IR 2025-06 conditional novelty 5.0 of 10

    For Arabic retrieval-augmented generation, sentence-aware chunking, BGE-M3 and Multilingual-E5-large embeddings, bge-reranker-v2-m3, and Aya-8B yield the highest RAGAS scores across six Arabic datasets.

  2. CVPD at QIAS 2025 Shared Task: An Efficient Encoder-Based Approach for Islamic Inheritance Reasoning

    cs.CL 2025-08 conditional novelty 4.0 of 10

    An encoder-based relevance-scoring system achieves 69.87% accuracy on Islamic inheritance multiple-choice questions, below Gemini's 87.60% but with far smaller compute.

  3. SHAMI-MT: A Syrian Arabic Dialect to Modern Standard Arabic Bidirectional Machine Translation System

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A bidirectional MSA-Syrian Arabic translation system built by fine-tuning AraT5v2 on the Nabra corpus; only the MSA-to-Shami direction is evaluated, with a GPT-4.1 score of 4.01/5.

  4. Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning

    cs.CL 2025-07 reject novelty 4.0 of 10

    Fine-tuning QARiB on Arabic offensive tweets reaches 92% macro F1, and a new transformer-based 'hate speech masking' system achieves 0.30 BLEU-1 without comparison to any baseline.

Pith tools