Pith. sign in

REVIEW 2 cited by

AraT5: Text-to-Text Transformers for Arabic Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.12068 v4 pith:6X3WK3CD submitted 2021-08-31 cs.CL

classification cs.CL
keywords languagearabicargenmodelstasksalthougharat5benchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transfer learning with a unified Transformer framework (T5) that converts all language problems into a text-to-text format was recently proposed as a simple and effective transfer learning approach. Although a multilingual version of the T5 model (mT5) was also introduced, it is not clear how well it can fare on non-English tasks involving diverse data. To investigate this question, we apply mT5 on a language with a wide variety of dialects--Arabic. For evaluation, we introduce a novel benchmark for ARabic language GENeration (ARGEN), covering seven important tasks. For model comparison, we pre-train three powerful Arabic T5-style models and evaluate them on ARGEN. Although pre-trained with ~49 less data, our new models perform significantly better than mT5 on all ARGEN tasks (in 52 out of 59 test sets) and set several new SOTAs. Our models also establish new SOTA on the recently-proposed, large Arabic language understanding evaluation benchmark ARLUE (Abdul-Mageed et al., 2021). Our new models are publicly available. We also link to ARGEN datasets through our repository: https://github.com/UBC-NLP/araT5.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model

    cs.CL 2025-05 reject novelty 5.0 of 10

    A compact 1.5B Arabic-English model beats GPT-4o mini only on the authors' own Tarjama-25 benchmark, while trailing large models on standard WMT24++ and IWSLT2017 tests.

  2. SHAMI-MT: A Syrian Arabic Dialect to Modern Standard Arabic Bidirectional Machine Translation System

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A bidirectional MSA-Syrian Arabic translation system built by fine-tuning AraT5v2 on the Nabra corpus; only the MSA-to-Shami direction is evaluated, with a GPT-4.1 score of 4.01/5.

Pith tools