Pith. sign in

Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Automatic diacritization of Arabic text involves adding diacritical marks (diacritics) to the text. This task poses a significant challenge with noteworthy implications for computational processing and comprehension. In this paper, we introduce PTCAD (Pre-FineTuned Token Classification for Arabic Diacritization, a novel two-phase approach for the Arabic Text Diacritization task. PTCAD comprises a pre-finetuning phase and a finetuning phase, treating Arabic Text Diacritization as a token classification task for pre-trained models. The effectiveness of PTCAD is demonstrated through evaluations on two benchmark datasets derived from the Tashkeela dataset, where it achieves state-of-the-art results, including a 20\% reduction in Word Error Rate (WER) compared to existing benchmarks and superior performance over GPT-4 in ATD tasks.

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Sadeed: Advancing Arabic Diacritization Through Small Language Model cs.CL · 2025-04-30 · reject · none · ref 14 · internal anchor

    The authors claim that Sadeed, a fine-tuned 1.5B Arabic SLM, reaches state-of-the-art word error rates on the Fadel benchmark and is competitive with proprietary models, while releasing a new benchmark and dataset.