REVIEW 2 cited by
CATT: Character-based Arabic Tashkeel Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Tashkeel, or Arabic Text Diacritization (ATD), greatly enhances the comprehension of Arabic text by removing ambiguity and minimizing the risk of misinterpretations caused by its absence. It plays a crucial role in improving Arabic text processing, particularly in applications such as text-to-speech and machine translation. This paper introduces a new approach to training ATD models. First, we finetuned two transformers, encoder-only and encoder-decoder, that were initialized from a pretrained character-based BERT. Then, we applied the Noisy-Student approach to boost the performance of the best model. We evaluated our models alongside 11 commercial and open-source models using two manually labeled benchmark datasets: WikiNews and our CATT dataset. Our findings show that our top model surpasses all evaluated models by relative Diacritic Error Rates (DERs) of 30.83\% and 35.21\% on WikiNews and CATT, respectively, achieving state-of-the-art in ATD. In addition, we show that our model outperforms GPT-4-turbo on CATT dataset by a relative DER of 9.36\%. We open-source our CATT models and benchmark dataset for the research community\footnote{https://github.com/abjadai/catt}.
Forward citations
Cited by 2 Pith papers
-
NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
The NADI 2025 shared task introduces a standardized speech benchmark for eight Arabic dialects and reports best results of 79.8% dialect ID accuracy, 35.68 WER for ASR, and 55 WER for diacritic restoration.
-
Sadeed: Advancing Arabic Diacritization Through Small Language Model
The authors claim that Sadeed, a fine-tuned 1.5B Arabic SLM, reaches state-of-the-art word error rates on the Fadel benchmark and is competitive with proprietary models, while releasing a new benchmark and dataset.
Discussion (0). Continue with ORCID to comment.