Pith. sign in

REVIEW 1 cited by

Take the Hint: Improving Arabic Diacritization with Partially-Diacritized Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03557 v2 pith:2KNG5UWJ submitted 2023-06-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords textarabicdiacriticsdiacritizationinputnon-diacritizedsupportwhile
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Automatic Arabic diacritization is useful in many applications, ranging from reading support for language learners to accurate pronunciation predictor for downstream tasks like speech synthesis. While most of the previous works focused on models that operate on raw non-diacritized text, production systems can gain accuracy by first letting humans partly annotate ambiguous words. In this paper, we propose 2SDiac, a multi-source model that can effectively support optional diacritics in input to inform all predictions. We also introduce Guided Learning, a training scheme to leverage given diacritics in input with different levels of random masking. We show that the provided hints during test affect more output positions than those annotated. Moreover, experiments on two common benchmarks show that our approach i) greatly outperforms the baseline also when evaluated on non-diacritized text; and ii) achieves state-of-the-art results while reducing the parameter count by over 60%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sadeed: Advancing Arabic Diacritization Through Small Language Model

    cs.CL 2025-04 reject novelty 5.0 of 10

    The authors claim that Sadeed, a fine-tuned 1.5B Arabic SLM, reaches state-of-the-art word error rates on the Fadel benchmark and is competitive with proprietary models, while releasing a new benchmark and dataset.

Pith tools