Pith. sign in

REVIEW 4 major objections 4 minor 60 references

This paper argues that Arabizi—Arabic written in Latin letters and numerals—has systematic, dialect-specific spelling patterns that speakers can recognize, even in sentences with no country-specific vocabulary.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 04:47 UTC pith:FT6HRO6Y

load-bearing objection Valuable new cross-dialectal Arabizi resources and credible word-level findings, but the sentence-recognition experiment is circular (same people who wrote the texts judged them), so the headline claim about dialect recognition is not yet supported. the 4 major comments →

arxiv 2608.02555 v1 pith:FT6HRO6Y submitted 2026-08-03 cs.CL

Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

classification cs.CL
keywords ArabiziRomanized Arabicdialect identificationspelling variationtransliterationArabic dialectssurvey study
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Working with 189 Arabic speakers from Algeria, Egypt, Lebanon, Morocco, and Tunisia, the paper collects opinions about Arabizi and transliterations of everyday words, then asks whether the Latin-script spellings are dialectally patterned. It reports two consistent findings: individual writers' transliterations of the same Arabic word are regular (a single form is usually preferred), and the preferred forms differ across dialects in ways that track known pronunciation differences—for example, Maghrebi writers render the letter qaf with 9 or g while Egyptians and Lebanese use 2 or q. At the sentence level, the same speakers identified around 90 percent of same-country transliterations as coming from their own dialect, even when the sentences contained no country-specific words. The paper concludes that Arabizi is not a temporary, makeshift orthography but a structured writing system whose spelling variation encodes dialect information worth modeling in natural language processing.

Core claim

The central discovery is that cross-dialectal spelling variation in Arabizi is systematic and perceptible. In a 210-word lexicon with about twenty transliterations per word, intra-dialectal regularity shows a preferred spelling for most words, while inter-dialectal differences surface in letter-to-script mappings: kha is 5 in Egypt, Lebanon, and Tunisia but kh in Algeria and Morocco; qaf is 9 mainly in the Maghreb and 2 or q in Egypt and Lebanon; and Algerian and Moroccan writers often double consonants to mark gemination. A parallel corpus of 189 Arabic sentences, each transliterated by two speakers per dialect, was judged by the ten participants who wrote them. Over 90 percent of same-coun

What carries the argument

The argument rests on two released resources and one alignment procedure. A character-level alignment algorithm maps Arabic letters to Latin-letter and numeral sequences across 1,390 unique transliterations, starting from a seed mapping table that the authors manually extend and validate; the resulting per-letter frequency table exposes contrasts like Maghrebi 9 versus Egyptian/Lebanese 2/q for qaf. A hand-built parallel corpus of 189 Arabic sentences—valid in at least two dialects, high in dialectality, free of country-specific words—tests whether spelling alone betrays the writer's dialect. Each transliteration was judged by all participants as 'My Country,' 'Not Sure,' or 'Another Country

Load-bearing premise

The sentence-level recognition result assumes the ten people who wrote the Arabizi sentences are unbiased judges of which dialect the spellings come from; if their near-perfect 'My Country' scores come from remembering their own writing instead of reading generalizable spelling cues, the central claim loses its support.

What would settle it

Recruit a fresh group of annotators from the same five countries who contributed none of the transliterations, give them the same 189 parallel sentences, and see whether same-country recognition stays above 90 percent. If fresh judges perform near chance—especially on sentences without dialect-specific words—the original scores reflect self-recognition rather than spelling-based dialect cues.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Arabizi carries enough dialectal signal to support dialect identification from orthography alone, even when texts contain no country-specific words.
  • NLP systems that transliterate, normalize, or identify dialects in Arabizi can exploit per-dialect grapheme preferences (5 vs kh, 9 vs 2/q, ch vs sh) rather than treating Arabizi as a single noisy script.
  • Because responses show decline—strongest in Egypt—and a persistent minority of heavy users, dialect-aware text processing remains relevant for moderation and user-facing chat systems.
  • The released 210-word lexicon and 189-sentence parallel corpus give researchers a shared benchmark for whether systems reproduce the same between-dialect distinctions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the recognizability claim would be to run the same annotation task with fresh judges who never wrote the sentences; we conjecture that same-country accuracy would drop, because the current annotators can rely on memory of their own spellings.
  • The paper suggests, without stating it, that Arabizi spelling norms could be codified into a first cross-dialect normalization standard (e.g., canonical mappings per dialect); we think the alignment table is the seed for such a standard.
  • The age skew toward younger participants, which the paper acknowledges, implies the Egypt decline might reflect cohort behavior as much as a general shift; extrapolating the trend to older speakers is not supported.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper reports a human-centered study of Arabizi (Romanized Arabic) across five dialects: Algerian, Egyptian, Lebanese, Moroccan, and Tunisian. It surveys 189 Arabic speakers about their perceptions and usage of Arabizi, releases a word-level transliteration lexicon with character-level Arabic–Latin alignments, and introduces a parallel corpus of Arabic-script sentences with multiple Arabizi transliterations. The central linguistic claim is that cross-dialectal spelling variation is systematic enough that Arabic speakers can identify their own dialect’s Arabizi and detect cues from other dialects, as evidenced by an annotation experiment in which participants judged the country of origin of transliterations. The word-level alignment and survey instruments are carefully described and largely reproduced, but the sentence-level recognition experiment has a load-bearing design flaw: the annotators are the same ten participants who wrote the transliterations, and no Arabic-script control or chance-level baseline is reported.

Significance. If the central claim were established, the paper would make a useful contribution: it would show that Arabizi is not merely an ad hoc transliteration but carries dialect-specific orthographic cues that NLP systems should model, and it would provide two reusable resources (a 210-word lexicon and a parallel sentence corpus). The strengths are the detailed manual validation of the alignment algorithm, the full reproduction of the survey questionnaire in three languages, the explicit attention to ethical considerations and positionality, and the release of the collected lexicons in the appendix. The perceptual survey results themselves are informative and, independent of the recognition experiment, support a nuanced picture of declining but persistent Arabizi use. However, the sentence-level recognition claim is currently supported only by a circular experimental design, and the comparative statement about Arabizi versus Arabic script is unsupported because no Arabic-script condition was run. The resource contribution is real, but the paper’s headline conclusion needs additional evidence before it can be accepted.

major comments (4)
  1. [§5.2, Table 4] The recognition experiment is circular. The paper states: “The same 10 participants who provided the transliterations acted as annotators for this experiment.” An annotator can label a transliteration with “My Country” because they recognize their own exact spelling from episodic memory, not because they have inferred a dialect-wide orthographic convention. The same holds for the one other same-country writer, whose idiolect may be recognized as an individual rather than as evidence of a country-level pattern. The high diagonal values in Table 4 therefore do not validate the §7 conclusion that “annotators near-perfect identification of same-country transliterations” establishes cross-dialectal spelling variation. The experiment should be redone with annotators who did not write the material and, ideally, with more than two writers per country or with held-out-writer evaluation.
  2. [§5.2, final paragraph] The paper claims that Arabic speakers are “often more capable of guessing the dialect of a sentence in Arabizi than in Arabic script,” but no Arabic-script-only condition is reported. Moreover, the annotation interface displays the Arabic-script sentence alongside the Arabizi forms. The annotators could therefore use lexical, morphological, or other dialectal cues from the Arabic script itself, independent of Arabizi-specific spelling variation. A factorial design (Arabizi only, Arabic script only, and both together), with appropriate controls, is needed before attributing the result to Arabizi orthography. As written, the comparative and causal claims in this paragraph are unsupported.
  3. [§5.2, Table 4 and bullets] No chance-level baseline or statistical test is reported. With three response categories (“My Country”, “Not Sure”, “Another Country”), the off-diagonal criterion of “>50% labeled Another Country” is not obviously above chance, especially if participants have a bias toward one category. The diagonal accuracy is high, but the paper’s additional claim that participants “generally recognize spelling cues of other dialects” rests on these un-baselined off-diagonal counts. Exact binomial tests against the observed label distribution, or a permutation baseline, should be provided. Inter-annotator agreement is also not reported, which matters given the small annotator pool.
  4. [§5.1 and sample size] Even taking the recognition result at face value, the generalization from ten annotators (two per country) to “Arabic speakers” is very weak. The paper itself notes variation between the two annotators from the same country in some cells. The sentence selection also relies on an ALDi threshold of 0.33 and on the authors’ prior MLADI/ALDi resources; this is a reasonable design choice, but it makes the sentence set dependent on the authors’ own annotation criteria, so an independent validation of the sentence-level result is especially important. I would not consider this a fatal flaw on its own, but it strengthens the need for a naive-annotator replication.
minor comments (4)
  1. [§5.1] Typo: “transliteartions” and “14 sentences that has minor typos.” Also, the sentence says the corpus “has a total of 189 sentences,” but the flow from 203 to 189 should be stated more precisely (13 expletive-filtered plus 14 discarded).
  2. [Table 4 caption] The color legend (“Blue cells... red otherwise”) is inaccessible in print or for color-blind readers. Please replace with textual markers (e.g., bold/underline/asterisks) and explicitly state the actual counts for the Tunisia (I,J) and Algeria (B) exceptions, since the text says “near-perfect” but the table shows those cells are not >90%.
  3. [§3.2] “Statistical significance is validated in Appendix C” is imprecise: the appendix reports pairwise Fisher exact tests with FDR correction. “Assessed” or “tested” would be more accurate.
  4. [§4] The Wikipedia “Arabic Chat” page is mentioned as the source of the seed mapping table but no URL or access date is given. Please add a citation or a link in the appendix.

Circularity Check

1 steps flagged

Same-participant annotation confound: the 10 annotators also wrote the transliterations, so same-country 'validation' may reflect self-recognition rather than generalizable dialect cues.

specific steps
  1. other [§5.2 'Annotating Parallel Arabizi Sentences'; also used in §7 conclusion]
    "The same 10 participants who provided the transliterations acted as annotators for this experiment."

    The annotation results are offered as validation that cross-dialectal spelling variation is systematic (§7: 'validated by annotators near-perfect identification of same-country transliterations'). But the annotators produced the very transliterations they judged. For same-country cells in Table 4, an annotator can answer 'My Country' from episodic memory of their own spelling choices rather than from abstract orthographic regularity; for the other writer from the same two-person country pool, familiarity with that individual's idiolect can produce the same result. The design has no control for self-recognition or writer familiarity, and only two writers per country, so the high diagonal accuracy and off-diagonal 'Another Country' labels are partly constructed from the annotators' own produ

full rationale

The descriptive core of the paper is largely self-contained: the §3 survey and §4 word/character transliteration data are independent participant-provided measurements, and the character-level alignment, while using a seed mapping from Wikipedia plus manual validation, is not a prediction derived from its own output. The use of the authors' prior ALDi/MLADI resources for sentence filtering is also not circular: those are external datasets, and the recognition conclusion does not reduce to their correctness. The central circularity is in the sentence-level recognition experiment: the paper explicitly states that the same 10 participants who wrote the transliterations also annotated them. This makes the 'near-perfect identification of same-country transliterations' in Table 4 and the §7 claim that the variation is 'validated by annotators' at least partly an artifact of self-recognition and acquaintance with the two writers per country. The paper's own limitation that annotators were not asked to explain their judgments only makes this confound harder to rule out. Separately, the comparative claim that speakers are 'more capable of guessing the dialect of a sentence in Arabizi than in Arabic script' is unsupported because no Arabic-script control is reported, though this is a missing control rather than circularity. Because the lexicon, spelling-variation analysis, and perceptual survey retain independent descriptive value, the paper is only partially circular, not a case where the entire derivation reduces to its inputs.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

No new theoretical entities are introduced. The paper's contributions are datasets and empirical findings. The main load-bearing assumptions are about annotator independence and representativeness of the convenience sample, both of which are only partially addressed.

free parameters (1)
  • ALDi threshold = 0.33
    Sentences with country-level average ALDi score >= 0.33 were selected for the parallel corpus, discarding MSA-like sentences. This hand-set cutoff shapes the dataset and the subsequent recognition experiment (Section 5.1).
axioms (4)
  • domain assumption Arabizi spelling conventions are stable and dialect-specific enough to be recognized from orthographic/phonological cues alone.
    Central premise of Section 5; if Arabizi is too noisy, the recognition judgments reflect lexical content or participant memory rather than spelling cues.
  • domain assumption The same people who produced the transliterations can serve as unbiased annotators of dialect origin.
    Section 5.2 states that the same 10 participants acted as annotators; this assumes no self-recognition bias, which is directly challenged by the experimental design.
  • domain assumption Snowball-recruited, mostly young participants are representative of Arabizi users in each country.
    The survey was disseminated via social media and NLP mailing lists, and the authors acknowledge the age skew in the Limitations section. This affects claims about usage decline and general perception patterns.
  • domain assumption The manually validated seed letter-mapping table used for alignment is correct.
    Section 4 relies on an alignment algorithm seeded by a Wikipedia-extracted mapping table that was manually edited for 222 of 1,390 transliterations; errors in this mapping propagate to Table E1 and the character-level frequency claims.

pith-pipeline@v1.3.0-daily-deepseek · 33999 in / 9042 out tokens · 97455 ms · 2026-08-04T04:47:42.449365+00:00 · methodology

0 comments
read the original abstract

Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most NLP research on Arabic texts treats it as a temporary phenomenon resulting from limited technological support for the Arabic script. In this work, we engage with Arabic speakers to collect insights on their perceptions and usage of Arabizi. We further examine writing norms among speakers of different dialects, focusing on Algerian, Egyptian, Lebanese, Moroccan, and Tunisian Arabic. To this end, we release two resources. First, a character-level alignment of Arabic words to study inter- and intra-dialectal variation across these five dialects, based on words transliterated by survey participants, finding systematic intra-dialectal regularity and inter-dialectal variation. Second, to study Arabic speakers' ability to identify this stylistic variation at the sentence-level, we build a manually curated parallel corpus of sentences written in Arabic script alongside multiple Arabizi transliterations, collected from speakers of the same five dialects. Our study presents the largest human-centered, cross-dialectal study of Arabizi's perceptions and practices to date.

Figures

Figures reproduced from arXiv: 2608.02555 by Ahmed Amine Ben Abdallah, Amr Keleg, Chadi Helwe, Imane Guellil, Nedjma Ousidhoum, Taha Yassine.

Figure 1
Figure 1. Figure 1: Overview of the paper’s components to study [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The distributions of the perception (2a) and usage of Arabizi (2b) across the five studied countries. The two aspects are complementary. For instance, Algeria and Egypt have similar perception patterns, but completely different usage patterns. Note: Age Group is marginalized in 2b, given the skew shown in 2a. A few participants provided free-form answers to Q7, hence, the number of responses reported for 2… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

60 extracted references · 15 canonical work pages

  1. [3]

    Wafia Adouane, Nasredine Semmar, and Richard Johansson. 2016. https://aclanthology.org/W16-4807/ R omanized B erber and R omanized A rabic automatic language identification using machine learning . In Proceedings of the Third Workshop on NLP for Similar Languages, Varieties and Dialects ( V ar D ial3) , pages 53--61, Osaka, Japan. The COLING 2016 Organizi...

  2. [9]

    Tim Buckwalter. 2004. Buckwalter A rabic morphological analyzer version 2.0 (ldc2004l02). Web Download. LDC catalog number LDC2004L02, ISBN 1-58563-324-0

  3. [10]

    Ryan Cotterell, Adithya Renduchintala, Naomi Saphra, and Chris Callison-Burch. 2014. An A lgerian A rabic- F rench code-switched corpus. In Workshop on free/open-source A rabic corpora and corpora processing tools , page 34

  4. [12]

    Ronald A. Fisher. 1935. The Principles of Experimentation, Illustrated by a Psycho-physical Experiment, chapter 2. Oliver and B oyd, Edinburgh

  5. [13]

    Chayma Fourati, Hatem Haddad, Abir Messaoudi, Moez BenHajhmida, Aymen Ben Elhaj Mabrouk, and Malek Naski. 2021. https://aclanthology.org/2021.wanlp-1.25/ Introducing a large T unisian A rabizi dialectal dataset for sentiment analysis . In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 226--230, Kyiv, Ukraine (Virtual). Associa...

  6. [16]

    Nizar Y Habash. 2010. Introduction to A rabic Natural Language Processing . Morgan & C laypool P ublishers

  7. [18]

    Samia A Jaran and Fawwaz Al-Abed Al-Haq. 2015. The use of hybrid terms and expressions in C olloquial A rabic among J ordanian college students: A sociolinguistic study. English L anguage T eaching , 8(12):86--97

  8. [19]

    Anna Kashina. 2020. Case study of language preferences in social media of T unisia. In Proceedings of the International Conference Digital Age: Traditions, Modernity and Innovations (ICDATMI 2020), pages 111--115. Atlantis P ress

  9. [23]

    Karen McNeil. 2022. https://doi.org/doi:10.1515/ijsl-2021-0126 ‘we don't speak the same language:’ language choice and identity on a T unisian internet forum . International Journal of the Sociology of Language, 2022(278):51--80

  10. [24]

    Leila Moudjari, Karima Akli-Astouati, and Farah Benamara. 2020. https://aclanthology.org/2020.lrec-1.151/ An A lgerian corpus and an annotation platform for opinion and emotion analysis . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1202--1210, Marseille, France. European Language Resources Association

  11. [27]

    Wael Salloum. 2018. Machine Translation of Arabic Dialects. Ph D thesis, Columbia University

  12. [30]

    Muhammad Wafa. 2024. Arabizi ( F ranco) in E gypt: A study of features, reasons, attitudes, and educational influence among youth in online communication. Master's thesis, The A merican University in C airo ( E gypt)

  13. [31]

    A rabizi

    Mohammad Ali Yaghan. 2008. http://www.jstor.org/stable/25224166 " A rabizi": A contemporary style of A rabic slang . Design Issues, 24(2):39--52

  14. [32]

    Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks , year =

    Graves, Alex and Fern\'. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks , year =. Proceedings of the 23rd International Conference on Machine Learning , pages =. doi:10.1145/1143844.1143891 , abstract =

  15. [33]

    Writing between languages: The case of

    Abu-Liel, Aula Khatteb and Eviatar, Zohar and Nir, Bracha , journal=. Writing between languages: The case of. 2019 , publisher=. doi:10.1080/17586801.2020.1814482 , URL =

  16. [34]

    Mohammad Ali Yaghan , journal =. "

  17. [35]

    Cotterell, Ryan and Renduchintala, Adithya and Saphra, Naomi and Callison-Burch, Chris , booktitle=. An

  18. [36]

    The Use of Hybrid Terms and Expressions in

    Jaran, Samia A and Al-Haq, Fawwaz Al-Abed , journal=. The Use of Hybrid Terms and Expressions in. 2015 , publisher=

  19. [37]

    Constructing Linguistic Resources for the

    Younes, Jihen and Achour, Hadhemi and Souissi, Emna , booktitle=. Constructing Linguistic Resources for the. 2015 , organization=

  20. [38]

    Arabizi Transliteration of

    Guellil, Imane and Azouaou, Fai. Arabizi Transliteration of. Social

  21. [39]

    Atar: Attention-based

    Talafha, Bashar and Abuammar, Analle and Al-Ayyoub, Mahmoud , journal=. Atar: Attention-based. 2021 , publisher=

  22. [40]

    Arabizi (

    Wafa, Muhammad , year=. Arabizi (

  23. [41]

    Case study of language preferences in social media of

    Kashina, Anna , booktitle=. Case study of language preferences in social media of. 2020 , organization=

  24. [42]

    Technology Literacies of the New Media: Phrasing the World in the `` A rab E asy'' (R)evolution

    Gonzalez-Quijano, Yves. Technology Literacies of the New Media: Phrasing the World in the `` A rab E asy'' (R)evolution. Media Evolution on the Eve of the Arab Spring. 2014. doi:10.1057/9781137403155_9

  25. [43]

    2004 , address =

    Buckwalter, Tim , title =. 2004 , address =

  26. [44]

    Shang, Guokan and Abdine, Hadi and Chamma, Ahmad and Mohamed, Amr and Anwar, Mohamed and Bounhar, Abdelaziz and Herraoui, Omar El and Nakov, Preslav and Vazirgiannis, Michalis and Xing, Eric , journal=. Nile-

  27. [45]

    Jais and

    Neha Sengupta and Sunil Kumar Sahu and Bokang Jia and Satheesh Katipomu and Haonan Li and Fajri Koto and William Marshall and Gurpreet Gosal and Cynthia Liu and Zhiming Chen and Osama Mohammed Afzal and Samta Kamboj and Onkar Pandit and Rahul Pal and Lalit Pradhan and Zain Muhammad Mujahid and Massa Baali and Xudong Han and Sondos Mahmoud Bsharat and Alha...

  28. [46]

    Alzahrani and Nouf M

    M Saiful Bari and Yazeed Alnumay and Norah A. Alzahrani and Nouf M. Alotaibi and Hisham A. Alyahya and Sultan AlRashed and Faisal A. Mirza and Shaykhah Z. Alsubaie and Hassan A. Alahmed and Ghadah Alabduljabbar and Raghad Alkhathran and Yousef Almushayqih and Raneem Alnajim and Salman Alsubaihi and Maryam Al Mansour and Majed Alrubaian and Ali Alammari an...

  29. [47]

    2025 , eprint=

    Fanar: An. 2025 , eprint=

  30. [48]

    Introduction to

    Habash, Nizar Y , year=. Introduction to

  31. [49]

    Machine Translation of Arabic Dialects , type =

    Salloum, Wael , month =. Machine Translation of Arabic Dialects , type =

  32. [50]

    A Deep Learning Approach for Sentiment and Emotional Analysis of L ebanese A rabizi Twitter Data

    Ra \"i dy, Maria and Harmanani, Haidar. A Deep Learning Approach for Sentiment and Emotional Analysis of L ebanese A rabizi Twitter Data. ITNG 2023 20th International Conference on Information Technology-New Generations. 2023

  33. [51]

    Moroccan

    Soufiane Hajbi and Omayma Amezian and Nawfal El Moukhi and Redouan Korchiyne and Younes Chihab , keywords =. Moroccan. Scientific. 2024 , issn =. doi:https://doi.org/10.1016/j.sciaf.2024.e02073 , url =

  34. [52]

    The Role of Transliteration in the Process of A rabizi Translation/Sentiment Analysis

    Guellil, Imane and Azouaou, Faical and Benali, Fodil and Hachani, Ala Eddine and Mendoza, Marcelo. The Role of Transliteration in the Process of A rabizi Translation/Sentiment Analysis. Recent Advances in NLP : The Case of A rabic Language. 2020. doi:10.1007/978-3-030-34614-0_6

  35. [53]

    2024 , publisher=

    Gaanoun, Kamel and Naira, Abdou Mohamed and Allak, Anass and Benelallam, Imade , journal=. 2024 , publisher=

  36. [54]

    2507.04569 , archivePrefix=

    Guokan Shang and Hadi Abdine and Ahmad Chamma and Amr Mohamed and Mohamed Anwar and Abdelaziz Bounhar and Omar El Herraoui and Preslav Nakov and Michalis Vazirgiannis and Eric Xing , year=. 2507.04569 , archivePrefix=

  37. [55]

    2505.18383 , archivePrefix=

    Abdellah El Mekki and Houdaifa Atou and Omer Nacar and Shady Shehata and Muhammad Abdul-Mageed , year=. 2505.18383 , archivePrefix=

  38. [56]

    Case Study of Language Preferences in Social Media of

    Anna Kashina , year=. Case Study of Language Preferences in Social Media of. Proceedings of the International Conference Digital Age: Traditions, Modernity and Innovations (ICDATMI 2020) , pages=. doi:10.2991/assehr.k.201212.025 , publisher=

  39. [57]

    International Journal of the Sociology of Language , doi =

    , author =. International Journal of the Sociology of Language , doi =. 2022 , lastchecked =

  40. [58]

    and Margatina, Katerina and Mosquera-Gomez, Rafael and Ciro, Juan and Bartolo, Max and Williams, Adina and He, He and Vidgen, Bertie and Hale, Scott , booktitle =

    Kirk, Hannah Rose and Whitefield, Alexander and Rottger, Paul and Bean, Andrew M. and Margatina, Katerina and Mosquera-Gomez, Rafael and Ciro, Juan and Bartolo, Max and Williams, Adina and He, He and Vidgen, Bertie and Hale, Scott , booktitle =. The

  41. [59]

    Muhammad Dehan Al Kautsar and Lucky Susanto and Derry Wijaya and Fajri Koto , year=. What Do. 2506.07506 , archivePrefix=

  42. [60]

    , title =

    Fisher, Ronald A. , title =. The Design of Experiments , year =

  43. [61]

    The Affinity Diagram , booktitle =

    Karen Holtzblatt and Hugh Beyer , keywords =. The Affinity Diagram , booktitle =. 2017 , series =. doi:https://doi.org/10.1016/B978-0-12-800894-2.00006-5 , url =

  44. [62]

    A rabizi Detection and Conversion to A rabic

    Darwish, Kareem. A rabizi Detection and Conversion to A rabic. Proceedings of the EMNLP 2014 Workshop on A rabic Natural Language Processing ( ANLP ). 2014. doi:10.3115/v1/W14-3629

  45. [63]

    Automatic Transliteration of R omanized Dialectal A rabic

    Al-Badrashiny, Mohamed and Eskander, Ramy and Habash, Nizar and Rambow, Owen. Automatic Transliteration of R omanized Dialectal A rabic. Proceedings of the Eighteenth Conference on Computational Natural Language Learning. 2014. doi:10.3115/v1/W14-1604

  46. [64]

    R omanized B erber and R omanized A rabic Automatic Language Identification Using Machine Learning

    Adouane, Wafia and Semmar, Nasredine and Johansson, Richard. R omanized B erber and R omanized A rabic Automatic Language Identification Using Machine Learning. Proceedings of the Third Workshop on NLP for Similar Languages, Varieties and Dialects ( V ar D ial3). 2016

  47. [65]

    A rabizi Identification in T witter Data

    Tobaili, Taha. A rabizi Identification in T witter Data. Proceedings of the ACL 2016 Student Research Workshop. 2016. doi:10.18653/v1/P16-3008

  48. [66]

    Transliteration of A rabizi into A rabic Orthography: Developing a Parallel Annotated A rabizi- A rabic Script SMS /Chat Corpus

    Bies, Ann and Song, Zhiyi and Maamouri, Mohamed and Grimes, Stephen and Lee, Haejoong and Wright, Jonathan and Strassel, Stephanie and Habash, Nizar and Eskander, Ramy and Rambow, Owen. Transliteration of A rabizi into A rabic Orthography: Developing a Parallel Annotated A rabizi- A rabic Script SMS /Chat Corpus. Proceedings of the EMNLP 2014 Workshop on ...

  49. [67]

    S en Z i: A Sentiment Analysis Lexicon for the Latinised A rabic ( A rabizi)

    Tobaili, Taha and Fernandez, Miriam and Alani, Harith and Sharafeddine, Sanaa and Hajj, Hazem and Glava s , Goran. S en Z i: A Sentiment Analysis Lexicon for the Latinised A rabic ( A rabizi). Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019). 2019. doi:10.26615/978-954-452-056-4_138

  50. [68]

    An A lgerian Corpus and an Annotation Platform for Opinion and Emotion Analysis

    Moudjari, Leila and Akli-Astouati, Karima and Benamara, Farah. An A lgerian Corpus and an Annotation Platform for Opinion and Emotion Analysis. Proceedings of the Twelfth Language Resources and Evaluation Conference. 2020

  51. [69]

    Introducing A large T unisian A rabizi Dialectal Dataset for Sentiment Analysis

    Fourati, Chayma and Haddad, Hatem and Messaoudi, Abir and BenHajhmida, Moez and Ben Elhaj Mabrouk, Aymen and Naski, Malek. Introducing A large T unisian A rabizi Dialectal Dataset for Sentiment Analysis. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  52. [70]

    Jailbreaking LLM s with A rabic Transliteration and A rabizi

    Al Ghanim, Mansour and Almohaimeed, Saleh and Zheng, Mengxin and Solihin, Yan and Lou, Qian. Jailbreaking LLM s with A rabic Transliteration and A rabizi. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.1034

  53. [71]

    A Conventional Orthography for A lgerian A rabic

    Saadane, Houda and Habash, Nizar. A Conventional Orthography for A lgerian A rabic. Proceedings of the Second Workshop on A rabic Natural Language Processing. 2015. doi:10.18653/v1/W15-3208

  54. [72]

    NADI 2024: The Fifth Nuanced A rabic Dialect Identification Shared Task

    Abdul-Mageed, Muhammad and Keleg, Amr and Elmadany, AbdelRahim and Zhang, Chiyu and Hamed, Injy and Magdy, Walid and Bouamor, Houda and Habash, Nizar. NADI 2024: The Fifth Nuanced A rabic Dialect Identification Shared Task. Proceedings of the Second Arabic Natural Language Processing Conference. 2024. doi:10.18653/v1/2024.arabicnlp-1.79

  55. [73]

    Revisiting Common Assumptions about A rabic Dialects in NLP

    Keleg, Amr and Goldwater, Sharon and Magdy, Walid. Revisiting Common Assumptions about A rabic Dialects in NLP. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.acl-long.166

  56. [74]

    ALD i: Quantifying the A rabic Level of Dialectness of Text

    Keleg, Amr and Goldwater, Sharon and Magdy, Walid. ALD i: Quantifying the A rabic Level of Dialectness of Text. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.655

  57. [75]

    LLM Alignment for the A rabs: A Homogenous Culture or Diverse Ones

    Keleg, Amr. LLM Alignment for the A rabs: A Homogenous Culture or Diverse Ones. Proceedings of the 3rd Workshop on Cross-Cultural Considerations in NLP (C3NLP 2025). 2025. doi:10.18653/v1/2025.c3nlp-1.1

  58. [76]

    What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for G erman Dialects

    Blaschke, Verena and Purschke, Christoph and Schuetze, Hinrich and Plank, Barbara. What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for G erman Dialects. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2024. doi:10.18653/v1/2024.acl-short.74

  59. [77]

    What Do Indonesians Really Need from Language Technology? A Nationwide Survey

    Al Kautsar, Muhammad Dehan and Susanto, Lucky and Wijaya, Derry Tanti and Koto, Fajri. What Do Indonesians Really Need from Language Technology? A Nationwide Survey. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.367

  60. [78]

    Enriching the NA rabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language

    Riabi, Arij and Mahamdi, Menel and Seddah, Djam \'e. Enriching the NA rabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language. Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII). 2023. doi:10.18653/v1/2023.law-1.26