REVIEW 4 major objections 4 minor 60 references
This paper argues that Arabizi—Arabic written in Latin letters and numerals—has systematic, dialect-specific spelling patterns that speakers can recognize, even in sentences with no country-specific vocabulary.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 04:47 UTC pith:FT6HRO6Y
load-bearing objection Valuable new cross-dialectal Arabizi resources and credible word-level findings, but the sentence-recognition experiment is circular (same people who wrote the texts judged them), so the headline claim about dialect recognition is not yet supported. the 4 major comments →
Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that cross-dialectal spelling variation in Arabizi is systematic and perceptible. In a 210-word lexicon with about twenty transliterations per word, intra-dialectal regularity shows a preferred spelling for most words, while inter-dialectal differences surface in letter-to-script mappings: kha is 5 in Egypt, Lebanon, and Tunisia but kh in Algeria and Morocco; qaf is 9 mainly in the Maghreb and 2 or q in Egypt and Lebanon; and Algerian and Moroccan writers often double consonants to mark gemination. A parallel corpus of 189 Arabic sentences, each transliterated by two speakers per dialect, was judged by the ten participants who wrote them. Over 90 percent of same-coun
What carries the argument
The argument rests on two released resources and one alignment procedure. A character-level alignment algorithm maps Arabic letters to Latin-letter and numeral sequences across 1,390 unique transliterations, starting from a seed mapping table that the authors manually extend and validate; the resulting per-letter frequency table exposes contrasts like Maghrebi 9 versus Egyptian/Lebanese 2/q for qaf. A hand-built parallel corpus of 189 Arabic sentences—valid in at least two dialects, high in dialectality, free of country-specific words—tests whether spelling alone betrays the writer's dialect. Each transliteration was judged by all participants as 'My Country,' 'Not Sure,' or 'Another Country
Load-bearing premise
The sentence-level recognition result assumes the ten people who wrote the Arabizi sentences are unbiased judges of which dialect the spellings come from; if their near-perfect 'My Country' scores come from remembering their own writing instead of reading generalizable spelling cues, the central claim loses its support.
What would settle it
Recruit a fresh group of annotators from the same five countries who contributed none of the transliterations, give them the same 189 parallel sentences, and see whether same-country recognition stays above 90 percent. If fresh judges perform near chance—especially on sentences without dialect-specific words—the original scores reflect self-recognition rather than spelling-based dialect cues.
If this is right
- Arabizi carries enough dialectal signal to support dialect identification from orthography alone, even when texts contain no country-specific words.
- NLP systems that transliterate, normalize, or identify dialects in Arabizi can exploit per-dialect grapheme preferences (5 vs kh, 9 vs 2/q, ch vs sh) rather than treating Arabizi as a single noisy script.
- Because responses show decline—strongest in Egypt—and a persistent minority of heavy users, dialect-aware text processing remains relevant for moderation and user-facing chat systems.
- The released 210-word lexicon and 189-sentence parallel corpus give researchers a shared benchmark for whether systems reproduce the same between-dialect distinctions.
Where Pith is reading between the lines
- A direct test of the recognizability claim would be to run the same annotation task with fresh judges who never wrote the sentences; we conjecture that same-country accuracy would drop, because the current annotators can rely on memory of their own spellings.
- The paper suggests, without stating it, that Arabizi spelling norms could be codified into a first cross-dialect normalization standard (e.g., canonical mappings per dialect); we think the alignment table is the seed for such a standard.
- The age skew toward younger participants, which the paper acknowledges, implies the Egypt decline might reflect cohort behavior as much as a general shift; extrapolating the trend to older speakers is not supported.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a human-centered study of Arabizi (Romanized Arabic) across five dialects: Algerian, Egyptian, Lebanese, Moroccan, and Tunisian. It surveys 189 Arabic speakers about their perceptions and usage of Arabizi, releases a word-level transliteration lexicon with character-level Arabic–Latin alignments, and introduces a parallel corpus of Arabic-script sentences with multiple Arabizi transliterations. The central linguistic claim is that cross-dialectal spelling variation is systematic enough that Arabic speakers can identify their own dialect’s Arabizi and detect cues from other dialects, as evidenced by an annotation experiment in which participants judged the country of origin of transliterations. The word-level alignment and survey instruments are carefully described and largely reproduced, but the sentence-level recognition experiment has a load-bearing design flaw: the annotators are the same ten participants who wrote the transliterations, and no Arabic-script control or chance-level baseline is reported.
Significance. If the central claim were established, the paper would make a useful contribution: it would show that Arabizi is not merely an ad hoc transliteration but carries dialect-specific orthographic cues that NLP systems should model, and it would provide two reusable resources (a 210-word lexicon and a parallel sentence corpus). The strengths are the detailed manual validation of the alignment algorithm, the full reproduction of the survey questionnaire in three languages, the explicit attention to ethical considerations and positionality, and the release of the collected lexicons in the appendix. The perceptual survey results themselves are informative and, independent of the recognition experiment, support a nuanced picture of declining but persistent Arabizi use. However, the sentence-level recognition claim is currently supported only by a circular experimental design, and the comparative statement about Arabizi versus Arabic script is unsupported because no Arabic-script condition was run. The resource contribution is real, but the paper’s headline conclusion needs additional evidence before it can be accepted.
major comments (4)
- [§5.2, Table 4] The recognition experiment is circular. The paper states: “The same 10 participants who provided the transliterations acted as annotators for this experiment.” An annotator can label a transliteration with “My Country” because they recognize their own exact spelling from episodic memory, not because they have inferred a dialect-wide orthographic convention. The same holds for the one other same-country writer, whose idiolect may be recognized as an individual rather than as evidence of a country-level pattern. The high diagonal values in Table 4 therefore do not validate the §7 conclusion that “annotators near-perfect identification of same-country transliterations” establishes cross-dialectal spelling variation. The experiment should be redone with annotators who did not write the material and, ideally, with more than two writers per country or with held-out-writer evaluation.
- [§5.2, final paragraph] The paper claims that Arabic speakers are “often more capable of guessing the dialect of a sentence in Arabizi than in Arabic script,” but no Arabic-script-only condition is reported. Moreover, the annotation interface displays the Arabic-script sentence alongside the Arabizi forms. The annotators could therefore use lexical, morphological, or other dialectal cues from the Arabic script itself, independent of Arabizi-specific spelling variation. A factorial design (Arabizi only, Arabic script only, and both together), with appropriate controls, is needed before attributing the result to Arabizi orthography. As written, the comparative and causal claims in this paragraph are unsupported.
- [§5.2, Table 4 and bullets] No chance-level baseline or statistical test is reported. With three response categories (“My Country”, “Not Sure”, “Another Country”), the off-diagonal criterion of “>50% labeled Another Country” is not obviously above chance, especially if participants have a bias toward one category. The diagonal accuracy is high, but the paper’s additional claim that participants “generally recognize spelling cues of other dialects” rests on these un-baselined off-diagonal counts. Exact binomial tests against the observed label distribution, or a permutation baseline, should be provided. Inter-annotator agreement is also not reported, which matters given the small annotator pool.
- [§5.1 and sample size] Even taking the recognition result at face value, the generalization from ten annotators (two per country) to “Arabic speakers” is very weak. The paper itself notes variation between the two annotators from the same country in some cells. The sentence selection also relies on an ALDi threshold of 0.33 and on the authors’ prior MLADI/ALDi resources; this is a reasonable design choice, but it makes the sentence set dependent on the authors’ own annotation criteria, so an independent validation of the sentence-level result is especially important. I would not consider this a fatal flaw on its own, but it strengthens the need for a naive-annotator replication.
minor comments (4)
- [§5.1] Typo: “transliteartions” and “14 sentences that has minor typos.” Also, the sentence says the corpus “has a total of 189 sentences,” but the flow from 203 to 189 should be stated more precisely (13 expletive-filtered plus 14 discarded).
- [Table 4 caption] The color legend (“Blue cells... red otherwise”) is inaccessible in print or for color-blind readers. Please replace with textual markers (e.g., bold/underline/asterisks) and explicitly state the actual counts for the Tunisia (I,J) and Algeria (B) exceptions, since the text says “near-perfect” but the table shows those cells are not >90%.
- [§3.2] “Statistical significance is validated in Appendix C” is imprecise: the appendix reports pairwise Fisher exact tests with FDR correction. “Assessed” or “tested” would be more accurate.
- [§4] The Wikipedia “Arabic Chat” page is mentioned as the source of the seed mapping table but no URL or access date is given. Please add a citation or a link in the appendix.
Circularity Check
Same-participant annotation confound: the 10 annotators also wrote the transliterations, so same-country 'validation' may reflect self-recognition rather than generalizable dialect cues.
specific steps
-
other
[§5.2 'Annotating Parallel Arabizi Sentences'; also used in §7 conclusion]
"The same 10 participants who provided the transliterations acted as annotators for this experiment."
The annotation results are offered as validation that cross-dialectal spelling variation is systematic (§7: 'validated by annotators near-perfect identification of same-country transliterations'). But the annotators produced the very transliterations they judged. For same-country cells in Table 4, an annotator can answer 'My Country' from episodic memory of their own spelling choices rather than from abstract orthographic regularity; for the other writer from the same two-person country pool, familiarity with that individual's idiolect can produce the same result. The design has no control for self-recognition or writer familiarity, and only two writers per country, so the high diagonal accuracy and off-diagonal 'Another Country' labels are partly constructed from the annotators' own produ
full rationale
The descriptive core of the paper is largely self-contained: the §3 survey and §4 word/character transliteration data are independent participant-provided measurements, and the character-level alignment, while using a seed mapping from Wikipedia plus manual validation, is not a prediction derived from its own output. The use of the authors' prior ALDi/MLADI resources for sentence filtering is also not circular: those are external datasets, and the recognition conclusion does not reduce to their correctness. The central circularity is in the sentence-level recognition experiment: the paper explicitly states that the same 10 participants who wrote the transliterations also annotated them. This makes the 'near-perfect identification of same-country transliterations' in Table 4 and the §7 claim that the variation is 'validated by annotators' at least partly an artifact of self-recognition and acquaintance with the two writers per country. The paper's own limitation that annotators were not asked to explain their judgments only makes this confound harder to rule out. Separately, the comparative claim that speakers are 'more capable of guessing the dialect of a sentence in Arabizi than in Arabic script' is unsupported because no Arabic-script control is reported, though this is a missing control rather than circularity. Because the lexicon, spelling-variation analysis, and perceptual survey retain independent descriptive value, the paper is only partially circular, not a case where the entire derivation reduces to its inputs.
Axiom & Free-Parameter Ledger
free parameters (1)
- ALDi threshold =
0.33
axioms (4)
- domain assumption Arabizi spelling conventions are stable and dialect-specific enough to be recognized from orthographic/phonological cues alone.
- domain assumption The same people who produced the transliterations can serve as unbiased annotators of dialect origin.
- domain assumption Snowball-recruited, mostly young participants are representative of Arabizi users in each country.
- domain assumption The manually validated seed letter-mapping table used for alignment is correct.
read the original abstract
Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most NLP research on Arabic texts treats it as a temporary phenomenon resulting from limited technological support for the Arabic script. In this work, we engage with Arabic speakers to collect insights on their perceptions and usage of Arabizi. We further examine writing norms among speakers of different dialects, focusing on Algerian, Egyptian, Lebanese, Moroccan, and Tunisian Arabic. To this end, we release two resources. First, a character-level alignment of Arabic words to study inter- and intra-dialectal variation across these five dialects, based on words transliterated by survey participants, finding systematic intra-dialectal regularity and inter-dialectal variation. Second, to study Arabic speakers' ability to identify this stylistic variation at the sentence-level, we build a manually curated parallel corpus of sentences written in Arabic script alongside multiple Arabizi transliterations, collected from speakers of the same five dialects. Our study presents the largest human-centered, cross-dialectal study of Arabizi's perceptions and practices to date.
Figures
Reference graph
Works this paper leans on
-
[3]
Wafia Adouane, Nasredine Semmar, and Richard Johansson. 2016. https://aclanthology.org/W16-4807/ R omanized B erber and R omanized A rabic automatic language identification using machine learning . In Proceedings of the Third Workshop on NLP for Similar Languages, Varieties and Dialects ( V ar D ial3) , pages 53--61, Osaka, Japan. The COLING 2016 Organizi...
2016
-
[9]
Tim Buckwalter. 2004. Buckwalter A rabic morphological analyzer version 2.0 (ldc2004l02). Web Download. LDC catalog number LDC2004L02, ISBN 1-58563-324-0
2004
-
[10]
Ryan Cotterell, Adithya Renduchintala, Naomi Saphra, and Chris Callison-Burch. 2014. An A lgerian A rabic- F rench code-switched corpus. In Workshop on free/open-source A rabic corpora and corpora processing tools , page 34
2014
-
[12]
Ronald A. Fisher. 1935. The Principles of Experimentation, Illustrated by a Psycho-physical Experiment, chapter 2. Oliver and B oyd, Edinburgh
1935
-
[13]
Chayma Fourati, Hatem Haddad, Abir Messaoudi, Moez BenHajhmida, Aymen Ben Elhaj Mabrouk, and Malek Naski. 2021. https://aclanthology.org/2021.wanlp-1.25/ Introducing a large T unisian A rabizi dialectal dataset for sentiment analysis . In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 226--230, Kyiv, Ukraine (Virtual). Associa...
2021
-
[16]
Nizar Y Habash. 2010. Introduction to A rabic Natural Language Processing . Morgan & C laypool P ublishers
2010
-
[18]
Samia A Jaran and Fawwaz Al-Abed Al-Haq. 2015. The use of hybrid terms and expressions in C olloquial A rabic among J ordanian college students: A sociolinguistic study. English L anguage T eaching , 8(12):86--97
2015
-
[19]
Anna Kashina. 2020. Case study of language preferences in social media of T unisia. In Proceedings of the International Conference Digital Age: Traditions, Modernity and Innovations (ICDATMI 2020), pages 111--115. Atlantis P ress
2020
-
[23]
Karen McNeil. 2022. https://doi.org/doi:10.1515/ijsl-2021-0126 ‘we don't speak the same language:’ language choice and identity on a T unisian internet forum . International Journal of the Sociology of Language, 2022(278):51--80
-
[24]
Leila Moudjari, Karima Akli-Astouati, and Farah Benamara. 2020. https://aclanthology.org/2020.lrec-1.151/ An A lgerian corpus and an annotation platform for opinion and emotion analysis . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1202--1210, Marseille, France. European Language Resources Association
2020
-
[27]
Wael Salloum. 2018. Machine Translation of Arabic Dialects. Ph D thesis, Columbia University
2018
-
[30]
Muhammad Wafa. 2024. Arabizi ( F ranco) in E gypt: A study of features, reasons, attitudes, and educational influence among youth in online communication. Master's thesis, The A merican University in C airo ( E gypt)
2024
- [31]
-
[32]
Graves, Alex and Fern\'. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks , year =. Proceedings of the 23rd International Conference on Machine Learning , pages =. doi:10.1145/1143844.1143891 , abstract =
-
[33]
Writing between languages: The case of
Abu-Liel, Aula Khatteb and Eviatar, Zohar and Nir, Bracha , journal=. Writing between languages: The case of. 2019 , publisher=. doi:10.1080/17586801.2020.1814482 , URL =
arXiv 2019
-
[34]
Mohammad Ali Yaghan , journal =. "
-
[35]
Cotterell, Ryan and Renduchintala, Adithya and Saphra, Naomi and Callison-Burch, Chris , booktitle=. An
-
[36]
The Use of Hybrid Terms and Expressions in
Jaran, Samia A and Al-Haq, Fawwaz Al-Abed , journal=. The Use of Hybrid Terms and Expressions in. 2015 , publisher=
2015
-
[37]
Constructing Linguistic Resources for the
Younes, Jihen and Achour, Hadhemi and Souissi, Emna , booktitle=. Constructing Linguistic Resources for the. 2015 , organization=
2015
-
[38]
Arabizi Transliteration of
Guellil, Imane and Azouaou, Fai. Arabizi Transliteration of. Social
-
[39]
Atar: Attention-based
Talafha, Bashar and Abuammar, Analle and Al-Ayyoub, Mahmoud , journal=. Atar: Attention-based. 2021 , publisher=
2021
-
[40]
Arabizi (
Wafa, Muhammad , year=. Arabizi (
-
[41]
Case study of language preferences in social media of
Kashina, Anna , booktitle=. Case study of language preferences in social media of. 2020 , organization=
2020
-
[42]
Technology Literacies of the New Media: Phrasing the World in the `` A rab E asy'' (R)evolution
Gonzalez-Quijano, Yves. Technology Literacies of the New Media: Phrasing the World in the `` A rab E asy'' (R)evolution. Media Evolution on the Eve of the Arab Spring. 2014. doi:10.1057/9781137403155_9
-
[43]
2004 , address =
Buckwalter, Tim , title =. 2004 , address =
2004
-
[44]
Shang, Guokan and Abdine, Hadi and Chamma, Ahmad and Mohamed, Amr and Anwar, Mohamed and Bounhar, Abdelaziz and Herraoui, Omar El and Nakov, Preslav and Vazirgiannis, Michalis and Xing, Eric , journal=. Nile-
-
[45]
Neha Sengupta and Sunil Kumar Sahu and Bokang Jia and Satheesh Katipomu and Haonan Li and Fajri Koto and William Marshall and Gurpreet Gosal and Cynthia Liu and Zhiming Chen and Osama Mohammed Afzal and Samta Kamboj and Onkar Pandit and Rahul Pal and Lalit Pradhan and Zain Muhammad Mujahid and Massa Baali and Xudong Han and Sondos Mahmoud Bsharat and Alha...
-
[46]
M Saiful Bari and Yazeed Alnumay and Norah A. Alzahrani and Nouf M. Alotaibi and Hisham A. Alyahya and Sultan AlRashed and Faisal A. Mirza and Shaykhah Z. Alsubaie and Hassan A. Alahmed and Ghadah Alabduljabbar and Raghad Alkhathran and Yousef Almushayqih and Raneem Alnajim and Salman Alsubaihi and Maryam Al Mansour and Majed Alrubaian and Ali Alammari an...
-
[47]
2025 , eprint=
Fanar: An. 2025 , eprint=
2025
-
[48]
Introduction to
Habash, Nizar Y , year=. Introduction to
-
[49]
Machine Translation of Arabic Dialects , type =
Salloum, Wael , month =. Machine Translation of Arabic Dialects , type =
-
[50]
A Deep Learning Approach for Sentiment and Emotional Analysis of L ebanese A rabizi Twitter Data
Ra \"i dy, Maria and Harmanani, Haidar. A Deep Learning Approach for Sentiment and Emotional Analysis of L ebanese A rabizi Twitter Data. ITNG 2023 20th International Conference on Information Technology-New Generations. 2023
2023
-
[51]
Soufiane Hajbi and Omayma Amezian and Nawfal El Moukhi and Redouan Korchiyne and Younes Chihab , keywords =. Moroccan. Scientific. 2024 , issn =. doi:https://doi.org/10.1016/j.sciaf.2024.e02073 , url =
-
[52]
The Role of Transliteration in the Process of A rabizi Translation/Sentiment Analysis
Guellil, Imane and Azouaou, Faical and Benali, Fodil and Hachani, Ala Eddine and Mendoza, Marcelo. The Role of Transliteration in the Process of A rabizi Translation/Sentiment Analysis. Recent Advances in NLP : The Case of A rabic Language. 2020. doi:10.1007/978-3-030-34614-0_6
-
[53]
2024 , publisher=
Gaanoun, Kamel and Naira, Abdou Mohamed and Allak, Anass and Benelallam, Imade , journal=. 2024 , publisher=
2024
-
[54]
Guokan Shang and Hadi Abdine and Ahmad Chamma and Amr Mohamed and Mohamed Anwar and Abdelaziz Bounhar and Omar El Herraoui and Preslav Nakov and Michalis Vazirgiannis and Eric Xing , year=. 2507.04569 , archivePrefix=
-
[55]
Abdellah El Mekki and Houdaifa Atou and Omer Nacar and Shady Shehata and Muhammad Abdul-Mageed , year=. 2505.18383 , archivePrefix=
-
[56]
Case Study of Language Preferences in Social Media of
Anna Kashina , year=. Case Study of Language Preferences in Social Media of. Proceedings of the International Conference Digital Age: Traditions, Modernity and Innovations (ICDATMI 2020) , pages=. doi:10.2991/assehr.k.201212.025 , publisher=
-
[57]
International Journal of the Sociology of Language , doi =
, author =. International Journal of the Sociology of Language , doi =. 2022 , lastchecked =
2022
-
[58]
and Margatina, Katerina and Mosquera-Gomez, Rafael and Ciro, Juan and Bartolo, Max and Williams, Adina and He, He and Vidgen, Bertie and Hale, Scott , booktitle =
Kirk, Hannah Rose and Whitefield, Alexander and Rottger, Paul and Bean, Andrew M. and Margatina, Katerina and Mosquera-Gomez, Rafael and Ciro, Juan and Bartolo, Max and Williams, Adina and He, He and Vidgen, Bertie and Hale, Scott , booktitle =. The
-
[59]
Muhammad Dehan Al Kautsar and Lucky Susanto and Derry Wijaya and Fajri Koto , year=. What Do. 2506.07506 , archivePrefix=
-
[60]
, title =
Fisher, Ronald A. , title =. The Design of Experiments , year =
-
[61]
The Affinity Diagram , booktitle =
Karen Holtzblatt and Hugh Beyer , keywords =. The Affinity Diagram , booktitle =. 2017 , series =. doi:https://doi.org/10.1016/B978-0-12-800894-2.00006-5 , url =
-
[62]
A rabizi Detection and Conversion to A rabic
Darwish, Kareem. A rabizi Detection and Conversion to A rabic. Proceedings of the EMNLP 2014 Workshop on A rabic Natural Language Processing ( ANLP ). 2014. doi:10.3115/v1/W14-3629
-
[63]
Automatic Transliteration of R omanized Dialectal A rabic
Al-Badrashiny, Mohamed and Eskander, Ramy and Habash, Nizar and Rambow, Owen. Automatic Transliteration of R omanized Dialectal A rabic. Proceedings of the Eighteenth Conference on Computational Natural Language Learning. 2014. doi:10.3115/v1/W14-1604
-
[64]
R omanized B erber and R omanized A rabic Automatic Language Identification Using Machine Learning
Adouane, Wafia and Semmar, Nasredine and Johansson, Richard. R omanized B erber and R omanized A rabic Automatic Language Identification Using Machine Learning. Proceedings of the Third Workshop on NLP for Similar Languages, Varieties and Dialects ( V ar D ial3). 2016
2016
-
[65]
A rabizi Identification in T witter Data
Tobaili, Taha. A rabizi Identification in T witter Data. Proceedings of the ACL 2016 Student Research Workshop. 2016. doi:10.18653/v1/P16-3008
-
[66]
Bies, Ann and Song, Zhiyi and Maamouri, Mohamed and Grimes, Stephen and Lee, Haejoong and Wright, Jonathan and Strassel, Stephanie and Habash, Nizar and Eskander, Ramy and Rambow, Owen. Transliteration of A rabizi into A rabic Orthography: Developing a Parallel Annotated A rabizi- A rabic Script SMS /Chat Corpus. Proceedings of the EMNLP 2014 Workshop on ...
-
[67]
S en Z i: A Sentiment Analysis Lexicon for the Latinised A rabic ( A rabizi)
Tobaili, Taha and Fernandez, Miriam and Alani, Harith and Sharafeddine, Sanaa and Hajj, Hazem and Glava s , Goran. S en Z i: A Sentiment Analysis Lexicon for the Latinised A rabic ( A rabizi). Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019). 2019. doi:10.26615/978-954-452-056-4_138
-
[68]
An A lgerian Corpus and an Annotation Platform for Opinion and Emotion Analysis
Moudjari, Leila and Akli-Astouati, Karima and Benamara, Farah. An A lgerian Corpus and an Annotation Platform for Opinion and Emotion Analysis. Proceedings of the Twelfth Language Resources and Evaluation Conference. 2020
2020
-
[69]
Introducing A large T unisian A rabizi Dialectal Dataset for Sentiment Analysis
Fourati, Chayma and Haddad, Hatem and Messaoudi, Abir and BenHajhmida, Moez and Ben Elhaj Mabrouk, Aymen and Naski, Malek. Introducing A large T unisian A rabizi Dialectal Dataset for Sentiment Analysis. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021
2021
-
[70]
Jailbreaking LLM s with A rabic Transliteration and A rabizi
Al Ghanim, Mansour and Almohaimeed, Saleh and Zheng, Mengxin and Solihin, Yan and Lou, Qian. Jailbreaking LLM s with A rabic Transliteration and A rabizi. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.1034
-
[71]
A Conventional Orthography for A lgerian A rabic
Saadane, Houda and Habash, Nizar. A Conventional Orthography for A lgerian A rabic. Proceedings of the Second Workshop on A rabic Natural Language Processing. 2015. doi:10.18653/v1/W15-3208
-
[72]
NADI 2024: The Fifth Nuanced A rabic Dialect Identification Shared Task
Abdul-Mageed, Muhammad and Keleg, Amr and Elmadany, AbdelRahim and Zhang, Chiyu and Hamed, Injy and Magdy, Walid and Bouamor, Houda and Habash, Nizar. NADI 2024: The Fifth Nuanced A rabic Dialect Identification Shared Task. Proceedings of the Second Arabic Natural Language Processing Conference. 2024. doi:10.18653/v1/2024.arabicnlp-1.79
-
[73]
Revisiting Common Assumptions about A rabic Dialects in NLP
Keleg, Amr and Goldwater, Sharon and Magdy, Walid. Revisiting Common Assumptions about A rabic Dialects in NLP. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.acl-long.166
-
[74]
ALD i: Quantifying the A rabic Level of Dialectness of Text
Keleg, Amr and Goldwater, Sharon and Magdy, Walid. ALD i: Quantifying the A rabic Level of Dialectness of Text. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.655
-
[75]
LLM Alignment for the A rabs: A Homogenous Culture or Diverse Ones
Keleg, Amr. LLM Alignment for the A rabs: A Homogenous Culture or Diverse Ones. Proceedings of the 3rd Workshop on Cross-Cultural Considerations in NLP (C3NLP 2025). 2025. doi:10.18653/v1/2025.c3nlp-1.1
-
[76]
Blaschke, Verena and Purschke, Christoph and Schuetze, Hinrich and Plank, Barbara. What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for G erman Dialects. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2024. doi:10.18653/v1/2024.acl-short.74
-
[77]
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
Al Kautsar, Muhammad Dehan and Susanto, Lucky and Wijaya, Derry Tanti and Koto, Fajri. What Do Indonesians Really Need from Language Technology? A Nationwide Survey. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.367
-
[78]
Enriching the NA rabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language
Riabi, Arij and Mahamdi, Menel and Seddah, Djam \'e. Enriching the NA rabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language. Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII). 2023. doi:10.18653/v1/2023.law-1.26
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.