{"id":"f95018dd-b055-4d4e-ad59-3086fb97e03c","arxiv_id":"2501.13419","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of code-switched Arabic NLP, categorizing the literature, quantifying task and corpus coverage, and listing research gaps and recommendations.","lead":"This preprint reviews research on code-switched Arabic, where people mix Modern Standard Arabic, dialects, and foreign languages in the same conversation. It maps the field, quantifies where data and tasks are missing, and recommends where to focus next.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's quantitative claims rest on an undocumented Google Scholar corpus; a reproducible search protocol is needed before the prevalence and gap statistics can be trusted as a complete map of the field.","rationale":"The reader's weakest-assumption analysis correctly identifies the undocumented literature corpus as the main load-bearing risk. My independent check of the paper's internal tables and appendices found the counts to be consistent, which strengthens confidence in the authors' annotation effort but not in the external completeness of the search. The paper is a survey, not a model or dataset contribution, so the correctness risk is concentrated in the representativeness of the corpus. Since the reader already assigned a CONDITIONAL verdict for exactly this reason, my stress-test does not move the verdict. I would keep CONDITIONAL: the survey is useful and internally coherent, but the quantitative claims cannot be fully verified or extended until the search protocol is disclosed or the set is independently reproduced. No stronger objection—such as a computational error in the tables or a contradiction in the narrative—survived scrutiny.","tokens_in":28306,"tokens_out":3562,"duration_ms":31506,"concrete_test":"Reconstruct the literature search: on a fixed date, run the stated Google Scholar keywords plus a pre-registered expanded query set (e.g., 'code-mixing Arabic', 'Arabizi', 'diglossic code-switching') and screen results against explicit inclusion/exclusion criteria; have a second annotator independently classify a random sample and report agreement. Then compare the resulting corpus to Appendix C and recompute the Section 5 statistics and Table 3 task counts on the union. If the expanded set adds papers in currently under-represented categories (e.g., speech tasks or MSA-DA pairs), the gap analysis and prevalence claims are materially incomplete.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution is a structured review whose value depends on the completeness and representativeness of the surveyed literature. Section 4 states only that papers were found on Google Scholar using keywords 'code-switch', 'code-mix', and 'Arabic', with no search date, no query formulation details, no inclusion/exclusion criteria, no deduplication procedure, and no inter-annotator agreement for the category labels. All quantitative statements in Section 5 and Table 3 inherit this uncertainty. For example, the '12 papers per year since 2014' average and the task-coverage gaps could shift if the search missed work indexed under alternative terms such as 'code-mixing', 'language alternation', 'Arabizi', or dialect-specific names, or if Google Scholar's ranking and coverage bias the set toward certain venues and language pairs. The Limitations section acknowledges scope limits but does not address the reproducibility of the corpus itself. A spot-check of Table 3 against Appendices B and C shows the internal counts are consistent, so the weakness is not arithmetic; it is the absence of a documented, auditable selection protocol for the corpus that underlies the headline statistics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys code-switched Arabic NLP by collecting papers from Google Scholar, categorizing them according to language pairs, methods, NLP tasks, and venue types, and then presenting descriptive statistics, resource and task overviews, research gaps, and future directions. It makes no claim of introducing a new model or dataset; its contribution is a structured, quantitative map of the existing literature and a set of recommendations for the community.","tokens_in":28478,"tokens_out":3795,"duration_ms":37358,"significance":"If the underlying literature collection is complete and representative, the survey would be a valuable reference for researchers working on Arabic code-switching: it consolidates information about language-pair coverage, resource availability, task popularity, and open problems that is currently scattered across venues. The manuscript is internally consistent in its descriptive statistics, as the appendix tables line up with the counts in Table 3, and it explicitly acknowledges scope limitations. The quantitative value of the survey, however, rests entirely on the undocumented corpus-selection procedure, so the reproducibility of that procedure is the main condition for the contribution to be trusted.","major_comments":[{"comment":"The corpus-collection procedure is not reproducible. The text states only that Google Scholar was searched using keywords involving 'code-switch,' 'code-mix,' and 'Arabic,' but it does not report the search date, the exact query formulations, the inclusion and exclusion criteria, the deduplication procedure, or the inter-annotator agreement for the category labels in Table 2. All prevalence statistics in Section 5 and the task-coverage counts in Table 3 inherit this uncertainty. Because the survey's central value is a complete and structured map of the field, the selection protocol must be documented in sufficient detail that another researcher can replicate it, and the text should discuss the risk that papers indexed under alternative terms (e.g., 'code-mixing,' 'language alternation,' 'Arabizi,' or dialect-specific names) are missing from the corpus.","section":"Section 4 (Paper Categorization Process)"},{"comment":"The claim of a '3-4 year delay' in adopting neural methods and pretrained models relative to Winata et al. (2023) is based on a comparison between two corpora constructed with different collection procedures, different keyword strategies, and possibly different time windows. The present text does not provide the comparative statistics needed to support this temporal claim, and the reader cannot tell whether the apparent delay is an artifact of the two surveys' different scopes or annotation schemes. The authors should either present the comparative data directly or soften the claim to an observation that, within the current corpus, adoption appears later than in the broader code-switching literature.","section":"Section 5.3 (Evolution in Methodology Methods)"},{"comment":"The counts in Table 3 are paper-task incidences rather than distinct resources or systems, but this is not stated explicitly. For example, a single resource paper can contribute to multiple rows if it supports several tasks, and the appendix tables list papers rather than concrete resources. The sentence in Section 6 that 'the task of language modeling is supported by all collected textual corpora' is difficult to verify from the table and appendices because Table 3 reports a count of 49 resource papers for language modeling while the appendix shows only a small subset of papers listed under that task. The authors should clarify the unit of counting (papers vs. corpora vs. task claims) and state how the rows in Table 3 relate to the entries in Appendices B and C.","section":"Section 6 (NLP Tasks' Coverage) and Table 3"}],"minor_comments":[{"comment":"The sentence 'on average, there are 12 papers per year since 2014' does not state whether shared-task papers are included in the average or how the start year was chosen; this should be specified for clarity.","section":"Section 5.1"},{"comment":"The phrase 'inline with their prevalent use' should be 'in line with their prevalent use.'","section":"Section 5.4"},{"comment":"The sentence 'where linguistic studies on morphological CSW patterns can providing valuable support' contains a typo: 'can providing' should be 'can provide.'","section":"Section 8 (Handling Morphological CSW)"},{"comment":"The language code table lists 'Libyian Arabic' instead of 'Libyan Arabic'; please correct the spelling.","section":"Appendix A, Table 4"},{"comment":"The list of dataset sources (social media, transcriptions, speech recordings, etc.) is informative, but the text would benefit from explicitly distinguishing text-based and speech-based sources when reporting these counts.","section":"Section 6 (Dataset Sources)"},{"comment":"The annotation categories in Table 2 are described as inspired by Winata et al. (2023), but the exact mapping from that prior guideline to the current taxonomy is not described; adding a short explanation would improve transparency.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The author team includes several active researchers in Arabic code-switching, and a large share of the surveyed papers are authored by members of the team. I did not find evidence of deliberate omission of competing work, but the future-directions section emphasizes topics that align closely with the authors' own research lines (e.g., user-adaptive models, evaluation metrics for ASR, data augmentation). The editor may want an independent check that the gap analysis gives due weight to research from outside the author network. The main technical issue remains the missing reproducible search protocol; once that is addressed, the survey could be a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. This is the first Arabic-focused survey of code-switched NLP that systematically organizes the literature by language pair, task, and resource. The tables in the appendices and the task-coverage table (Table 3) are genuinely useful: if you work on CSW Arabic, you can quickly see which dialect pairs have corpora, which tasks are empirically studied, and where the gaps are. The paper also does a decent job situating the work relative to the general CSW surveys (Winata, Sitaram, Dogruoz) and noting the 3-4 year lag in adopting neural methods. No new models or datasets, and the paper doesn't claim any. That's fine for a survey.\n\nThe soft spot is the one the reader flagged. Section 4 says papers were found on Google Scholar with keywords 'code-switch', 'code-mix', and 'Arabic', and then 'we categorize.' There's no search date, no query strings, no inclusion/exclusion criteria, no deduplication procedure, and no agreement numbers on the annotations. Every quantitative claim in Section 5 and Table 3 inherits that uncertainty. The '12 papers per year' and the task-coverage gaps could shift if the search missed work indexed under 'code-mixing,' 'language alternation,' 'Arabizi,' or dialect-specific terms. I spot-checked the appendix tables against the main text; the arithmetic is internally consistent, so this isn't a data error. It's an auditable-protocol gap. The Limitations section is honest about scope but doesn't address reproducibility of the corpus itself.\n\nI don't think this is fatal. Surveys in NLP often get away with less rigor, and the qualitative parts—the discussion of LID, ASR, MT, evaluation metrics, morphological CSW—are well organized and well cited. The authors are also central figures in this subfield; the overlap between cited work and their own publications is heavy, but that's expected here and not a red flag by itself. The gap recommendations are reasonable, though they do lean toward the authors' own research interests.\n\nWho is this for? New researchers in Arabic NLP, and anyone building multilingual systems for Arab region languages. It's a map, not a method paper. I'd bring it to reading group for the table alone.\n\nRecommendation: send it to peer review. A competent referee can ask the authors to document the search protocol, specify the search date, list inclusion/exclusion decisions, and ideally add inter-annotator agreement on the categorization. None of that requires new experiments; it makes the existing survey auditable. With that revision, I'd be happy to see it published.","headline":"A competent, Arabic-specific survey with real map value, but the undocumented Google Scholar corpus makes the headline statistics less auditable than they should be.","tokens_in":29013,"tokens_out":2027,"would_cite":true,"duration_ms":18323,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A decade of code-switched Arabic NLP, mapped task by task, shows research clustered on a few language pairs.","keywords":["code-switching","Arabic NLP","survey","language identification","diglossia","Arabic dialects","machine translation","speech recognition"],"falsifier":"Re-run the survey with a documented, reproducible search strategy (query strings, dates, databases, and inclusion criteria) and compare the resulting corpus size and task-coverage table against the paper's counts; finding many more than the reported average of 12 papers per year, or locating datasets for tasks the paper reports as unsupported, would show the gap analysis is incomplete.","tokens_in":28116,"feed_emoji":"🗣️","tokens_out":4900,"duration_ms":40644,"temperature":0.7,"pith_summary":"This survey maps the field of code-switched Arabic natural language processing by compiling roughly a decade of published work and classifying it by language pair, NLP task, method, and empirical versus resource contribution. It finds that research clusters on a narrow set of settings—Egyptian Arabic with English, Algerian with French, and MSA-dialect switching—while most NLP tasks remain nearly untouched. The paper's contribution is the map itself: a structured picture of what resources exist, where models have been tested, and which gaps block progress. A sympathetic reader would use this as the entry point for choosing research directions in this area.","feed_headline":"Code-switched Arabic NLP: research clustered, gaps wide open","feed_subtitle":"Survey of ~10 years finds Egyptian-English and Algerian-French dominate; most NLP tasks remain unexplored.","key_machinery":"The organizing instrument is the paper's annotation scheme, shown in its Table 2. Each collected paper is tagged for year, venue, language pair (among MSA-DA, MSA-Foreign, DA-Foreign, Arabic-Foreign, and MSA-DA-Foreign), methodology (rule-based, statistical, neural, or pretrained), and NLP task, with empirical papers separated from resource papers. This scheme turns a heterogeneous literature into a countable matrix from which prevalence statistics, the task-coverage table, and the gap analysis are read directly.","core_discovery":"The paper's central claim is that the code-switched Arabic NLP literature, while growing at an average of 12 papers per year since 2014, is concentrated in a narrow band of language pairs and tasks. Word-level language identification, automatic speech recognition, named entity recognition, and machine translation dominate empirical work; only half of the tasks the authors tabulate have been explored at all. On the resource side, social media supplies most text data, and language modeling is the only task every textual corpus supports. The paper derives a gap list: benchmarks covering more language pairs and tasks, evaluations of pretrained models' code-switching ability, user-facing applications, personalized code-switched text generation, morphological code-switching handling, privacy and ethics, and several entirely unstudied high-level tasks.","pith_inferences":["The near-absence of Arabic from code-switching benchmarks suggests an implicit transfer argument: datasets built for Egyptian-English could serve as seed data for other Arabic-foreign pairs, though the paper does not test this.","The dominance of social media text, where intra-word switching is discouraged by script differences, implies that speech corpora are the more faithful evidence about morphological code-switching; the paper notes this tension but does not draw the sampling conclusion.","Given that human annotators reach only fair agreement on the naturalness of generated code-switched text, automatic naturalness metrics will likely need to be personal rather than global; that is an extension the paper leaves open.","The paper's dynamic view of code-switching—affected by topic, channel, demographics, and personality—implies user-adaptive models as the eventual goal, but it stops at listing factors; operationalizing those factors into features is a direct next step."],"forward_implications":["Researchers entering the area can use the paper's tables to pick tasks and language pairs with little competition; for example, only a handful of papers address speech translation or sentence-level language identification.","Because Arabic appears in the LinCE benchmark for only language identification and named entity recognition, building a broader code-switched Arabic benchmark covering more tasks and language pairs is the paper's most concrete next step.","The reported headroom—ASR word error rates of 28% to 54% and inconsistent large-language-model translation scores—means pretrained models are not yet reliable for production code-switched Arabic systems.","The authors recommend reporting evaluation results separately for morphological code-switching, since it behaves differently from sentence- or word-level switching and degrades machine translation and ASR performance.","Only half of the tabulated NLP tasks have any empirical work, so resource creation for tasks such as question answering, text-to-speech, and speech translation is a clear gap."],"supporting_citations":[{"why":"Supplies the categorization guidelines and the general code-switching survey that this Arabic-focused review narrows and builds upon.","marker":"Winata et al. (2023)"},{"why":"First shared task on language identification in code-switched data; drives the 2014 activity peak and the prominence of the LID task.","marker":"Solorio et al. (2014)"},{"why":"Introduces the LinCE benchmark, the main benchmark that includes Arabic code-switching, for LID and NER.","marker":"Aguilar et al. (2020)"},{"why":"Introduces the GLUECoS code-switching benchmark, which the paper contrasts with Arabic's sparse benchmark coverage.","marker":"Khanuja et al. (2020)"},{"why":"Defines the inter-sentential, extra-sentential, and intra-sentential code-switching typology that structures the paper's Section 3.","marker":"Poplack (1980)"},{"why":"The CALCS 2021 shared task on machine translation for code-switched data; underlies the paper's MT coverage claims.","marker":"Chen et al. (2021)"},{"why":"Supplies the terms 'diglossic' and 'bilingual' code-switching used throughout the survey.","marker":"Adouane et al. (2018a)"},{"why":"Motivates the call for user-facing code-switching applications and provides a cross-linguistic survey the paper extends to Arabic.","marker":"Do˘gruöz et al. (2021)"}],"fun_headline_variants":["Arabic code-switching NLP: narrow focus, big gaps","Survey: most Arabic code-switching tasks unexplored","Code-switched Arabic NLP: more gaps than coverage","Arabic code-switching research: clustered, not comprehensive"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's quantitative picture rests on a Google Scholar search with no documented inclusion criteria, search date, or agreement between annotators; if that search missed a substantial slice of the literature, the prevalence statistics and gap analysis would be incomplete.","fun_headline_variants_meta":{"raw":{"variants":["Arabic code-switching NLP: narrow focus, big gaps","Survey: most Arabic code-switching tasks unexplored","Code-switched Arabic NLP: more gaps than coverage","Arabic code-switching research: clustered, not comprehensive"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000566,"raw_usage":{"total_tokens":2609,"prompt_tokens":798,"completion_tokens":1811,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":1747}},"tokens_in":414,"tokens_out":1811,"duration_ms":11752,"temperature":1.0,"reasoning_tokens":1747,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:57:30.673892+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the survey with a documented, reproducible search strategy (query strings, dates, databases, and inclusion criteria) and compare the resulting corpus size and task-coverage table against the paper's counts; finding many more than the reported average of 12 papers per year, or locating datasets for tasks the paper reports as unsupported, would show the gap analysis is incomplete.","supporting_citations":[],"review_version":1}