Pith. sign in

REVIEW 3 major objections 6 minor 50 references

Arabic dialects are causally steerable in LLMs through sparse late-layer neurons and, more reliably, distributed residual activation directions—without any fine-tuning.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 22:52 UTC pith:QS3AIOUT

load-bearing objection Solid first demonstration that Arabic dialects are causally steerable at inference time; vector steering works better than sparse neurons, and the residual-coverage analysis is the real interpretability payoff. the 3 major comments →

arxiv 2607.03936 v1 pith:QS3AIOUT submitted 2026-07-04 cs.CL

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

classification cs.CL
keywords Arabic dialectsmechanistic interpretabilityactivation steeringlanguage-specific neuronsLAPEresidual streaminference-time controlModern Standard Arabic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Arabic LLMs overproduce Modern Standard Arabic because dialect data is scarce. This paper asks where dialect features live inside the model and whether they can be used to control generation at inference time. It finds dialect-associated MLP neurons that concentrate in late, generation-facing layers and that are more shared among spoken dialects than with MSA. Amplifying those neurons can push outputs toward a target dialect, but because the features are only partly localized, a second method works better: extract a residual-stream direction as the mean difference between dialect and MSA response activations, then add a scaled version of that vector during decoding. Together the two probes show that dialect identity is both sparse and distributed, and that vector steering can improve dialect authenticity even when the prompt itself is in MSA.

Core claim

Dialectal information in Arabic-centric LLMs is neither fully modular nor fully diffuse: sparse LAPE-selected neurons in late layers encode dialect-selective features and can be rescaled to reinforce a target dialect, yet they capture only a partial projection of a broader residual-space dialect direction. Injecting that full contrastive direction at inference time yields more reliable dialect control than neuron rescaling alone, including when overriding MSA prompts.

What carries the argument

Two complementary inference-time interventions: (1) LAPE-selected dialect-associated MLP neurons, rescaled by amplification/suppression factors during decoding; (2) response-side dialect-minus-MSA mean activation difference vectors injected into a chosen residual layer for a limited token budget.

Load-bearing premise

That neurons and residual directions recovered from parallel MADAR dialect–MSA pairs truly isolate dialect identity rather than register, length, or corpus style, and that ADI2 plus an LLM judge plus small human ratings measure authentic dialect well enough to support the causal claims.

What would settle it

If residual dialect-minus-MSA vectors extracted from matched non-dialect style pairs (formal vs informal MSA, or length-matched paraphrases) produced equal ADI2 and authenticity gains, or if human raters from the target dialect communities scored vector-steered outputs as no more authentic than unsteered baselines, the claim that the directions specifically encode dialect would fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Dialect generation quality can be improved at decode time without collecting large dialect fine-tuning sets or changing model weights.
  • Late-layer MLP neurons and residual directions become concrete control knobs for dialect style in Arabic agents and content tools.
  • MSA-prompted generation can still be shifted toward a spoken dialect by residual injection, decoupling prompt register from output variety.
  • High neuron overlap among geographically close dialects suggests regional subcircuits that could support transfer between related varieties.
  • Because sparse neurons cover only a minority of the residual dialect direction, distributed vector methods should be preferred when robust control is the goal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same contrastive residual recipe may transfer to other diglossic or closely related language pairs where standard and colloquial varieties share vocabulary and orthography.
  • If dialect directions sit in residual space alongside persona or formality directions, compositional multi-attribute steering becomes a natural next experiment.
  • Weak Gulf results and incomplete neuron coverage imply that some varieties may need denser or multi-layer interventions rather than single-layer mean-difference vectors.
  • Community-grounded human evaluation across more speaker groups would be the cleanest way to pressure-test whether automatic gains track sociolinguistic authenticity.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper asks whether Arabic dialects, despite high lexical/syntactic overlap with MSA and each other, can be steered at inference time like distinct languages. On ALLaM-7B and Fanar-1-9B it identifies sparse LAPE dialect-associated MLP neurons (concentrated in late layers, with MSA largely separated from spoken varieties and regional sharing among dialects) and shows that rescaling them can reinforce dialectal generation. Motivated by incomplete localization, it also extracts residual-stream dialect–MSA mean-difference vectors from MADAR response tokens and injects them during decoding. Across mono-dialect and MSA-prompt settings, with ADI2/macro-ADI2, Gemini-as-judge, and blinded human ratings, vector steering is more reliable than neuron rescaling and can override MSA prompts; residual-subspace coverage shows LAPE neurons capture a meaningful but minority share of the dialect direction (with random-exclusion significance tests).

Significance. If the results hold, this is a useful first mechanistic account of intra-Arabic dialect encoding and a practical, training-free control recipe for under-resourced dialect generation. Strengths include dual models, multi-dialect MADAR extraction, systematic layer/coefficient/token ablations, residual-coverage tests against matched random neuron subsets, an MSA-prompt override experiment, human evaluation with agreement statistics, and released code. The sparse-plus-distributed geometry claim is a concrete contribution beyond pure engineering, and the finding that sparse neurons only partially span residual dialect directions explains the performance gap between the two interventions.

major comments (3)
  1. [§2.2, Appendix B.1] §2.2 and Appendix B.1: Steering vectors are mean response-token differences h(sk)−h(sMSA) on MADAR parallel pairs, and LAPE neurons are selected from the same dialect corpora. MADAR MSA counterparts are systematically more formal than the short colloquial dialect utterances; the paper’s own PCA notes a residual axis that may reflect register/formality rather than geography, and MSA formality is the metric that moves most cleanly under steering. Without a control that holds register fixed (e.g., colloquial MSA vs dialect, or subtracting a pure formality direction estimated from formal vs informal MSA), the causal claim that dialect identity—not colloquial-vs-formal register—is the steerable factor remains under-supported. ADI2’s dialect-specific term and geographic neuron/PCA structure partially mitigate this, but do not replace an explicit formality control.
  2. [§3.5–3.6, Table 1] §3.5–3.6 and Table 1: Layer, coefficient, and token ablations (and best α for neuron steering) are performed only on Egyptian and Moroccan Arabic, then transferred to Saudi and Syrian without retuning. For ALLaM on SAU/SYR, explicit prompting retains a judge advantage over vector steering; for Fanar, SAU/SYR ADI2 gains under vector steering are near floor. The central multi-dialect claim therefore rests on unvalidated transfer of free parameters. Either ablate on at least one Gulf and one Levantine variety, or clearly restrict the strong claims to the ablated dialects and treat the others as exploratory.
  3. [§4.2, Table 3] §4.2 and Table 3: Neuron steering fails entirely under MSA prompts (authenticity collapses, ADI2 → 0), while vector steering still induces dialect signal. The abstract and introduction frame both methods as complementary causal probes of dialect encoding; the MSA-prompt result shows sparse LAPE interventions can reinforce but not induce dialectal generation. The paper should state this limitation more sharply in the main claims and discuss what it implies for the “sparse neurons encode dialect-specific features” interpretation—i.e., that the selected neurons may be more diagnostic of dialect-conditioned generation than sufficient causal drivers of dialect induction.
minor comments (6)
  1. [Figure 1, §1] Figure 1 caption and §1 claim vector steering produces the “most colloquial” output; quantify this with the same metrics used elsewhere rather than a single qualitative example.
  2. [§3.5] LAPE selection uses top 5% activation and bottom 1% LAPE (§3.5) with little sensitivity analysis; a brief sweep or justification relative to Tang et al. (2024) would help.
  3. [Table 3] Table 3 reports identical ADI2/macro numbers for ALLaM and Fanar across several dialects (e.g., EGY 0.215/0.221); verify this is not a copy-paste error and, if real, discuss the coincidence.
  4. [§3.2, Appendix D.1] Appendix D.1: Gemini 2.5 Flash as dialect authenticity judge for Arabic varieties is a potential bias source; the human agreement (Table 8) is reassuring but should be flagged earlier in the main evaluation section.
  5. [§2.1–2.2] Notation: both neuron amplification and vector injection use α; disambiguate (e.g., α_n vs α_v) to avoid confusion when comparing methods.
  6. [§6] Related work could more explicitly contrast with Elshabrawy et al. (2025) on MSA–dialect entanglement and subspace decoupling, which is closely related to the residual-direction story.

Circularity Check

0 steps flagged

No circular derivation: interventions are extracted from MADAR contrasts and evaluated on independent generation metrics; residual coverage is a geometric comparison, not a tautology.

full rationale

The paper’s load-bearing chain is empirical, not definitional. LAPE neurons are selected from activation probabilities and entropy on dialect corpora (Sec. 2.1); steering multiplies those activations by free hyperparameters α/γ (ablated, not fitted to force the reported scores). Vector directions are mean residual differences v_k_ℓ = (1/N) Σ (h_ℓ(s_k) − h_ℓ(s_MSA)) from response tokens (Sec. 2.2), then injected at inference; success is measured by held-out ADI2/macro-ADI2, Gemini-as-judge dimensions, and blinded human ratings on mono-dialect and MSA-prompt generation (Sec. 3–4, Tables 1–3). Residual-subspace coverage ρ_l = ||Proj_{S_l}(v_l)||² / ||v_l||² (App. C) compares two independently constructed objects (LAPE down-projection span vs. contrastive residual direction) and is further tested against random-exclusion baselines—it does not redefine either quantity as the other. Citations for LAPE and activation addition (Tang et al. 2024; Turner et al. 2024; Rimsky et al. 2024) are external methods, not self-authored uniqueness theorems that force the dialect-control claim. No equation equates a “prediction” to a fitted input by construction; confounds such as register vs. dialect identity are validity concerns, not circularity. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

The paper imports standard activation-steering and LAPE machinery and adds only operational definitions (dialect-associated neurons, residual dialect vectors) plus a handful of selection and scaling hyperparameters chosen by ablation. No new physical or mathematical entities are postulated; the free parameters are the usual inference-time knobs of the steering literature.

free parameters (5)
  • target-dialect amplification α = 2.0 / 4.0
    Multiplicative scale on selected neurons; ablated and set to 2.0 (ALLaM) / 4.0 (Fanar).
  • LAPE selection percentiles (top n% activation, bottom m% LAPE) = n=5, m=1
    Neuron selection thresholds fixed at top 5% activation and bottom 1% LAPE without exhaustive search.
  • vector steering coefficient α and layer = layer/coeff pairs from ablation
    Chosen by grid search on Egyptian/Moroccan then transferred; final values model- and dialect-dependent (e.g., layer 20, α≈1–3).
  • number of steered tokens N = 30
    Token budget ablated and fixed at 30 for final runs.
  • MSA / competitor suppression γ, γ_comp = ≈1.0 (final)
    Tested but ultimately set near 1 (no suppression) after ablations showed limited benefit.
axioms (4)
  • domain assumption Low-LAPE, high-activation-probability MLP neurons are causally linked to dialect identity rather than generic Arabic or register features.
    Core premise of the neuron-steering pipeline (Section 2.1); supported by overlap statistics but not exhaustively controlled for confounds.
  • domain assumption Mean residual difference between dialect and MSA response activations is a valid linear steering direction for generation.
    Standard activation-addition assumption (Section 2.2); justified by PCA geometry and prior work but still an empirical modeling choice.
  • domain assumption ADI2, Gemini-2.5-Flash judge scores, and small human panels are adequate proxies for dialect authenticity and quality.
    Evaluation backbone (Sections 3.2–3.4); human–LLM agreement is reported but remains imperfect.
  • domain assumption MADAR city-level parallel sentences are representative of the broader dialect groups used at evaluation.
    Data mapping (Cairo→Egyptian, Rabat→Moroccan, etc.) is stated but not validated against other corpora.

pith-pipeline@v1.1.0-grok45 · 30833 in / 2836 out tokens · 25744 ms · 2026-07-11T22:52:26.913585+00:00 · methodology

0 comments
read the original abstract

A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLMs to overproduce MSA and struggle with dialectally accurate generation. From an interpretability perspective, this raises a fundamental question: where and how are dialectal features encoded within model internals, and can these representations be leveraged to improve dialect generation without fine-tuning? This study investigates two complementary inference-time approaches that serve simultaneously as interpretability probes and control mechanisms. First, we conduct a neuron-level analysis, identifying sparse neuron populations that encode dialect-specific features and showing that amplifying or suppressing these neurons can steer model outputs toward target dialects. Second, motivated by the entanglement of dialectal features at the single-neuron level, we apply a vector-steering approach that extracts dialect-specific activation directions and injects them during inference. Together, these methods illuminate the geometry of dialectal knowledge in Arabic LLMs and offer a principled, interpretability-grounded framework for dialect control without requiring dialect-specific fine-tuning.

Figures

Figures reproduced from arXiv: 2607.03936 by Fahim Dalvi, Kareem Elozeiri, Kentaro Inui, Mervat Abassy, Nadir Durrani, Omar Kallas, Preslav Nakov.

Figure 1
Figure 1. Figure 1: Example of unsteered, neuron-steered, and vector [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Dialect-specific neuron analysis for ALLaM and [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Layer ablation on Egyptian and Moroccan Arabic [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Residual-subspace coverage of vector-steering di [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Neuron-steering target-dialect amplification factor ablation results for [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Neuron-steering MSA suppression factor ablation results for [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Neuron-steering non-target dialect suppression factor ablation results for [PITH_FULL_IMAGE:figures/full_fig_p013_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: PCA 3D projections of mean response-side hidden [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Steering coefficient ablation heatmaps for AL [PITH_FULL_IMAGE:figures/full_fig_p014_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Token budget ablation for vector steering across [PITH_FULL_IMAGE:figures/full_fig_p015_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Sample-size sensitivity of the Cairo steering vector. The top panels show cosine similarity between vectors estimated [PITH_FULL_IMAGE:figures/full_fig_p016_11.png] view at source ↗
Figure 13
Figure 13. Figure 13: Explicit system prompt for the explicit-prompt [PITH_FULL_IMAGE:figures/full_fig_p018_13.png] view at source ↗
Figure 12
Figure 12. Figure 12: Prompt template used for LLM-as-a-judge evaluation. [PITH_FULL_IMAGE:figures/full_fig_p020_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 5 canonical work pages

  1. [1]

    ALD i: Quantifying the A rabic Level of Dialectness of Text

    Keleg, Amr and Goldwater, Sharon and Magdy, Walid. ALD i: Quantifying the A rabic Level of Dialectness of Text. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.655

  2. [2]

    NADI 2024: The Fifth Nuanced A rabic Dialect Identification Shared Task

    Abdul-Mageed, Muhammad and Keleg, Amr and Elmadany, AbdelRahim and Zhang, Chiyu and Hamed, Injy and Magdy, Walid and Bouamor, Houda and Habash, Nizar. NADI 2024: The Fifth Nuanced A rabic Dialect Identification Shared Task. Proceedings of the Second Arabic Natural Language Processing Conference. 2024. doi:10.18653/v1/2024.arabicnlp-1.79

  3. [3]

    The F lores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation

    Goyal, Naman and Gao, Cynthia and Chaudhary, Vishrav and Chen, Peng-Jen and Wenzek, Guillaume and Ju, Da and Krishnan, Sanjana and Ranzato, Marc ' Aurelio and Guzm \'a n, Francisco and Fan, Angela. The F lores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation. Transactions of the Association for Computational Linguistics. 2022...

  4. [4]

    chr F : character n-gram F -score for automatic MT evaluation

    Popovi \'c , Maja. chr F : character n-gram F -score for automatic MT evaluation. Proceedings of the Tenth Workshop on Statistical Machine Translation. 2015. doi:10.18653/v1/W15-3049

  5. [5]

    Representation Engineering: A Top-Down Approach to

    Zou, Andy and Phan, Long and Chen, Sarah and Campbell, James and Guo, Phillip and Ren, Richard and Pan, Alexander and Yin, Xuwang and Mazeika, Mantas and Dombrowski, Ann-Kathrin and others , journal =. Representation Engineering: A Top-Down Approach to. 2025 , url =

  6. [9]

    A rabic POS Tagging: Don ' t Abandon Feature Engineering Just Yet

    Darwish, Kareem and Mubarak, Hamdy and Abdelali, Ahmed and Eldesouki, Mohamed. A rabic POS Tagging: Don ' t Abandon Feature Engineering Just Yet. Proceedings of the Third A rabic Natural Language Processing Workshop. 2017. doi:10.18653/v1/W17-1316

  7. [10]

    arXiv preprint arXiv:2508.13130 , year =

    MuDRiC: Multi‐Dialect Reasoning for Arabic Commonsense Validation , author =. arXiv preprint arXiv:2508.13130 , year =

  8. [11]

    arXiv preprint arXiv:2502.11614 , year =

    Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI , author =. arXiv preprint arXiv:2502.11614 , year =

  9. [12]

    The MADAR A rabic Dialect Corpus and Lexicon

    Bouamor, Houda and Habash, Nizar and Salameh, Mohammad and Zaghouani, Wajdi and Rambow, Owen and Abdulrahim, Dana and Obeid, Ossama and Khalifa, Salam and Eryani, Fadhl and Erdmann, Alexander and Oflazer, Kemal. The MADAR A rabic Dialect Corpus and Lexicon. Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC 20...

  10. [16]

    Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , year=

    NeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks , author=. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , year=

  11. [18]

    The Interplay of Variant, Size, and Task Type in A rabic Pre-trained Language Models

    Inoue, Go and Alhafni, Bashar and Baimukan, Nurpeiis and Bouamor, Houda and Habash, Nizar. The Interplay of Variant, Size, and Task Type in A rabic Pre-trained Language Models. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  12. [19]

    Arid and Hasanain, Maram and Kabbani, Tameem and Dalvi, Fahim and Chowdhury, Shammur Absar and Alam, Firoj

    Mousi, Basel and Durrani, Nadir and Ahmad, Fatema and Hasan, Md. Arid and Hasanain, Maram and Kabbani, Tameem and Dalvi, Fahim and Chowdhury, Shammur Absar and Alam, Firoj. A ra D i CE : Benchmarks for Dialectal and Cultural Capabilities in LLM s. Proceedings of the 31st International Conference on Computational Linguistics. 2025

  13. [21]

    ALLaM: Large Language Models for Arabic and English , url =

    Bari, M Saiful and Alnumay, Yazeed and Alzahrani, Norah and Alotaibi, Nouf and Alyahya, Hisham and AlRashed, AlRashed and Mirza, Faisal and Alsubaie, Shaykhah and Alahmed, Hassan and Alabduljabbar, Ghadah and Alkhathran, Raghad and Almushayqih, Yousef and Alnajim, Raneem and Alsubaihi, Salman I and Al Mansour, Maryam and Hassan, Saad and Alrubaian, Majed ...

  14. [22]

    2025 , eprint=

    UI-Level Evaluation of ALLaM 34B: Measuring an Arabic-Centric LLM via HUMAIN Chat , author=. 2025 , eprint=

  15. [23]

    and Nacar, Omer and Nagoudi, El Moatez Billah and Abdel-Salam, Reem and Atwany, Hanin and Nafea, Youssef and Yahya, Abdulfattah Mohammed and Alhamouri, Rahaf and Alsayadi, Hamzah A

    Alwajih, Fakhraddin and El Mekki, Abdellah and Magdy, Samar Mohamed and Elmadany, AbdelRahim A. and Nacar, Omer and Nagoudi, El Moatez Billah and Abdel-Salam, Reem and Atwany, Hanin and Nafea, Youssef and Yahya, Abdulfattah Mohammed and Alhamouri, Rahaf and Alsayadi, Hamzah A. and Zayed, Hiba and Shatnawi, Sara and Sibaee, Serry and Ech-chammakhy, Yasir a...

  16. [24]

    D ial2 MSA -Verified: A Multi-Dialect A rabic Social Media Dataset for Neural Machine Translation to M odern S tandard A rabic

    Khered, Abdullah and Benkhedda, Youcef and Batista-Navarro, Riza. D ial2 MSA -Verified: A Multi-Dialect A rabic Social Media Dataset for Neural Machine Translation to M odern S tandard A rabic. Proceedings of the 4th Workshop on Arabic Corpus Linguistics (WACL-4). 2025

  17. [25]

    ARBERT & MARBERT : Deep Bidirectional Transformers for A rabic

    Abdul-Mageed, Muhammad and Elmadany, AbdelRahim and Nagoudi, El Moatez Billah. ARBERT & MARBERT : Deep Bidirectional Transformers for A rabic. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. doi:10.18653/v1/2021...

  18. [26]

    2023 , eprint=

    Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models , author=. 2023 , eprint=

  19. [27]

    2025 , url=

    Fanar: An Arabic-Centric Multimodal Generative AI Platform , author=. 2025 , url=

  20. [28]

    2025 , eprint=

    Qwen3 Technical Report , author=. 2025 , eprint=

  21. [29]

    2024 , eprint=

    GPT-4o System Card , author=. 2024 , eprint=

  22. [30]

    2025 , eprint=

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities , author=. 2025 , eprint=

  23. [31]

    2022 , eprint=

    No Language Left Behind: Scaling Human-Centered Machine Translation , author=. 2022 , eprint=

  24. [32]

    Habibi - a multi Dialect multi National A rabic Song Lyrics Corpus

    El-Haj, Mahmoud. Habibi - a multi Dialect multi National A rabic Song Lyrics Corpus. Proceedings of the Twelfth Language Resources and Evaluation Conference. 2020

  25. [33]

    Understanding and Mitigating Language Confusion in

    Marchisio, Kelly and Ko, Wei-Yin and B. Understanding and Mitigating Language Confusion in. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , month = nov, year =

  26. [36]

    2026 , eprint=

    Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs , author=. 2026 , eprint=

  27. [37]

    Anwar, Mohamed and Freihat, Abdelhakim and Ibrahim, George and Awad, Mostafa and Sadallah, Abdelrahman Atef Mohamed Ali and Gosal, Gurpreet and Ramakrishnan, Gokul and Chandran, Sarath and Mishra, Biswajit and Joshi, Rituraj and Frikha, Ahmed and Goffinet, Etienne and Maiti, Abhishek and El Filali, Ali and Al Barri, Sarah and Ghosh, Samujjwal and Pal, Rah...

  28. [39]

    2025 , eprint=

    When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models , author=. 2025 , eprint=

  29. [40]

    Ahmed Abdelali, Nadir Durrani, Fahim Dalvi, and Hassan Sajjad. 2022. https://doi.org/10.18653/v1/2022.blackboxnlp-1.8 Post-hoc analysis of A rabic transformer models . In Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 91--103, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computationa...

  30. [41]

    Muhammad Abdul-Mageed, AbdelRahim Elmadany, Chiyu Zhang, El Moatez Billah Nagoudi, Houda Bouamor, and Nizar Habash. 2023. https://doi.org/10.18653/v1/2023.arabicnlp-1.62 NADI 2023: The fourth nuanced A rabic dialect identification shared task . In Proceedings of ArabicNLP 2023, pages 600--613, Singapore (Hybrid). Association for Computational Linguistics

  31. [42]

    Krishak Aneja, Manas Mittal, Anmol Goel, Ponnurangam Kumaraguru, and Vamshi Krishna Bonagiri. 2026. https://arxiv.org/abs/2605.10633 Intrinsic guardrails: How semantic geometry of personality interacts with emergent misalignment in llms . Preprint, arXiv:2605.10633

  32. [43]

    M Saiful Bari, Yazeed Alnumay, Norah Alzahrani, Nouf Alotaibi, Hisham Alyahya, AlRashed AlRashed, Faisal Mirza, Shaykhah Alsubaie, Hassan Alahmed, Ghadah Alabduljabbar, Raghad Alkhathran, Yousef Almushayqih, Raneem Alnajim, Salman I Alsubaihi, Maryam Al Mansour, Saad Hassan, Majed Alrubaian, Ali Alammari, Zaki Alawami, and 7 others. 2025. https://proceedi...

  33. [44]

    Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, and Kemal Oflazer. 2018. https://aclanthology.org/L18-1535/ The MADAR A rabic dialect corpus and lexicon . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (...

  34. [45]

    Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey. 2025. https://arxiv.org/abs/2507.21509 Persona vectors: Monitoring and controlling character traits in language models . arXiv preprint arXiv:2507.21509

  35. [46]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, and 3416 others. 2025. https://arxiv.org/abs/2507.06261 Gemini 2.5: Pus...

  36. [47]

    Mahmoud El-Haj. 2020. https://aclanthology.org/2020.lrec-1.165/ Habibi - a multi dialect multi national A rabic song lyrics corpus . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1318--1326, Marseille, France. European Language Resources Association

  37. [48]

    Ahmed Elshabrawy, Hour Kaing, Haiyue Song, Alham Fikri Aji, Hideki Tanaka, Masao Utiyama, and Raj Dabre. 2025. https://arxiv.org/abs/2508.12803 When alignment hurts: Decoupling representational spaces in multilingual models . Preprint, arXiv:2508.12803

  38. [49]

    Fanar Team , Ummar Abbas, Mohammad Shahmeer Ahmad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf, Sabri Boughorbel, Sanjay Chawla, Shammur Chowdhury, Fahim Dalvi, Kareem Darwish, Nadir Durrani, Mohamed Elfeky, Ahmed Elmagarmid, Mohamed Eltabakh, Masoomali Fatehkia, Anastasios Fragkopoulos, Maram Hasanain, and 23 others. 2025. https://arxiv....

  39. [50]

    Daniil Gurgurov, Katharina Trinley, Yusser Al Ghussin, Tanja Baeumel, Josef Van Genabith, and Simon Ostermann. 2025. https://doi.org/10.18653/v1/2025.ijcnlp-long.156 Language arithmetics: Towards systematic language neuron identification and manipulation . In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th...

  40. [51]

    Nizar Y. Habash. 2010. https://doi.org/10.2200/S00277ED1V01Y201008HLT010 Introduction to Arabic Natural Language Processing , volume 3 of Synthesis Lectures on Human Language Technologies. Morgan & Claypool Publishers

  41. [52]

    Abdulmuizz Khalak, Abderrahmane Issam, and Gerasimos Spanakis. 2026. https://doi.org/10.18653/v1/2026.vardial-1.16 From F us H a to folk: Exploring cross-lingual transfer in A rabic language models . In Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects , pages 196--209, Rabat, Morocco. Association for Computational Linguistics

  42. [53]

    Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, and Firoj Alam

    Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, and Firoj Alam. 2025. https://aclanthology.org/2025.coling-main.283/ A ra D i CE : Benchmarks for dialectal and cultural capabilities in LLM s . In Proceedings of the 31st International Conference on Computational Linguistics, pa...

  43. [54]

    Omer Nacar. 2025. https://arxiv.org/abs/2508.17378 Ui-level evaluation of allam 34b: Measuring an arabic-centric llm via humain chat . Preprint, arXiv:2508.17378

  44. [55]

    Inaya Rahmanisa, Lyzander Marciano Andrylie, Mahardika Krisna Ihsani, Alfan Farizki Wicaksono, Haryo Akbarianto Wibowo, and Alham Fikri Aji. 2025. https://doi.org/10.18653/v1/2025.findings-ijcnlp.55 Unveiling the influence of amplifying language-specific neurons . In Proceedings of the 14th International Joint Conference on Natural Language Processing and...

  45. [56]

    Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. 2024. https://doi.org/10.18653/v1/2024.acl-long.828 Steering llama 2 via contrastive activation addition . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15504--15522, Bangkok, Thailand. Assoc...

  46. [57]

    Nathaniel Romney Robinson, Shahd Abdelmoneim, Kelly Marchisio, and Sebastian Ruder. 2025. https://doi.org/10.18653/v1/2025.findings-acl.1137 AL - QASIDA : Analyzing LLM quality and accuracy systematically in dialectal A rabic . In Findings of the Association for Computational Linguistics: ACL 2025, pages 22048--22065, Vienna, Austria. Association for Comp...

  47. [58]

    Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024. https://doi.org/10.18653/v1/2024.acl-long.309 Language-specific neurons: The key to multilingual capabilities in large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

  48. [59]

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2022. https://arxiv.org/abs/2207.04672 No language...

  49. [60]

    Vazquez, Ulisse Mini, and Monte MacDiarmid

    Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. 2024. https://arxiv.org/abs/2308.10248 Steering language models with activation engineering . arXiv preprint arXiv:2308.10248

  50. [61]

    Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, and 1 others. 2025. https://arxiv.org/abs/2310.01405 Representation engineering: A top-down approach to AI transparency . arXiv preprint arXiv:2310.01405