Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that small monolingual models for lower- and medium-resourced languages are the most vulnerable to adversarial attacks, whereas larger multilingual models are generally more secure, and multilinguality alone does not…

desk verdict A valuable first security benchmark for low-resource languages, with a headline claim about model size that outruns the confounded comparison. read the letter →

arxiv 2507.03473 v1 pith:6XLGRU54 submitted 2025-07-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords adversarialattackslow-resourcelanguagesmultilinguallanguagemodelsTextFoolerNLPsecurityround-tripmachinetranslationmodelsizeclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether language models deployed for lower- and medium-resourced languages can withstand adversarial attacks, and which model choices make them more resilient. It extends the popular TextFooler attack to 70 languages using FastText for synonym lookup, and compares attack success rates across small monolingual models (Goldfish), moderately multilingual models, and large multilingual models (mBERT, XLM-R, Glot500). The central finding is that the smallest monolingual models are the most vulnerable, while larger multilingual models are generally the most secure; moderately multilingual models do not consistently beat similarly sized monolingual ones. The paper concludes that monolingual models for these languages are often too small to ensure sound security and that multilinguality helps but does not guarantee improved security.

What carries the argument

The machinery is a multilingual extension of TextFooler, the synonym-substitution adversarial attack, adapted to work without the resources TextFooler normally assumes. Instead of the English-only counter-fitted embeddings, synonym candidates come from FastText, which covers all 70 languages; the similarity threshold is lowered from 0.7 to 0.6; and stop-word filtering, POS tagging, and final sentence-encoder quality checks are dropped because they are unavailable for most of these languages. A second attack, round-trip machine translation through NLLB-200 via Zulu, serves as a baseline that mirrors common MT-based evaluations. The attack success rate of these two attacks on fine-tuned classifiers is what carries the paper's comparison of monolingual versus multilingual models.

What would settle it

A controlled experiment that trains a BERT-style monolingual model for a few lower-resourced languages on the same pretraining corpus sizes and domains as the Goldfish models, then re-runs the multilingual TextFooler attack across all model types; if the ordering of attack success rates by model size and multilinguality does not reproduce, the paper's central conclusion about monolingual insecurity would be weakened.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is an empirical ordering rather than a new mechanism: across 70 languages, the smallest Goldfish models (39M parameters) suffer the highest attack success rates under multilingual TextFooler, often above 0.80; the larger multilingual models (XLM-R large, Glot500) are generally the most secure, with ASR around 0.3–0.6; and moderately multilingual models such as IndicBERT, SlavicBERT, and CINO do not consistently improve over monolingual models of similar size. The paper also finds that the round-trip machine translation attack yields much lower ASR than TextFooler, suggesting that MT-based evaluations can underestimate the threat level. These results are presented as the first study of LM security for lower- and medium-resourced languages in their own right, rather than as a vehicle for attacking them.

Load-bearing premise

The comparison assumes that differences in attack success rate are explained by model size and multilinguality, but the models also differ in architecture, pretraining data size, and data domain—a confound the paper acknowledges as a lack of an apples-to-apples comparison.

Editorial extensions

If this is right

  • If the findings hold, deployers of language models for lower-resourced communities cannot treat small monolingual models as acceptable security baselines; they are the most easily flipped by synonym substitution.
  • Larger multilingual models such as XLM-R large and Glot500 are the safer default, though the paper stresses that even they retain substantial attack success rates for LRLs.
  • Increasing multilinguality alone, as in moderately multilingual models, is not a reliable security fix; language coverage and model size interact in ways that vary by language.
  • Security evaluations that rely only on round-trip machine translation may understate the real threat, since targeted attacks like TextFooler succeed at much higher rates.
  • Because most LRLs and MRLs lack the resources to train larger monolingual models, the practical route to better security runs through improving multilingual models and building defenses for the languages they serve.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's comparison is confounded: model size, pretraining corpus size, architecture, and domain vary together across the compared models, so its ordering is suggestive rather than causal; a controlled comparison that varies one factor at a time could confirm or overturn the size-versus-multilinguality story.
  • If the ordering survives such controls, then model size and language coverage could serve as cheap proxies when assessing the security of a new or unseen language model, before running expensive attacks.
  • The multilingual TextFooler pipeline could be extended to other attack families, such as character-level perturbations tailored to non-Latin scripts, to see whether the same model-level ordering appears.
  • The paper's framing that multilingual models are only as secure as their weakest language suggests a model-level risk score that aggregates the worst-performing language, which would be a natural next step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents a security evaluation of monolingual and multilingual encoder-only language models for 70 lower-, medium-, and higher-resourced languages. The authors extend the black-box TextFooler attack to a multilingual setting by replacing English synonym resources with FastText and dropping stop-word, POS-tagger, and sentence-encoder constraints, and they add a round-trip machine-translation baseline through Zulu. They fine-tune classifiers from four Goldfish monolingual models and four multilingual models (mBERT, XLM-R base/large, Glot500) on Taxi1500 and SIB200, then report attack success rates by resource level for the 45 languages with even model coverage, with full results in the appendix. A human evaluation of 80 adversarial samples in 8 languages supports the label fidelity of the generated attacks. The paper concludes that monolingual models are often too small to ensure sound security and that multilinguality helps but does not guarantee improved security.

Significance. The empirical contribution is substantial if read descriptively: it is, to my knowledge, the first systematic adversarial security audit spanning 70 languages with resource-level stratification, it demonstrates a feasible multilingual adaptation of a word-level attack in the absence of POS taggers and sentence encoders, and it reports results from over 3000 classifiers with multiple seeds and a human validation of sample quality. The descriptive finding that currently deployed monolingual options for LRLs and MRLs are much more vulnerable than the largest multilingual models is actionable for practitioners. The paper's causal interpretation of this ordering in terms of parameter count and multilinguality, however, is currently over-stated because the compared model families differ along several other axes; the revision should either add matched comparisons or restrict the claims to the models actually evaluated.

major comments (3)
  1. [Abstract; §5, "Monolingual vs Multilingual"; §B Limitations] The central claim that monolingual models are "often too small in total number of parameters to ensure sound security" is not established by the ASR ordering in Table 3 because the monolingual and multilingual families differ simultaneously in architecture, tokenizer, pretraining data scale, and pretraining domain. All monolingual models are GPT-2-style causal LMs with a 50k vocabulary and 5MB-1GB corpora, whereas all multilingual models are masked-LM BERT/RoBERTa variants with 110k-401k vocabularies and much larger, different-domain pretraining data; the Limitations section acknowledges the lack of an "apples-to-apples" comparison, so the abstract and conclusion should be reworded to a descriptive comparison of currently available models, or supplemented with matched controls (e.g., a BERT-style monolingual model of comparable parameter count).
  2. [§5, "Monolingual vs Multilingual"; Table 3] The statement that "the goldfish100mb and 1000mb models are more secure than mBERT despite being smaller in model size" is not consistently supported by the reported numbers: on Taxi1500 the ASR values are 0.62/0.65 vs 0.63 for LRLs, 0.73/0.73 vs 0.72 for MRLs, and 0.70/0.70 vs 0.68 for HRLs, so mBERT is equal or more secure in those cells, while the ordering reverses on SIB200. A per-language paired analysis with confidence intervals or statistical tests is needed before drawing any conclusion about the relative security of same-scale monolingual and multilingual models.
  3. [§5, "Does Multilinguality Guarantee Improved Security?"; Figure 3; Tables 2 and 4] The claim that moderately multilingual models "do not consistently outperform" similarly-sized monolingual models rests on three model families (Indic, Slavic, Chinese minority) with one model each, no reported statistical comparison, and size matches that are only approximate (IndicBERT v2 is 278M parameters against 125M for the largest Goldfish models, and SlavicBERT/CINO are 180M/190M). The observed pattern may hold, but the analysis as presented does not isolate model size, multilinguality, or language family; it should be quantified (e.g., paired per-language differences with variance) or explicitly framed as a case study.
minor comments (6)
  1. [Section 1, contribution bullet] The phrase "ensuring our work can is also comparable" should read "can also be compared".
  2. [Section 6, Conclusion] The phrase "as the the lack of resources" duplicates "the" and should be corrected.
  3. [Section B, Limitations] The phrase "scales for for HRLs" duplicates "for" and should be corrected.
  4. [Section 3, Multilingual TextFooler] The "average poisoning rate of 0.14 ± 0.07" is not defined; please state what quantity this rate measures (e.g., fraction of candidate substitutions that flip the label) and how it was computed.
  5. [Table 6] "norwegian" and "norwegian bokmål" appear as two rows with the same ISO code (nob); merge them or explicitly explain the intended distinction.
  6. [Section 2, Related Work] The statement that "no previous works in NLP Security have carefully examined LRLs or MRLs" is too strong given the Amharic and Arabic studies cited later; suggest softening to "systematically evaluated a broad set of LRLs and MRLs."

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the security rankings are measured attack success rates, not values derived from fitted inputs or self-citations.

full rationale

The paper's central findings are empirical attack success rates (ASR) obtained by running two black-box attacks, multilingual TextFooler and round-trip MT, over fine-tuned classifiers. ASR is defined as 1 minus the model's accuracy on the generated adversarial samples, and no parameter is fitted to the reported ASR values; the cosine similarity threshold (delta = 0.6) is fixed before the experiments and is not tuned to produce the observed ordering. The conclusion that small monolingual models are often too small to ensure sound security is an interpretation of the measured ordering, not a quantity derived from the inputs by construction. Self-citations (e.g., Lent et al. 2022, 2025; Chen, Lent, and Bjerva 2024) appear only as motivational, ethical, or contextual support and do not carry the load of the security ranking. The acknowledged lack of an apples-to-apples model comparison, stated in Limitations B, is a genuine validity concern about confounded architecture, pretraining data, and parameter count, but it is not circularity: the reported numbers would still be measured facts even if the attribution to parameter size were mistaken. No equation, fitted parameter, or load-bearing self-citation reduces the conclusions to their inputs.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The paper introduces no new theoretical entities or fitted predictive models. Its central dependency is on a set of domain assumptions about attack validity and model comparability. The only hand-tuned free parameters are the FastText similarity threshold, the top-k candidate pool, and standard fine-tuning hyperparameters. The absence of code and the small human evaluation are the main reproducibility constraints.

free parameters (3)
  • TextFooler cosine similarity threshold δ = 0.6
    Set by hand, lowered from the original 0.7 because FastText is not trained for synonym lookup. This threshold directly controls which candidate words qualify for replacement and therefore affects ASR.
  • Top-k synonym candidates = 50
    The top 50 nearest FastText neighbors per word are pulled as candidate synonyms, following the original TextFooler. The size of this pool influences attack success and sample quality.
  • Fine-tuning hyperparameters (learning rate, epochs, batch size) = 2e-5, 15 epochs, batch 32
    Chosen during initial hyperparameter tuning and applied uniformly across all models. These affect classifier accuracy and ASR, but they are not fitted to the security outcome per language.
assumptions (6)
  • domain assumption FastText cosine similarity is a sufficient proxy for synonymy across 70 languages, including low-resource ones.
    Section 3 replaces the English synonym embedding with FastText and does not provide per-language validation beyond the small human evaluation of 80 samples from 8 languages.
  • domain assumption Round-trip translation through Zulu preserves topic labels while producing useful adversarial inputs.
    Section 3 adopts the Zulu round-trip design from prior jailbreak work, without validating per-language that the translated sentences remain on-topic; the low RT-MT ASR suggests limited corruption.
  • domain assumption Classifiers that beat random weighted guessing are sufficiently trained for fair security comparisons.
    Section 4 uses this as an inclusion criterion and excludes mT5 when it fails, which could bias model coverage toward configurations that happen to train well.
  • domain assumption Model differences in ASR are attributable to model size and multilinguality rather than pretraining data domains, tokenizer, or architecture.
    Section 5 and the Limitations state that no 'apples to apples' comparison is possible, yet the main conclusions are framed in terms of size and multilinguality.
  • domain assumption Adversarial samples generated without POS taggers or sentence encoders are valid for measuring security.
    Section 5 Human Evaluation uses 80 samples across 8 languages with one annotator per language to support validity, which is a thin basis for generalizing to all 70 languages and all models.
  • domain assumption Evaluation datasets are not contaminated in a way that biases security comparisons.
    The Limitations section explicitly notes that data contamination is likely for Goldfish, Glot500, and XLM-R because their pretraining data overlaps with Bible and Wikipedia domains, but no contamination check is performed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right." pith.science (2026). https://pith.science/paper/6XLGRU54

@misc{pith2026250703473,
  author       = {Pith},
  title        = {Pith review of: Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6XLGRU54}},
  note         = {Machine review of arXiv:2507.03473}
}
read the original abstract

Despite mounting evidence that multilinguality can be easily weaponized against language models (LMs), works across NLP Security remain overwhelmingly English-centric. In terms of securing LMs, the NLP norm of "English first" collides with standard procedure in cybersecurity, whereby practitioners are expected to anticipate and prepare for worst-case outcomes. To mitigate worst-case outcomes in NLP Security, researchers must be willing to engage with the weakest links in LM security: lower-resourced languages. Accordingly, this work examines the security of LMs for lower- and medium-resourced languages. We extend existing adversarial attacks for up to 70 languages to evaluate the security of monolingual and multilingual LMs for these languages. Through our analysis, we find that monolingual models are often too small in total number of parameters to ensure sound security, and that while multilinguality is helpful, it does not always guarantee improved security either. Ultimately, these findings highlight important considerations for more secure deployment of LMs, for communities of lower-resourced languages.

Figures

Figures reproduced from arXiv: 2507.03473 by the authors.

Figure 1
Figure 1. We investigate LM security for 70 lower, medium, [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Examples of multilingual TextFooler applied to sentences from the SIB200 dataset. For each pair, the original inputs [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. We compare the best Mono(lingual) and Multi(lingual) models to the relevant Mod(erately) multilingual model. On [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. QQ: A Language Metadata Toolkit for Multilingual NLP

    cs.CL 2026-02 conditional novelty 6.0 of 10

    QQ unifies ISO, BCP-47, Glottocode, Wikidata, and other language identifiers into a traversable graph, and a HuggingFace audit shows deprecated codes like eml remain common.

Reference graph

Works this paper leans on

73 extracted references · 57 canonical work pages · cited by 1 Pith paper

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    G.; and Cevher, V

    Abad Rocamora, E.; Wu, Y.; Liu, F.; Chrysos, G. G.; and Cevher, V. 2024. Revisiting Character-level Adversarial Attacks for Language Models. In International Conference on Machine Learning (ICML)

  4. [4]

    Abdalla, M.; and Abdalla, M. 2020. The Grey Hoodie Project: Big Tobacco, Big Tech, and the threat on academic integrity. CoRR, abs/2009.13676

  5. [5]

    Adelani, D.; Liu, H.; Shen, X.; Vassilyev, N.; Alabi, J.; Mao, Y.; Gao, H.; and Lee, E.-S. 2024 a . SIB -200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects. In Graham, Y.; and Purver, M., eds., Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguisti...

  6. [6]

    I.; Do g ru \"o z, A

    Adelani, D. I.; Do g ru \"o z, A. S.; Coneglian, A.; and Ojha, A. K. 2024 b . Comparing LLM prompting with Cross-lingual transfer performance on Indigenous and Low-resource B razilian Languages. In Mager, M.; Ebrahimi, A.; Rijhwani, S.; Oncevay, A.; Chiruzzo, L.; Pugh, R.; and von der Wense, K., eds., Proceedings of the 4th Workshop on Natural Language Pr...

  7. [7]

    Akhtar, N.; and Mian, A. S. 2018. Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey. IEEE Access, 6: 14410--14430

  8. [8]

    Alshahrani, N.; Alshahrani, S.; Wali, E.; and Matthews, J. 2024. A rabic Synonym BERT -based Adversarial Examples for Text Classification. In Falk, N.; Papi, S.; and Zhang, M., eds., Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, 137--147. St. Julian ' s, Malta: Assoc...

Show all 73 references
  1. [9]

    Alslman, Y.; and Awajan, A. 2025. ArabicTextFool: A Framework for Arabic Adversarial Attacks. In 2025 International Conference on New Trends in Computing Sciences (ICTCS), 164--169

  2. [10]

    Alzantot, M.; Sharma, Y.; Elgohary, A.; Ho, B.-J.; Srivastava, M.; and Chang, K.-W. 2018. Generating Natural Language Adversarial Examples. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Conference on Empirical Methods in Natural Lang...

  3. [11]

    Antoun, W.; Baly, F.; and Hajj, H. 2020. A ra BERT : Transformer-based Model for A rabic Language Understanding. In Al-Khalifa, H.; Magdy, W.; Darwish, K.; Elsayed, T.; and Mubarak, H., eds., Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, w...

  4. [12]

    Arkhipov, M.; Trofimova, M.; Kuratov, Y.; and Sorokin, A. 2019. Tuning Multilingual Transformers for Language-Specific Named Entity Recognition. In Erjavec, T.; Marci \'n czuk, M.; Nakov, P.; Piskorski, J.; Pivovarova, L.; S najder, J.; Steinberger, J.; and Yangarber, R., eds....

  5. [13]

    Artetxe, M.; Labaka, G.; and Agirre, E. 2020. Translation Artifacts in Cross-lingual Transfer Learning. In Webber, B.; Cohn, T.; He, Y.; and Liu, Y., eds., Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 7674--7684. Online: Assoc...

  6. [14]

    Bird, S. 2022. Local Languages, Third Spaces, and other High-Resource Scenarios. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds., Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7817--7829. Dublin, Ireland...

  7. [15]

    Blasi, D.; Anastasopoulos, A.; and Neubig, G. 2022. Systematic Inequalities in Language Technology Performance across the World ' s Languages. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds., Proceedings of the 60th Annual Meeting of the Association for Computational Li...

  8. [16]

    Bojanowski, P.; Grave, E.; Joulin, A.; and Mikolov, T. 2016. Enriching Word Vectors with Subword Information. arXiv preprint arXiv:1607.04606

  9. [17]

    Cao, X.; Dawa, D.; Qun, N.; and Nyima, T. 2023. Pay Attention to the Robustness of C hinese Minority Language Models! Syllable-level Textual Adversarial Attack on T ibetan Script. In Ovalle, A.; Chang, K.-W.; Mehrabi, N.; Pruksachatkun, Y.; Galystan, A.; Dhamala, J.; Verma, A....

  10. [18]

    A.; Arnett, C.; Tu, Z.; and Bergen, B

    Chang, T. A.; Arnett, C.; Tu, Z.; and Bergen, B. 2024 a . When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages. In Al-Onaizan, Y.; Bansal, M.; and Chen, Y.-N., eds., Proceedings of the 2024 Conference on Empirical Methods in Natural Langu...

  11. [19]

    A.; Arnett, C.; Tu, Z.; and Bergen, B

    Chang, T. A.; Arnett, C.; Tu, Z.; and Bergen, B. K. 2024 b . Goldfish: Monolingual Language Models for 350 Languages. arXiv:2408.10441

  12. [20]

    Chen, Y.; Biswas, R.; Lent, H.; and Bjerva, J. 2025. Against All Odds: Overcoming Typology, Script, and Language Confusion in Multilingual Embedding Inversion Attacks. Proceedings of the AAAI Conference on Artificial Intelligence, 39(22): 23632--23641

  13. [21]

    Chen, Y.; Lent, H.; and Bjerva, J. 2024. Text Embedding Inversion Security for Multilingual Language Models. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7808...

  14. [22]

    Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzm \'a n, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; and Stoyanov, V. 2020. Unsupervised Cross-lingual Representation Learning at Scale. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J., eds., Proceed...

  15. [23]

    J.; and Bing, L

    Deng, Y.; Zhang, W.; Pan, S. J.; and Bing, L. 2024. Multilingual Jailbreak Challenges in Large Language Models. arXiv:2310.06474

  16. [24]

    Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North A merican Chapter of the Association...

  17. [25]

    M.; Kunchukuttan, A.; and Kumar, P

    Doddapaneni, S.; Aralikatte, R.; Ramesh, G.; Goyal, S.; Khapra, M. M.; Kunchukuttan, A.; and Kumar, P. 2023. Towards Leaving No I ndic Language Behind: Building Monolingual Corpora, Benchmark and Models for I ndic Languages. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds...

  18. [26]

    Dyrmishi, S.; Ghamizi, S.; and Cordy, M. 2023. How do humans perceive adversarial text? A reality check on the validity and naturalness of word-based adversarial attacks. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Assoc...

  19. [27]

    Fang, X.; Cheng, S.; Liu, Y.; and Wang, W. 2023. Modeling Adversarial Attack on Pre-trained Language Models as Sequential Decision Making. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Findings of the Association for Computational Linguistics: ACL 2023, 7322--7336. To...

  20. [28]

    Foo, J.; and Khoo, S. 2025. L ion G uard: A Contextualized Moderation Classifier to Tackle Localized Unsafe Content. In Rambow, O.; Wanner, L.; Apidianaki, M.; Al-Khalifa, H.; Eugenio, B. D.; Schockaert, S.; Darwish, K.; and Agarwal, A., eds., Proceedings of the 31st Internati...

  21. [29]

    L.; and Qi, Y

    Gao, J.; Lanchantin, J.; Soffa, M. L.; and Qi, Y. 2018. Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers. In 2018 IEEE Security and Privacy Workshops (SPW), 50--56

  22. [30]

    Goswami, D.; Zampieri, M.; North, K.; Malmasi, S.; and Anastasopoulos, A. 2025. Multilingual Native Language Identification with Large Language Models. In Ebrahimi, A.; Haider, S.; Liu, E.; Haider, S.; Leonor Pacheco, M.; and Wein, S., eds., Proceedings of the 2025 Conference ...

  23. [31]

    Graham, Y.; Haddow, B.; and Koehn, P. 2020. Statistical Power and Translationese in Machine Translation Evaluation. In Webber, B.; Cohn, T.; He, Y.; and Liu, Y., eds., Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 72--81. Onlin...

  24. [32]

    He, X.; Wang, J.; Xu, Q.; Minervini, P.; Stenetorp, P.; Rubinstein, B. I. P.; and Cohn, T. 2024. Transferring Troubles: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning. arXiv:2404.19597

  25. [33]

    H.; Severini, S.; Jalili Sabet, M.; Kassner, N.; Ma, C.; Schmid, H.; Martins, A.; Yvon, F.; and Sch \"u tze, H

    Imani, A.; Lin, P.; Kargaran, A. H.; Severini, S.; Jalili Sabet, M.; Kassner, N.; Ma, C.; Schmid, H.; Martins, A.; Yvon, F.; and Sch \"u tze, H. 2023. Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., e...

  26. [34]

    Jha, P.; Arora, A.; and Ganesh, V. 2024. LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs. arXiv:2411.08862

  27. [35]

    T.; and Szolovits, P

    Jin, D.; Jin, Z.; Zhou, J. T.; and Szolovits, P. 2020. Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment. arXiv:1907.11932

  28. [36]

    Joshi, P.; Santy, S.; Budhiraja, A.; Bali, K.; and Choudhury, M. 2020. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J., eds., Proceedings of the 58th Annual Meeting of the Association for Com...

  29. [37]

    Kogkalidis, K.; and Chatzikyriakidis, S. 2024. On Tables with Numbers, with Numbers. arXiv:2408.06062

  30. [38]

    J.; and Morreale, P

    Kumar, Y.; Paredes, C.; Yang, G.; Li, J. J.; and Morreale, P. 2024. Adversarial Testing of LLMs Across Multiple Languages. 2024 International Symposium on Networks, Computers and Communications (ISNCC), 1--6

  31. [39]

    Lei, Y.; Cao, Y.; Li, D.; Zhou, T.; Fang, M.; and Pechenizkiy, M. 2022. Phrase-level Textual Adversarial Attack with Label Preservation. In Carpuat, M.; de Marneffe, M.-C.; and Meza Ruiz, I. V., eds., Findings of the Association for Computational Linguistics: NAACL 2022, 1095-...

  32. [40]

    Lembersky, G.; Ordan, N.; and Wintner, S. 2012. Language Models for Machine Translation: Original vs. Translated Texts. Computational Linguistics, 38(4): 799--825

  33. [41]

    M.; Derczynski, L.; and Bjerva, J

    Lent, H.; Galinkin, E.; Chen, Y.; Pedersen, J. M.; Derczynski, L.; and Bjerva, J. 2025. NLP Security and Ethics, in the Wild. arXiv:2504.06669

  34. [42]

    Lent, H.; Ogueji, K.; de Lhoneux, M.; Ahia, O.; and S gaard, A. 2022. What a Creole Wants, What a Creole Needs. In Calzolari, N.; B \'e chet, F.; Blache, P.; Choukri, K.; Cieri, C.; Declerck, T.; Goggi, S.; Isahara, H.; Maegaard, B.; Mariani, J.; Mazo, H.; Odijk, J.; and Piper...

  35. [43]

    Li, L.; Ma, R.; Guo, Q.; Xue, X.; and Qiu, X. 2020. BERT - ATTACK : Adversarial Attack Against BERT Using BERT . In Webber, B.; Cohn, T.; He, Y.; and Liu, Y., eds., Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 6193--6202. Onli...

  36. [44]

    K.-W.; Chua, W.; Goh, J

    Lim, I.; Khoo, S.; Lee, R. K.-W.; Chua, W.; Goh, J. Y.; and Foo, J. 2025. Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study. arXiv:2502.12485

  37. [45]

    Läubli, S.; Sennrich, R.; and Volk, M. 2018. Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation. arXiv:1808.07048

  38. [46]

    Ma, C.; ImaniGooghari, A.; Ye, H.; Asgari, E.; and Schütze, H. 2023. Taxi1500: A Multilingual Dataset for Text Classification in 1500 Languages. arXiv:2305.08487

  39. [47]

    Mager, M.; Mager, E.; Kann, K.; and Vu, N. T. 2023. Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association for Comp...

  40. [48]

    Mayer, T.; and Cysouw, M. 2014. Creating a massively parallel B ible corpus. In Calzolari, N.; Choukri, K.; Declerck, T.; Loftsson, H.; Maegaard, B.; Mariani, J.; Moreno, A.; Odijk, J.; and Piperidis, S., eds., Proceedings of the Ninth International Conference on Language Reso...

  41. [49]

    J.; Cotterell, R.; Gorman, K.; Roark, B.; and Eisner, J

    Mielke, S. J.; Cotterell, R.; Gorman, K.; Roark, B.; and Eisner, J. 2019. What Kind of Language Is Hard to Language-Model? In Korhonen, A.; Traum, D.; and M \`a rquez, L., eds., Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4975--4989...

  42. [50]

    Y.; Grigsby, J.; Jin, D.; and Qi, Y

    Morris, J.; Lifland, E.; Yoo, J. Y.; Grigsby, J.; Jin, D.; and Qi, Y. 2020. TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System De...

  43. [51]

    M.; Su, P.-H.; Vandyke, D.; Wen, T.-H.; and Young, S

    Mrk s i \'c , N.; \'O S \'e aghdha, D.; Thomson, B.; Ga s i \'c , M.; Rojas-Barahona, L. M.; Su, P.-H.; Vandyke, D.; Wen, T.-H.; and Young, S. 2016. Counter-fitting Word Vectors to Linguistic Constraints. In Knight, K.; Nenkova, A.; and Rambow, O., eds., Proceedings of the 201...

  44. [52]

    Nakhleh, S.; Qasaimeh, M.; and Qasaimeh, A. 2024. Character-level Adversarial Attacks Evaluation for AraBERT’s. 2024 15th International Conference on Information and Communication Systems (ICICS), 1--6

  45. [53]

    H.; and Raji, I

    Nigatu, H. H.; and Raji, I. D. 2024. “I Searched for a Religious Song in Amharic and Got Sexual Content Instead’’: Investigating Online Harm in Low-Resourced Languages on YouTube. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT '2...

  46. [54]

    NLLB Team ; Costa-jussà, M. R.; Cross, J.; Çelebi, O.; Elbayad, M.; Heafield, K.; Heffernan, K.; Kalbassi, E.; Lam, J.; Licht, D.; Maillard, J.; Sun, A.; Wang, S.; Wenzek, G.; Youngblood, A.; Akula, B.; Barrault, L.; Gonzalez, G. M.; Hansanti, P.; Hoffman, J.; Jarrett, S.; Sad...

  47. [55]

    Pfeiffer, J.; Goyal, N.; Lin, X.; Li, X.; Cross, J.; Riedel, S.; and Artetxe, M. 2022. Lifting the Curse of Multilinguality by Pre-training Modular Transformers. In Carpuat, M.; de Marneffe, M.-C.; and Meza Ruiz, I. V., eds., Proceedings of the 2022 Conference of the North Ame...

  48. [56]

    Ploeger, E.; Poelman, W.; de Lhoneux, M.; and Bjerva, J. 2024 a . What is Typological Diversity in NLP ? In Al-Onaizan, Y.; Bansal, M.; and Chen, Y.-N., eds., Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 5681--5700. Miami, Florida, US...

  49. [57]

    Ploeger, E.; Poelman, W.; H eg-Petersen, A.; Schlichtkrull, A.; de Lhoneux , M.; and Bjerva, J. 2024 b . A Principled Framework for Evaluating on Typologically Diverse Languages

  50. [58]

    Pruthi, D.; Dhingra, B.; and Lipton, Z. C. 2019. Combating Adversarial Misspellings with Robust Word Recognition. In Korhonen, A.; Traum, D.; and M \`a rquez, L., eds., Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 5582--5591. Florenc...

  51. [59]

    Qi, F.; Chen, Y.; Zhang, X.; Li, M.; Liu, Z.; and Sun, M. 2021. Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer. In Moens, M.-F.; Huang, X.; Specia, L.; and Yih, S. W.-t., eds., Proceedings of the 2021 Conference on Empirical Methods in Na...

  52. [60]

    Ranathunga, S.; and de Silva, N. 2022. Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World. In He, Y.; Ji, H.; Li, S.; Liu, Y.; and Chang, C.-H., eds., Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Ass...

  53. [61]

    Reimers, N.; and Gurevych, I. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics

  54. [62]

    Ren, S.; Deng, Y.; He, K.; and Che, W. 2019. Generating Natural Language Adversarial Examples through Probability Weighted Word Saliency. In Korhonen, A.; Traum, D.; and M \`a rquez, L., eds., Proceedings of the 57th Annual Meeting of the Association for Computational Linguist...

  55. [63]

    A.; Muhammad, S

    Sani, S. A.; Muhammad, S. H.; and Jarvis, D. 2025. Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in H ausa Language Using A fri BERT a. In Hettiarachchi, H.; Ranasinghe, T.; Rayson, P.; Mitkov, R.; Gaber, M.; Premasiri, D.; Tan, F. A.; and Uya...

  56. [64]

    Toral, A.; Castilho, S.; Hu, K.; and Way, A. 2018. Attaining the Unattainable? Reassessing Claims of Human Parity in Neural Machine Translation. In Bojar, O.; Chatterjee, R.; Federmann, C.; Fishel, M.; Graham, Y.; Haddow, B.; Huck, M.; Yepes, A. J.; Koehn, P.; Monz, C.; Negri,...

  57. [65]

    Upadhayay, B.; and Behzadan, V. 2024. Sandwich attack: Multi-language Mixture Adaptive Attack on LLM s. In Ovalle, A.; Chang, K.-W.; Cao, Y. T.; Mehrabi, N.; Zhao, J.; Galstyan, A.; Dhamala, J.; Kumar, A.; and Gupta, R., eds., Proceedings of the 4th Workshop on Trustworthy Nat...

  58. [66]

    Wang, J.; Xu, Q.; He, X.; Rubinstein, B.; and Cohn, T. 2024 a . Backdoor Attacks on Multilingual Machine Translation. In Duh, K.; Gomez, H.; and Bethard, S., eds., Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics...

  59. [67]

    Wang, W.; Tu, Z.; Chen, C.; Yuan, Y.; Huang, J.-t.; Jiao, W.; and Lyu, M. 2024 b . All Languages Matter: On the Multilingual Safety of LLM s. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Association for Computational Linguistics ACL 2024, 5865--5877. Bang...

  60. [68]

    Yang, Z.; Xu, Z.; Cui, Y.; Wang, B.; Lin, M.; Wu, D.; and Chen, Z. 2022. CINO : A C hinese Minority Pre-trained Language Model. In Calzolari, N.; Huang, C.-R.; Kim, H.; Pustejovsky, J.; Wanner, L.; Choi, K.-S.; Ryu, P.-M.; Chen, H.-H.; Donatelli, L.; Ji, H.; Kurohashi, S.; Pag...

  61. [69]

    Yong, Z.-X.; Menghini, C.; and Bach, S. H. 2024. Low-Resource Languages Jailbreak GPT-4. arXiv:2310.02446

  62. [70]

    Yoo, H.; Yang, Y.; and Lee, H. 2024. Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding. arXiv:2406.15481

  63. [71]

    Zang, Y.; Qi, F.; Yang, C.; Liu, Z.; Zhang, M.; Liu, Q.; and Sun, M. 2020. Word-level Textual Adversarial Attacking as Combinatorial Optimization. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J., eds., Proceedings of the 58th Annual Meeting of the Association for Co...

  64. [72]

    Zhou, H.; Wang, Z.; Wang, H.; Chen, D.; Mu, W.; and Zhang, F. 2024. Evaluating the Validity of Word-level Adversarial Attacks with Large Language Models. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Association for Computational Linguistics ACL 2024, 4902...

  65. [73]

    Üstün, A.; Aryabumi, V.; Yong, Z.-X.; Ko, W.-Y.; D'souza, D.; Onilude, G.; Bhandari, N.; Singh, S.; Ooi, H.-L.; Kayid, A.; Vargus, F.; Blunsom, P.; Longpre, S.; Muennighoff, N.; Fadaee, M.; Kreutzer, J.; and Hooker, S. 2024. Aya Model: An Instruction Finetuned Open-Access Mult...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.