Pith. sign in

Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Large language models (LLMs) have demonstrated impressive capabilities across various natural language processing (NLP) tasks in recent years. However, their susceptibility to jailbreaks and perturbations necessitates additional evaluations. Many LLMs are multilingual, but safety-related training data contains mainly high-resource languages like English. This can leave them vulnerable to perturbations in low-resource languages such as Polish. We show how surprisingly strong attacks can be cheaply created by altering just a few characters and using a small proxy model for word importance calculation. We find that these character and word-level attacks drastically alter the predictions of different LLMs, suggesting a potential vulnerability that can be used to circumvent their internal safety mechanisms. We validate our attack construction methodology on Polish, a low-resource language, and find potential vulnerabilities of LLMs in this language. Additionally, we show how it can be extended to other languages. We release the created datasets and code for further research.

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

representative citing papers

PL-Guard: Benchmarking Language Model Safety for Polish

cs.CL · 2025-06-19 · reject · novelty 6.0

A small Polish BERT classifier proved more robust than larger fine-tuned LLMs at classifying safe versus unsafe Polish content, including under character-level adversarial perturbations.

citing papers explorer

Showing 1 of 1 citing paper.

  • PL-Guard: Benchmarking Language Model Safety for Polish cs.CL · 2025-06-19 · reject · none · ref 2025 · internal anchor

    A small Polish BERT classifier proved more robust than larger fine-tuned LLMs at classifying safe versus unsafe Polish content, including under character-level adversarial perturbations.