REVIEW 4 major objections 28 references
Low-resource multilingual jailbreaks leave the refusal mechanism intact but untriggered by routing harm through a misaligned residual subspace.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 14:12 UTC pith:3OVPV6DC
load-bearing objection Solid multi-perturbation multilingual safety benchmark with a coherent geometric story; the T2–T3 cliff and Tier-4 ASRs rest on a soft judge pipeline, but the attack-type profiles and subspace misalignment findings still hold up. the 4 major comments →
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Low-resource multilingual jailbreaks succeed mainly by routing harmful content through a geometrically misaligned residual subspace that projects insufficiently onto the refusal direction, leaving the refusal mechanism intact but untriggered (subthreshold activation). Each of four attack types yields a distinct, cross-model vulnerability profile, with a sharp refuse-to-comply regime transition between resource Tiers 2 and 3.
What carries the argument
A two-stage geometric account of residual-stream safety: harmfulness subspace (from linear probes) versus refusal direction (from English refuse–comply contrast), with failure typed as upstream semantic-recovery collapse or downstream subthreshold activation when the harm signal fails to clear the refusal threshold. Principal angles between non-English and English harm subspaces quantify the misalignment.
Load-bearing premise
Non-English attack success is assumed to be reliably measured by back-translating model answers into English and scoring them with an English safety judge, even when that back-translation is weak for the lowest-resource languages.
What would settle it
Have native speakers judge the original non-English model outputs (no back-translation) for the same three models and four attack types; if the Tier 2–3 refuse-to-comply cliff or the script-conditioned transliteration pattern disappears or reverses, the reported profiles and geometric story fail.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Minionese, a multilingual jailbreak benchmark of 55,440 prompts spanning 18 languages, four resource tiers, and four meaning-preserving perturbations (standard translation, code-switching, transliteration, translationese), built from AdvBench harmful–harmless pairs. Evaluating Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Aya-Expanse-8B, it reports distinct, cross-model vulnerability profiles: transliteration ASR is script-conditioned, code-switching remains effective through Tier 4, and a sharp refuse-to-comply transition appears between Tiers 2 and 3. A geometric analysis of residual-stream activations then argues that low-resource jailbreaks primarily succeed via subthreshold activation—harmful content is often linearly decodable but projects insufficiently onto an intact refusal direction—while Tier-4 and non-Latin transliteration can also collapse the harm representation itself. A multi-dimensional refusal-cone analysis further claims that only the first cone basis direction transfers cross-lingually.
Significance. If the behavioral and geometric claims hold, the work is a substantial contribution to multilingual LLM safety: it upgrades translation-only benchmarks by treating perturbation type as an independent variable, supplies paired harmful–harmless prompts for mechanistic analysis, and offers a falsifiable two-stage account (semantic recovery failure vs. subthreshold activation) that is consistent across three architectures. The released benchmark and analysis code, the script-conditioned transliteration pattern, and the effective one-dimensionality of cross-lingual refusal are concrete, reusable results for evaluation design and geometry-aware interventions. These strengths matter for governance and deployment beyond English-centric red-teaming.
major comments (4)
- §4.1.3 and Limitations F.1: Non-English ASR is defined by back-translation to English then WildGuard scoring. The only human validation is a spot-check claiming 100% agreement on an unspecified sample, with no per-tier, per-perturbation, or per-model agreement rates and no second judge. Because the sharp Tier 2–3 regime transition and the tier gradients are load-bearing for the central claim, the paper needs either (i) stratified human labels (or dual-judge agreement) on a non-trivial Tier-3/4 sample, or (ii) a sensitivity analysis showing that the reported cliffs survive plausible mislabel rates. Without this, the behavioral foundation of the mechanistic narrative remains under-anchored.
- §4.2 / Eq. (1): ASR is reported as point estimates with no confidence intervals, bootstrap errors, or significance tests across languages, tiers, or models. The claimed “sharp” Tier 2–3 transition and cross-model rank-order stability of attack types cannot be assessed for sampling variability. Adding per-cell uncertainty (and, where appropriate, simple tests of tier contrasts) is necessary for the regime-transition claim to be scientifically usable.
- §5.1–5.2: Failure-type classification (upstream / mixed / downstream) and “subthreshold activation” rest on free analysis choices—τ^(ℓ)=0.95, harmfulness subspace rank k (≥95% spectral mass), and the English-only mean-difference refusal direction—without reported sensitivity. The dominant “mixed” classification and the claim that refusal is never suppressed once adequately triggered should be shown to be robust under reasonable variation of τ and k, or the threshold should be calibrated (e.g., from English ROC) rather than fixed at 0.95.
- §5 and F.1: Causal language (“routing,” “starving,” “leaving the refusal mechanism intact but untriggered”) is stronger than the evidence. The geometric results are correlational; the manuscript itself defers activation/path patching to future work. Either run a minimal causal check (e.g., projecting Tier-3 harmful activations onto ˆr and measuring refusal recovery) or systematically rephrase claims as geometric associations consistent with, but not yet proving, the two-stage mechanism.
Circularity Check
No significant circularity: empirical ASR measurements and observational residual-stream geometry are independent of each other; failure labels restate measured projections, not redefine the behavioral outcomes.
full rationale
Minionese is an empirical benchmark plus post-hoc geometric analysis, not a first-principles derivation. Attack success rates are measured externally via WildGuard (with back-translation) on held-out harmful prompts; resource tiers and the four perturbation types are design choices that partition the data, not quantities defined from the ASR they are used to report. The geometric pipeline (logistic probes for harmfulness subspaces, mean-difference English refusal direction, silhouette/AUC/principal-angle statistics, and the optional multi-dimensional cone fit with free weights λ_cone, λ_xl) observes residual-stream structure and correlates it with those independently measured ASRs. Calling a language “subthreshold” when harm is linearly decodable but ˆr⊤h fails a fixed τ=0.95 is a labeling of the same geometric measurements, not a prediction forced by fitting ASR. There is no self-citation load-bearing chain (Arditi, Zhao, Wollschläger, Wang et al. are external), no uniqueness theorem imported from the authors, and no ansatz smuggled in as a theorem. Limitations correctly flag that causal attribution remains correlational. Measurement-validity concerns about WildGuard-on-back-translation affect correctness risk, not circularity of the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (5)
- harmfulness subspace rank k =
≥95% spectral mass
- refusal decision threshold τ^(ℓ) =
0.95
- cone optimization hyperparameters =
η=0.05, λcone=5, λxl=1, N=5, 400 steps
- coherence filter cutoffs =
conf<0.3, unicode<0.8, len<20
- resource tier language assignment =
4 tiers / 18 languages as listed in §3
axioms (6)
- domain assumption Refusal in aligned LLMs is mediated by a low-dimensional direction (or cone) in residual activation space that can be extracted by behavioral contrast.
- domain assumption Harmfulness detection and refusal activation are separable linear components at instruction vs post-instruction tokens.
- domain assumption Linear probes and principal angles on residual-stream activations are adequate diagnostics of multilingual safety geometry.
- ad hoc to paper Google Translate (and romanization/Cyrillic conversion) preserves enough semantic content that ASR differences reflect safety failure rather than meaning collapse.
- ad hoc to paper WildGuard judgments on English back-translations are a valid primary ASR oracle across all 18 languages.
- standard math Standard linear algebra / SVD / logistic regression / silhouette score definitions.
invented entities (4)
-
Minionese benchmark (4 perturbation types × 18 languages × paired harmful/harmless)
no independent evidence
-
Subthreshold activation failure
no independent evidence
-
Semantic recovery failure
no independent evidence
-
Cross-Lingual Representational Independence (CL-RepInd)
no independent evidence
read the original abstract
Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each attack type produces a distinct vulnerability profile: transliteration vulnerability is mediated by script identity, code-switching maintains effectiveness through the lowest-resource tier, and a sharp safety regime transition between Tiers 2 and 3 is consistent across all models. Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometrically misaligned subspace that projects insufficiently onto the refusal directions, leaving the refusal mechanism intact but untriggered. These findings show that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage. The benchmark and analysis code is at https://github.com/Brentkong/Minionese-Comprehensive-Benchmark-and-Mechanistic-Study-of-Multilingual-LLM-Safety.git.
Figures
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2310.06474 , year=
Multilingual Jailbreak Challenges in Large Language Models , author=. arXiv preprint arXiv:2310.06474 , year=. 2310.06474 , archivePrefix=
-
[2]
arXiv preprint arXiv:2404.01318 , year=
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models , author=. arXiv preprint arXiv:2404.01318 , year=. 2404.01318 , archivePrefix=
-
[3]
arXiv preprint arXiv:2307.15043 , year=
Universal and Transferable Adversarial Attacks on Aligned Language Models , author=. arXiv preprint arXiv:2307.15043 , year=. 2307.15043 , archivePrefix=
-
[4]
arXiv preprint arXiv:2505.17306 , year=
Refusal Direction is Universal Across Safety-Aligned Languages , author=. arXiv preprint arXiv:2505.17306 , year=. 2505.17306 , archivePrefix=
-
[5]
Upadhayay, Bibek and Behzadan, Vahid , booktitle =. Tongue-Tied: Breaking. 2025 , month = may, pages =. doi:10.18653/v1/2025.calcs-1.5 , url =
-
[6]
Towards Understanding the Fragility of Multilingual
Poppi, Samuele and Yong, Zheng Xin and He, Yifei and Chern, Bobbie and Zhao, Han and Yang, Aobo and Chi, Jianfeng , booktitle =. Towards Understanding the Fragility of Multilingual. 2025 , month = apr, pages =. doi:10.18653/v1/2025.findings-naacl.126 , url =
-
[7]
arXiv preprint arXiv:2406.11717 , year=
Refusal in Language Models Is Mediated by a Single Direction , author=. arXiv preprint arXiv:2406.11717 , year=. 2406.11717 , archivePrefix=
-
[8]
arXiv preprint arXiv:2505.23556 , year=
Understanding Refusal in Language Models with Sparse Autoencoders , author=. arXiv preprint arXiv:2505.23556 , year=. 2505.23556 , archivePrefix=
-
[9]
arXiv preprint arXiv:2507.11878 , year=
LLMs Encode Harmfulness and Refusal Separately , author=. arXiv preprint arXiv:2507.11878 , year=. 2507.11878 , archivePrefix=
-
[12]
Zhao, Weixiang and Guo, Jiahe and Hu, Yulin and Deng, Yang and Zhang, An and Sui, Xingyu and Han, Xinyang and Zhao, Yanyan and Qin, Bing and Chua, Tat-Seng and Liu, Ting , booktitle =. AdaSteer: Your Aligned. 2025 , month = nov, pages =. doi:10.18653/v1/2025.emnlp-main.1248 , url =
-
[13]
arXiv preprint arXiv:2602.22554 , year=
Multilingual Safety Alignment Via Sparse Weight Editing , author=. arXiv preprint arXiv:2602.22554 , year=. 2602.22554 , archivePrefix=
-
[14]
2025 , month = apr, howpublished =
Mind the (Language) Gap: Mapping the Challenges of LLM Development in Low-Resource Language Contexts , author =. 2025 , month = apr, howpublished =
2025
-
[15]
Icml.cc , author=
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence , url=. Icml.cc , author=. 2025 , month=
2025
-
[16]
arXiv.org , author=
Understanding Refusal in Language Models with Sparse Autoencoders , url=. arXiv.org , author=
-
[17]
arXiv.org , author=
LLMs Encode Harmfulness and Refusal Separately , url=. arXiv.org , author=
-
[18]
Zhang, Zekai and Guo, Yiduo and Lin, Jiuheng and Quan, Shanghaoran and Zhang, Huishuai and Zhao, Dongyan , year=. English as Defense Proxy: Mitigating Multilingual Jailbreak via Eliciting English Safety Knowledge , url=. doi:https://doi.org/10.18653/v1/2025.findings-emnlp.62 , journal=
-
[19]
Wu, Jiaqi and Chen, Chen and Hou, Chunyan and Yuan, Xiaojie , year=. SafeInt: Shielding Large Language Models from Jailbreak Attacks via Safety-Aware Representation Intervention , url=. doi:https://doi.org/10.18653/v1/2025.findings-emnlp.450 , journal=
-
[20]
Multilingual Safety Alignment Via Sparse Weight Editing , doi =
Liang, Jiaming and Wang, Zhaoxin and Wang, Handing , year =. Multilingual Safety Alignment Via Sparse Weight Editing , doi =
-
[21]
LSSF : Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
Zhou, Guanghao and Qiu, Panjia and Chen, Cen and Li, Hongyu and Chu, Jason and Zhang, Xin and Zhou, Jun. LSSF : Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.acl-long.1479
-
[22]
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier , doi =
Dang, John and Singh, Shivalika and D'souza, Daniel and Ahmadian, Arash and Salamanca, Alejandro and Smith, Madeline and Peppin, Aidan and Hong, Sungjin and Govindassamy, Manoj and Zhao, Terrence and Kublik, Sandra and Amer, Meor and Aryabumi, Viraat and Campos, Jon and Tan, Yi-Chern and Kocmi, Tom and Strub, Florian and Grinsztajn, Nathan and Flet-Berlia...
-
[23]
arXiv.org , author=
Qwen2.5 Technical Report , url=. arXiv.org , author=
-
[24]
arXiv.org , author=
The Llama 3 Herd of Models , url=. arXiv.org , author=
-
[25]
Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =
Han, Seungju and Rao, Kavel and Ettinger, Allyson and Jiang, Liwei and Lin, Bill Yuchen and Lambert, Nathan and Choi, Yejin and Dziri, Nouha , title =. Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =. 2024 , isbn =
2024
-
[26]
Peter J. Rousseeuw , keywords =. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis , journal =. 1987 , issn =. doi:https://doi.org/10.1016/0377-0427(87)90125-7 , url =
-
[27]
The effective rank: A measure of effective dimensionality , year=
Roy, Olivier and Vetterli, Martin , booktitle=. The effective rank: A measure of effective dimensionality , year=
-
[28]
Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=
Efficient Memory Management for Large Language Model Serving with PagedAttention , author=. Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=
-
[29]
Koishekenov, Yeskendir and Berard, Alexandre and Nikoulina, Vassilina. Memory-efficient NLLB -200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.198
-
[30]
2024 , eprint=
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal , author=. 2024 , eprint=
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.