Pith. sign in

REVIEW 4 major objections 28 references

Low-resource multilingual jailbreaks leave the refusal mechanism intact but untriggered by routing harm through a misaligned residual subspace.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:12 UTC pith:3OVPV6DC

load-bearing objection Solid multi-perturbation multilingual safety benchmark with a coherent geometric story; the T2–T3 cliff and Tier-4 ASRs rest on a soft judge pipeline, but the attack-type profiles and subspace misalignment findings still hold up. the 4 major comments →

arxiv 2607.10112 v1 pith:3OVPV6DC submitted 2026-07-11 cs.CR cs.AI

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

classification cs.CR cs.AI
keywords multilingual jailbreakLLM safetyrefusal directionresidual stream geometrycode-switchingtransliterationlow-resource languagesMinionese
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that LLM safety that works in English fails in structured, measurable ways once language, script, and surface form change. It introduces Minionese: 18 languages in four resource tiers, paired harmful and harmless prompts, and four meaning-preserving attack styles—full translation, code-switching, transliteration, and translationese. Across three instruction-tuned models, each attack type has its own vulnerability profile: transliteration risk tracks script (not resource alone), code-switching stays effective into the lowest tier, and models flip from mostly refusing to mostly complying between mid-high and mid-low resource tiers. The mechanistic claim is geometric. Harmful content in lower-resource settings often still lives in a detectable harm representation, but that representation sits in a subspace poorly aligned with English and projects too weakly onto the refusal direction, so refusal never fires. English-only safety tests therefore miss risks that depend on script family, perturbation type, and per-language alignment coverage.

Core claim

Low-resource multilingual jailbreaks succeed mainly by routing harmful content through a geometrically misaligned residual subspace that projects insufficiently onto the refusal direction, leaving the refusal mechanism intact but untriggered (subthreshold activation). Each of four attack types yields a distinct, cross-model vulnerability profile, with a sharp refuse-to-comply regime transition between resource Tiers 2 and 3.

What carries the argument

A two-stage geometric account of residual-stream safety: harmfulness subspace (from linear probes) versus refusal direction (from English refuse–comply contrast), with failure typed as upstream semantic-recovery collapse or downstream subthreshold activation when the harm signal fails to clear the refusal threshold. Principal angles between non-English and English harm subspaces quantify the misalignment.

Load-bearing premise

Non-English attack success is assumed to be reliably measured by back-translating model answers into English and scoring them with an English safety judge, even when that back-translation is weak for the lowest-resource languages.

What would settle it

Have native speakers judge the original non-English model outputs (no back-translation) for the same three models and four attack types; if the Tier 2–3 refuse-to-comply cliff or the script-conditioned transliteration pattern disappears or reverses, the reported profiles and geometric story fail.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 0 minor

Summary. The paper introduces Minionese, a multilingual jailbreak benchmark of 55,440 prompts spanning 18 languages, four resource tiers, and four meaning-preserving perturbations (standard translation, code-switching, transliteration, translationese), built from AdvBench harmful–harmless pairs. Evaluating Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Aya-Expanse-8B, it reports distinct, cross-model vulnerability profiles: transliteration ASR is script-conditioned, code-switching remains effective through Tier 4, and a sharp refuse-to-comply transition appears between Tiers 2 and 3. A geometric analysis of residual-stream activations then argues that low-resource jailbreaks primarily succeed via subthreshold activation—harmful content is often linearly decodable but projects insufficiently onto an intact refusal direction—while Tier-4 and non-Latin transliteration can also collapse the harm representation itself. A multi-dimensional refusal-cone analysis further claims that only the first cone basis direction transfers cross-lingually.

Significance. If the behavioral and geometric claims hold, the work is a substantial contribution to multilingual LLM safety: it upgrades translation-only benchmarks by treating perturbation type as an independent variable, supplies paired harmful–harmless prompts for mechanistic analysis, and offers a falsifiable two-stage account (semantic recovery failure vs. subthreshold activation) that is consistent across three architectures. The released benchmark and analysis code, the script-conditioned transliteration pattern, and the effective one-dimensionality of cross-lingual refusal are concrete, reusable results for evaluation design and geometry-aware interventions. These strengths matter for governance and deployment beyond English-centric red-teaming.

major comments (4)
  1. §4.1.3 and Limitations F.1: Non-English ASR is defined by back-translation to English then WildGuard scoring. The only human validation is a spot-check claiming 100% agreement on an unspecified sample, with no per-tier, per-perturbation, or per-model agreement rates and no second judge. Because the sharp Tier 2–3 regime transition and the tier gradients are load-bearing for the central claim, the paper needs either (i) stratified human labels (or dual-judge agreement) on a non-trivial Tier-3/4 sample, or (ii) a sensitivity analysis showing that the reported cliffs survive plausible mislabel rates. Without this, the behavioral foundation of the mechanistic narrative remains under-anchored.
  2. §4.2 / Eq. (1): ASR is reported as point estimates with no confidence intervals, bootstrap errors, or significance tests across languages, tiers, or models. The claimed “sharp” Tier 2–3 transition and cross-model rank-order stability of attack types cannot be assessed for sampling variability. Adding per-cell uncertainty (and, where appropriate, simple tests of tier contrasts) is necessary for the regime-transition claim to be scientifically usable.
  3. §5.1–5.2: Failure-type classification (upstream / mixed / downstream) and “subthreshold activation” rest on free analysis choices—τ^(ℓ)=0.95, harmfulness subspace rank k (≥95% spectral mass), and the English-only mean-difference refusal direction—without reported sensitivity. The dominant “mixed” classification and the claim that refusal is never suppressed once adequately triggered should be shown to be robust under reasonable variation of τ and k, or the threshold should be calibrated (e.g., from English ROC) rather than fixed at 0.95.
  4. §5 and F.1: Causal language (“routing,” “starving,” “leaving the refusal mechanism intact but untriggered”) is stronger than the evidence. The geometric results are correlational; the manuscript itself defers activation/path patching to future work. Either run a minimal causal check (e.g., projecting Tier-3 harmful activations onto ˆr and measuring refusal recovery) or systematically rephrase claims as geometric associations consistent with, but not yet proving, the two-stage mechanism.

Circularity Check

0 steps flagged

No significant circularity: empirical ASR measurements and observational residual-stream geometry are independent of each other; failure labels restate measured projections, not redefine the behavioral outcomes.

full rationale

Minionese is an empirical benchmark plus post-hoc geometric analysis, not a first-principles derivation. Attack success rates are measured externally via WildGuard (with back-translation) on held-out harmful prompts; resource tiers and the four perturbation types are design choices that partition the data, not quantities defined from the ASR they are used to report. The geometric pipeline (logistic probes for harmfulness subspaces, mean-difference English refusal direction, silhouette/AUC/principal-angle statistics, and the optional multi-dimensional cone fit with free weights λ_cone, λ_xl) observes residual-stream structure and correlates it with those independently measured ASRs. Calling a language “subthreshold” when harm is linearly decodable but ˆr⊤h fails a fixed τ=0.95 is a labeling of the same geometric measurements, not a prediction forced by fitting ASR. There is no self-citation load-bearing chain (Arditi, Zhao, Wollschläger, Wang et al. are external), no uniqueness theorem imported from the authors, and no ansatz smuggled in as a theorem. Limitations correctly flag that causal attribution remains correlational. Measurement-validity concerns about WildGuard-on-back-translation affect correctness risk, not circularity of the derivation chain.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 4 invented entities

The work rests on standard linear-representation assumptions from prior refusal literature, an operational resource-tier taxonomy, and several analysis hyperparameters (subspace rank, refusal threshold, cone loss weights). Invented constructs are diagnostic labels and metrics (subthreshold activation, CL-RepInd, Minionese attack taxonomy), not new physical entities. Free parameters affect mechanistic figures more than raw ASR tables.

free parameters (5)
  • harmfulness subspace rank k = ≥95% spectral mass
    Chosen so top-k right singular vectors of stacked probe weights explain ≥95% spectral mass; controls principal-angle and silhouette analyses.
  • refusal decision threshold τ^(ℓ) = 0.95
    Subthreshold activation is defined as harmful activations failing to exceed τ=0.95 despite nontrivial harm signal; directly labels the dominant failure mode.
  • cone optimization hyperparameters = η=0.05, λcone=5, λxl=1, N=5, 400 steps
    Stiefel-manifold ascent for 400 steps with η=0.05, λcone=5, λxl=1, canonical N=5; shapes which secondary directions are declared non-transferring.
  • coherence filter cutoffs = conf<0.3, unicode<0.8, len<20
    langdetect confidence <0.3, Unicode validity <0.8, length <20 chars exclude responses from ASR; can bias per-language rates if generation quality correlates with tier.
  • resource tier language assignment = 4 tiers / 18 languages as listed in §3
    18 languages hand-grouped into 4 tiers by ‘training-data availability in contemporary LLMs’; the Tier 2–3 regime transition depends on this grouping.
axioms (6)
  • domain assumption Refusal in aligned LLMs is mediated by a low-dimensional direction (or cone) in residual activation space that can be extracted by behavioral contrast.
    Imported from Arditi et al. (2024) and Wollschläger et al. (2025); used throughout §5 to define r̂ and the cone basis.
  • domain assumption Harmfulness detection and refusal activation are separable linear components at instruction vs post-instruction tokens.
    From Zhao et al. (2025); underpins upstream vs downstream vs mixed failure classification.
  • domain assumption Linear probes and principal angles on residual-stream activations are adequate diagnostics of multilingual safety geometry.
    Assumed in §5; authors note nonlinear mechanisms may be missed (Limitations F.1).
  • ad hoc to paper Google Translate (and romanization/Cyrillic conversion) preserves enough semantic content that ASR differences reflect safety failure rather than meaning collapse.
    All four attack types are built via Google Translate API (§3); quality variation is acknowledged but not measured.
  • ad hoc to paper WildGuard judgments on English back-translations are a valid primary ASR oracle across all 18 languages.
    §4.1.3; spot-check claims 100% human agreement on a random sample, but sample size and tier coverage are not detailed.
  • standard math Standard linear algebra / SVD / logistic regression / silhouette score definitions.
    Used for subspaces, probes, and separation metrics without modification.
invented entities (4)
  • Minionese benchmark (4 perturbation types × 18 languages × paired harmful/harmless) no independent evidence
    purpose: Enable attack-type-aware multilingual jailbreak evaluation and mechanistic analysis beyond translation-only corpora.
    New dataset construction; independent evidence is the public release claim and construction recipe, not external prior measurement.
  • Subthreshold activation failure no independent evidence
    purpose: Name the case where harm is linearly decodable but projects below the refusal threshold.
    Diagnostic label defined from probe AUC + refusal projection; falsifiable via activation patching (not yet done).
  • Semantic recovery failure no independent evidence
    purpose: Name upstream collapse of the harm representation (esp. transliteration / Tier-4).
    Defined from near-chance probe AUC and collapsed cone margins; correlational in this paper.
  • Cross-Lingual Representational Independence (CL-RepInd) no independent evidence
    purpose: Measure whether refusal cone basis directions exploit shared vs distinct cross-lingual circuits via margin-vector correlation.
    Paper-introduced metric; no external validation yet.

pith-pipeline@v1.1.0-grok45 · 20946 in / 4346 out tokens · 40275 ms · 2026-07-14T14:12:41.751794+00:00 · methodology

0 comments
read the original abstract

Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each attack type produces a distinct vulnerability profile: transliteration vulnerability is mediated by script identity, code-switching maintains effectiveness through the lowest-resource tier, and a sharp safety regime transition between Tiers 2 and 3 is consistent across all models. Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometrically misaligned subspace that projects insufficiently onto the refusal directions, leaving the refusal mechanism intact but untriggered. These findings show that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage. The benchmark and analysis code is at https://github.com/Brentkong/Minionese-Comprehensive-Benchmark-and-Mechanistic-Study-of-Multilingual-LLM-Safety.git.

Figures

Figures reproduced from arXiv: 2607.10112 by Ayushi Mehrotra, Brent Kong, Chigozirim Ifebi.

Figure 1
Figure 1. Figure 1: ASR by language and perturbation type, Llama￾3.1-8B-Instruct. Heatmaps for Aya and Qwen are in Ap￾pendix B. at Swahili, which reaches 95%. Translationese con￾sistently exceeds standard translation ASR at mid￾resource languages, with the gap largest for Korean (14% vs. 51% on Aya) and Hindi (15% vs. 48% on Llama). This is consistent with the hypothesis that round-trip translation shifts phrasing toward natu… view at source ↗
Figure 2
Figure 2. Figure 2: Harmful-harmless separation in activation space (silhouette score) across languages and layers, all three models using standard_translation attack type. 5.1. Preliminaries Let M be a transformer with L layers and hidden dimension d. We extract residual stream activations h (ℓ) (x) ∈ R d at the last post-instruction token position t ∗ for each input x. For language λ and layer ℓ, we de￾fine harmful and harm… view at source ↗
Figure 3
Figure 3. Figure 3: Linear probe AUC for harmfulness detection across all layers using standard_translation attack type (category all; chance = 0.5). maintain high and stable separation across the full layer range on all three models, indicating tightly clus￾tered, well-separated harmful and harmless represen￾tations. Separation declines from Tier 3 onward, with the pattern consistent across all three architectures. The most … view at source ↗
Figure 5
Figure 5. Figure 5: Cone quality vs. dimensionality and canonical N = 5 basis margins (Aya and Qwen). Llama in Ap￾pendix D. 6. Discussion Our results establish that multilingual jailbreak vul￾nerability is a structured set of distinct failure modes whose character depends on the type of linguistic ma￾nipulation applied. Safety failures arise either be￾cause harmful content is not represented in a geo￾metrically separable form… view at source ↗
Figure 7
Figure 7. Figure 7: ASR by language and perturbation type, Qwen2.5-7B-Instruct. Std. Translation Translationese Code Switching Transliteration Perturbation Type German English Spanish French Chinese Arabic Japanese Korean Russian Hindi Indonesian Swahili Turkish Scottish Gaelic Guaraní Javanese Yoruba Zulu Language 11% 11% 15% 28% 2% 2% 2% 0% 8% 8% 17% 11% 8% 9% 13% 0% 11% 24% 19% 86% 18% 42% 29% 84% 16% 24% 38% 78% 14% 51% 4… view at source ↗
Figure 8
Figure 8. Figure 8: ASR by language and perturbation type, Aya-Expanse-8b. 12 [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Disentanglement of harmfulness and refusal representations across all three models using standard_translation attack type. Rows from top: (1) harm vs. refusal component norms per language; (2) contrastive harm signal at tinst vs. refusal signal at tpost; (3) layer-wise failure type distribution per language; (4) tier-averaged harm and refusal signal trajectories across normalized layer depth. 13 [PITH_FUL… view at source ↗
Figure 10
Figure 10. Figure 10: Cone quality vs. dimensionality and canonical N = 5 basis margins, Llama-3.1-8B-Instruct [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Per-language cone margins, N = 5 basis, Llama-3.1-8B-Instruct. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

28 extracted references · 2 canonical work pages

  1. [1]

    arXiv preprint arXiv:2310.06474 , year=

    Multilingual Jailbreak Challenges in Large Language Models , author=. arXiv preprint arXiv:2310.06474 , year=. 2310.06474 , archivePrefix=

  2. [2]

    arXiv preprint arXiv:2404.01318 , year=

    JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models , author=. arXiv preprint arXiv:2404.01318 , year=. 2404.01318 , archivePrefix=

  3. [3]

    arXiv preprint arXiv:2307.15043 , year=

    Universal and Transferable Adversarial Attacks on Aligned Language Models , author=. arXiv preprint arXiv:2307.15043 , year=. 2307.15043 , archivePrefix=

  4. [4]

    arXiv preprint arXiv:2505.17306 , year=

    Refusal Direction is Universal Across Safety-Aligned Languages , author=. arXiv preprint arXiv:2505.17306 , year=. 2505.17306 , archivePrefix=

  5. [5]

    Tongue-Tied: Breaking

    Upadhayay, Bibek and Behzadan, Vahid , booktitle =. Tongue-Tied: Breaking. 2025 , month = may, pages =. doi:10.18653/v1/2025.calcs-1.5 , url =

  6. [6]

    Towards Understanding the Fragility of Multilingual

    Poppi, Samuele and Yong, Zheng Xin and He, Yifei and Chern, Bobbie and Zhao, Han and Yang, Aobo and Chi, Jianfeng , booktitle =. Towards Understanding the Fragility of Multilingual. 2025 , month = apr, pages =. doi:10.18653/v1/2025.findings-naacl.126 , url =

  7. [7]

    arXiv preprint arXiv:2406.11717 , year=

    Refusal in Language Models Is Mediated by a Single Direction , author=. arXiv preprint arXiv:2406.11717 , year=. 2406.11717 , archivePrefix=

  8. [8]

    arXiv preprint arXiv:2505.23556 , year=

    Understanding Refusal in Language Models with Sparse Autoencoders , author=. arXiv preprint arXiv:2505.23556 , year=. 2505.23556 , archivePrefix=

  9. [9]

    arXiv preprint arXiv:2507.11878 , year=

    LLMs Encode Harmfulness and Refusal Separately , author=. arXiv preprint arXiv:2507.11878 , year=. 2507.11878 , archivePrefix=

  10. [12]

    AdaSteer: Your Aligned

    Zhao, Weixiang and Guo, Jiahe and Hu, Yulin and Deng, Yang and Zhang, An and Sui, Xingyu and Han, Xinyang and Zhao, Yanyan and Qin, Bing and Chua, Tat-Seng and Liu, Ting , booktitle =. AdaSteer: Your Aligned. 2025 , month = nov, pages =. doi:10.18653/v1/2025.emnlp-main.1248 , url =

  11. [13]

    arXiv preprint arXiv:2602.22554 , year=

    Multilingual Safety Alignment Via Sparse Weight Editing , author=. arXiv preprint arXiv:2602.22554 , year=. 2602.22554 , archivePrefix=

  12. [14]

    2025 , month = apr, howpublished =

    Mind the (Language) Gap: Mapping the Challenges of LLM Development in Low-Resource Language Contexts , author =. 2025 , month = apr, howpublished =

  13. [15]

    Icml.cc , author=

    The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence , url=. Icml.cc , author=. 2025 , month=

  14. [16]

    arXiv.org , author=

    Understanding Refusal in Language Models with Sparse Autoencoders , url=. arXiv.org , author=

  15. [17]

    arXiv.org , author=

    LLMs Encode Harmfulness and Refusal Separately , url=. arXiv.org , author=

  16. [18]

    English as Defense Proxy: Mitigating Multilingual Jailbreak via Eliciting English Safety Knowledge , url=

    Zhang, Zekai and Guo, Yiduo and Lin, Jiuheng and Quan, Shanghaoran and Zhang, Huishuai and Zhao, Dongyan , year=. English as Defense Proxy: Mitigating Multilingual Jailbreak via Eliciting English Safety Knowledge , url=. doi:https://doi.org/10.18653/v1/2025.findings-emnlp.62 , journal=

  17. [19]

    SafeInt: Shielding Large Language Models from Jailbreak Attacks via Safety-Aware Representation Intervention , url=

    Wu, Jiaqi and Chen, Chen and Hou, Chunyan and Yuan, Xiaojie , year=. SafeInt: Shielding Large Language Models from Jailbreak Attacks via Safety-Aware Representation Intervention , url=. doi:https://doi.org/10.18653/v1/2025.findings-emnlp.450 , journal=

  18. [20]

    Multilingual Safety Alignment Via Sparse Weight Editing , doi =

    Liang, Jiaming and Wang, Zhaoxin and Wang, Handing , year =. Multilingual Safety Alignment Via Sparse Weight Editing , doi =

  19. [21]

    LSSF : Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion

    Zhou, Guanghao and Qiu, Panjia and Chen, Cen and Li, Hongyu and Chu, Jason and Zhang, Xin and Zhou, Jun. LSSF : Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.acl-long.1479

  20. [22]

    Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier , doi =

    Dang, John and Singh, Shivalika and D'souza, Daniel and Ahmadian, Arash and Salamanca, Alejandro and Smith, Madeline and Peppin, Aidan and Hong, Sungjin and Govindassamy, Manoj and Zhao, Terrence and Kublik, Sandra and Amer, Meor and Aryabumi, Viraat and Campos, Jon and Tan, Yi-Chern and Kocmi, Tom and Strub, Florian and Grinsztajn, Nathan and Flet-Berlia...

  21. [23]

    arXiv.org , author=

    Qwen2.5 Technical Report , url=. arXiv.org , author=

  22. [24]

    arXiv.org , author=

    The Llama 3 Herd of Models , url=. arXiv.org , author=

  23. [25]

    Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =

    Han, Seungju and Rao, Kavel and Ettinger, Allyson and Jiang, Liwei and Lin, Bill Yuchen and Lambert, Nathan and Choi, Yejin and Dziri, Nouha , title =. Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =. 2024 , isbn =

  24. [26]

    Rousseeuw , keywords =

    Peter J. Rousseeuw , keywords =. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis , journal =. 1987 , issn =. doi:https://doi.org/10.1016/0377-0427(87)90125-7 , url =

  25. [27]

    The effective rank: A measure of effective dimensionality , year=

    Roy, Olivier and Vetterli, Martin , booktitle=. The effective rank: A measure of effective dimensionality , year=

  26. [28]

    Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

    Efficient Memory Management for Large Language Model Serving with PagedAttention , author=. Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

  27. [29]

    Memory-efficient NLLB -200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model

    Koishekenov, Yeskendir and Berard, Alexandre and Nikoulina, Vassilina. Memory-efficient NLLB -200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.198

  28. [30]

    2024 , eprint=

    HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal , author=. 2024 , eprint=