Pith. sign in

hub

Multilingual jailbreak challenges in large language models

25 Pith papers cite this work, alongside 16 external citations. Polarity classification is still indexing.

25 Pith papers citing it
16 external citations · external index

hub tools

citation-role summary

background 3

citation-polarity summary

roles

background 3

polarities

background 2 support 1

representative citing papers

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

cs.LG · 2026-04-10 · unverdicted · novelty 6.0

DACO curates a 15,000-concept dictionary from 400K image-caption pairs and uses it to initialize an SAE that enables granular, concept-specific steering of MLLM activations, raising safety scores on MM-SafetyBench and JailBreakV while preserving general capabilities.

Low-Resource Languages Jailbreak GPT-4

cs.CL · 2023-10-03 · conditional · novelty 6.0

Translating unsafe inputs to low-resource languages jailbreaks GPT-4 at rates on par with or exceeding state-of-the-art attacks.

Multilingual jailbreaking of LLMs using low-resource languages

cs.CL · 2026-05-18 · unverdicted · novelty 5.0

Multi-turn prompts in Afrikaans, Kiswahili, isiXhosa and isiZulu achieve 52-83% harmful response rates across GPT, Claude, Gemini and others, rising further with native-speaker red-teaming, showing translation quality limits jailbreak success.

citing papers explorer

Showing 25 of 25 citing papers.