A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.
Proceedings of the AAAI Conference on Artificial Intelligence 36, 10581–10589
1 Pith paper cite this work, alongside 22 external citations. Polarity classification is still indexing.
1
Pith paper citing it
22
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Tailored untruths: How personalisation challenges LLM safeguards
A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.