The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—while the abstract claims a much larger 'alignment veto' study absent from the body
Personas as a way to model truthfulness in language models
1 Pith paper cite this work, alongside 2 external citations. Polarity classification is still indexing.
1
Pith paper citing it
2
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—while the abstract claims a much larger 'alignment veto' study absent from the body