Pith. sign in

Personas as a way to model truthfulness in language models

1 Pith paper cite this work, alongside 2 external citations. Polarity classification is still indexing.

1 Pith paper citing it
2 external citations · OpenAlex

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

representative citing papers

The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs

cs.CL · 2025-10-15 · reject · novelty 6.0

The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—while the abstract claims a much larger 'alignment veto' study absent from the body

citing papers explorer

Showing 1 of 1 citing paper.

  • The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs cs.CL · 2025-10-15 · reject · none · ref 31

    The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—while the abstract claims a much larger 'alignment veto' study absent from the body