RL-trained persuader agents flip LLMs away from correct answers in a single message, reaching 93.7% success on the training-time target and transferring to unseen open-weight and frontier models.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
RL-trained persuader agents flip LLMs away from correct answers in a single message, reaching 93.7% success on the training-time target and transferring to unseen open-weight and frontier models.