Preference Tree Optimization uses look-ahead simulations scored by an AI oracle to generate DPO preference data, and the resulting Motivational Interviewing agent scores higher on that same oracle than the base model.
Broaden your scope! efficient multi-turn conversation planning for llms with semantic space
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
Preference Tree Optimization uses look-ahead simulations scored by an AI oracle to generate DPO preference data, and the resulting Motivational Interviewing agent scores higher on that same oracle than the base model.