LDPP automatically discovers latent dialogue policies from raw records and uses offline hierarchical reinforcement learning to plan in that latent space, outperforming strong baselines on proactive dialogue benchmarks.
Controllable Mixed-Initiative Dialogue Generation through Prompting
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Mixed-initiative dialogue tasks involve repeated exchanges of information and conversational control. Conversational agents gain control by generating responses that follow particular dialogue intents or strategies, prescribed by a policy planner. The standard approach has been fine-tuning pre-trained language models to perform generation conditioned on these intents. However, these supervised generation models are limited by the cost and quality of data annotation. We instead prompt large language models as a drop-in replacement to fine-tuning on conditional generation. We formalize prompt construction for controllable mixed-initiative dialogue. Our findings show improvements over fine-tuning and ground truth responses according to human evaluation and automatic metrics for two tasks: PersuasionForGood and Emotional Support Conversations.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues
LDPP automatically discovers latent dialogue policies from raw records and uses offline hierarchical reinforcement learning to plan in that latent space, outperforming strong baselines on proactive dialogue benchmarks.