A DQN chatbot that selects among 100 clustered reply types and is rewarded for picking true human responses learns on training dialogues but generalizes poorly to unseen dialogues.
A study on dialogue reward prediction for open-ended conversational agents,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2019 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Deep Reinforcement Learning for Chatbots Using Clustered Actions and Human-Likeness Rewards
A DQN chatbot that selects among 100 clustered reply types and is rewarded for picking true human responses learns on training dialogues but generalizes poorly to unseen dialogues.