REVIEW 2 major objections 7 minor 65 references
A Multi-Turn Emotionally Engaging Dialog Model
T0 review · 2 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Emotion-tracking chatbot beats two baselines in human tests
desk verdict A plausible incremental architecture whose main empirical claim is undermined by a likely train/test overlap in the DailyDialog-based human evaluation; worth engaging as a fixable, conditional accept. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the emotion context vector $e$, produced by a unidirectional GRU that reads, in conversation order, a six-dimensional LIWC indicator vector for each history utterance (positive, negative, anxious, angry, sad, or neutral), each embedded by a dense sigmoid layer. At every decoding step, $e$ is concatenated with the decoder's language context vector before the softmax, so the emotional trajectory of the conversation directly biases word choice. This sits on an HRAN-style hierarchical attention encoder for semantics, so the model claims to track both what was said and how it felt.
What would settle it
Check whether any of the 100 test dialogs (or their first three utterances) appear in the 46,797 DailyDialog training pairs; if they do, re-run the human evaluation on a set provably unseen by MEED and see whether the emotional-appropriateness advantage over HRAN survives.
Extended reading notes
Core claim
The central claim is that explicitly modeling the emotional flow of a multi-turn conversation, as a separate recurrent encoder over utterance-level emotion indicators, produces responses that human raters find emotionally more appropriate and better tied to the context than a plain sequence-to-sequence model or a hierarchical attention model. In the paper's own experiments, MEED scored an average of 0.917 on emotional appropriateness and 0.990 on contextual coherence versus 0.895 and 0.958 for HRAN and 0.688 and 0.713 for S2S, with the MEED-over-S2S differences reported as statistically significant. The authors further show that the emotion layer's weight vectors cluster positive and negative words in t-SNE, evidence that the encoder is tracking affect rather than merely copying context semantics. The claim is that this data-driven, rule-free emotion tracking is what makes the difference.
Load-bearing premise
The load-bearing premise is that the 100 human-evaluated dialogs were held out from training, even though 50 come from the same DailyDialog corpus used for fine-tuning.
Editorial extensions
If this is right
- MEED can be deployed without any emotion label at response time: the only affective input is the LIWC tags of the history, so the model chooses the emotional stance itself from context.
- Perplexity and BLEU do not decide which model feels most human: HRAN ranks below S2S on perplexity yet above it in human judgments, so response quality needs human evaluation.
- The emotion encoder's learned weights separate positive from negative words, suggesting the same architecture could distinguish finer affect categories if given richer labels.
- Replacing the emotion classifier (LIWC) with a finer or domain-specific affect recognizer should slot into the same encoder without architectural change.
Reading between the lines
- If the 50 DailyDialog test dialogs were not held out, part of MEED's emotional-appropriateness edge could reflect memorization of DailyDialog's emotion patterns rather than generalizable emotion tracking; a held-out test would settle this.
- MEED's emotion encoder is trained end-to-end, so the LIWC indicator vectors act as a fixed-feature front end; a learned emotion classifier trained jointly might capture affect words LIWC misses, especially in domain slang.
- The architecture suggests a natural control knob for chatbot personality: scaling or shifting the emotion context vector $e$ could let the system amplify or dampen the emotional intensity of its replies, a step the paper does not explore.
- Because the model only mirrors the emotional trajectory of the context, it may reproduce the interlocutor's negativity in a 'miserable' conversation; a separate objective to de-escalate or re-engage could be added on top of $e$.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes MEED, a multi-turn open-domain dialog model that extends a hierarchical recurrent attention network with an emotion encoder. For each context utterance, LIWC2015 produces a six-dimensional binary emotion indicator; a dense layer embeds the indicator, and a GRU encodes the sequence of emotion embeddings into a single emotion context vector e, which is concatenated with the decoder hidden state before the output softmax. The model is pre-trained on the Cornell Movie Dialogs Corpus and fine-tuned on DailyDialog, then compared with a vanilla seq2seq model (S2S) and with HRAN. The paper reports lower perplexity and higher BLEU for MEED on the validation and test sets, and human evaluations on 100 four-turn dialogs (50 positive, 50 negative) in which MEED receives the highest average scores for grammatical correctness, contextual coherence, and emotional appropriateness. The authors also provide a t-SNE visualization of output-layer weights and a case study.
Significance. If the empirical claims hold, the paper makes a modest but useful contribution to affect-aware response generation: it demonstrates a simple data-driven way to condition a multi-turn dialog model on a compressed emotion trajectory rather than on a hand-selected target emotion, and it documents a human-evaluation design that balances positive and negative test dialogs. The authors release source code, which aids reproducibility. The main value is the combination of hierarchical context encoding with an emotion-tracking signal and the attempt to evaluate emotional appropriateness with a balanced test set. The strength of these claims, however, depends entirely on the integrity of the held-out test set and on the statistical comparison with HRAN, both of which need clarification.
major comments (2)
- [Section 4, 'Preparation of Natural Dialog Test Set'; Tables 2, 4, 5] The test set used for both the automatic and human evaluations is constructed from DailyDialog, but the paper never states that these dialogs were excluded from the training or validation pairs. The model is fine-tuned on 46,797 DailyDialog context-response pairs, and the training construction in Section 4 creates a context-response pair for every position in every dialog; therefore a four-turn test dialog's context (u1,u2,u3) and response u4 will appear as a training pair if that dialog is in the training split. Since the test set is built by selecting 50 positive dialogs from 78 DailyDialog dialogs and the negative pool can include the 14 negative DailyDialog dialogs, up to 64 of 100 test contexts may be DailyDialog-sourced. The paper does not state that these dialogs were held out, and the test set is not released ('due to privacy concerns, we do not plan to release this dataset'). If any of these contexts were seen during training, the perplexity/BLEU advantage in Table 2 and the human-judged coherence and emotional-appropriateness scores in Tables 4 and 5 would measure memorization rather than generalization. Please state explicitly how the test dialogs were held out from training and validation, and release the DailyDialog dialog IDs or an overlap check.
- [Section 4, 'Human Evaluation Results'; Tables 4, 5; Section 1 Contributions] The paper's contribution statement claims MEED produces 'emotionally more appropriate responses than both baselines,' but the only significance test reported for the human evaluation is MEED versus S2S (p<0.01). For MEED versus HRAN, the average scores are 0.990 versus 0.958 for contextual coherence and 0.917 versus 0.895 for emotional appropriateness; these differences are not tested, and with four raters and 100 items they may well be within noise. Either add pairwise significance tests for MEED versus HRAN, or revise the claim to say that MEED improves over S2S and is at least comparable to HRAN on these dimensions.
minor comments (7)
- [Section 4, 'Preparation of Natural Dialog Test Set'] The description of worker-created negative dialogs is ambiguous: it is unclear whether each of the two workers wrote five dialogs per topic (50 total) or five dialogs across all topics (25 total); the final selection of 50 negative dialogs implies the former, but the text should state this explicitly.
- [Abstract and Section 4, 'Automatic Evaluation Results'] The significance test for perplexity is reported only for the two validation sets; the statement that MEED 'performs the best' on the test set in Table 2 should be phrased descriptively, or accompanied by a significance test on the test set.
- [Section 4, 'Human Evaluation Results'] The reported Fleiss' kappa values (0.327-0.389) indicate only fair inter-rater agreement; this should be noted in the text when interpreting the strength of the human judgments.
- [Section 2 and Section 5] The paper repeatedly describes the approach as 'free of human-defined heuristic rules' and 'completely data-driven', but the emotion signal is obtained from LIWC, a human-curated lexicon with hand-selected categories; please soften or qualify this claim.
- [Section 4, 'Visualization of Output Layer Weights'] The description of the t-SNE analysis does not state how the '100 most frequent positive words' and '100 most frequent negative words' are identified; please specify the source of the sentiment labels.
- [References] The reference [10] (Fleiss and Cohen 1973) is not the original source of Fleiss' kappa; please cite the appropriate Fleiss (1971) article.
- [Throughout] Typos such as 'Wokrshops' in the ACM Reference Format, 'gurantee' in Section 5, 'mechansim' in Section 4, 'begain' in Section 5, and 'U/t_terance-level' in Figure 1 should be corrected.
Circularity Check
No circularity: the evaluation is an empirical benchmark and no prediction reduces to its inputs by construction.
full rationale
The paper's claims are empirical: MEED is compared with S2S and HRAN by perplexity, BLEU, and human ratings. There is no derivation chain in which an output quantity is definitionally equal to an input quantity. The emotion context vector e is computed from LIWC indicators over the source context utterances (Equations 11-12), and the decoder conditions on the concatenation of language and emotion vectors (Equations 13-16); the measured perplexity and human scores are outcomes of trained models, not renamings of training objectives. The baselines are external (sequence-to-sequence and HRAN), and no load-bearing uniqueness theorem or prior result by the same authors is invoked. The only notable risk is data hygiene rather than circularity: 50 of the 100 human-evaluation dialogs are sampled from DailyDialog, on which MEED was fine-tuned, and the paper does not explicitly state that those test dialogs were excluded from training (Section 4, Datasets and Preparation of Natural Dialog Test Set). If overlap exists, the generalization claims would be weakened, but that is a data-leakage and correctness concern, not a reduction of the claimed result to its own inputs. The paper also states that the test set is not released for privacy reasons, which hinders replication, but this is not circularity. Accordingly, no circular step is identified and the circularity score is 0.
Assumptions & free parameters
free parameters (6)
- LIWC binary utterance emotion indicator =
threshold: any word in category -> entry 1, else neutral
- Emotion context vector e =
final hidden state of emotion GRU
- Number of preceding context utterances =
up to 4 (si = max(1, i-4))
- Beam width =
256
- Hidden and embedding sizes =
256 word embedding, 256 RNN hidden, 128 utterance attention
- Neutral category entry =
sixth dimension set to 1 when no emotion word detected
assumptions (4)
- domain assumption LIWC2015's five emotion categories (positive, negative, anxious, angry, sad) adequately capture the emotionally salient content of dialog utterances.
- domain assumption Human raters can reliably judge emotional appropriateness of short responses on a 0-2 scale.
- domain assumption The test dialogs were not used in training or validation.
- standard math Bidirectional GRU and hierarchical attention mechanisms provide adequate semantic context modeling.
invented entities (1)
-
Emotion context vector e
Cite this review
Pith. "Pith review of A Multi-Turn Emotionally Engaging Dialog Model." pith.science (2026). https://pith.science/paper/QY7LEDSU
@misc{pith2026190807816,
author = {Pith},
title = {Pith review of: A Multi-Turn Emotionally Engaging Dialog Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/QY7LEDSU}},
note = {Machine review of arXiv:1908.07816}
}
read the original abstract
Open-domain dialog systems (also known as chatbots) have increasingly drawn attention in natural language processing. Some of the recent work aims at incorporating affect information into sequence-to-sequence neural dialog modeling, making the response emotionally richer, while others use hand-crafted rules to determine the desired emotion response. However, they do not explicitly learn the subtle emotional interactions captured in human dialogs. In this paper, we propose a multi-turn dialog system aimed at learning and generating emotional responses that so far only humans know how to do. Compared with two baseline models, offline experiments show that our method performs the best in perplexity scores. Further human evaluations confirm that our chatbot can keep track of the conversation context and generate emotionally more appropriate responses while performing equally well on grammar.
Figures
Reference graph
Works this paper leans on
-
[1]
Nabiha Asghar, Pascal Poupart, Jesse Hoey, Xin Jiang, and Lili Mou
-
[2]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR abs/1409.0473 (2014). arXiv:1409.0473 http://arxiv.org/abs/1409.0473
arXiv 2014
-
[3]
https://doi.org/10.1007/978-3-319-76941-7_12
154–166. https://doi.org/10.1007/978-3-319-76941-7_12
-
[4]
Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017. A Survey on Dialogue Systems: Recent Advances and New Frontiers. SIGKDD Explorations 19, 2 (2017), 25–35. https://doi.org/10.1145/ 3166054.3166058
arXiv 2017
-
[5]
Timothy W. Bickmore and Rosalind W. Picard. 2005. Establishing and Maintaining Long-Term Human-Computer Relationships. ACM Trans. Comput.-Hum. Interact. 12, 2 (2005), 293–327. https://doi.org/10.1145/ 1067860.1067867
arXiv 2005
-
[6]
Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011. Chameleons in Imagined Conversations: A New Approach to Understanding Coor- dination of Linguistic Style in Dialogs. In Proceedings of CMCL@ACL
work page 2011
-
[7]
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. A Multi-Turn Emotionally Engaging Dialog Model IUI ’20 Workshops, March 17, 2020, Cagliari, Italy Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation. In Proceedings of EMNLP 2014 . 1724–
work page 2014
-
[8]
Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, and Sune Lehmann. 2017. Using Millions of Emoji Occurrences to Learn Any- Domain Representations for Detecting Sentiment, Emotion and Sar- casm. In Proceedings of EMNLP 2017 . 1615–1625. https://aclanthology. info/papers/D17-1169/d17-1169
work page 2017
Show all 65 references
-
[9]
Robert H Finn. 1970. A Note on Estimating the Reliability of Categorical Data. Educational and Psychological Measurement 30, 1 (1970), 71–76
1970
-
[10]
Joseph L Fleiss and Jacob Cohen. 1973. The Equivalence of Weighted kappa and the Intraclass Correlation Coefficient as Measures of Relia- bility. Educational and psychological measurement 33, 3 (1973), 613– 619
1973
-
[11]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova
-
[12]
2016.Fundamental Statistics for the Behavioral Sciences
David C Howell. 2016.Fundamental Statistics for the Behavioral Sciences. Nelson Education
2016
-
[13]
George Hripcsak and Daniel F. Heitjan. 2002. Measuring Agreement in Medical Informatics Reliability Studies. Journal of Biomedical Informat- ics 35, 2 (2002), 99–110. https://doi.org/10.1016/S1532-0464(02)00500-2
2002 doi
-
[14]
Tianran Hu, Anbang Xu, Zhe Liu, Quanzeng You, Yufan Guo, Vibha Sinha, Jiebo Luo, and Rama Akkiraju. 2018. Touch Your Heart: A Tone- aware Chatbot for Customer Care on Social Media. In Proceedings of CHI 2018. 415. https://doi.org/10.1145/3173574.3173989
2018
-
[15]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Sto- chastic Optimization. CoRR abs/1412.6980 (2014). arXiv:1412.6980 http://arxiv.org/abs/1412.6980
2014 arXiv
-
[16]
Jonathan Klein, Youngme Moon, and Rosalind W. Picard. 2001. This Computer Responds to User Frustration: Theory, Design, and Results. Interacting with Computers 14, 2 (2001), 119–140. https://doi.org/10. 1016/S0953-5438(01)00053-4
2001
-
[17]
Sayan Ghosh, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, and Stefan Scherer. 2017. Affect-LM: A Neural Language Model for Customizable Affective Text Generation. In Proceedings of ACL 2017. 634–642. https://doi.org/10.18653/v1/P17-1059
2017 doi
-
[18]
Spithourakis, Jian- feng Gao, and William B
Jiwei Li, Michel Galley, Chris Brockett, Georgios P. Spithourakis, Jian- feng Gao, and William B. Dolan. 2016. A Persona-Based Neural Conver- sation Model. InProceedings of ACL 2016. http://aclweb.org/anthology/ P/P16/P16-1094.pdf
2016
-
[19]
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016. Deep Reinforcement Learning for Dialogue Gen- eration. In Proceedings of EMNLP 2016 . 1192–1202. http://aclweb.org/ anthology/D/D16/D16-1127.pdf
2016
-
[20]
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu
-
[21]
Pierre Lison and Jörg Tiedemann. 2016. OpenSubtitles2016: Extract- ing Large Parallel Corpora from Movie and TV Subtitles. In Proceed- ings of LREC 2016 . http://www.lrec-conf.org/proceedings/lrec2016/ summaries/947.html
2016
-
[22]
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. In Proceedings of EMNLP 2016 . 2122–
2016
-
[23]
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan
-
[24]
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard- Gontier, Yoshua Bengio, and Joelle Pineau. 2017. Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses. In Proceedings ACL 2017. 1116–1126. https://doi.org/10.18653/v1/P17-1103
2017 doi
-
[25]
Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, Nov (2008), 2579–2605
2008
-
[26]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Ef- ficient Estimation of Word Representations in Vector Space. CoRR abs/1301.3781 (2013). arXiv:1301.3781 http://arxiv.org/abs/1301.3781
2013 arXiv
-
[27]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of ACL 2002 . 311–318. http://www.aclweb.org/anthology/ P02-1040.pdf
2002
-
[28]
James W Pennebaker, Martha E Francis, and Roger J Booth. 2001. Linguistic Inquiry and Word Count: LIWC 2001. Mahway: Lawrence Erlbaum Associates 71, 2001 (2001), 2001
2001
-
[29]
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau
-
[30]
Alan Ritter, Colin Cherry, and William B. Dolan. 2011. Data-Driven Response Generation in Social Media. In Proceedings of EMNLP 2011 . 583–593. http://www.aclweb.org/anthology/D11-1054
2011
-
[31]
Klaus R Scherer. 2005. What Are Emotions? And How Can They Be Measured? Social science information 44, 4 (2005), 695–729
2005
-
[32]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov
-
[33]
CoRR abs/1907.11692 (2019)
RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/ 1907.11692
2019 arXiv
-
[34]
Lifeng Shang, Zhengdong Lu, and Hang Li. 2015. Neural Responding Machine for Short-Text Conversation. In Proceedings of ACL-IJCNLP
2015
-
[35]
Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi, Christina Li- oma, Jakob Grue Simonsen, and Jian-Yun Nie. 2015. A Hierarchical Recurrent Encoder-Decoder for Generative Context-Aware Query Sug- gestion. In Proceedings of CIKM 2015. 553–562. https://doi.org/10.1145/ 2806416.2806493
2015
-
[36]
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan. 2015. A Neural Network Approach to Context-Sensitive Gen- eration of Conversational Responses. In Proceedings of NAACL-HLT
2015
-
[37]
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to Sequence Learning with Neural Networks. In Proceedings of NIPS 2014 . 3104–3112. http://papers.nips.cc/paper/5346-sequence-to-sequence- IUI ’20 Workshops, March 17, 2020, Cagliari, Italy Xie, et al. learning-with...
2014
-
[38]
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018. RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems. In Proceedings of AAAI 2018 . 722–729. https://www. aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16179
2018
-
[39]
Christoph Tillmann and Hermann Ney. 2003. Word Reordering and a Dynamic Programming Beam Search Algorithm for Statistical Machine Translation. Computational Linguistics 29, 1 (2003), 97–133. https: //doi.org/10.1162/089120103321337458
2003 doi
-
[40]
In Proceedings of ACL 2019
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset. In Proceedings of ACL 2019. 5370–5381. https://www.aclweb.org/anthology/P19-1534/
2019
-
[41]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Proceedings of NIPS 2017 . 5998–6008. http://papers.nips.cc/paper/7181-attention-is-all-you-need
2017
-
[42]
Oriol Vinyals and Quoc V. Le. 2015. A Neural Conversational Model. CoRR abs/1506.05869 (2015). arXiv:1506.05869 http://arxiv.org/abs/ 1506.05869
2015 arXiv
-
[43]
Courville, and Joelle Pineau
Iulian Vlad Serban, Alessandro Sordoni, Yoshua Bengio, Aaron C. Courville, and Joelle Pineau. 2016. Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models. In Proceedings of AAAI 2016 . 3776–3784. http://www.aaai.org/ocs/index. php/AAAI/AAAI16...
2016
-
[44]
Courville, and Yoshua Bengio
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Char- lin, Joelle Pineau, Aaron C. Courville, and Yoshua Bengio. 2017. A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues. In Proceedings of AAAI 2017 . 3295–3301. http://aaai.org/ ocs/index....
2017
-
[45]
Chen Xing, Yu Wu, Wei Wu, Yalou Huang, and Ming Zhou. 2018. Hierarchical Recurrent Attention Network for Response Generation. In Proceedings of AAAI 2018 . 5610–5617. https://www.aaai.org/ocs/ index.php/AAAI/AAAI18/paper/view/16510
2018
-
[46]
Anbang Xu, Zhe Liu, Yufan Guo, Vibha Sinha, and Rama Akkiraju. 2017. A New Chatbot for Customer Service on Social Media. In Proceedings of CHI 2017. 3506–3510. https://doi.org/10.1145/3025453.3025496
2017
-
[47]
Jennifer Zamora. 2017. I’m Sorry, Dave, I’m Afraid I Can’t Do That: Chatbot Perception and Expectations. In Proceedings of HAI 2017 . 253–
2017
-
[48]
Peixiang Zhong, Di Wang, and Chunyan Miao. 2019. An Affect-Rich Neural Conversational Model with Biased Attention and Weighted Cross-Entropy Loss. In Proceedings of AAAI 2019 . 7492–7500. https: //aaai.org/ojs/index.php/AAAI/article/view/4740
2019
-
[49]
http://aclweb.org/anthology/N/N15/N15-1020.pdf
196–205. http://aclweb.org/anthology/N/N15/N15-1020.pdf
-
[50]
Xianda Zhou and William Yang Wang. 2018. MojiTalk: Generating Emotional Responses at Scale. In Proceedings of ACL 2018 . 1128–1137. https://doi.org/10.18653/v1/P18-1104
2018 doi
-
[53]
Howard E Tinsley and David J Weiss. 1975. Interrater Reliability and Agreement of Subjective Judgments. Journal of Counseling Psychology 22, 4 (1975), 358
1975
-
[56]
Amy Beth Warriner, Victor Kuperman, and Marc Brysbaert. 2013. Norms of Valence, Arousal, and Dominance for 13,915 English Lemmas. Behavior research methods 45, 4 (2013), 1191–1207
2013
-
[57]
Chen Xing, Wei Wu, Yu Wu, Jie Liu, Yalou Huang, Ming Zhou, and Wei-Ying Ma. 2017. Topic Aware Neural Response Generation. In Proceedings of AAAI 2017 . 3351–3357. http://aaai.org/ocs/index.php/ AAAI/AAAI17/paper/view/14563
2017
-
[63]
Hao Zhou, Minlie Huang, Tianyang Zhang, Xiaoyan Zhu, and Bing Liu. 2018. Emotional Chatting Machine: Emotional Conversation Generation with Internal and External Memory. InProceedings of AAAI
2018
-
[64]
https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/ 16455
-
[260]
https://doi.org/10.1145/3125739.3125766
-
[1734]
http://aclweb.org/anthology/D/D14/D14-1179.pdf
-
[2011]
https://aclanthology.info/papers/W11-0609/w11-0609
76–87. https://aclanthology.info/papers/W11-0609/w11-0609
-
[2015]
http://aclweb.org/anthology/P/P15/P15-1152.pdf
1577–1586. http://aclweb.org/anthology/P/P15/P15-1152.pdf
-
[2016]
In Proceedings of NAACL-HLT 2016
A Diversity-Promoting Objective Function for Neural Conver- sation Models. In Proceedings of NAACL-HLT 2016 . 110–119. http: //aclweb.org/anthology/N/N16/N16-1014.pdf
2016
-
[2017]
In Proceedings of IJCNLP 2017
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In Proceedings of IJCNLP 2017
2017
-
[2018]
In Proceedings of ECIR
Affective Neural Response Generation. In Proceedings of ECIR
-
[2019]
In Proceedings of NAACL-HLT 2019
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT 2019. 4171–
2019
-
[2132]
http://aclweb.org/anthology/D/D16/D16-1230.pdf
-
[4186]
https://aclweb.org/anthology/papers/N/N19/N19-1423/
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.