Pith. sign in

REVIEW 2 major objections 7 minor 65 references

A Multi-Turn Emotionally Engaging Dialog Model

T0 review · 2 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Emotion-tracking chatbot beats two baselines in human tests

desk verdict A plausible incremental architecture whose main empirical claim is undermined by a likely train/test overlap in the DailyDialog-based human evaluation; worth engaging as a fixable, conditional accept. read the letter →

arxiv 1908.07816 v3 pith:QY7LEDSU submitted 2019-08-15 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords multi-turndialogemotionrecognitionencoderhierarchicalattentionLIWCchatbotaffectivecomputinghumanevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes MEED, a multi-turn chatbot that learns emotional exchanges directly from human dialogs rather than relying on hand-crafted rules for choosing an emotion. The model adds an emotion encoder to a hierarchical attention network: each context utterance is tagged with coarse emotion indicators from the LIWC lexicon, and a recurrent network turns that sequence into an emotion context vector that shapes response generation. Offline tests on the Cornell movie and DailyDialog corpora show MEED reaches the lowest perplexity and highest BLEU among the three models tested, and human raters judged its responses more contextually coherent and emotionally appropriate on average than the two baselines while matching them on grammar, with the gap over the sequence-to-sequence baseline reported as significant. The paper argues this is a step toward chatbots that reproduce the social-emotional intelligence humans display in conversation.

What carries the argument

The load-bearing object is the emotion context vector $e$, produced by a unidirectional GRU that reads, in conversation order, a six-dimensional LIWC indicator vector for each history utterance (positive, negative, anxious, angry, sad, or neutral), each embedded by a dense sigmoid layer. At every decoding step, $e$ is concatenated with the decoder's language context vector before the softmax, so the emotional trajectory of the conversation directly biases word choice. This sits on an HRAN-style hierarchical attention encoder for semantics, so the model claims to track both what was said and how it felt.

What would settle it

Check whether any of the 100 test dialogs (or their first three utterances) appear in the 46,797 DailyDialog training pairs; if they do, re-run the human evaluation on a set provably unseen by MEED and see whether the emotional-appropriateness advantage over HRAN survives.

Watch

Extended reading notes

Core claim

The central claim is that explicitly modeling the emotional flow of a multi-turn conversation, as a separate recurrent encoder over utterance-level emotion indicators, produces responses that human raters find emotionally more appropriate and better tied to the context than a plain sequence-to-sequence model or a hierarchical attention model. In the paper's own experiments, MEED scored an average of 0.917 on emotional appropriateness and 0.990 on contextual coherence versus 0.895 and 0.958 for HRAN and 0.688 and 0.713 for S2S, with the MEED-over-S2S differences reported as statistically significant. The authors further show that the emotion layer's weight vectors cluster positive and negative words in t-SNE, evidence that the encoder is tracking affect rather than merely copying context semantics. The claim is that this data-driven, rule-free emotion tracking is what makes the difference.

Load-bearing premise

The load-bearing premise is that the 100 human-evaluated dialogs were held out from training, even though 50 come from the same DailyDialog corpus used for fine-tuning.

Editorial extensions

If this is right

  • MEED can be deployed without any emotion label at response time: the only affective input is the LIWC tags of the history, so the model chooses the emotional stance itself from context.
  • Perplexity and BLEU do not decide which model feels most human: HRAN ranks below S2S on perplexity yet above it in human judgments, so response quality needs human evaluation.
  • The emotion encoder's learned weights separate positive from negative words, suggesting the same architecture could distinguish finer affect categories if given richer labels.
  • Replacing the emotion classifier (LIWC) with a finer or domain-specific affect recognizer should slot into the same encoder without architectural change.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 50 DailyDialog test dialogs were not held out, part of MEED's emotional-appropriateness edge could reflect memorization of DailyDialog's emotion patterns rather than generalizable emotion tracking; a held-out test would settle this.
  • MEED's emotion encoder is trained end-to-end, so the LIWC indicator vectors act as a fixed-feature front end; a learned emotion classifier trained jointly might capture affect words LIWC misses, especially in domain slang.
  • The architecture suggests a natural control knob for chatbot personality: scaling or shifting the emotion context vector $e$ could let the system amplify or dampen the emotional intensity of its replies, a step the paper does not explore.
  • Because the model only mirrors the emotional trajectory of the context, it may reproduce the interlocutor's negativity in a 'miserable' conversation; a separate objective to de-escalate or re-engage could be added on top of $e$.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. This paper proposes MEED, a multi-turn open-domain dialog model that extends a hierarchical recurrent attention network with an emotion encoder. For each context utterance, LIWC2015 produces a six-dimensional binary emotion indicator; a dense layer embeds the indicator, and a GRU encodes the sequence of emotion embeddings into a single emotion context vector e, which is concatenated with the decoder hidden state before the output softmax. The model is pre-trained on the Cornell Movie Dialogs Corpus and fine-tuned on DailyDialog, then compared with a vanilla seq2seq model (S2S) and with HRAN. The paper reports lower perplexity and higher BLEU for MEED on the validation and test sets, and human evaluations on 100 four-turn dialogs (50 positive, 50 negative) in which MEED receives the highest average scores for grammatical correctness, contextual coherence, and emotional appropriateness. The authors also provide a t-SNE visualization of output-layer weights and a case study.

Significance. If the empirical claims hold, the paper makes a modest but useful contribution to affect-aware response generation: it demonstrates a simple data-driven way to condition a multi-turn dialog model on a compressed emotion trajectory rather than on a hand-selected target emotion, and it documents a human-evaluation design that balances positive and negative test dialogs. The authors release source code, which aids reproducibility. The main value is the combination of hierarchical context encoding with an emotion-tracking signal and the attempt to evaluate emotional appropriateness with a balanced test set. The strength of these claims, however, depends entirely on the integrity of the held-out test set and on the statistical comparison with HRAN, both of which need clarification.

major comments (2)
  1. [Section 4, 'Preparation of Natural Dialog Test Set'; Tables 2, 4, 5] The test set used for both the automatic and human evaluations is constructed from DailyDialog, but the paper never states that these dialogs were excluded from the training or validation pairs. The model is fine-tuned on 46,797 DailyDialog context-response pairs, and the training construction in Section 4 creates a context-response pair for every position in every dialog; therefore a four-turn test dialog's context (u1,u2,u3) and response u4 will appear as a training pair if that dialog is in the training split. Since the test set is built by selecting 50 positive dialogs from 78 DailyDialog dialogs and the negative pool can include the 14 negative DailyDialog dialogs, up to 64 of 100 test contexts may be DailyDialog-sourced. The paper does not state that these dialogs were held out, and the test set is not released ('due to privacy concerns, we do not plan to release this dataset'). If any of these contexts were seen during training, the perplexity/BLEU advantage in Table 2 and the human-judged coherence and emotional-appropriateness scores in Tables 4 and 5 would measure memorization rather than generalization. Please state explicitly how the test dialogs were held out from training and validation, and release the DailyDialog dialog IDs or an overlap check.
  2. [Section 4, 'Human Evaluation Results'; Tables 4, 5; Section 1 Contributions] The paper's contribution statement claims MEED produces 'emotionally more appropriate responses than both baselines,' but the only significance test reported for the human evaluation is MEED versus S2S (p<0.01). For MEED versus HRAN, the average scores are 0.990 versus 0.958 for contextual coherence and 0.917 versus 0.895 for emotional appropriateness; these differences are not tested, and with four raters and 100 items they may well be within noise. Either add pairwise significance tests for MEED versus HRAN, or revise the claim to say that MEED improves over S2S and is at least comparable to HRAN on these dimensions.
minor comments (7)
  1. [Section 4, 'Preparation of Natural Dialog Test Set'] The description of worker-created negative dialogs is ambiguous: it is unclear whether each of the two workers wrote five dialogs per topic (50 total) or five dialogs across all topics (25 total); the final selection of 50 negative dialogs implies the former, but the text should state this explicitly.
  2. [Abstract and Section 4, 'Automatic Evaluation Results'] The significance test for perplexity is reported only for the two validation sets; the statement that MEED 'performs the best' on the test set in Table 2 should be phrased descriptively, or accompanied by a significance test on the test set.
  3. [Section 4, 'Human Evaluation Results'] The reported Fleiss' kappa values (0.327-0.389) indicate only fair inter-rater agreement; this should be noted in the text when interpreting the strength of the human judgments.
  4. [Section 2 and Section 5] The paper repeatedly describes the approach as 'free of human-defined heuristic rules' and 'completely data-driven', but the emotion signal is obtained from LIWC, a human-curated lexicon with hand-selected categories; please soften or qualify this claim.
  5. [Section 4, 'Visualization of Output Layer Weights'] The description of the t-SNE analysis does not state how the '100 most frequent positive words' and '100 most frequent negative words' are identified; please specify the source of the sentiment labels.
  6. [References] The reference [10] (Fleiss and Cohen 1973) is not the original source of Fleiss' kappa; please cite the appropriate Fleiss (1971) article.
  7. [Throughout] Typos such as 'Wokrshops' in the ACM Reference Format, 'gurantee' in Section 5, 'mechansim' in Section 4, 'begain' in Section 5, and 'U/t_terance-level' in Figure 1 should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the evaluation is an empirical benchmark and no prediction reduces to its inputs by construction.

full rationale

The paper's claims are empirical: MEED is compared with S2S and HRAN by perplexity, BLEU, and human ratings. There is no derivation chain in which an output quantity is definitionally equal to an input quantity. The emotion context vector e is computed from LIWC indicators over the source context utterances (Equations 11-12), and the decoder conditions on the concatenation of language and emotion vectors (Equations 13-16); the measured perplexity and human scores are outcomes of trained models, not renamings of training objectives. The baselines are external (sequence-to-sequence and HRAN), and no load-bearing uniqueness theorem or prior result by the same authors is invoked. The only notable risk is data hygiene rather than circularity: 50 of the 100 human-evaluation dialogs are sampled from DailyDialog, on which MEED was fine-tuned, and the paper does not explicitly state that those test dialogs were excluded from training (Section 4, Datasets and Preparation of Natural Dialog Test Set). If overlap exists, the generalization claims would be weakened, but that is a data-leakage and correctness concern, not a reduction of the claimed result to its own inputs. The paper also states that the test set is not released for privacy reasons, which hinders replication, but this is not circularity. Accordingly, no circular step is identified and the circularity score is 0.

Assumptions & free parameters 6 free parameters · 4 assumptions · 1 invented entities

The paper's contribution is empirical. The only invented entity is the emotion context vector, a learned latent summary. The main assumptions are the adequacy of LIWC emotion labels, the reliability of four human raters, and the implied independence of the test set from the training data. The free-parameter list captures the hand-chosen design decisions that shape the emotion channel and decoding.

free parameters (6)
  • LIWC binary utterance emotion indicator = threshold: any word in category -> entry 1, else neutral
    Section 3, Emotion Encoder. Hand-chosen binarization loses intensity and frequency information, and it defines the emotion signal the encoder sees.
  • Emotion context vector e = final hidden state of emotion GRU
    Section 3, Decoding. The entire emotional history is compressed into one vector, a design choice that limits the model to a single fixed emotion summary.
  • Number of preceding context utterances = up to 4 (si = max(1, i-4))
    Section 4, Datasets. Chosen for computational efficiency; the central claim about multi-turn emotion tracking depends on this window.
  • Beam width = 256
    Section 4, Baselines and Implementation. Decoding parameter that directly affects generated response quality.
  • Hidden and embedding sizes = 256 word embedding, 256 RNN hidden, 128 utterance attention
    Section 4, Baselines and Implementation. Standard but arbitrary choices.
  • Neutral category entry = sixth dimension set to 1 when no emotion word detected
    Section 3, Emotion Encoder. Introduces a subjective definition of neutrality for every utterance without emotion words.
assumptions (4)
  • domain assumption LIWC2015's five emotion categories (positive, negative, anxious, angry, sad) adequately capture the emotionally salient content of dialog utterances.
    Section 3, Emotion Encoder. The entire emotion channel is derived from these categories; if LIWC labels are noisy or incomplete, the emotion vector e cannot track emotional interactions accurately.
  • domain assumption Human raters can reliably judge emotional appropriateness of short responses on a 0-2 scale.
    Section 4, Human Evaluation. The main evidence for the central claim rests on four raters' subjective scores; Fleiss' kappa values around 0.33-0.39 indicate only fair agreement.
  • domain assumption The test dialogs were not used in training or validation.
    Section 4, Datasets and Preparation of Natural Dialog Test Set. Both the fine-tuning set and the test set are built from DailyDialog, but no explicit exclusion is stated, so the comparative evaluation assumes independence.
  • standard math Bidirectional GRU and hierarchical attention mechanisms provide adequate semantic context modeling.
    Section 3, Hierarchical Attention. Standard sequence modeling assumptions inherited from HRAN, not re-derived.
invented entities (1)
  • Emotion context vector e
    purpose: Summarizes the emotional trajectory of the dialog history and is concatenated with the decoder hidden state at every step to bias word choice.
    An internal latent variable with no falsifiable handle outside the paper; its value is never directly validated, only indirectly through human ratings of the final response.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Multi-Turn Emotionally Engaging Dialog Model." pith.science (2026). https://pith.science/paper/QY7LEDSU

@misc{pith2026190807816,
  author       = {Pith},
  title        = {Pith review of: A Multi-Turn Emotionally Engaging Dialog Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QY7LEDSU}},
  note         = {Machine review of arXiv:1908.07816}
}
read the original abstract

Open-domain dialog systems (also known as chatbots) have increasingly drawn attention in natural language processing. Some of the recent work aims at incorporating affect information into sequence-to-sequence neural dialog modeling, making the response emotionally richer, while others use hand-crafted rules to determine the desired emotion response. However, they do not explicitly learn the subtle emotional interactions captured in human dialogs. In this paper, we propose a multi-turn dialog system aimed at learning and generating emotional responses that so far only humans know how to do. Compared with two baseline models, offline experiments show that our method performs the best in perplexity scores. Further human evaluations confirm that our chatbot can keep track of the conversation context and generate emotionally more appropriate responses while performing equally well on grammar.

Figures

Figures reproduced from arXiv: 1908.07816 by the authors.

Figure 1
Figure 1. The overall architecture of our model. encoder output. More specifically, at decoding step t, the summary of utterance xj is a linear combination of hjk , for k = 1, 2, . . . ,nj , r t j = Õnj k=1 α t jkhjk . (5) Here α t jk is the word-level attention score placed on hjk , and can be calculated as a t jk = v T a tanh(Uast−1 +Vaℓ t j+1 +Wahjk ), (6) α t jk = exp(a t jk ) Ínj k ′=1 exp(a t jk′ ) , (7) where st−1 is t… view at source ↗
Figure 2
Figure 2. t-SNE visualization of the output layer weights in HRAN and MEED. 100 most frequent positive words and 100 most [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 49 canonical work pages

  1. [1]

    Nabiha Asghar, Pascal Poupart, Jesse Hoey, Xin Jiang, and Lili Mou

  2. [2]

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR abs/1409.0473 (2014). arXiv:1409.0473 http://arxiv.org/abs/1409.0473

  3. [3]

    https://doi.org/10.1007/978-3-319-76941-7_12

    154–166. https://doi.org/10.1007/978-3-319-76941-7_12

  4. [4]

    Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017. A Survey on Dialogue Systems: Recent Advances and New Frontiers. SIGKDD Explorations 19, 2 (2017), 25–35. https://doi.org/10.1145/ 3166054.3166058

  5. [5]

    Bickmore and Rosalind W

    Timothy W. Bickmore and Rosalind W. Picard. 2005. Establishing and Maintaining Long-Term Human-Computer Relationships. ACM Trans. Comput.-Hum. Interact. 12, 2 (2005), 293–327. https://doi.org/10.1145/ 1067860.1067867

  6. [6]

    Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011. Chameleons in Imagined Conversations: A New Approach to Understanding Coor- dination of Linguistic Style in Dialogs. In Proceedings of CMCL@ACL

  7. [7]

    Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. A Multi-Turn Emotionally Engaging Dialog Model IUI ’20 Workshops, March 17, 2020, Cagliari, Italy Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation. In Proceedings of EMNLP 2014 . 1724–

  8. [8]

    Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, and Sune Lehmann. 2017. Using Millions of Emoji Occurrences to Learn Any- Domain Representations for Detecting Sentiment, Emotion and Sar- casm. In Proceedings of EMNLP 2017 . 1615–1625. https://aclanthology. info/papers/D17-1169/d17-1169

Show all 65 references
  1. [9]

    Robert H Finn. 1970. A Note on Estimating the Reliability of Categorical Data. Educational and Psychological Measurement 30, 1 (1970), 71–76

  2. [10]

    Joseph L Fleiss and Jacob Cohen. 1973. The Equivalence of Weighted kappa and the Intraclass Correlation Coefficient as Measures of Relia- bility. Educational and psychological measurement 33, 3 (1973), 613– 619

  3. [11]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova

  4. [12]

    2016.Fundamental Statistics for the Behavioral Sciences

    David C Howell. 2016.Fundamental Statistics for the Behavioral Sciences. Nelson Education

  5. [13]

    George Hripcsak and Daniel F. Heitjan. 2002. Measuring Agreement in Medical Informatics Reliability Studies. Journal of Biomedical Informat- ics 35, 2 (2002), 99–110. https://doi.org/10.1016/S1532-0464(02)00500-2

  6. [14]

    Tianran Hu, Anbang Xu, Zhe Liu, Quanzeng You, Yufan Guo, Vibha Sinha, Jiebo Luo, and Rama Akkiraju. 2018. Touch Your Heart: A Tone- aware Chatbot for Customer Care on Social Media. In Proceedings of CHI 2018. 415. https://doi.org/10.1145/3173574.3173989

  7. [15]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Sto- chastic Optimization. CoRR abs/1412.6980 (2014). arXiv:1412.6980 http://arxiv.org/abs/1412.6980

  8. [16]

    Jonathan Klein, Youngme Moon, and Rosalind W. Picard. 2001. This Computer Responds to User Frustration: Theory, Design, and Results. Interacting with Computers 14, 2 (2001), 119–140. https://doi.org/10. 1016/S0953-5438(01)00053-4

  9. [17]

    Sayan Ghosh, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, and Stefan Scherer. 2017. Affect-LM: A Neural Language Model for Customizable Affective Text Generation. In Proceedings of ACL 2017. 634–642. https://doi.org/10.18653/v1/P17-1059

  10. [18]

    Spithourakis, Jian- feng Gao, and William B

    Jiwei Li, Michel Galley, Chris Brockett, Georgios P. Spithourakis, Jian- feng Gao, and William B. Dolan. 2016. A Persona-Based Neural Conver- sation Model. InProceedings of ACL 2016. http://aclweb.org/anthology/ P/P16/P16-1094.pdf

  11. [19]

    Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016. Deep Reinforcement Learning for Dialogue Gen- eration. In Proceedings of EMNLP 2016 . 1192–1202. http://aclweb.org/ anthology/D/D16/D16-1127.pdf

  12. [20]

    Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu

  13. [21]

    Pierre Lison and Jörg Tiedemann. 2016. OpenSubtitles2016: Extract- ing Large Parallel Corpora from Movie and TV Subtitles. In Proceed- ings of LREC 2016 . http://www.lrec-conf.org/proceedings/lrec2016/ summaries/947.html

  14. [22]

    Chia-Wei Liu, Ryan Lowe, Iulian Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. In Proceedings of EMNLP 2016 . 2122–

  15. [23]

    Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan

  16. [24]

    Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard- Gontier, Yoshua Bengio, and Joelle Pineau. 2017. Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses. In Proceedings ACL 2017. 1116–1126. https://doi.org/10.18653/v1/P17-1103

  17. [25]

    Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, Nov (2008), 2579–2605

  18. [26]

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Ef- ficient Estimation of Word Representations in Vector Space. CoRR abs/1301.3781 (2013). arXiv:1301.3781 http://arxiv.org/abs/1301.3781

  19. [27]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of ACL 2002 . 311–318. http://www.aclweb.org/anthology/ P02-1040.pdf

  20. [28]

    James W Pennebaker, Martha E Francis, and Roger J Booth. 2001. Linguistic Inquiry and Word Count: LIWC 2001. Mahway: Lawrence Erlbaum Associates 71, 2001 (2001), 2001

  21. [29]

    Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau

  22. [30]

    Alan Ritter, Colin Cherry, and William B. Dolan. 2011. Data-Driven Response Generation in Social Media. In Proceedings of EMNLP 2011 . 583–593. http://www.aclweb.org/anthology/D11-1054

  23. [31]

    Klaus R Scherer. 2005. What Are Emotions? And How Can They Be Measured? Social science information 44, 4 (2005), 695–729

  24. [32]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov

  25. [33]

    CoRR abs/1907.11692 (2019)

    RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/ 1907.11692

  26. [34]

    Lifeng Shang, Zhengdong Lu, and Hang Li. 2015. Neural Responding Machine for Short-Text Conversation. In Proceedings of ACL-IJCNLP

  27. [35]

    Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi, Christina Li- oma, Jakob Grue Simonsen, and Jian-Yun Nie. 2015. A Hierarchical Recurrent Encoder-Decoder for Generative Context-Aware Query Sug- gestion. In Proceedings of CIKM 2015. 553–562. https://doi.org/10.1145/ 2806416.2806493

  28. [36]

    Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan. 2015. A Neural Network Approach to Context-Sensitive Gen- eration of Conversational Responses. In Proceedings of NAACL-HLT

  29. [37]

    Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to Sequence Learning with Neural Networks. In Proceedings of NIPS 2014 . 3104–3112. http://papers.nips.cc/paper/5346-sequence-to-sequence- IUI ’20 Workshops, March 17, 2020, Cagliari, Italy Xie, et al. learning-with...

  30. [38]

    Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018. RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems. In Proceedings of AAAI 2018 . 722–729. https://www. aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16179

  31. [39]

    Christoph Tillmann and Hermann Ney. 2003. Word Reordering and a Dynamic Programming Beam Search Algorithm for Statistical Machine Translation. Computational Linguistics 29, 1 (2003), 97–133. https: //doi.org/10.1162/089120103321337458

  32. [40]

    In Proceedings of ACL 2019

    Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset. In Proceedings of ACL 2019. 5370–5381. https://www.aclweb.org/anthology/P19-1534/

  33. [41]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Proceedings of NIPS 2017 . 5998–6008. http://papers.nips.cc/paper/7181-attention-is-all-you-need

  34. [42]

    Oriol Vinyals and Quoc V. Le. 2015. A Neural Conversational Model. CoRR abs/1506.05869 (2015). arXiv:1506.05869 http://arxiv.org/abs/ 1506.05869

  35. [43]

    Courville, and Joelle Pineau

    Iulian Vlad Serban, Alessandro Sordoni, Yoshua Bengio, Aaron C. Courville, and Joelle Pineau. 2016. Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models. In Proceedings of AAAI 2016 . 3776–3784. http://www.aaai.org/ocs/index. php/AAAI/AAAI16...

  36. [44]

    Courville, and Yoshua Bengio

    Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Char- lin, Joelle Pineau, Aaron C. Courville, and Yoshua Bengio. 2017. A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues. In Proceedings of AAAI 2017 . 3295–3301. http://aaai.org/ ocs/index....

  37. [45]

    Chen Xing, Yu Wu, Wei Wu, Yalou Huang, and Ming Zhou. 2018. Hierarchical Recurrent Attention Network for Response Generation. In Proceedings of AAAI 2018 . 5610–5617. https://www.aaai.org/ocs/ index.php/AAAI/AAAI18/paper/view/16510

  38. [46]

    Anbang Xu, Zhe Liu, Yufan Guo, Vibha Sinha, and Rama Akkiraju. 2017. A New Chatbot for Customer Service on Social Media. In Proceedings of CHI 2017. 3506–3510. https://doi.org/10.1145/3025453.3025496

  39. [47]

    Jennifer Zamora. 2017. I’m Sorry, Dave, I’m Afraid I Can’t Do That: Chatbot Perception and Expectations. In Proceedings of HAI 2017 . 253–

  40. [48]

    Peixiang Zhong, Di Wang, and Chunyan Miao. 2019. An Affect-Rich Neural Conversational Model with Biased Attention and Weighted Cross-Entropy Loss. In Proceedings of AAAI 2019 . 7492–7500. https: //aaai.org/ojs/index.php/AAAI/article/view/4740

  41. [49]

    http://aclweb.org/anthology/N/N15/N15-1020.pdf

    196–205. http://aclweb.org/anthology/N/N15/N15-1020.pdf

  42. [50]

    Xianda Zhou and William Yang Wang. 2018. MojiTalk: Generating Emotional Responses at Scale. In Proceedings of ACL 2018 . 1128–1137. https://doi.org/10.18653/v1/P18-1104

  43. [53]

    Howard E Tinsley and David J Weiss. 1975. Interrater Reliability and Agreement of Subjective Judgments. Journal of Counseling Psychology 22, 4 (1975), 358

  44. [56]

    Amy Beth Warriner, Victor Kuperman, and Marc Brysbaert. 2013. Norms of Valence, Arousal, and Dominance for 13,915 English Lemmas. Behavior research methods 45, 4 (2013), 1191–1207

  45. [57]

    Chen Xing, Wei Wu, Yu Wu, Jie Liu, Yalou Huang, Ming Zhou, and Wei-Ying Ma. 2017. Topic Aware Neural Response Generation. In Proceedings of AAAI 2017 . 3351–3357. http://aaai.org/ocs/index.php/ AAAI/AAAI17/paper/view/14563

  46. [63]

    Hao Zhou, Minlie Huang, Tianyang Zhang, Xiaoyan Zhu, and Bing Liu. 2018. Emotional Chatting Machine: Emotional Conversation Generation with Internal and External Memory. InProceedings of AAAI

  47. [64]

    https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/ 16455

  48. [260]

    https://doi.org/10.1145/3125739.3125766

  49. [1734]

    http://aclweb.org/anthology/D/D14/D14-1179.pdf

  50. [2011]

    https://aclanthology.info/papers/W11-0609/w11-0609

    76–87. https://aclanthology.info/papers/W11-0609/w11-0609

  51. [2015]

    http://aclweb.org/anthology/P/P15/P15-1152.pdf

    1577–1586. http://aclweb.org/anthology/P/P15/P15-1152.pdf

  52. [2016]

    In Proceedings of NAACL-HLT 2016

    A Diversity-Promoting Objective Function for Neural Conver- sation Models. In Proceedings of NAACL-HLT 2016 . 110–119. http: //aclweb.org/anthology/N/N16/N16-1014.pdf

  53. [2017]

    In Proceedings of IJCNLP 2017

    DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In Proceedings of IJCNLP 2017

  54. [2018]

    In Proceedings of ECIR

    Affective Neural Response Generation. In Proceedings of ECIR

  55. [2019]

    In Proceedings of NAACL-HLT 2019

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT 2019. 4171–

  56. [2132]

    http://aclweb.org/anthology/D/D16/D16-1230.pdf

  57. [4186]

    https://aclweb.org/anthology/papers/N/N19/N19-1423/

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.