REVIEW 3 major objections 6 minor 1 cited by
Entertaining and Opinionated but Too Controlling: A Large-Scale User Study of an Open Domain Alexa Prize System
T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read In a field trial of over 10,000 conversations, users rated a social bot's storytelling and games higher than its search and general chit-chat, while also finding the system too controlling in open chit-chat.
desk verdict A valuable large-scale field study of an open-domain social bot whose module-level comparisons are real but observational—worth reviewing, with the search-fallback confound noted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is SlugBot (SB), an open-domain social dialogue system whose dialogue manager logs a 'signature' for every system turn, labeling which of four high-level modules produced it: Search, Topic-oriented Chit-Chat, Interactive Games, or Storytelling. Those signatures let the authors retrospectively group the 10,000 field conversations by module and compare ratings, turn counts, and durations. The supporting mechanism is content sourcing: crowd-sourced games and hypothetical questions, public fables and personal narratives, and topic-indexed trivia and news retrieved by Elasticsearch. The signature logging is what turns a messy public deployment into a per-module user study, while the playful, opinion-driven content is what the authors claim drives the positive ratings.
What would settle it
Randomly assign users to four otherwise identical bot conditions that differ only in module type, search, chit-chat, games, and storytelling, and compare ratings; if the storytelling advantage over search falls below significance or reverses, the paper's central module-preference conclusion is an artifact of user self-selection.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that module design determines perceived conversational success: conversations involving the Storytelling module received the highest average user rating (3.62, versus 3.12 for chit-chat, 3.20 for games, and 3.01 for search), and search-driven interactions were significantly worse across rating, turns, and duration. In the authors' interpretation, users do not want a talking encyclopedia; factual information is welcomed only when it serves another activity, such as supporting an opinion or extending a story. The qualitative data further show that structured activities with clear turn-taking, such as stories told in installments and games with predictable question formats, let users understand their role, reduce coverage ambiguity, and lead users to attribute intelligence and personality to the system. The same structure is experienced as over-control in open chit-chat, where users wanted more influence over the agenda. The paper's prescriptive conclusion is that future social bots should have their own opinions and personal stories to share, with SlugBot as a working example.
Load-bearing premise
The comparisons assume that users who chose to engage with stories or games are comparable to those who engaged with search or chit-chat, but since users picked their own modules, the rating differences could reflect the types of users each module attracted rather than the module's quality.
Editorial extensions
If this is right
- Search-style question answering, on its own, shortens conversations and lowers ratings, so open-domain social bots should not be built primarily around factual retrieval.
- Stories and games can carry long, highly rated conversations even when the underlying understanding is imperfect, because their predictable structure lowers ambiguity for the user.
- Users attribute intelligence and personality to a bot that expresses and defends opinions, so opinion exchange is a usable design lever for perceived humanness.
- System-led control is acceptable inside a clearly signaled activity, but in free chit-chat it reads as domination; future designs need to hand agenda control to the user.
- Initial system prompts that promise broad coverage raise expectations that hurt later ratings, so expectation-setting language should be conservative.
Reading between the lines
- The per-module rating gaps are correlational because users self-selected modules; a randomized assignment that controls user type could confirm whether the storytelling advantage is causal.
- If the mechanism is opinion exchange plus predictable structure, then other structured activities, such as debates, collaborative storytelling, or shared hypotheticals, should produce similar engagement without needing more factual content.
- The 'too controlling' complaint suggests quantifiable design targets, such as a user-initiated topic-switch success rate or the fraction of system turns that ask the user to choose a direction, which could be measured and optimized.
- The finding that users impute intelligence from opinionated responses implies that perceived competence may be manipulable by style of response rather than by underlying knowledge coverage.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes the design and large-scale field evaluation of SlugBot, an open-domain social conversational system that participated in the 2018 Alexa Prize. SlugBot combines topic-oriented chit-chat, interactive games, storytelling, and search-based fallback strategies, with content sourced from public corpora and crowdsourcing. The reported deployment collected over 10,000 conversations in August 2018, with user ratings and system logs; an in-lab qualitative study of 16 participants was also conducted. The central claims are that storytelling and games receive higher user ratings and lead to longer conversations than search and general chit-chat, and that users find the system too controlling in general chit-chat. From this, the authors draw design implications favoring conversational systems that express opinions and share personal stories over systems that primarily provide factual information.
Significance. If the module-level comparison is taken at face value, the paper makes a useful empirical contribution to the design of open-domain social chatbots, providing evidence at a scale (over 10,000 conversations) that is rare in this area. The use of nonparametric Mann-Whitney tests for the rating comparisons and the grounding of qualitative claims in participant quotes are strengths. However, the central quantitative comparison is confounded by user self-selection into modules and by the system's use of search as a fallback after coverage failures, which the authors themselves acknowledge in the Limitations section. As a result, the paper's headline design implication---that users do not want factual information provision---is not fully supported by the reported data. The paper is honest about these limitations, but the Abstract and Discussion state the implication more strongly than the evidence warrants.
major comments (3)
- [Section 4.1 / Table 2] The comparison of ratings across dialogue modules is confounded by the system's fallback behavior and by user self-selection. Section 3.2 states that search is used 'to fill in gaps in our database of content' and is a 'fall-back strategy,' while Section 4.1 notes that 'users were able to choose which modules they interacted with.' Therefore, a conversation labeled Combined Search may have been a conversation in which the system failed to understand the user or lacked topic content, triggering a search-based response, rather than a conversation in which the user deliberately sought factual information. The lower ratings and shorter durations for Combined Search in Table 2 could be driven by these system failures. The Mann-Whitney tests show that the differences are statistically significant, but they do not establish that the cause is a user preference against factual information. This confound directly affects the central design implication stated in the Abstract and Section 5.
- [Section 5] The Discussion contains an internal tension. It acknowledges, 'search is technically difficult and it may have been that it returned irrelevant results - a possibility that we intend to evaluate more systematically,' yet it also states, 'users did not seem to want Information Provision, as evidenced by low ratings for Search modules.' If the low search ratings may reflect a failure artifact, they cannot simultaneously serve as evidence that users are averse to factual information. The paper should either re-analyze the log data to separate user-initiated search from system-initiated fallback search (or at least remove conversations with known system errors), or substantially temper the conclusion to say that the current implementation of search underperformed. As written, the key design implication is not supported.
- [Section 4.2] The qualitative study is described as confirming the quantitative results from Table 2, but this is not an independent confirmation. The 16 participants were instructed to engage in general conversation first and then directly interact with specific modules, so the module exposure was not self-selected but still not randomized, and the sample is small and homogeneous (mean age 22.3). The qualitative reactions are valuable as complementary evidence, but they should be presented as illustrative rather than as confirming the deployment comparison. Additionally, the participant demographics sentence states that 'Seven owned an Alexa device while 7 described themselves as having limited or no Alexa experience'; these two groups sum to 14, leaving 2 participants unaccounted for, and the overlap between the two categories is unclear.
minor comments (6)
- [Section 3.2] There is a typo in the sentence 'Overall, this approach is imrpoves the ability of SB to stay on topic'; 'imrpoves' should be 'improves', and the phrase 'is improves' should be corrected.
- [Section 4.1] In the sentence reporting the Mann-Whitney tests, 'topic-oriented Chat-Chat' should be 'topic-oriented Chit-Chat'.
- [Table 2] Table 2 reports mean and median ratings, turns, and time for each module but does not report the number of conversations assigned to each module group. These counts are necessary to assess the Mann-Whitney U values and to understand the composition of each group, especially because many conversations involve multiple modules.
- [Figure 7 caption] The caption states 'Higher ratings are significantly more likely with longer conversations,' but the reported Pearson correlation is only r = 0.18. 'More likely' is imprecise for a continuous rating variable; consider phrasing such as 'longer conversations tend to receive higher ratings.'
- [References] References [55] and [56] appear to be the same paper (Thorne, Korobov, and Morgan 2007) and are listed with identical bibliographic details. One should be removed or the citations should be merged.
- [Section 4.1] The paper reports 'over 10,000 individual conversations' and 'over 290,000 user turns' but does not provide exact counts or the number of unique users. Providing exact numbers would strengthen the scale claims and facilitate comparison with other deployments.
Circularity Check
No significant circularity: the user-study conclusions are drawn from external field and lab observations, not from the paper's own definitions or cited prior work.
full rationale
This paper contains no mathematical derivation, fitted model, or prediction that reduces by construction to its inputs. The central claims—that Storytelling and Games received higher ratings than Search and Chit-Chat, and that users found the system too controlling—are empirical findings from an Alexa Prize field deployment ("We collected data for the month of August 2018, resulting in over 10,000 individual conversations") and from a follow-up qualitative study. The module comparisons in Table 2 are produced by associating logged turn signatures with user ratings, an external benchmark, not by a formula that presupposes the conclusion. The self-citations to prior SlugBot system work [5] and to corpora [19, 29] are engineering and resource background; they do not supply the evaluative results. The paper's own Limitations section candidly notes that the deployment was not a controlled experiment and that results are "less definitive than those from a controlled deployment," which is an internal-validity caveat about self-selection and confounding, not circularity. The skeptic's concern that Search ratings may reflect fallback failures rather than dislike of factual information is a plausible alternative interpretation of the empirical correlation, but it does not show that any derivation is equivalent to its inputs. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The per-turn 'signature' logs correctly identify which dialogue module (Search, Chit-Chat, Games, Stories) produced each system turn.
- domain assumption User ratings on the 1-5 scale are a valid measurement of engagement and success at the conversation level.
- domain assumption The 16-participant in-lab sample is informative about the broader population of Alexa Prize users.
- domain assumption Random assignment of users to Alexa Prize systems by Amazon holds, so the user population for SB is comparable to other systems.
Cite this review
Pith. "Pith review of Entertaining and Opinionated but Too Controlling: A Large-Scale User Study of an Open Domain Alexa Prize System." pith.science (2026). https://pith.science/paper/2TNRMQ6Z
@misc{pith2026190804832,
author = {Pith},
title = {Pith review of: Entertaining and Opinionated but Too Controlling: A Large-Scale User Study of an Open Domain Alexa Prize System},
year = {2026},
howpublished = {\url{https://pith.science/paper/2TNRMQ6Z}},
note = {Machine review of arXiv:1908.04832}
}
read the original abstract
Conversational systems typically focus on functional tasks such as scheduling appointments or creating todo lists. Instead we design and evaluate SlugBot (SB), one of 8 semifinalists in the 2018 AlexaPrize, whose goal is to support casual open-domain social inter-action. This novel application requires both broad topic coverage and engaging interactive skills. We developed a new technical approach to meet this demanding situation by crowd-sourcing novel content and introducing playful conversational strategies based on storytelling and games. We collected over 10,000 conversations during August 2018 as part of the Alexa Prize competition. We also conducted an in-lab follow-up qualitative evaluation. Over-all users found SB moderately engaging; conversations averaged 3.6 minutes and involved 26 user turns. However, users reacted very differently to different conversation subtypes. Storytelling and games were evaluated positively; these were seen as entertaining with predictable interactive structure. They also led users to impute personality and intelligence to SB. In contrast, search and general Chit-Chat induced coverage problems; here users found it hard to infer what topics SB could understand, with these conversations seen as being too system-driven. Theoretical and design implications suggest a move away from conversational systems that simply provide factual information. Future systems should be designed to have their own opinions with personal stories to share, and SB provides an example of how we might achieve this.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
LLMs assisting cyber threat intelligence fail mainly due to spurious correlations, contradictory knowledge, and constrained generalization that stem from the threat landscape itself.
Reference graph
Works this paper leans on
-
[1]
David Ameixa, Luisa Coheur, Pedro Fialho, and Paulo Quaresma. 2014. Luke, I am Your Father: Dealing with Out-of-Domain Requests by Using Movies Subtitles . Springer, 13–21
work page 2014
-
[2]
Jaime Arguello, Filip Radlinski, Hideo Joho, Damiano Spina, and Julia Kisel- eva. 2018. Second International Workshop on Conversational Approaches to Information Retrieval (CAIR’18). Proceedings of SIGIR’18 (2018), 32–41
work page 2018
-
[3]
Rafael E Banchs and Haizhou Li. 2012. IRIS: A Chat-oriented Dialogue System Based on the Vector Space Model. In Proc of the ACL 2012 System Demonstrations (ACL ’12). 37–42
work page 2012
-
[4]
Jerome R Bellegarda. 2013. Large-scale personal assistant technology deployment: the siri experience.. In INTERSPEECH. 2029–2033
work page 2013
-
[5]
Kevin K Bowden, Jiaqi Wu, Wen Cui, Juraj Juraska, Vrindavan Harrison, Brian Schwarzmann, Nick Santer, and Marilyn Walker. 2018. SlugBot: Developing a Computational Model and Framework of a Novel Dialogue Genre. Alexa Prize Proceedings (2018)
work page 2018
-
[6]
Kevin Burton, Akshay Java, and Ian Soboroff. 2009. The ICWSM 2009 Spinn3r dataset. In Proc. of the Annual Conference on Weblogs and Social Media (ICWSM)
work page 2009
-
[7]
Mikhail Burtsev, Aleksandr Chuklin, Julia Kiseleva, and Alexey Borisov. 2017. Search-oriented conversational AI (SCAI). In Proceedings of the ACM SIGIR Inter- national Conference on Theory of Information Retrieval . ACM, 333–334
work page 2017
-
[8]
Justine Cassell and Kristinn R Thorisson. 1999. The power of a nod and a glance: Envelope vs. emotional feedback in animated conversational agents. Applied Artificial Intelligence 13, 4-5 (1999), 519–538
work page 1999
Show all 61 references
-
[9]
Chun-Yen Chen, Dian Yu, Weiming Wen, Yi Mang Yang, Jiaping Zhang, Mingyang Zhou, Kevin Jesse, Chau Austin, Antara Bhowmick, Shreenath Iyer, Giritheja Sreenivasulu, Runxiang Cheng, Ashwin Bhandare, and Zhou Yu. 2018. Gunrock: Building A Human-Like Social Bot By Leveraging Large...
2018
-
[10]
Aleksandr Chuklin, Ilya Markov, and Maarten de Rijke. 2015. Click models for web search. Synthesis Lectures on Information Concepts, Retrieval, and Services 7, 3 (2015), 1–115
2015
-
[11]
Amanda Cercas Curry, Ioannis Papaiannou, Alessandro Suglia, Shubham Agar- wal, Igor Shalyminov, Xinnuo Xu, Ondrej Dusek, Arash Eshghi, Ioannis Konstas, Verena Rieser, and Oliver Lemon. 2018. Alana v2: Entertaining and Informative Open-domain Social Dialogue using Ontologies an...
2018
-
[12]
Ansgar E Depping, Regan L Mandryk, Colby Johanson, Jason T Bowey, and Shelby C Thomson. 2016. Trust me: social games are better than social icebreakers at building trust. In Proceedings of the 2016 Annual Symposium on Computer- Human Interaction in Play . ACM, 116–129
2016
-
[13]
Guillaume Dubuisson Duplessis, Vincent Letard, Anne-Laure Ligozat, and Sophie Rosset. 2016. Purely Corpus-based Automatic Conversation Authoring. In 10th edition of the Language Resources and Evaluation Conference (LREC)
2016
-
[14]
David K. Elson. 2012. DramaBank: Annotating Agency in Narrative Discourse. In Proc. of the Eighth International Conference on Language Resources and Evaluation (LREC 2012)
2012
-
[15]
Hao Fang, Hao Cheng, Elizabeth Clark, Ariel Holtzman, Maarten Sap, Mari Ostendorf, Yejin Choi, and Noah A Smith. 2017. Sounding board–university of washington’s alexa prize submission. Alexa Prize Proceedings (2017)
2017
-
[16]
Matthew Henderson, Blaise Thomson, and Jason Williams. 2014. The Second Dialog State Tracking Challenge. In Proceedings of SIGDIAL. ACL - Association for Computational Linguistics. https://www.microsoft.com/en-us/research/ publication/the-second-dialog-state-tracking-challenge/
2014
-
[17]
Ryuichiro Higashinaka, Kenji Imamura, Toyomi Meguro, Chiaki Miyazaki, No- zomi Kobayashi, Hiroaki Sugiyama, Toru Hirano, Toshiro Makino, and Yoshihiro Matsuo. 2014. Towards an open-domain conversational system fully based on natural language processing. InProceedings of COLING...
2014
-
[18]
Lynette Hirschman. 2000. Evaluating spoken language interaction: Experiences from the darpa spoken language program 1990–1995. Spoken Language Discourse. MIT Press, Cambridge, Mass (2000)
2000
-
[19]
Fox Tree, and Marilyn Walker
Zhichao Hu, Michelle Dick, Johnnie Chang, Kevin Bowden, Michael Neff, Jean E. Fox Tree, and Marilyn Walker. 2016. A Corpus of Human-Generated Dialogs from Personal Narratives with Gesture Annotations. In Language Resources and Evaluation Conference, LREC 2016 . 3447–3454
2016
-
[20]
Katherine Isbister and Clifford Nass. 2000. Consistency of personality in interac- tive characters: verbal cues, non-verbal cues, and user characteristics. Interna- tional Journal of Human-Computer Studies 53, 2 (2000), 251 – 267
2000
-
[21]
Katherine Isbister and Clifford Nass. 2000. Consistency of personality in interac- tive characters: verbal cues, non-verbal cues, and user characteristics. Interna- tional journal of human-computer studies 53, 2 (2000), 251–267
2000
-
[22]
Chandra Khatri, Behnam Hedayatnia, Anu Venkatesh, Jeff Nunn, Yi Pan, Qing Liu, Han Song, Anna Gottardi, Sanjeev Kwatra, Sanju Pancholi, Ming Cheng, Chen Qinglang, Lauren Stubel, Karthik Gopalakrishnan, Kate Bland, Raefer Gabriel, Arindam Mandal, Dilek Hakkani-Tur, Gene Hwang, ...
2018
-
[23]
Julia Kiseleva, Kyle Williams, Ahmed Hassan Awadallah, Aidan C Crook, Imed Zitouni, and Tasos Anastasakos. 2016. Predicting user satisfaction with intelli- gent assistants. In Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information...
2016
-
[24]
William Labov and David Fanshel. 1977. Therapeutic discourse : psychotherapy as conversation. Academic Press
1977
-
[25]
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A Persona-Based Neural Conversation Model. arXiv preprint arXiv:1603.06155 (2016)
2016 arXiv
-
[26]
Pierre Lison and Jörg Tiedemann. 2016. Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles. (2016)
2016
-
[27]
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015. The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems. In Proceedings of the SIGDIAL 2015 Conference, The 16th Annual Meeting of the Special Interest Group on Discours...
2015
-
[28]
Like Having a Really Bad PA
Ewa Luger and Abigail Sellen. 2016. "Like Having a Really Bad PA": The Gulf Between User Expectation and Experience of Conversational Agents. In Proceed- ings of the 2016 CHI Conference on Human Factors in Computing Systems (CHI ’16) . ACM, New York, NY, USA, 5286–5297. https:...
2016
-
[29]
Stephanie Lukin, Kevin Bowden, Casey Barackman, and Marilyn Walker. 2016. PersonaBank: A Corpus of Personal Narratives and Their Story Intention Graphs. In Language Resources and Evaluation Conference, LREC2016 . 1026–1033
2016
-
[30]
Chelsea Myers, Anushay Furqan, Jessica Nebolsky, Karina Caro, and Jichen Zhu
-
[31]
Tien T Nguyen, Duyen T Nguyen, Shamsi T Iqbal, and Eyal Ofek. 2015. The known stranger: Supporting conversations between strangers with personalized topic suggestions. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems . ACM, 555–564
2015
-
[32]
Lasguido Nio, Sakriani Sakti, Graham Neubig, Tomoki Toda, Mirna Adriani, and Satoshi Nakamura. 2014. Developing Non-goal Dialog System Based on Examples of Drama Television. Springer, 355–361
2014
-
[33]
Neal R Norrick. 2000. Conversational narrative: Storytelling in everyday talk . Vol. 203. John Benjamins Publishing
2000
-
[34]
Steffi Paepcke and Leila Takayama. 2010. Judging a Bot by Its Cover: An Ex- periment on Expectation Setting for Personal Robots. In Proceedings of the 5th ACM/IEEE International Conference on Human-robot Interaction (HRI ’10) . IEEE Press, Piscataway, NJ, USA, 45–52. http://dl...
2010
-
[35]
Monisha Pasupathi and Timothy Hoyt. 2009. The development of narrative identity in late adolescence and emergent adulthood: The continued importance of listeners. Developmental Psychology 45, 2 (2009), 558–574
2009
-
[36]
Jan Pichl, Petr Marek, Jakub Konrád, Martin Matulík, and Jan Šediv`y. 2018. Alquist 2.0: Alexa Prize Socialbot Based on Sub-Dialogue Model. Alexa Prize Proceedings (2018)
2018
-
[37]
Livia Polanyi. 1989. Telling the American Story: A Structural and Cultural Analysis of Conversational Storytelling. MIT Press
1989
-
[38]
Fischer, Stuart Reeves, and Sarah Sharples
Martin Porcheron, Joel E. Fischer, Stuart Reeves, and Sarah Sharples. 2018. Voice Interfaces in Everyday Life. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18) . ACM, New York, NY, USA, Article 640, 12 pages. https://doi.org/10.1145/317...
2018
-
[39]
Patti Price, Lynette Hirschman, Elizabeth Shriberg, and Elizabeth Wade. 1992. Subject-based evaluation measures for interactive spoken language systems. In Proceedings of the workshop on Speech and Natural Language . Association for Computational Linguistics, 34–39
1992
-
[40]
Filip Radlinski and Nick Craswell. 2017. A Theoretical Framework for Con- versational Search. In Proceedings of the 2017 Conference on Conference Human Information Interaction and Retrieval (CHIIR ’17) . ACM, New York, NY, USA, 117–126. https://doi.org/10.1145/3020165.3020183
2017
-
[41]
Ashwin Ram, Rohit Prasad, Chandra Khatri, Anu Venkatesh, Raefer Gabriel, Qing Liu, Jeff Nunn, Behnam Hedayatnia, Ming Cheng, Ashish Nagar, et al. 2017. Conversational ai: The science behind the alexa prize. Alexa Prize Proceedings (2017)
2017
-
[42]
Byron Reeves and Clifford Nass. 1996. The Media Equation. University of Chicago Press
1996
-
[43]
Alexander I Rudnicky, Eric Thayer, Paul Constantinides, Chris Tchou, R Shern, Kevin Lenzo, Wei Xu, and Alice Oh. 1999. Creating natural dialogs in the Carnegie Mellon Communicator system. In Eurospeech. 1531–1534
1999
-
[44]
Pashupati Sapkota and Junita Sharma. 1996. Participatory interactions with children in Nepal. PLA notes 25 (1996), 61–64
1996
-
[45]
Schegloff
Emanuel A. Schegloff. 1990. On the Organization of Sequences as a Source of Co- herence in Talk-In-Interaction. In Conversational Coherence and Its Development , CUI 2019, August 22–23, 2019, Dublin, Ireland Bowden and Wu, et al. Bruce Dorval (Ed.). Ablex, Norwood, N.J., 51–77
1990
-
[46]
Christopher Schmandt. 1985. Voice Communication with computers. InAdvances in Human-computer interaction, H. Rex Hartson (Ed.). Vol. 1. 133–159
1985
-
[47]
Stephanie Seneff and Joseph Polifroni. 2000. Dialogue management in the Mer- cury flight reservation system. In Proceedings of the 2000 ANLP/NAACL Workshop on Conversational systems-Volume 3. Association for Computational Linguistics, 11–16
2000
-
[48]
Iulian V Serban, Chinnadhurai Sankar, Mathieu Germain, Saizheng Zhang, Zhouhan Lin, Sandeep Subramanian, Taesup Kim, Michael Pieper, Sarath Chan- dar, Nan Rosemary Ke, et al. 2017. A deep reinforcement learning chatbot. arXiv preprint arXiv:1709.02349 (2017)
2017 arXiv
-
[49]
Pararth Shah, Dilek Hakkani-Tur, Bing Liu, and Gokhan Tur. 2018. Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning. In Proceedings of the 2018 Conference of the North Ameri- can Chapter of the Association for ...
2018
-
[50]
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan. 2015. A Neural Network Approach to Context-Sensitive Generation of Conversational Responses. arXiv preprint arXiv:1506.06714 (2015)
2015 arXiv
-
[51]
David Stallard. 1998. Talk’n’Travel: A Conversational System for Air Travel Planning. In In Proc. of the 6th Applied Natural Language Processing Conference and the 1st Meeting of the North American Chapter of the Association for Computational Linguistics (ANLP-NAACL 2000). 68–75
1998
-
[52]
A. Stent. 2001. Dialogue Systems as Conversational Partners: Applying conversation acts theory to natural language generation for task-oriented mixed-initiative spoken dialogue. Ph.D. Dissertation. University of Rochester
2001
-
[53]
Hiroaki Sugiyama, Toyomi Meguro, Ryuichiro Higashinaka, and Yasuhiro Mi- nami. 2013. Open-domain utterance generation for conversational dialogue systems using web-scale dependency structures. In Proceedings of the SIGDIAL 2013 Conference. 334–338
2013
-
[54]
Deborah Tannen. 2007. Talking voices: Repetition, dialogue, and imagery in con- versational discourse. Vol. 26. Cambridge University Press
2007
-
[56]
Avril Thorne, Neill Korobov, and Elizabeth M Morgan. 2007. Channeling identity: A study of storytelling in conversations between introverted and extraverted friends. Journal of research in personality 41, 5 (2007), 1008–1031
2007
-
[57]
Pirros Tsiakoulis, Milica Gasic, Matthew Henderson, Joaquin Plannels, Jorge Prombonas, Blaise Thomson, Kai Yu, Steve Young, and Eli Tzirkel. 2012. Statistical methods for building robust spoken dialogue systems in an automobile. In in 4th International Conference on Applied Hu...
2012
-
[58]
Oriol Vinyals and Quoc Le. 2015. A Neural Conversational Model. arXiv preprint arXiv:1506.05869 (2015)
2015 arXiv
-
[59]
Marilyn A Walker, Diane J Litman, Candace A Kamm, and Alicia Abella. 1997. PARADISE: A framework for evaluating spoken dialogue agents. In Proceedings of the eighth conference on European chapter of the Association for Computational Linguistics. Association for Computational L...
1997
-
[60]
Marilyn A Walker, Rebecca Passonneau, and Julie E Boland. 2001. Quantitative and qualitative evaluation of DARPA Communicator spoken dialogue systems. In Proceedings of the 39th Annual Meeting on Association for Computational Lin- guistics. Association for Computational Lingui...
2001
-
[61]
Steve Whittaker and Phil Stenton. 1989. User Studies and the Design of Natural Language Systems. 116–123
1989
-
[2018]
In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18)
Patterns for How Users Overcome Obstacles in Voice User Interfaces. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18). ACM, New York, NY, USA, Article 6, 7 pages. https://doi.org/10.1145/ 3173574.3173580
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.