REVIEW 4 major objections 4 minor 3 cited by
Proactive Conversational Agents with Inner Thoughts
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read An AI that continuously forms and scores its own inner thoughts can join group conversations proactively, outperforming next-speaker prediction on anthropomorphism, coherence, intelligence, and turn-taking appropriateness.
desk verdict A solid framework with an honest but unproven comparative claim: the evaluation doesn't yet separate the inner-thought mechanism from the underlying LLM capability. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the thought reservoir: a continuously updated set of covert candidate contributions that the AI forms in parallel with the overt conversation. Each cycle runs five stages—trigger, retrieval, thought formation, evaluation, and participation—where an incoming message or a ten-second pause triggers retrieval of relevant long-term-memory items by embedding saliency, formation of one fast 'System 1' and two deliberate 'System 2' thoughts per batch, and an LLM evaluation that assigns each thought a 1–5 intrinsic-motivation score, adjusted upward the longer the AI has been silent. Participation is then gated by interpretable thresholds: an imThreshold for open turns, system1Prob for low-motivation fill-in responses, and an interruptThreshold that lets the AI override turn allocation when a thought is urgent. This machinery is what carries the argument, because it replaces 'who is externally most likely to speak next' with 'which internal thought is most worth saying now'.
What would settle it
A controlled comparison that keeps the response-generation step identical and varies only the participation decision—for example, replacing the intrinsic-motivation gate with random thought selection at the same speaking rate—would settle whether the thought-evaluation mechanism carries the effect; the paper reports no such ablation.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that AI proactivity in multi-party conversation is better modelled as an internal, motivation-driven process than as a prediction problem. The paper shows that even strong LLMs predict the next speaker at near-chance levels in self-selection cases, and that relabelling those cases as 'anyone' improves the task but sidesteps the real question of who should actually contribute. The Inner Thoughts framework answers that question by generating thoughts in parallel, scoring each thought on relevance, information gap, expected impact, urgency, coherence, originality, balance, and dynamics, and letting a threshold decide whether to speak, remain silent, or even interrupt an allocated turn. In the technical evaluation, all seven metrics—anthropomorphism, coherence, engagement, intelligence, turn appropriateness, initiative, and adaptability—were rated significantly higher for Inner Thoughts conversations than for the baseline, with a stated 82% preference for the Inner Thoughts condition.
Load-bearing premise
The load-bearing premise is that the next-speaker-prediction-plus-persona system is a fair and representative baseline; if it is weaker than typical proactive systems, the reported advantage may come from better response generation or extra LLM calls rather than from the inner-thought mechanism.
Editorial extensions
If this is right
- Group-chat assistants could engage without being named, contributing when a thought clears the motivation bar.
- Next-speaker prediction may be better reframed as a two-level task: detect whether a turn is open or allocated, then let content motivation decide participation.
- Adjustable proactivity parameters give users direct control over how talkative, assertive, or reserved an AI is.
- Because the evaluation criteria are explicit and score-based, the same machinery can double as an automatic, explainable benchmark for proactive dialogue quality.
- The framework's scope extends to task-oriented settings such as brainstorming and negotiation by swapping in goal-aligned evaluation criteria.
Reading between the lines
- The 82% preference should be read with caution: the baseline's response generation was tied to fixed personas, so part of the gap may come from weaker content rather than from the thought-gating mechanism itself.
- A testable extension would hold the response-generation step fixed and vary only the participation decision, isolating the contribution of intrinsic-motivation scoring.
- If the thought reservoir were surfaced to users, it could act as a transparency interface, letting people see and veto the AI's unspoken reasoning—a direction the paper raises but does not test.
- The near-chance self-selection results suggest a domain ceiling for history-based turn-taking models; any gains from richer multimodal cues may still need an internal motivation model to convert attention into appropriate speech.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the Inner Thoughts framework, in which a conversational AI continuously generates a covert stream of thoughts in parallel to an ongoing multi-party conversation, retrieves relevant long-term memories, evaluates each thought's 'intrinsic motivation' to be expressed using LLM-based scoring over heuristics derived from a 24-participant think-aloud study, and then decides when and how to participate on the basis of thresholds and turn-taking type. The authors instantiate this framework in a playground web app and a Slackbot, report a technical evaluation of simulated conversations against a next-speaker-prediction-plus-persona baseline, and report a user study comparing three proactivity settings of the framework.
Significance. The paper addresses an important and underexplored problem: how a conversational agent should decide to self-select in multi-party social conversations, where next-speaker prediction is known to be difficult. The think-aloud study and the resulting ten heuristics are a useful qualitative contribution, and the two open-sourced, deployed implementations make the framework concrete and reproducible. The human raters' preference for the Inner Thoughts conversations over the baseline in the simulated evaluation is encouraging evidence for the approach. However, the central comparative claim that the framework 'significantly surpasses existing baselines' currently rests on a single simulated comparison in which the two arms are not matched for model capability, number of LLM calls, or parameter tuning. The paper is therefore best read as presenting a promising framework with an initial proof-of-concept evaluation rather than a fully established superiority claim.
major comments (4)
- [§6.2, Table 2, Figure 7] The baseline condition is specified as a fine-tuned GPT-3.5 model for next-speaker prediction with persona-based response generation, while the Inner Thoughts condition is described only as 'the framework described in section 5' with no LLM backend stated. If condition 2 runs on a stronger model (e.g., GPT-4-class) or makes many more LLM calls per turn, then the observed advantages in anthropomorphism, coherence, intelligence, and turn-taking appropriateness, as well as the 82% preference, could be driven by generation quality or prompt engineering rather than by the inner-thought mechanism. This is a load-bearing point for the paper's headline claim; the authors should either report the exact backend used for condition 2 and run a matched-model ablation, or temper the superiority claim accordingly.
- [§6.3 and §5.5] The simulation parameters listed in §6.3 ('Overt proactivity = 3.95, Covert proactivity = 0.1') are inconsistent with the definitions in §5.5, where overt proactivity is the system1Prob parameter (0-1) and covert proactivity is the imThreshold parameter (1-5). The values appear to be swapped, and 3.95 cannot be a valid system1Prob. This suggests that the simulation configuration needs an audit. Since participation behavior depends directly on these thresholds and the authors state in §5 that the hyperparameters are chosen empirically, the paper should report the exact parameter settings used in the simulations and ideally include a sensitivity analysis.
- [§5.4 and §6.6] The thought-evaluation stage uses an LLM to score each thought's intrinsic motivation with criteria and five-level definitions that were themselves derived from a think-aloud study, and the same class of LLM also generates the thoughts. This creates a partial self-consistency loop: the model's judgments about its own thoughts may be coherent without corresponding to human judgments of appropriate participation. Although the human preference data in §6.6 provide independent evidence, the paper does not report whether human raters agree with the LLM's intrinsic-motivation scores. A small validation study comparing LLM motivation scores with human ratings of the same thoughts would substantially strengthen this component.
- [§7 and Abstract] The user study in §7 compares three proactivity settings of the Inner Thoughts framework (Non-stop Chatter, Active Contributor, Selective Participant) and does not include the next-speaker-prediction baseline. Consequently, the abstract's statement that the framework 'significantly surpasses existing baselines' is supported only by the simulated evaluation in §6. The authors should either add a baseline condition to the user study or revise the abstract and introduction to specify that the superiority claim is based on the technical evaluation of simulated conversations, while the user study demonstrates that different proactivity settings are perceptible and that the Active Contributor setting is preferred over the other two.
minor comments (4)
- [§5.4] The word 'calculaed' appears in the sentence describing the score formula; it should be 'calculated'.
- [§6.6 and Figure 7] The stacked bar plot in Figure 7 would benefit from error bars or a point plot showing the distribution of ratings, because the Mann-Whitney U tests are reported on the underlying Likert responses but the figure does not convey variance.
- [§7.2] The identification accuracy rates in §7.3.2 are reported against a 33.3% chance baseline, but with 12 participants the confidence intervals are wide; the authors should be careful in describing these as evidence that users can reliably distinguish the three styles.
- [§5.4] The notation 'd_p = lambda(t - tau_p)' is introduced twice in the same section; once for memory decay and once for the silence-based motivation increase, with different decay-rate values. Using distinct symbols or names would avoid confusion.
Circularity Check
No circularity: the headline comparison rests on independent human ratings; tuning and baseline gaps are limitations, not construction-level circularity.
full rationale
The paper's central comparative claim is an empirical head-to-head: Inner Thoughts (condition 2) versus next-speaker prediction plus persona (condition 1), judged by human raters on seven Likert items and a forced preference choice. No equation in the paper defines the outcome metric as a formal function of the framework's inputs. The intrinsic-motivation score in Section 5.4 is a weighted combination of LLM output-token probabilities and a silence factor, but the human ratings in Table 3 are collected independently from the simulation logs, so the reported significance is not a formal consequence of the score definition. The heuristics (relevance, coherence, etc.) are used both to decide participation and to define what natural conversation looks like, which creates a design tautology—the system is optimized for relevant/coherent contributions and is then rated as relevant/coherent—but this is not a circular derivation because the LLM evaluator's ratings are not the same measurement as the human Likert ratings, and the baseline is not constrained to use those heuristics. Section 3's next-speaker-prediction failure is a standalone experiment on MPC with reported accuracy numbers, not a fitted parameter reused as the evaluation outcome. Section 6.2's baseline uses the fine-tuned GPT-3.5 from Section 3, but the predicted quantity is the next speaker, not the human quality ratings; using that model as a baseline can be criticized as weak or unmatched (condition 2's backend is unspecified), but that is a validity confound, not circularity. The paper's own limitation statements—Section 5: 'we chose the hyperparameters described in the following paragraphs empirically' and 'We recognize conducting formal ablation studies an important direction for future work'; Section 8.4: 'thresholds are adjusted through trial and error'—concern parameter tuning and missing ablations, not a result that is true by construction. Section 7 compares three proactivity styles without the headline baseline, so it does not create a circular claim. Thus no circular step meeting the quoted-reduction bar is present.
Assumptions & free parameters
free parameters (9)
- saliency threshold =
0.3
- retrieval decay rate lambda =
0.95
- motivation increase rate lambda (dp) =
1.02
- pause trigger duration =
10 seconds
- system1Prob =
0.7 / 0.2 / 0 across personas; unclear in simulation settings
- imThreshold =
4.49, 3.59, 4.09 across personas
- interruptThreshold =
4.8 / 5.0
- number of thoughts per trigger =
1 System 1, 2 System 2
- evaluation response count =
5
assumptions (6)
- domain assumption Sacks et al.'s turn-taking systematics: turn-allocation, self-selection, and continuance rules govern when humans speak.
- domain assumption Dual-process theory (System 1 fast/automatic, System 2 slow/deliberate) describes human thought formation.
- ad hoc to paper The 10 heuristics derived from the 24-participant think-aloud study are a valid model of intrinsic motivation to speak.
- ad hoc to paper An LLM's self-evaluated 'intrinsic motivation' score is a reliable proxy for human judgments of appropriate participation.
- domain assumption MPC corpus communicative link annotations distinguish turn-allocation from self-selection.
- ad hoc to paper Longer silence monotonically increases motivation to speak.
invented entities (2)
-
Inner thought
-
Thought reservoir
Cite this review
Pith. "Pith review of Proactive Conversational Agents with Inner Thoughts." pith.science (2026). https://pith.science/paper/EQBTO5RQ
@misc{pith2026250100383,
author = {Pith},
title = {Pith review of: Proactive Conversational Agents with Inner Thoughts},
year = {2026},
howpublished = {\url{https://pith.science/paper/EQBTO5RQ}},
note = {Machine review of arXiv:2501.00383}
}
read the original abstract
One of the long-standing aspirations in conversational AI is to allow them to autonomously take initiatives in conversations, i.e., being proactive. This is especially challenging for multi-party conversations. Prior NLP research focused mainly on predicting the next speaker from contexts like preceding conversations. In this paper, we demonstrate the limitations of such methods and rethink what it means for AI to be proactive in multi-party, human-AI conversations. We propose that just like humans, rather than merely reacting to turn-taking cues, a proactive AI formulates its own inner thoughts during a conversation, and seeks the right moment to contribute. Through a formative study with 24 participants and inspiration from linguistics and cognitive psychology, we introduce the Inner Thoughts framework. Our framework equips AI with a continuous, covert train of thoughts in parallel to the overt communication process, which enables it to proactively engage by modeling its intrinsic motivation to express these thoughts. We instantiated this framework into two real-time systems: an AI playground web app and a chatbot. Through a technical evaluation and user studies with human participants, our framework significantly surpasses existing baselines on aspects like anthropomorphism, coherence, intelligence, and turn-taking appropriateness.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 3 Pith papers
-
ProactiveVA: Proactive Visual Analytics with LLM-Based UI Agent
An LLM-based UI agent monitors visual analytics interactions, detects when users struggle, infers their intent, and proactively provides context-aware guidance.
-
RetroChat: Designing for the Preservation of Past Digital Experiences
An LLM chat agent prompted with 2000-2010 Chinese BBS dialogue and deployed in a restored MSN window elicited nostalgic memory flashbacks and language adaptation from participants familiar with that era.
-
Amplifying Minority Voices: AI-Mediated Devil's Advocate System for Inclusive Group Decision-Making
An LLM-powered devil's advocate that paraphrases minority members' private dissents as its own messages could reduce social pressure and increase opinion diversity in group decisions, but the paper provides no user st...
Reference graph
Works this paper leans on
-
[1]
James E Allen, Curry I Guinn, and Eric Horvtz. 1999. Mixed-initiative interaction. IEEE Intelligent Systems and their Applications 14, 5 (1999), 14–23
work page 1999
-
[2]
Salvatore Andolina, Valeria Orso, Hendrik Schneider, Khalil Klouche, Tuukka Ruotsalo, Luciano Gamberini, and Giulio Jacucci. 2018. Investigating Proactive Search Support in Conversations. In Proceedings of the 2018 Designing Interactive Systems Conference (Hong Kong, China) (DIS ’18) . ACM, 1295–1307. https: //doi.org/10.1145/3196709.3196734
arXiv 2018
-
[3]
Salvatore Andolina, Valeria Orso, Hendrik Schneider, Khalil Klouche, Tuukka Ruotsalo, Luciano Gamberini, and Giulio Jacucci. 2018. SearchBot: Supporting Voice Conversations With Proactive Search. In Companion of the 2018 ACM Conference on Computer Supported Cooperative Work and Social Computing(Jersey City, NJ, USA) (CSCW ’18). ACM, 9–12. https://doi.org/...
-
[4]
Christoph Bartneck, Dana Kulić, Elizabeth Croft, and Susana Zoghbi. 2009. Mea- surement instruments for the anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety of robots. International journal of social robotics 1 (2009), 71–81
work page 2009
-
[5]
Daniel E. Berlyne. 1960. Conflict, Arousal, and Curiosity. McGraw-Hill
work page 1960
-
[6]
Keping Bi, Qingyao Ai, and W Bruce Croft. 2021. Asking clarifying questions based on negative feedback in conversational search. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval . 157–166
work page 2021
-
[7]
Dan Bohus and Eric Horvitz. 2009. Learning to predict engagement with a spoken dialog system in open-world settings. In Proceedings of the SIGDIAL 2009 Conference. 244–252
work page 2009
-
[8]
Dan Bohus and Eric Horvitz. 2009. Models for multiparty engagement in open- world dialog. In Proceedings of the SIGDIAL 2009 conference, the 10th annual meeting of the special interest group on discourse and dialogue . 10
work page 2009
Show all 76 references
-
[9]
Dan Bohus and Eric Horvitz. 2011. Multiparty turn taking in situated dialog: Study, lessons, and directions. In Proceedings of the SIGDIAL 2011 Conference . 98–109
2011
-
[10]
Simone Borsci, Alessio Malizia, Martin Schmettow, Frank Van Der Velde, Gunay Tariverdiyeva, Divyaa Balaji, and Alan Chamberlain. 2022. The chatbot usability scale: the design and pilot of a usability scale for interaction with AI-based conversational agents. Personal and ubiqu...
2022
-
[11]
Paul T Brady. 1968. A statistical analysis of on-off patterns in 16 conversations. Bell System Technical Journal 47, 1 (1968), 73–91
1968
-
[12]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101
2006
-
[13]
Penelope Brown. 1987. Politeness: Some universals in language usage . Vol. 4. Cambridge university press
1987
-
[14]
Shuo-yiin Chang, Bo Li, Tara N Sainath, Chao Zhang, Trevor Strohman, Qiao Liang, and Yanzhang He. 2022. Turn-taking prediction for natural conversational speech. arXiv preprint arXiv:2208.13321 (2022)
2022 arXiv
-
[15]
Maira Gatti de Bayser, Paulo Cavalin, Claudio Pinhanez, and Bianca Zadrozny
-
[16]
Yang Deng, Wenqiang Lei, Wai Lam, and Tat-Seng Chua. 2023. A survey on proactive dialogue systems: Problems, methods, and prospects. arXiv preprint arXiv:2305.02750 (2023)
2023 arXiv
-
[17]
Starkey Duncan. 1972. Some signals and rules for taking speaking turns in conversations. Journal of personality and social psychology 23, 2 (1972), 283
1972
-
[18]
Starkey Duncan Jr and George Niederehe. 1974. On signalling that it’s your turn to speak. Journal of experimental social psychology 10, 3 (1974), 234–247
1974
-
[19]
Paul Ekman and Wallace V Friesen. 1969. The repertoire of nonverbal behavior: Categories, origins, usage, and coding. semiotica 1, 1 (1969), 49–98
1969
-
[20]
Erik Ekstedt and Gabriel Skantze. 2020. Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog.arXiv preprint arXiv:2010.10874 (2020)
2020 arXiv
-
[21]
Jonathan St BT Evans. 2008. Dual-processing accounts of reasoning, judgment, and social cognition. Annu. Rev. Psychol. 59, 1 (2008), 255–278
2008
-
[22]
Cecilia E Ford and Sandra A Thompson. 1996. Interactional units in conversation: Syntactic, intonational, and pragmatic resources for the management of turns. Studies in interactional sociolinguistics 13 (1996), 134–184
1996
-
[23]
Jianfeng Gao, Michel Galley, and Lihong Li. 2019. Neural approaches to conversa- tional AI: Question answering, task-oriented dialogues and social chatbots . Now Foundations and Trends
2019
-
[24]
Erving Goffman. 1967. Interaction Ritual: Essays on Face-to-Face Behavior . Dou- bleday
1967
-
[25]
H. P. Grice. 1975. Logic and conversation. In Syntax and semantics . Vol. 3. Academic Press, 41–58
1975
-
[26]
Zhiwei Guan, Shirley Lee, Elisabeth Cuddihy, and Judith Ramey. 2006. The validity of the stimulated retrospective think-aloud method as measured by eye tracking. In Proceedings of the SIGCHI conference on Human Factors in computing systems. 1253–1262
2006
-
[27]
Charles T Hemphill, John J Godfrey, and George R Doddington. 1990. The ATIS spoken language systems pilot corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990
1990
-
[28]
Karen Holtzblatt and Hugh Beyer. 1997. Contextual design: defining customer- centered systems. Elsevier
1997
-
[29]
Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. https://doi.org/10.1145...
1999
-
[30]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021)
2021 arXiv
-
[31]
who will be next speaker and when?
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, and Junji Yamato. 2014. Analysis of respiration for prediction of" who will be next speaker and when?" in multi- party meetings. In Proceedings of the 16th international conference on multimodal interaction. 18–25
2014
-
[32]
Donghyun Kim, Youbin Ahn, Wongyu Kim, Chanhee Lee, Kyungchan Lee, Kyong- Ho Lee, Jeonguk Kim, Donghoon Shin, and Yeonsoo Lee. 2023. Persona expansion with commonsense knowledge for diverse and consistent response generation. In Proceedings of the 17th Conference of the Europea...
2023
-
[33]
Rachna Konigari, Saurabh Ramola, Vijay Vardhan Alluri, and Manish Shrivastava
-
[34]
Kenichi Kumatani, Sankaran Panchapagesan, Minhua Wu, Minjae Kim, Nikko Strom, Gautam Tiwari, and Arindam Mandai. 2017. Direct modeling of raw audio with dnns for wake word detection. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 252–257
2017
-
[35]
John E Laird. 2019. The Soar cognitive architecture . MIT press
2019
-
[36]
Lizi Liao, Grace Hui Yang, and Chirag Shah. 2023. Proactive conversational agents in the post-chatgpt world. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 3452–3455
2023
-
[37]
Bruce" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang
Xingyu "Bruce" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang "Anthony" Chen, and Ruofei Du. 2023. Visual Captions: Augmenting Verbal Communication with On-the-fly Visuals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamb...
2023
-
[38]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634 (2023)
2023 arXiv
-
[39]
Zeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che, and Ting Liu. 2020. Towards conversational recommendation over multi-type dialogs. arXiv preprint arXiv:2005.03954 (2020)
2020 arXiv
-
[40]
John K Local, John Kelly, and William HG Wells. 1986. Towards a phonology of conversation: turn-taking in Tyneside English1. Journal of Linguistics 22, 2 (1986), 411–437
1986
-
[41]
David H McFarland. 2001. Respiratory markers of conversational interaction. (2001)
2001
-
[42]
John P Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa, Pang Wei Koh, Vaishaal Shankar, Percy Liang, Yair Carmon, and Ludwig Schmidt. 2021. Ac- curacy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In International ...
2021
-
[43]
OpenAI. 2024. Learning to Reason with LLMs. (September 2024). https://openai. com/index/learning-to-reason-with-llms/
2024
-
[44]
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22
2023
-
[45]
Christopher Peters. 2005. Direction of attention perception for conversation initiation in virtual environments. In International Workshop on Intelligent Virtual Agents. Springer, 215–228
2005
-
[46]
Christopher Peters, Catherine Pelachaud, Elisabetta Bevacqua, Maurizio Mancini, and Isabella Poggi. 2005. A model of attention and interest using gaze behavior. In International Workshop on Intelligent Virtual Agents . Springer, 229–240
2005
-
[47]
Xuhui Ren, Hongzhi Yin, Tong Chen, Hao Wang, Zi Huang, and Kai Zheng
-
[48]
Bradley Rhodes and Thad Starner. 1996. Remembrance Agent: A Continuously Running Automated Information Retrieval System. In The Proceedings of the First International Conference on the Practical Application of Intelligent Agents and Multi Agent Technology, Vol. 1. ACM, 487–495
1996
-
[49]
Frank E Ritter, Farnaz Tehranchi, and Jacob D Oury. 2019. ACT-R: A cognitive architecture for modeling cognition. Wiley Interdisciplinary Reviews: Cognitive Science 10, 3 (2019), e1488
2019
-
[50]
In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval
Learning to ask appropriate questions in conversational recommendation. In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval. 808–817
-
[51]
Jurgen Ruesch, Gregory Bateson, Eve C Pinsker, and Gene Combs. 2017. Com- munication: The social matrix of psychiatry . Routledge
2017
-
[52]
Harvey Sacks, Emanuel A Schegloff, and Gail Jefferson. 1974. A simplest system- atics for the organization of turn-taking for conversation. language 50, 4 (1974), 696–735
1974
-
[53]
Sonia Roccas, Lilach Sagiv, Shalom H Schwartz, and Ariel Knafo. 2002. The big five personality factors and personal values. Personality and social psychology bulletin 28, 6 (2002), 789–801. CHI ’25, April 26-May 1, 2025, Yokohama, Japan Xingyu Bruce Liu, Shitao Fang, Weiyan Sh...
2002
-
[54]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[55]
Gabriel Skantze. 2021. Turn-taking in conversational systems and human-robot interaction: a review. Computer Speech & Language 67 (2021), 101178
2021
-
[56]
Samira Shaikh, Tomek Strzalkowski, George Aaron Broadwell, Jennifer Stromer- Galley, Sarah M Taylor, and Nick Webb. 2010. MPC: A Multi-Party Chat Corpus for Modeling Social Phenomena in Discourse.. In LREC. Citeseer
2010
-
[57]
Jianheng Tang, Tiancheng Zhao, Chenyan Xiong, Xiaodan Liang, Eric P Xing, and Zhiting Hu. 2019. Target-guided open-domain conversation. arXiv preprint arXiv:1905.11553 (2019)
2019 arXiv
-
[58]
Louis Ten Bosch, Nelleke Oostdijk, and Lou Boves. 2005. On temporal aspects of turn taking in conversational dialogues. Speech Communication 47, 1-2 (2005), 80–86
2005
-
[59]
Lucille Alice Suchman. 1987. Plans and situated actions: The problem of human- machine communication. Cambridge university press
1987
-
[60]
David Traum, Antonio Roque, Anton Leuski, Panayiotis Georgiou, Jillian Gerten, Bilyana Martinovski, Shrikanth Narayanan, Susan Robinson, and Ashish Vaswani
-
[61]
Marilyn Walker and Steve Whittaker. 1995. Mixed initiative in dialogue: An investigation into discourse segmentation. arXiv preprint cmp-lg/9504007 (1995)
1995 arXiv
-
[62]
Naftali Tishby, Fernando C Pereira, and William Bialek. 2000. The information bottleneck method. arXiv preprint physics/0004057 (2000)
2000 arXiv
-
[63]
Jimmy Wei, Kurt Shuster, Arthur Szlam, Jason Weston, Jack Urbanek, and Mojtaba Komeili. 2023. Multi-party chat: Conversational agents in group settings with humans and models. arXiv preprint arXiv:2304.13835 (2023)
2023 arXiv
-
[64]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837
2022
-
[65]
Joseph Weizenbaum. 1966. ELIZA—a Computer Program for the Study of Natural Language Communication between Man and Machine. Commun. ACM 9, 1 (jan 1966), 36–45. https://doi.org/10.1145/365153.365168
1966
-
[66]
Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel. 2021. Towards mutual theory of mind in human-ai interaction: How language reflects what students perceive about a virtual teaching assistant. In Proceedings of the 2021 CHI conference on human factors in co...
2021
-
[67]
Jun Xu, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che, and Ting Liu
-
[68]
Sanae Yamashita, Koji Inoue, Ao Guo, Shota Mochizuki, Tatsuya Kawahara, and Ryuichiro Higashinaka. 2023. Realpersonachat: A realistic persona chat corpus with interlocutors’ own personalities. In Proceedings of the 37th Pacific Asia Conference on Language, Information and Comp...
2023
-
[69]
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[70]
Allison Woodruff and Paul M Aoki. 2003. How push-to-talk makes talk less pushy. In Proceedings of the 2003 ACM International Conference on Supporting Group Work. 170–179
2003
-
[71]
Saizheng Zhang. 2018. Personalizing dialogue agents: I have a dog, do you have pets too. arXiv preprint arXiv:1801.07243 (2018). 10 Appendix 10.1 System Figure 9 shows the Inner Thoughts playground settings panel. 10.2 User Evaluation 10.2.1 Proactivity Settings. (1) Non-stop ...
2018 arXiv
-
[75]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629 (2022)
2022 arXiv
-
[2007]
In Proceedings of the 8th SIGdial Workshop on Discourse and Dialogue
Hassan: A virtual human for tactical questioning. In Proceedings of the 8th SIGdial Workshop on Discourse and Dialogue . 71–74
-
[2019]
arXiv preprint arXiv:1907.02090 (2019)
Learning multi-party turn-taking models from dialogue logs. arXiv preprint arXiv:1907.02090 (2019)
2019 arXiv
-
[2020]
In Proceedings of the 58th annual meeting of the association for computational linguistics
Conversational graph grounded policy learning for open-domain conver- sation generation. In Proceedings of the 58th annual meeting of the association for computational linguistics. 1835–1845
-
[2021]
InProceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue
Topic shift detection for mixed initiative response. InProceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue . 161–166
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.