REVIEW 4 major objections 6 minor 115 references
Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper reports that in VR conversations with LLM-powered agents, response delays above four seconds degrade users' perceived response time and experience, while natural conversational fillers—thinking gestures plus filler…
desk verdict A well-run, fully interactive VR study of latency and fillers whose headline claims are a bit broader than the pairwise data will bear—still worth a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing intervention is the natural conversational filler: a multimodal delay-mitigation cue in which the agent performs a "thinking" gesture (head turn, chin touch, subtle breathing) and speaks a filler voice line during the response delay. It works by occupying the wait with social cues that mimic human deliberation, and the paper measures its effect through a six-item post-condition survey whose first item asks whether the agent "was quick to start responding meaningfully."
What would settle it
Run the same study with a natural filler that uses only the thinking gesture and no voice line; if the perceived-response-time advantage disappears, the effect came from participants counting the filler phrase as the start of the answer rather than from a genuine reduction in perceived waiting.
Extended reading notes
Core claim
This paper claims that response delay is a first-class usability problem for free-form, LLM-powered embodied conversational agents, and that the right interface-level remedy is a natural conversational filler rather than a conventional loading indicator. In a within-subjects VR study with 54 participants conversing with nine agents across three scenarios, delay degraded perceived response time and broader perception metrics, with effects becoming pronounced at 4.0s and 6.5s. Natural fillers—a thinking gesture plus a filler phrase such as "Hmm, let's see..."—significantly improved perceived response time at both of those delay levels compared with no filler, supporting H2a; artificial wait indicators did not produce significant improvement, and neither filler type rescued the broader experience metrics.
Load-bearing premise
The main measurement assumption is that the survey question about the agent being "quick to start responding meaningfully" measures perceived latency and not the fact that, in the Natural condition, the agent audibly speaks during the wait, so participants may treat the filler phrase itself as the start of the reply.
Editorial extensions
If this is right
- Designers of LLM-powered conversational agents should target system response times under four seconds, because longer delays significantly degrade perceived response time, engagement, impression, competence, and willingness to interact again.
- When delays of four to 6.5 seconds are unavoidable, natural conversational fillers (gesture plus filler phrase) can recover perceived response time, though they do not restore the broader perception metrics.
- Artificial loading icons and processing sounds should not be relied on as a latency-mitigation strategy, since the study found no significant benefit over no filler.
- Studies using LLM-based free-form agents should report response latency, because high delays bias user perceptions and limit comparability across results.
- The open-source deployment pipeline makes the system reproducible for future VR agent studies.
Reading between the lines
- The gesture-versus-voice split in participant preferences suggests a personalization opportunity: a single filler design may not serve all users, and offering selectable filler modality could strengthen mitigation.
- The null result for artificial wait indicators may be specific to generic spinners; more communicative indicators that signal content (for example, showing that the agent is checking a record or that a response stage is underway) might behave differently, and the paper's recommendation to explore such indicators is a natural next test.
- The four-second threshold is measured in a leisurely task-guided VR setting; time-pressure or high-stakes contexts such as emergency or medical conversation may compress tolerance, so the threshold is likely an upper bound rather than a universal constant.
- In non-embodied channels such as phone voice assistants or text chatbots, the natural filler's gestural component is absent, so the voice-line alone may be the transferable part; the paper's embodied results do not directly establish that transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a within-subjects VR experiment (N=54) in which participants held free-form, task-guided conversations with nine LLM-powered embodied agents under three response-latency levels (1.5s, 4.0s, 6.5s) and three filler conditions (None, Artificial wait indicator, Natural conversational filler). The authors use ART ANOVA with Holm-Bonferroni-corrected ART-C contrasts and find that latency significantly degrades perceived response time and several broader perception ratings, that Natural fillers significantly improve the single-item perceived-response-time rating at Medium and High latency (H2a), and that Artificial wait indicators do not produce significant effects (H3a/H3b unsupported). The paper interprets these results as showing that response delays above 4 seconds degrade quality of experience, that natural fillers mitigate perceived delay, and it contributes an open-source Unity/ASR/LLM/TTS pipeline for deploying conversational agents in VR.
Significance. The study is a well-powered full-factorial experiment in a realistic LLM-based conversational pipeline, and the open-source release is a practical contribution to VR conversational-agent research. Strengths include a power analysis, balanced-Latin-square counterbalancing, appropriate nonparametric ART ANOVA with Holm-Bonferroni correction, and a credible interactive system with a measured 1.5s SRT. If the H2a finding survives a cleaner measure of perceived latency, it would usefully extend prior WoZ and pre-recorded filler results to real-time LLM-driven free-form conversation. However, the central interpretation depends on a single questionnaire item whose wording does not fully separate filler speech from response speech, and the 'above 4 seconds' threshold claim goes beyond the sampled latency grid. The paper also makes summary claims in the Discussion and Conclusion that are broader than the significant results.
major comments (4)
- [§3.2.1, §3.1.2, §5.2] The main positive result (H2a) is measured by a single item, Q1: 'From the moment I stopped talking, the agent was quick to start responding meaningfully.' In the Natural condition the agent utters a semantically interpretable filler line ('Hmm, let me think about that...') and performs a thinking gesture immediately after the user stops talking, so the objective time to the first audible speech is shorter in Natural than in None or Artificial. The word 'meaningfully' was intended to prevent participants from counting filler speech, but no manipulation check in §3.2 or §4.2 verifies that they did so; PSQ6 asks whether fillers were noticed but is not used to test this interpretation. As a result, the Medium/High Q1 advantage of Natural over None (p < 0.01 and p < 0.0001 in §5.2) is also consistent with participants treating the filler utterance as the start of the response, making the effect partly definitional rather than a genuine improvement in perceived latency. Please add a manipulation check (for example, a post-condition item that distinguishes 'the agent said something while thinking' from 'the agent answered my question') or re-analyze with a measure that isolates the final response onset, and adjust the H2a claim accordingly.
- [§5.3, Figure 5] The abstract, §5.3, and the design recommendations state that response delays 'above 4 seconds' degrade quality of experience, but the experiment samples only 1.5s, 4.0s, and 6.5s. Q1 shows all three pairwise delay differences, but the broader perception metrics (Q2-Q6) show significant degradation only between Low and High, with Q2 the only broader metric distinguishing Medium from High and no broader metric distinguishing Low from Medium. The data therefore establish that 6.5s is worse than 1.5s and that 4.0s and 6.5s are perceived as slower than 1.5s on Q1, but they do not locate a threshold 'above 4 seconds.' Please soften the threshold claim to 'the two longer delays tested' or add a latency grid that brackets the boundary.
- [§5.2, §7] The Discussion and Conclusion go beyond the supported effects. The only significant filler effects are on Q1 at Medium and High latency; no filler effect reaches significance on Q2-Q6, and the paper appropriately reports that H2b, H3a, and H3b are unsupported. Yet §7 concludes that Natural fillers 'enhance VR user experience' and 'reduce latency's negative repercussions,' and §5.2 states that fillers 'improve tolerance for delayed responses.' Please align these summary claims with the measured single-item Q1 effect and the broader null results, and avoid implying that the study demonstrates improvement in global user experience.
- [§4.2.5] The claim that conversational fillers increase willingness to use the slowest system rests on a comparison of PSQ5 and PSQ10, which are repeated measures from the same 54 participants. The reported chi-square test treats the two sets of responses as independent; a paired analysis (e.g., a McNemar test on collapsed agree/disagree categories or a paired ordinal model) is needed because the same participants answer both questions. Please re-analyze these data and adjust §5.2 if the conclusion changes.
minor comments (6)
- [§1] The phrase 'free-from conversation' appears to be a typo for 'free-form conversation'; please correct it.
- [§4.1] The set of 36 Holm-Bonferroni-adjusted contrasts should be defined explicitly in the text or supplementary material; currently the reader cannot tell which family of comparisons the correction controls.
- [§3.2, §6] The custom questionnaires are appropriately acknowledged in Section 6 as unvalidated, but the single-item Q1 and the aggregated RoSAS-based Q4/Q5 would benefit from a brief psychometric justification or a pilot validation, especially because Q1 carries the main result.
- [Figures 4-8] In the copy under review, the text inside Figures 4-8 appears as long unreadable token sequences (e.g., '/uni000000...' strings); please ensure the final PDF contains clean figure graphics with readable labels.
- [§4.2.6] The inductive analysis of text justifications is described as preliminary, but no inter-rater reliability or coding agreement is reported; a brief statement on the coding process would help readers assess the thematic counts.
- [Table 3] The observed filler effect sizes are small (η²p = 0.017-0.081), whereas the power analysis targeted a medium-to-large effect; the null results for Q2-Q6 should be explicitly framed as 'not supported' rather than as evidence of no effect, which the text mostly does but could state more consistently.
Circularity Check
Q1's 'responding meaningfully' does not fully separate the Natural filler's spoken utterance from the start of a meaningful response, so the H2a result is partly definitional.
-
self definitional
[§3.2.1 (Q1), §3.1.2 (Natural filler), §5.2 (H2a result)]
"Q1 assessed perceived response latency and included the word "meaningfully" to avoid bias toward Natural fillers that included speech. (Q1) Response Time: From the moment I stopped talking, the agent was quick to start responding meaningfully. In the Natural condition, each agent randomly selected from three 'thinking' gestures and six filler voice lines at runtime."
The load-bearing H2a result rests on Q1, which asks whether the agent was quick to start responding meaningfully. In the Natural condition, the agent speaks a filler voice line immediately at the onset of the wait, before the delayed LLM response. These filler lines are not semantically empty: "Hmm, let me think about that..." acknowledges the user and signals deliberation. Thus the objective time from the participant stopping talking to the agent's first meaningful speech is shorter in the Natural condition than in None or Artificial by construction. Participants who count the filler as the beginning of a meaningful response will rate Q1 higher regardless of perceived waiting time.
full rationale
This is an empirical VR user study with no fitted parameters, no uniqueness theorems, and no mathematical derivation chain, so most circularity patterns do not apply. Self-citations are present (e.g., the authors' prior pilot study and the VALID avatar library) but are not load-bearing: the experiment is self-contained, and the main latency results are statistical comparisons from Likert responses. The one genuine circularity risk is measurement-level and concerns H2a. The outcome Q1 asks whether "the agent was quick to start responding meaningfully," and the Natural filler condition makes the agent speak a semantically meaningful filler line immediately when the wait begins. Because filler lines like "Hmm, let me think about that..." acknowledge the user and signal deliberation, participants may count the filler as the start of a meaningful response, making the Natural condition objectively shorter on the very construct Q1 measures. The paper added the word "meaningfully" to avoid this bias, showing awareness, but it reports no manipulation check confirming that participants excluded filler speech. The Q1 advantage of Natural fillers at Medium and High latency is thus at least partly definitional rather than pure evidence of perceived-latency mitigation. This does not undermine H1a or the delay main effects, which do not involve fillers, and the broader perception results (H2b not supported) are unaffected because those questions do not reference response onset. Overall, one partial measurement-level circularity in the central filler claim warrants a score of 4.
Assumptions & free parameters
assumptions (3)
- domain assumption The 1.5s, 4.0s, and 6.5s delay levels span natural, comfortable, and excessive conversational pauses.
- ad hoc to paper The Q1 wording 'quick to start responding meaningfully' separates the filler utterance from the actual response.
- standard math Aligned Rank Transform ANOVA is appropriate for these five-point Likert responses.
Cite this review
Pith. "Pith review of Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents." pith.science (2026). https://pith.science/paper/WGDWWY6U
@misc{pith2026250722352,
author = {Pith},
title = {Pith review of: Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/WGDWWY6U}},
note = {Machine review of arXiv:2507.22352}
}
read the original abstract
We investigated the challenges of mitigating response delays in free-form conversations with virtual agents powered by Large Language Models (LLMs) within Virtual Reality (VR). For this, we used conversational fillers, such as gestures and verbal cues, to bridge delays between user input and system responses and evaluate their effectiveness across various latency levels and interaction scenarios. We found that latency above 4 seconds degrades quality of experience, while natural conversational fillers improve perceived response time, especially in high-delay conditions. Our findings provide insights for practitioners and researchers to optimize user engagement whenever conversational systems' responses are delayed by network limitations or slow hardware. We also contribute an open-source pipeline that streamlines deploying conversational agents in virtual environments.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Mohammad Rafayet Ali, Seyedeh Zahra Razavi, Raina Langevin, Abdullah Al Mamun, Benjamin Kane, Reza Rawassizadeh, Lenhart K. Schubert, and Ehsan Hoque. 2020. A Virtual Conversational Agent for Teens with Autism Spectrum Disorder: Experimental Results and Design Lessons. In Proceed- ings of the 20th ACM International Conference on Intelligent Virtual Agents...
arXiv 2020
-
[2]
T. S. Amer and Todd L. Johnson. 2016. Information Technology Progress Indi- cators: Temporal Expectancy, User Preference, and the Perception of Process Duration. International Journal of Technology and Human Interaction (IJTHI) 12, 4 (Oct. 2016), 1—-14. https://doi.org/10.4018/IJTHI.2016100101
-
[4]
Jana Appel, Astrid von der Pütten, Nicole C. Krämer, and Jonathan Gratch. 2012. Does Humanity Matter? Analyzing the Importance of Social Cues and Perceived Agency of a Computer System for the Emergence of Social Reactions during Human-Computer Interaction. Advances in Human-Computer Interaction 2012, 1 (2012), 324694. https://doi.org/10.1155/2012/324694
-
[5]
Bailey, Jeremy N
Jakki O. Bailey, Jeremy N. Bailenson, Jelena Obradović, and Naomi R. Aguiar
-
[6]
Janet Beavin Bavelas and Nicole Chovil. 1997. Faces in Dialogue. In The Psychology of Facial Expression , James A. Russell and José Miguel Fernández- Dols (Eds.). Cambridge University Press, Cambridge, 334–346. https://doi.org/ 10.1017/CBO9780511659911.017
-
[7]
Rojin Bayat, Elios De Maio, Jacopo Fiorenza, Massimo Migliorini, and Fab- rizio Lamberti. 2024. Exploring Methodologies to Create a Unified VR User- Experience in the Field of Virtual Museum Experiences. In 2024 IEEE Gaming, Entertainment, and Media Conference (GEM) . IEEE, Turin, Italy, 1–4. https: //doi.org/10.1109/GEM61861.2024.10585452 ISSN: 2766-6530
arXiv 2024
-
[8]
Baylor and Rinat B
Amy L. Baylor and Rinat B. Rosenberg-Kima. 2006. Interface agents to allevi- ate online frustration, In Proceedings of the 7th International Conference on Learning Sciences (Bloomington, Indiana). Proceedings of ICLS 2006 1, 30–36. https://repository.isls.org//handle/1/3514
2006
-
[9]
Štefan Beňuš and Marián Trnka. 2014. Prosody, Voice Assimilation, and Conversational Fillers. In Speech Prosody 2014 . ISCA, Online, 75–79. https: //doi.org/10.21437/SpeechProsody.2014-3 Conversational Agents’ Response Latency Mitigation CUI ’25, July 8–10, 2025, Waterloo, ON, Canada
Show all 115 references
-
[10]
Andrea J. Bingham. 2023. From Data Management to Actionable Findings: A Five-Phase Process of Qualitative Data Analysis. International Journal of Quali- tative Methods 22 (Jan. 2023), 1–11. https://doi.org/10.1177/16094069231183620 Publisher: SAGE Publications Inc
2023 doi
-
[11]
Heather Bortfeld, Silvia D Leon, Jonathan E Bloom, Michael F Schober, and Susan E Brennan. 2001. Disfluency Rates in Conversation: Effects of Age, Relationship, Topic, Role, and Gender. Language and Speech 44, 2 (2001), 123–
2001
-
[12]
Halim-Antoine Boukaram, Micheline Ziadee, and Majd F Sakr. 2021. Mitigating the Effects of Delayed Virtual Agent Response Time Using Conversational Fillers. In Proceedings of the 9th International Conference on Human-Agent Interaction (HAI ’21). Association for Computing Machi...
2021
-
[13]
Russell J Branaghan and Christopher A Sanchez. 2009. Feedback Preferences and Impressions of Waiting. Human Factors 51, 4 (2009), 528–538. https: //doi.org/10.1177/0018720809345684 PMID: 19899362
2009 doi
-
[14]
Carpinella, Alisa B
Colleen M. Carpinella, Alisa B. Wyman, Michael A. Perez, and Steven J. Stroess- ner. 2017. The Robotic Social Attributes Scale (RoSAS): Development and Validation. In Proceedings of the 2017 ACM/IEEE International Conference on Human-Robot Interaction (HRI ’17) . Association f...
2017
-
[15]
Llogari Casas, Samantha Hannah, and Kenny Mitchell. 2024. MoodFlow: Orches- trating Conversations with Emotionally Intelligent Avatars in Mixed Reality. In 2024 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW). IEEE, Orlando, FL, USA, 86–...
2024
-
[16]
Florian Charlier, Marc Weber, Dariusz Izak, Emerson Harkin, Marcin Magnus, Joseph Lalli, Louison Fresnais, Matt Chan, Nikolay Markov, Oren Amsalem, Sebastian Proost, Agamemnon Krasoulis, getzze, and Stefan Repplinger. 2022. Statannotations. Statannotations Contributors. https:...
2022 doi
-
[17]
Anping Cheng, Dongming Ma, Hao Qian, and Younghwan Pan. 2024. The Effects of Mobile Applications’ Passive and Interactive Loading Screen Types on Waiting Experience. Behaviour & Information Technology 43, 8 (2024), 1652–1663. https://doi.org/10.1080/0144929X.2023.2224901
2024
-
[18]
Vuthea Chheang, Shayla Sharmin, Rommy Márquez-Hernández, Megha Pa- tel, Danush Rajasekaran, Gavin Caulfield, Behdokht Kiafar, Jicheng Li, Pinar Kullu, and Roghayeh Leila Barmaki. 2024. Towards Anatomy Education with Generative AI-based Virtual Assistants in Immersive Virtual R...
2024
-
[19]
Nicholas Christenfeld. 1995. Does It Hurt to Say Um? Journal of Nonverbal Behavior 19 (1995), 171–186. https://doi.org/10.1007/BF02175503
1995 doi
-
[20]
Herbert H Clark. 1996. Using language. Cambridge University Press, Cambridge, UK. https://doi.org/10.1017/CBO9780511620539
1996 doi
-
[21]
Frederick G Conrad, Mick P Couper, Roger Tourangeau, and Andy Peytchev
-
[22]
Creem-Regehr, Jeanine K
Sarah H. Creem-Regehr, Jeanine K. Stefanucci, and Bobby Bodenheimer. 2022. Perceiving Distance in Virtual Reality: Theoretical Insights from Contemporary Technologies. Philosophical Transactions of the Royal Society B: Biological Sciences 378, 1869 (Dec. 2022), 1–12. https://d...
2022
-
[23]
de Melo, Kangsoo Kim, Nahal Norouzi, Gerd Bruder, and Gregory Welch
Celso M. de Melo, Kangsoo Kim, Nahal Norouzi, Gerd Bruder, and Gregory Welch. 2020. Reducing Cognitive Load and Improving Warfighter Problem Solving With Intelligent Virtual Assistants. Frontiers in Psychology 11 (2020), 12 pages. https://doi.org/10.3389/fpsyg.2020.554706
2020
-
[24]
Divekar*, Jaimie Drozdal*, Samuel Chabot*, Yalun Zhou, Hui Su, Yue Chen, Houming Zhu, James A
Rahul R. Divekar*, Jaimie Drozdal*, Samuel Chabot*, Yalun Zhou, Hui Su, Yue Chen, Houming Zhu, James A. Hendler, and Jonas Braasch. 2022. Foreign Language Acquisition via Artificial Intelligence and Extended Reality: Design and Evaluation. Computer Assisted Language Learning 3...
2022
-
[25]
Do, Steve Zelenty, Mar Gonzalez-Franco, and Ryan P
Tiffany D. Do, Steve Zelenty, Mar Gonzalez-Franco, and Ryan P. McMahan
-
[26]
Filippo Domaneschi, Marcello Passarelli, and Carlo Chiorri. 2017. Facial Expres- sions and Speech Acts: Experimental Evidences on the Role of the Upper Face as an Illocutionary Force Indicating Device in Language Comprehension.Cognitive Processing 18, 3 (Aug. 2017), 285–306. h...
2017 doi
-
[27]
Elkin, Matthew Kay, James J
Lisa A. Elkin, Matthew Kay, James J. Higgins, and Jacob O. Wobbrock. 2021. An Aligned Rank Transform Procedure for Multifactor Contrast Tests. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association for Computing ...
2021
-
[28]
Franz Faul, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009. Sta- tistical Power Analyses Using G*Power 3.1: Tests for Correlation and Re- gression Analyses. Behavior Research Methods 41, 4 (Nov 2009), 1149–1160. https://doi.org/10.3758/BRM.41.4.1149
2009 doi
-
[29]
Kotaro Funakoshi, Kazuki Kobayashi, Mikio Nakano, Seiji Yamada, Yasuhiko Kitamura, and Hiroshi Tsujino. 2008. Smoothing Human-Robot Speech In- teractions by Using a Blinking-light as Subtle Expression. In Proceedings of the 10th international conference on Multimodal interface...
2008
-
[30]
Markus Funk, Carie Cunningham, Duygu Kanver, Christopher Saikalis, and Rohan Pansare. 2020. Usable and Acceptable Response Delays of Conversa- tional Agents in Automotive User Interfaces. In 12th International Conference on Automotive User Interfaces and Interactive Vehicular ...
2020
-
[32]
Ulrich Gnewuch, Stefan Morana, Marc Adam, and Alexander Maedche. 2018. Faster is Not Always Better: Understanding the Effect of Dynamic Response Delays in Human-Chatbot Interaction. In European Conference on Information Systems. AIS Electronic Library (AISeL), Portsmouth, UK, ...
2018
-
[33]
The Chatbot is typing
Ulrich Gnewuch, Stefan Morana, Marc Adam, and Alexander Maedche. 2018. “The Chatbot is typing ... ” – The Role of Typing Indicators in Human-Chatbot Interaction. In SIGHCI 2018 Proceedings . AIS Electronic Library (AISeL), San Francisco, CA, 1–5. https://aisel.aisnet.org/sighci2018/14
2018
-
[34]
Sanchez-Vives, Jeremy Bailenson, Mel Slater, and Jaron Lanier
Mar Gonzalez-Franco, Eyal Ofek, Ye Pan, Angus Antley, Anthony Steed, Bern- hard Spanlang, Antonella Maselli, Domna Banakou, Nuria Pelechano, Sergio Orts-Escolano, Veronica Orvalho, Laura Trutoiu, Markus Wojcik, Maria V. Sanchez-Vives, Jeremy Bailenson, Mel Slater, and Jaron La...
2020
-
[35]
Chris Harrison, Brian Amento, Stacey Kuznetsov, and Robert Bell. 2007. Re- thinking the Progress Bar. In Proceedings of the 20th annual ACM symposium on User interface software and technology (UIST ’07). Association for Computing Ma- chinery, New York, NY, USA, 115–118. https:...
2007
-
[36]
Chris Harrison, Zhiquan Yeo, and Scott E. Hudson. 2010. Faster Progress Bars: Manipulating Perceived Duration with Visual Augmentations. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’10) . Association for Computing Machinery, New York, NY,...
2010
-
[37]
Masum Hasan, Cengiz Ozel, Sammy Potter, and Ehsan Hoque. 2023. SAPIEN: Affective Virtual Agents Powered by Large Language Models*. In 2023 11th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW) . IEEE, Cambridge, MA, USA, 1...
2023
-
[38]
Mattias Heldner and Jens Edlund. 2010. Pauses, Gaps and Overlaps in Conver- sations. Journal of Phonetics 38, 4 (Oct. 2010), 555–568. https://doi.org/10.1016/ j.wocn.2010.08.002
2010
-
[39]
Jess Hohenstein, Hani Khan, Kramer Canfield, Samuel Tung, and Rocio Perez Cano. 2016. Shorter Wait Times: The Effects of Various Loading Screens on Perceived Performance. In Proceedings of the 2016 CHI Confer- ence Extended Abstracts on Human Factors in Computing Systems (CHI ...
2016
-
[40]
T. M. Holtgraves, S. J. Ross, C. R. Weywadt, and T. L. Han. 2007. Perceiving Artificial Social Agents. Computers in Human Behavior 23, 5 (Sept. 2007), 2163–
2007
-
[41]
Kate Hone. 2006. Empathic Agents to Reduce User Frustration: The Effects of Varying Agent Characteristics. Interacting with computers 18, 2 (2006), 227–245. https://doi.org/10.1016/j.intcom.2005.05.003
2006 doi
-
[42]
Jacob Hornik. 1984. Subjective vs. Objective Time Measures: A Note on the Perception of Time in Consumer Behavior. Journal of consumer research 11, 1 (1984), 615–618. https://doi.org/10.1086/208998
1984 doi
-
[43]
John Hoxmeier and Chris DiCesare. 2000. System Response Time and User Satisfaction: An Experimental Study of Browser-based Applications. , 6 pages. https://aisel.aisnet.org/amcis2000/347
2000
-
[44]
Qian Hu and Zhao Pan. 2024. Is Cute AI More Forgivable? The Impact of Informal Language Styles and Relationship Norms of Conversational Agents on Service Recovery. Electronic Commerce Research and Applications 65 (May 2024), 1–14. https://doi.org/10.1016/j.elerap.2024.101398
2024
-
[45]
Zuwen Huang and Ada Lo. 2025. Human vs. Robot Service Provider Agents in Service Failures: Comparing Customer Dissatisfaction and The Mediating Role of Forgiveness and Service Recovery Expectation. Information Technology & Tourism 27 (Feb. 2025), 1–32. https://doi.org/10.1007/...
2025 doi
-
[46]
Sun Young Hwang, Negar Khojasteh, and Susan R. Fussell. 2019. When Delayed in a Hurry: Interpretations of Response Delays in Time-Sensitive Instant Mes- saging. Proceedings of the ACM on Human-Computer Interaction 3, GROUP (Dec. 2019), 1–20. https://doi.org/10.1145/3361115
2019 doi
-
[47]
Together But Not To- gether
Zainab Iftikhar, Yumeng Ma, and Jeff Huang. 2023. “Together But Not To- gether”: Evaluating Typing Indicators for Interaction-Rich Communication. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Compu...
2023
-
[48]
Joseph Jaffe and Stanley Feldstein. 1970. Rhythms of Dialogue. Academic Press, New York, NY, USA. https://cir.nii.ac.jp/crid/1130282269381241984
1970
-
[49]
Yuin Jeong, Juho Lee, and Younah Kang. 2019. Exploring Effects of Conver- sational Fillers on User Perception of Conversational Agents. In Extended Ab- stracts of the 2019 CHI Conference on Human Factors in Computing Systems (CHI EA ’19). Association for Computing Machinery, N...
2019
-
[50]
Takayuki Kanda, Masayuki Kamasima, Michita Imai, Tetsuo Ono, Daisuke Sakamoto, Hiroshi Ishiguro, and Yuichiro Anzai. 2007. A Humanoid Robot that Pretends to Listen to Route Guidance from a Human. Auton. Robots 22, 1 (Jan. 2007), 87–100. https://doi.org/10.1007/s10514-006-9007-6
2007 doi
-
[51]
Woojoo Kim, Shuping Xiong, and Zhuoqian Liang. 2017. Effect of Loading Symbol of Online Video on Perception of Waiting Time. International Journal of Human–Computer Interaction 33, 12 (2017), 1001–1009. https://doi.org/10. 1080/10447318.2017.1305051
2017
-
[52]
Takanori Komatsu, Chenxi Xie, and Seiji Yamada. 2024. Waiting Time Percep- tions for Faster Count-downs/ups Are More Sensitive Than Slower Ones: Exper- imental Investigation and Its Application. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’2...
2024
-
[53]
Junyeong Kum and Myungho Lee. 2022. Can Gestural Filler Reduce User- Perceived Latency in Conversation with Digital Humans? Applied Sciences 12, 21 (Jan. 2022), 10972. https://doi.org/10.3390/app122110972
2022 doi
-
[54]
Bailenson, and Caitlin Mills
Vishal Kiran Kuvar, Jeremy N. Bailenson, and Caitlin Mills. 2024. A novel quantitative assessment of engagement in virtual reality: Task-unrelated thought is reduced compared to 2D videos.Computers & Education 209 (Feb. 2024), 104959. https://doi.org/10.1016/j.compedu.2023.104959
2024
-
[55]
Christos Kyrlitsias and Despina Michael-Grigoriou. 2022. Social Interaction With Agents and Avatars in Immersive Virtual Environments: A Survey. Frontiers in Virtual Reality 2 (Jan. 2022), 13 pages. https://doi.org/10.3389/frvir.2021.786665 Publisher: Frontiers
2022
-
[56]
Lafferty, P.M
J.C. Lafferty, P.M. Eady, A.W. Pond, and Human Synergistics. 1974. The Desert Survival Problem: A Group Decision Making Experience for Examin- ing and Increasing Individual and Team Effectiveness: Manual . Experimen- tal Learning Methods, Plymouth, Michigan: Experimental Learn...
1974
-
[57]
Asif Ali Laghari, Hui He, Muhammad Shafiq, and Asiya Khan. 2019. Application of Quality of Experience in Networked Services: Review, Trend & Perspectives. Systemic Practice and Action Research 32, 5 (Oct. 2019), 501–519. https://doi. org/10.1007/s11213-018-9471-x
2019 doi
-
[58]
Levinson and Francisco Torreira
Stephen C. Levinson and Francisco Torreira. 2015. Timing in Turn-taking and its Implications for Processing Models of Language. Frontiers in Psychology 6 (2015), 731. https://doi.org/10.3389/fpsyg.2015.00731
2015
-
[59]
Zijian Lew, Joseph B Walther, Augustine Pang, and Wonsun Shin. 2018. Interac- tivity in Online Chat: Conversational Contingency and Response Latency in Computer-mediated Communication. Journal of Computer-Mediated Communi- cation 23, 4 (July 2018), 201–221. https://doi.org/10....
2018 doi
-
[60]
Shasha Li and Chien-Hsiung Chen. 2019. The Effect of Progress Indicator Speeds on Users’ Time Perceptions and Experience of a Smartphone User Interface. In Human-Computer Interaction. Recognition and Interaction Technologies: Thematic Area, HCI 2019, Held as Part of the 21st H...
2019 doi
-
[61]
Shasha Li and Chien-Hsiung Chen. 2019. The Effects of Visual Feedback Designs on Long Wait Time of Mobile Application User Interface. Interacting with Computers 31, 1 (Jan. 2019), 1–12. https://doi.org/10.1093/iwc/iwz001
2019 doi
-
[62]
Jose Llanes-Jurado, Lucía Gómez-Zaragozá, Maria Eleonora Minissi, Mariano Alcañiz, and Javier Marín-Morales. 2024. Developing Conversational Virtual Humans for Social Emotion Elicitation Based on Large Language Models.Expert Systems with Applications 246 (Jul 2024), 123261. ht...
2024
-
[63]
Soledad López Gambino, Sina Zarrieß, and David Schlangen. 2017. Beyond On-hold Messages: Conversational Time-buying in Task-oriented Dialogue. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue , Kristiina Jokinen, Manfred Stede, David DeVault, and Ann...
2017 doi
-
[64]
Soledad López Gambino, Sina Zarrieß, and David Schlangen. 2019. Testing Strategies For Bridging Time-To-Content In Spoken Dialogue Systems. In 9th International Workshop on Spoken Dialogue System Technology , Luis Fernando D’Haro, Rafael E. Banchs, and Haizhou Li (Eds.). Sprin...
2019 doi
-
[65]
Redowan Mahmud, Satish Narayana Srirama, Kotagiri Ramamohanarao, and Rajkumar Buyya. 2019. Quality of Experience (QoE)-aware Placement of Appli- cations in Fog Computing Environments. J. Parallel and Distrib. Comput. 132 (Oct. 2019), 190–203. https://doi.org/10.1016/j.jpdc.2018.03.004
2019 doi
-
[66]
LaViola Jr
Mykola Maslych, Christian Pumarada, Amirpouya Ghasemaghaei, and Joseph J. LaViola Jr. 2024. Takeaways from Applying LLM Capabilities to Multiple Conversational Avatars in a VR Pilot Study. arXiv:2501.00168 [cs.HC]
2024 arXiv
-
[68]
David Matsumoto and Hyisung C. Hwang. 2018. Microexpressions Differentiate Truths From Lies About Future Malicious Intent. Frontiers in Psychology 9 (Dec. 2018), 2545. https://doi.org/10.3389/fpsyg.2018.02545
2018
-
[69]
Thomas McWilliams, Bryan Reimer, Bruce Mehler, Jonathan Dobres, and Hale McAnulty. 2015. A Secondary Assessment of the Impact of Voice Interface Turn Delays on Driver Attention and Arousal in Field Conditions.Driving Assessment Conference 8, 2015 (June 2015), 408–414. https://...
2015
-
[70]
Robert B. Miller. 1968. Response Time in Man-computer Conversational Trans- actions. In Proceedings of the December 9-11, 1968, fall joint computer conference, part I on - AFIPS ’68 (Fall, part I) . ACM Press, San Francisco, California, 267. https://doi.org/10.1145/1476589.1476628
1968
-
[71]
Brad A Myers. 1985. The Importance of Percent-Done Progress Indicators for Computer-Human Interfaces. ACM SIGCHI Bulletin 16, 4 (1985), 11–17. https://doi.org/10.1145/1165385.317459
1985
-
[72]
Matthias Müller-Brockhausen, Giulio Barbero, and Mike Preuss. 2023. Chatter Generation through Language Models. In 2023 IEEE Conference on Games (CoG) . IEEE, Boston, MA, USA, 6. https://doi.org/10.1109/CoG57401.2023.10333244 ISSN: 2325-4289
2023
-
[73]
Andreea Niculescu, Betsy van Dijk, Anton Nijholt, Dilip Kumar Limbu, Swee Lan See, and Alvin Hong Yee Wong. 2010. Socializing with Olivia, the Youngest Robot Receptionist Outside the Lab. In Social Robotics, Shuzhi Sam Ge, Haizhou Li, John-John Cabibihan, and Yeow Kee Tan (Eds...
2010 doi
-
[74]
Naoki Ohshima, Keita Kimijima, Junji Yamato, and Naoki Mukawa. 2015. A Con- versational Robot with Vocal and Bodily Fillers for Recovering from Awkward Silence at Turn-takings. In 2015 24th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN) . IE...
2015
-
[75]
Xueni Pan and Antonia F. de C. Hamilton. 2018. Why and How to Use Virtual Reality to Study Human Social Interaction: The Challenges of Exploring a New Research Landscape. British Journal of Psychology 109, 3 (2018), 395–417. https://doi.org/10.1111/bjop.12290
2018 doi
-
[76]
O’Brien, Carrie J
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442 [cs.HC]
2023 arXiv
-
[77]
Gustav Bøg Petersen, Aske Mottelson, and Guido Makransky. 2021. Pedagogical Agents in Educational VR: An in the Wild Study. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21) . Association for Computing Machinery, New York, NY, USA, 1–12....
2021
-
[78]
Fischer, Stuart Reeves, and Sarah Sharples
Martin Porcheron, Joel E. Fischer, Stuart Reeves, and Sarah Sharples. 2018. Voice Interfaces in Everyday Life. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . ACM, Montreal QC Canada, 1–12. https://doi.org/10.1145/3173574.3174214
2018
-
[79]
Hua Xuan Qin, Shan Jin, Ze Gao, Mingming Fan, and Pan Hui. 2024. Char- acterMeet: Supporting Creative Writers’ Entire Story Character Construc- tion Processes Through Conversation with LLM-Powered Chatbot Avatars. In Proceedings of the CHI Conference on Human Factors in Comput...
2024
-
[80]
Amir Bani Saeed, Zahra Moussavi, and Bruce Hardy. 2024. Developing an Avatar in Virtual Reality for Mental Health Treatment. CMBES Proceedings 46 (June 2024), 1–1. https://proceedings.cmbes.ca/index.php/proceedings/article/view/ 1107
2024
-
[81]
Junaid Shaikh, Markus Fiedler, and Denis Collange. 2010. Quality of Experience from User and Network Perspectives. Annals of Telecommunications - Annales Des Télécommunications 65, 1 (Feb. 2010), 47–57. https://doi.org/10.1007/s12243- 009-0142-x
2010 doi
-
[82]
Horowitz
Nicole Shechtman and Leonard M. Horowitz. 2003. Media Inequality in Conver- sation: How People Behave Differently When Interacting with Computers and Conversational Agents’ Response Latency Mitigation CUI ’25, July 8–10, 2025, Waterloo, ON, Canada People. In Proceedings of the...
2003
-
[83]
Toshiyuki Shiwa, Takayuki Kanda, Michita Imai, Hiroshi Ishiguro, and Norihiro Hagita. 2009. How Quickly Should a Communication Robot Respond? Delaying Strategies and Habituation Effects. International Journal of Social Robotics 1, 2 (April 2009), 141–155. https://doi.org/10.10...
2009 doi
-
[84]
Alon Shoa, Ramon Oliva, Mel Slater, and Doron Friedman. 2023. Sushi with Einstein: Enhancing Hybrid Live Events with LLM-Based Virtual Humans. In Proceedings of the 23rd ACM International Conference on Intelligent Virtual Agents (IV A ’23). Association for Computing Machinery,...
2023
-
[85]
Gabriel Skantze and Anna Hjalmarsson. 2013. Towards Incremental Speech Generation in Conversational Systems. Computer Speech & Language 27, 1 (2013), 243–262. https://doi.org/10.1016/j.csl.2012.05.004
2013 doi
-
[86]
Sruti Srinidhi, Edward Lu, and Anthony Rowe. 2024. XaiR: An XR Platform that Integrates Large Language Models with the Physical World. In 2024 IEEE Inter- national Symposium on Mixed and Augmented Reality (ISMAR) . IEEE, Bellevue, WA, USA, 759–767. https://doi.org/10.1109/ISMA...
2024
-
[87]
Ian Steenstra, Farnaz Nouraei, Mehdi Arjmand, and Timothy Bickmore. 2024. Virtual Agents for Alcohol Use Counseling: Exploring LLM-Powered Moti- vational Interviewing. In Proceedings of the ACM International Conference on Intelligent Virtual Agents . ACM, GLASGOW United Kingdo...
2024
-
[88]
Tanya Stivers, N. J. Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter de Ruiter, Kyung-Eun Yoon, and Stephen C. Levinson. 2009. Universals and Cultural Variation in Turn-taking in Conversation. Proceedings...
2009 doi
-
[89]
Jan Svartvik (Ed.). 1990. The London–Lund Corpus of Spoken English: Description and Research. Lund Studies in English, Vol. 82. Lund University Press, Lund, Sweden. https://lup.lub.lu.se/record/c9ccd3ca-4a6e-4885-9ca5-1f939baa977f
1990
-
[90]
Marc Swerts. 1998. Filled Pauses as Markers of Discourse Structure. Journal of pragmatics 30, 4 (1998), 485–496. https://doi.org/10.1016/S0378-2166(98)00014-9
1998 doi
-
[91]
Maite Taboada. 2006. Spontaneous and Non-spontaneous Turn-taking. Prag- matics. Quarterly Publication of the International Pragmatics Association (IPrA) 16, 2-3 (2006), 329–360. https://doi.org/10.1075/prag.16.2-3.04tab
2006 doi
-
[92]
Templeton, Luke J
Emma M. Templeton, Luke J. Chang, Elizabeth A. Reynolds, Marie D. Cone LeBeaumont, and Thalia Wheatley. 2022. Fast Response Times Sig- nal Social Connection in Conversation. Proceedings of the National Acad- emy of Sciences of the United States of America 119, 4 (Jan. 2022), e...
2022 doi
-
[93]
Oguzhan Topsakal and Elif Topsakal. 2022. Framework for A Foreign Language Teaching Software for Children Utilizing AR, Voicebots and ChatGPT (Large Language Models). The Journal of Cognitive Systems 7, 2 (Dec 2022), 33–38. https://doi.org/10.52876/jcs.1227392
2022 doi
-
[94]
Vivian Tsai, Timo Baumann, Florian Pecune, and Justine Cassell. 2019. Faster Re- sponses Are Better Responses: Introducing Incrementality into Sociable Virtual Personal Assistants. In 9th International Workshop on Spoken Dialogue System Technology, Luis Fernando D’Haro, Rafael...
2019 doi
-
[95]
Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana, Noor Ahsan, Nevasini Sasikumar, Omkar Thawakar, Henok Biadglign Ademtew, Yahya Hmaiti, Aman- deep Kumar, and Kartik Kuckreja et al. 2024. All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages. arXiv:2411...
2024 arXiv
-
[96]
Ana Villar, Mario Callegaro, and Yongwei Yang. 2013. Where am I? A Meta- analysis of Experiments on the Effects of Progress Indicators for Web Surveys. Social Science Computer Review 31, 6 (2013), 744–762. https://doi.org/10.1177/ 0894439313497468
2013
-
[97]
Hongyu Wan, Jinda Zhang, Abdulaziz Arif Suria, Bingsheng Yao, Dakuo Wang, Yvonne Coady, and Mirjana Prpa. 2024. Building LLM-based AI Agents in Social Virtual Reality. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA ’24) . Associa...
2024
-
[98]
Zhan Wang, Lin-Ping Yuan, Liangwei Wang, Bingchuan Jiang, and Wei Zeng
-
[99]
Joseph Weizenbaum. 1966. ELIZA—a computer program for the study of natural language communication between man and machine. Commun. ACM 9, 1 (Jan. 1966), 36–45. https://doi.org/10.1145/365153.365168
1966
-
[100]
Chatbots
David Westerman, Aaron C. Cross, and Peter G. Lindmark. 2019. I Believe in a Thing Called Bot: Perceptions of the Humanness of “Chatbots”. Communica- tion Studies 70, 3 (May 2019), 295–312. https://doi.org/10.1080/10510974.2018. 1557233
2019
-
[101]
Noel Wigdor, Joachim de Greeff, Rosemarijn Looije, and Mark A Neerincx
-
[102]
Wobbrock, Leah Findlater, Darren Gergle, and James J
Jacob O. Wobbrock, Leah Findlater, Darren Gergle, and James J. Higgins
-
[103]
Charley Wu, Eric Schulz, Timothy Pleskac, and Maarten Speekenbrink. 2022. Time pressure changes how people explore and respond to uncertainty.Scientific Reports 12 (March 2022), 4122. https://doi.org/10.1038/s41598-022-07901-1
2022 doi
-
[104]
Takato Yamazaki, Tomoya Mizumoto, Katsumasa Yoshikawa, Masaya Ohagi, Toshiki Kawamoto, and Toshinori Sato. 2023. An Open-Domain Avatar Chatbot by Exploiting a Large Language Model. InProceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue...
2023 doi
-
[105]
Dorneich
Euijung Yang and Michael C. Dorneich. 2015. The Effect of Time Delay on Emotion, Arousal, and Satisfaction in Human-Robot Interaction. Proceedings of the Human Factors and Ergonomics Society Annual Meeting 59, 1 (Sept. 2015), 443–
2015
-
[106]
Fu-Chia Yang, Kevin Duque, and Christos Mousas. 2024. The Effects of Depth of Knowledge of a Virtual Agent. IEEE Transactions on Visualization and Computer Graphics 30, 11 (Nov. 2024), 7140–7151. https://doi.org/10.1109/TVCG.2024. 3456148
2024 doi
-
[107]
Difeng Yu, Qiushi Zhou, Benjamin Tag, Tilman Dingler, Eduardo Velloso, and Jorge Goncalves. 2020. Engaging Participants during Selection Studies in Virtual Reality. In 2020 IEEE Conference on Virtual Reality and 3D User Interfaces (VR) . IEEE, Atlanta, GA, USA, 500–509. https:...
2020
-
[108]
Jiarui Zhu, Radha Kumaran, Chengyuan Xu, and Tobias Höllerer. 2023. Free- form Conversation with Human and Symbolic Avatars in Mixed Reality. In2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR) . IEEE, Sydney, Australia, 751–760. https://doi.org/10.1109/...
2023
-
[147]
https://doi.org/10.1177/00238309010440020101 PMID: 11575901
-
[447]
https://doi.org/10.1177/1541931215591094 Publisher: SAGE Publications Inc
-
[2010]
Interacting with computers 22, 5 (2010), 417–427
The Impact of Progress Indicators on Task Completion. Interacting with computers 22, 5 (2010), 417–427. https://doi.org/10.1016/j.intcom.2010.03.001
2010 doi
-
[2011]
In Proceedings of the SIGCHI Conference on Hu- man Factors in Computing Systems (Vancouver, BC, Canada) (CHI ’11)
The aligned rank transform for nonparametric factorial analyses us- ing only anova procedures. In Proceedings of the SIGCHI Conference on Hu- man Factors in Computing Systems (Vancouver, BC, Canada) (CHI ’11) . As- sociation for Computing Machinery, New York, NY, USA, 143–146....
-
[2016]
In 2016 25th IEEE international symposium on robot and human interactive communication (RO-MAN)
How to Improve Human-robot Interaction with Conversational Fillers. In 2016 25th IEEE international symposium on robot and human interactive communication (RO-MAN). IEEE, IEEE, New York, NY, USA, 219–224. https: //doi.org/10.1109/ROMAN.2016.7745134
2016
-
[2019]
Journal of Applied Developmental Psychology 64 (July 2019), 101052
Virtual Reality’s Effect on Children’s Inhibitory Control, Social Compli- ance, and Sharing. Journal of Applied Developmental Psychology 64 (July 2019), 101052. https://doi.org/10.1016/j.appdev.2019.101052
2019
-
[2023]
Frontiers in Virtual Reality 4 (Nov
VALID: A Perceptually Validated Virtual Avatar Library for Inclusion and Diversity. Frontiers in Virtual Reality 4 (Nov. 2023), 15 pages. https: //doi.org/10.3389/frvir.2023.1248915
2023
-
[2024]
In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24)
VirtuWander: Enhancing Multi-modal Interaction for Virtual Tour Guid- ance through Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24) . Association for Computing Machinery, New York, NY, USA, 1–20. https://doi.org/10.114...
-
[2174]
https://doi.org/10.1016/j.chb.2006.02.017
2006 doi
-
[2360]
https://doi.org/10.1080/09588221.2021.1879162
2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.