REVIEW 3 major objections 5 minor 39 references
Behavioral and Symbolic Fillers as Delay Mitigation for Embodied Conversational Agents in Virtual Reality
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read In a VR job-interview study, a thinking animation with a vocal "Hhm" made a delayed conversational agent feel more human and natural than progress-bar indicators.
desk verdict Useful VR comparison of delay fillers with a real confound between filler type and character identity that the authors acknowledge but cannot fully rule out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the multimodal behavioral filler: a set of nine "thinking" animations (gaze aversion, filler gestures) combined with a non-lexical verbal filler, "Hhm," spoken with six pitches and occurring every three questions. This filler runs during the delay between the user's question and the agent's response. The comparison conditions are a base idle animation and two symbolic progress indicators (a badge-embedded bar and a floating thinking-bubble bar). The experimental machinery is a within-subject VR job interview with MetaHuman agents, pseudo-random delays of 2, 4, and 8 seconds, questionnaire measures of perceived response time, presence, and agent impression, plus continuo
What would settle it
Run the same study with each filler type paired with every character in a counterbalanced design while keeping delays identical; if the significant advantages of the behavioral filler disappear or shift across character pairings, the effect is character-driven, not filler-driven.
Extended reading notes
Core claim
The paper reports a within-subject VR study with 24 participants who held simulated job interviews with four embodied conversational agents, one per condition. During scripted response delays of 2, 4, or 8 seconds, the agent either played idle animations (BASE), performed nine thinking animations plus a non-lexical 'Hhm' vocalization with six pitches (BEHAVIORAL), showed a green progress bar embedded in a visitor badge (EMBEDDED), or showed a floating thinking bubble with a progress bar (EXTERNAL). The central claim is that the BEHAVIORAL filler made the perceived response time significantly more appropriate, raised parasocial interaction, engagement, and social realism, and improved humanli
Load-bearing premise
The four filler types were always paired with the same MetaHuman character each, so differences in character appearance or voice, rather than the filler itself, could drive the reported perceptual differences.
Editorial extensions
If this is right
- If the behavioral filler's effect holds, VR agents backed by LLMs can mask multi-second computation delays with low-cost animations and a brief vocal filler rather than loading UI.
- Symbolic progress indicators, even embedded ones, did not help users predict when the response would arrive; the gaze data suggests they redirect attention away from the agent's face, which may make response onset feel abrupt.
- The absence of significant differences between embedded and external progress bars implies that the placement of a symbolic indicator matters less than whether a symbolic indicator is used at all.
- For real deployments with unpredictable response times, behavioral fillers need an animation state machine that can extend or finish naturally, a design problem the paper explicitly identifies.
- For longer delays, behavioral fillers may lose their advantage because they provide no estimate of remaining wait time, so symbolic indicators or hybrids may be needed.
Reading between the lines
- The effect may be partly driven by the specific MetaHuman character paired with the behavioral filler, since filler type and character were not crossed in the design; a replication with rotated character-filler pairings would isolate the filler itself.
- If behavioral fillers work because they signal imminent speech, agents could deliberately time a finishing gesture to the arrival of the response, making even unpredictable delays feel natural.
- In high-stakes or information-retrieval tasks, symbolic indicators might be more appropriate than the job-interview scenario suggests; the paper itself speculates along these lines, but future work would need to test it directly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a within-subject VR experiment (N=24) in which participants interviewed four MetaHuman ECAs while the agent displayed one of four delay-mitigation strategies during a 2/4/8-s thinking pause: BASE idle animation, BEHAVIORAL thinking animations with non-lexical "Hhm" fillers, EMBEDDED progress bar in a badge, and EXTERNAL progress bar in a thinking bubble. The authors measure perceived response-time appropriateness/expectedness, presence subscales, social presence, impression-of-agent scales (humanlikeness, intelligence, likeability, willingness, competence, naturalness, assuredness), gaze on face, and final hiring preference. They report significant benefits of BEHAVIORAL over all other conditions on appropriateness, parasocial interaction, engagement, humanlikeness, and naturalness, and over the two symbolic conditions on social realism and face gaze; 16/24 participants preferred BEHAVIORAL. H2 (symbolic indicators improve expectedness) and H3 (embedded better than external) were not confirmed.
Significance. If the causal interpretation were valid, the result would be practically useful for designers of LLM-based ECAs in VR, showing that naturalistic behavioral fillers can mask response delays more effectively than progress bars. The work extends prior screen-based findings to immersive VR and directly compares behavioral and symbolic strategies. The manuscript is generally clear, uses appropriate nonparametric tests with post-hoc corrections, and makes data and code available, which aids reproducibility. However, the central causal claim is weakened by the full confounding of filler type with character identity, so the study currently supports a conditional rather than a general causal conclusion.
major comments (3)
- [§3.3, §5 (Limitations)] The four filler conditions are perfectly confounded with character identity: Section 3.3 states that two male and two female MetaHumans were created, and Section 5 acknowledges that 'each filler type was always displayed with the same character.' Since the dependent variables include humanlikeness, naturalness, parasocial interaction, and hiring preference, the significant BEHAVIORAL advantages reported in Tables 2–3 and the 16/24 preference could in principle be caused by the specific character's appearance, voice, or perceived attractiveness rather than by the filler. Counterbalancing the presentation order does not break this confound because it does not randomize the character–filler mapping. The authors' acknowledgment that appearance 'might have influenced our results' understates the problem: with the current design, the two explanations are indistinguishable. I recommend either a
- [§3.1, Table 2] The manuscript reports 13 subjective dependent variables and several eye-gaze measures without any correction for multiple testing across the family of measures. Some of the reported significant effects (e.g., parasocial interaction p=0.014, engagement p=0.016, naturalness p=0.013) would not survive a very conservative Bonferroni correction, although they would survive FDR. The consistency of the pattern mitigates the concern, but the number of 'significant' outcomes should be interpreted with caution. An explicit multiple-comparison strategy or pre-registered analysis plan would strengthen the claims; at minimum, report adjusted p-values or justify the unadjusted family-wise error rate.
- [§4.1, §5] The self-reported expectedness measure was not significant (p=0.066), but the authors use the gaze-at-response-onset results to argue that 'participants expected the response more naturally with the BEHAVIORAL fillers.' This is a post hoc interpretation, not a planned test. The gaze measure was not validated as an expectedness measure, and the claim goes beyond the data. Please label this as exploratory and remove it from the summary of hypothesis support.
minor comments (5)
- [§4.5] The preference counts are inconsistent: 16 (66.7%) + 3 (12.5%) + 1 (4.2%) + 2 (8.3%) = 22, not 24. Two participants are unaccounted for. Please clarify whether some participants did not state a preference or whether the counts/percentages contain an error.
- [Figure 4, Tables 2/4/5] Minor typos: 'Humalikeness' in Figure 4 and 'd f1 = d fe f f ect' in Table headers should read 'Humanlikeness' and 'df1 = df_effect'.
- [§3.3] Replication would benefit from more detail about the ECAs' voices (e.g., TTS engine, language/accent, whether all agents used the same voice) and about the animation state machine controlling the transition from filler to response. These details are also relevant to the character-confound concern.
- [§4.4] The relative gaze time on the face is reported as a proportion in Table 3 (e.g., 0.66) but as a percentage in Figure 3(g). Please use a consistent scale. Also clarify whether the 'cumulative average' smoothing affects the RGTF values reported.
- [§3.1] No sample size justification or power analysis is provided. Given the number of dependent variables, reporting confidence intervals for the main effect sizes would aid interpretation.
Circularity Check
No circularity: the paper is an empirical user study with outcome measures independent of the experimental manipulations; no claim reduces to its inputs by construction.
full rationale
This paper does not contain a derivation chain to audit for circularity. It reports a within-subjects VR user study (n=24) comparing four delay-mitigation conditions (BASE, BEHAVIORAL, EMBEDDED, EXTERNAL) on perceived response time, presence subscales, agent impression, gaze, and preference. The outcome measures are collected through questionnaires, eye-tracking, and preference choices; they are not computed from the experimental conditions, and no parameter is fitted to the data and then renamed as a prediction. Hypotheses H1-H3 are framed as predictions, but they are tested against independent Likert-scale responses, gaze ratios, and hiring/preference choices, so the conclusions do not reduce to the inputs by definition or construction. The paper's self-citations (e.g., prior work by Kum and Lee, Elfleet and Chollet) are used only as motivation and comparison, not as load-bearing evidence that forces the results. The most substantive weakness is a validity/confound issue explicitly acknowledged by the authors: each filler type was always paired with the same MetaHuman character, so character appearance cannot be excluded as an alternative explanation (Section 5). That is a legitimate experimental-design concern, but it is not circularity as defined here: the measured outcomes are still independent of the manipulation, and the causal claim, even if insecure, is not equivalent to its inputs by construction. Accordingly, no circular steps are identified and the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The Temple Presence Inventory (TPI) subscales and Godspeed questionnaire items validly measure presence, humanlikeness, and naturalness in VR ECA interactions.
- domain assumption The fixed pseudo-random delays (2, 4, 8 seconds) are representative of realistic LLM response delays in conversational VR agents.
- domain assumption The four MetaHuman characters are perceptually equivalent enough that condition-to-character assignment does not confound results.
- domain assumption The scripted interview responses were neutral and similar enough that response quality did not confound the filler effects.
Cite this review
Pith. "Pith review of Behavioral and Symbolic Fillers as Delay Mitigation for Embodied Conversational Agents in Virtual Reality." pith.science (2026). https://pith.science/paper/63QIIXOO
@misc{pith2026250811781,
author = {Pith},
title = {Pith review of: Behavioral and Symbolic Fillers as Delay Mitigation for Embodied Conversational Agents in Virtual Reality},
year = {2026},
howpublished = {\url{https://pith.science/paper/63QIIXOO}},
note = {Machine review of arXiv:2508.11781}
}
read the original abstract
When communicating with embodied conversational agents (ECAs) in virtual reality, there might be delays in the responses of the agents lasting several seconds, for example, due to more extensive computations of the answers when large language models are used. Such delays might lead to unnatural or frustrating interactions. In this paper, we investigate filler types to mitigate these effects and lead to a more positive experience and perception of the agent. In a within-subject study, we asked 24 participants to communicate with ECAs in virtual reality, comparing four strategies displayed during the delays: a multimodal behavioral filler consisting of conversational and gestural fillers, a base condition with only idle motions, and two symbolic indicators with progress bars, one embedded as a badge on the agent, the other one external and visualized as a thinking bubble. Our results indicate that the behavioral filler improved perceived response time, three subscales of presence, humanlikeness, and naturalness. Participants looked away from the face more often when symbolic indicators were displayed, but the visualizations did not lead to a more positive impression of the agent or to increased presence. The majority of participants preferred the behavioral fillers, only 12.5% and 4.2% favored the symbolic embedded and external conditions, respectively.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Adkins, A. Normoyle, L. Lin, Y . Sun, Y . Ye, M. Di Luca, and S. Jörg. How important are detailed hand motions for communication for a virtual character through the lens of charades? ACM Trans. Graph., 42(3), article no. 27, 16 pages, May 2023. doi: 10.1145/3578575 2
-
[2]
C. Bartneck, D. Kuli ´c, E. Croft, and S. Zoghbi. Measurement in- struments for the anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety of robots. International Journal of Social Robotics, 1(1):71–81, 2009. doi: 10.1007/s12369-008-0001-3 3
-
[3]
M. J. Blanca, J. Arnau, F. J. García-Castro, R. Alarcón, and R. Bono. Non-normal data in repeated measures ANOV A: Impact on type I error and power. Psicothema, 35(1):21–29, 2023. doi: 10.7334/ psicothema2022.292 6
work page 2023
-
[4]
H.-A. Boukaram, M. Ziadee, and M. F. Sakr. Mitigating the effects of delayed virtual agent response time using conversational fillers. In Proceedings of the 9th International Conference on Human-Agent In- teraction, HAI ’21, 9 pages, pp. 130–138, 2021. doi: 10.1145/3472307 .3484181 2, 3
-
[5]
J. Cassell. Embodied conversational agents: Representation and intelli- gence in user interfaces. AI magazine, 22(4):67–67, 2001. 1
work page 2001
-
[6]
R. B. Church and S. Goldin-Meadow. The mismatch between gesture and speech as an index of transitional knowledge. Cognition, 23(1):43– 71, 1986. 2
work page 1986
-
[7]
E. Coppa and I. Finocchi. On data skewness, stragglers, and mapreduce progress indicators. In Proceedings of the Sixth ACM Symposium on Cloud Computing, pp. 139–152, 2015. 8
work page 2015
-
[8]
J. Ehret, A. Bönsch, L. Aspöck, C. T. Röhr, S. Baumann, M. Grice, J. Fels, and T. W. Kuhlen. Do prosody and embodiment influence the perceived naturalness of conversational agents’ speech? ACM Trans. Appl. Percept., 18(4), article no. 21, 15 pages, Oct. 2021. doi: 10.1145/3486580 2
Show all 39 references
-
[9]
Elfleet and M
M. Elfleet and M. Chollet. Investigating the impact of multimodal feedback on user-perceived latency and immersion with llm-powered embodied conversational agents in virtual reality. In Proceedings of 8 Authors’ version. To appear in IEEE Transactions on Visualization and Comp...
2024
-
[10]
S. L. Gambino, S. Zarrieß, and D. Schlangen. Testing strategies for bridging time-to-content in spoken dialogue systems. In Proceedings of the Ninth International Workshop on Spoken Dialogue Systems Technology, pp. 1–7, 2018. 2
2018
-
[11]
Gnewuch, S
U. Gnewuch, S. Morana, M. Adam, and A. Maedche. Faster is not always better: Understanding the effect of dynamic response delays in human-chatbot interaction. 2018. 2, 8
2018
-
[12]
The chatbot is typing
U. Gnewuch, S. Morana, M. T. Adam, and A. Maedche. “The chatbot is typing...”–the role of typing indicators in human-chatbot interaction. In Proceedings of the Sixteenth Annual Pre-ICIS Workshop on HCI Research, 2018. 2
2018
-
[13]
Goble and C
H. Goble and C. E. and. A robot that communicates with vocal fillers has . . . uhhh . . . greater social presence. Communication Research Reports, 35(3):256–260, 2018. doi: 10.1080/08824096.2018.1447454 2
2018
-
[14]
Gronier and A
G. Gronier and A. Baudet. Does progress bars’ behavior influence the user experience in human-computer interaction. Psychology and Cognitive Sciences–Open Journal, 5(1):6–13, 2019. 2
2019
-
[15]
Jaffe, B
J. Jaffe, B. Beebe, S. Feldstein, C. L. Crown, M. D. Jasnow, P. Rochat, and D. N. Stern. Rhythms of dialogue in infancy: Coordinated timing in development. Monographs of the society for research in child development, pp. i–149, 2001. 2
2001
-
[16]
Jeong, J
Y . Jeong, J. Lee, and Y . Kang. Exploring effects of conversational fillers on user perception of conversational agents. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems , CHI EA ’19, 6 pages, pp. 1—-6, 2019. doi: 10.1145/3290607.3312913 1, 2
2019
-
[17]
K. Kim, L. Boelling, S. Haesler, J. Bailenson, G. Bruder, and G. F. Welch. Does a digital assistant need a body? The influence of visual embodiment and social behavior on the perception of intelligent virtual agents in AR. In 2018 IEEE International Symposium on Mixed and Augm...
2018
-
[18]
Kum and M
J. Kum and M. Lee. Can gestural filler reduce user-perceived latency in conversation with digital humans? Applied Sciences, 12(21):10972,
-
[19]
M. E. Latoschik, F. Kern, J.-P. Stauffert, A. Bartl, M. Botsch, and J.-L. Lugrin. Not alone here?! Scalability and user experience of embodied ambient crowds in distributed social virtual reality. IEEE Transactions on Visualization and Computer Graphics, 25(5):2134–2144, 2019. 2
2019
-
[20]
G. Liu, M. W. Liu, and Q. Zhu. Hmm, the effect of AI conversational fillers on consumer purchase intentions. Marketing Letters, pp. 1–13,
-
[21]
Lombard, T
M. Lombard, T. B. Ditton, and L. Weinstein. Measuring presence: The temple presence inventory. In Proceedings of the 12th International Workshop on Presence (PRESENCE’09), 01 2009. 3
2009
-
[22]
Lombard, L
M. Lombard, L. Weinstein, and T. B. Ditton. Measuring telepresence: The validity of the temple presence inventory (TPI) in a gaming context. In Proceedings of the 13th International Workshop on Presence , 2011. 3
2011
-
[23]
M. L. Lupetti, E. Hagens, W. Van Der Maden, R. Steegers-Theunissen, and M. Rousian. Trustworthy embodied conversational agents for healthcare: A design exploration of embodied conversational agents for the periconception period at Erasmus MC. In Proceedings of the 5th Internat...
2023
-
[24]
Maslych, C
M. Maslych, C. Pumarada, A. Ghasemaghaei, and J. J. LaViola Jr. Takeaways from applying LLM capabilities to multiple conversational avatars in a VR pilot study. arXiv preprint arXiv:2501.00168, 2025. 1, 2
2025 arXiv
-
[25]
McCloud and M
S. McCloud and M. Martin. Understanding comics: The invisible art , vol. 106. Kitchen sink press Northampton, MA, 1993. 2
1993
-
[26]
Mukawa, H
N. Mukawa, H. Sasaki, and A. Kimura. How do verbal/bodily fillers ease embarrassing situations during silences in conversations? In The 23rd IEEE International Symposium on Robot and Human Interactive Communication, pp. 30–35. IEEE, 2014. 2
2014
-
[27]
B. A. Myers. The importance of percent-done progress indicators for computer-human interfaces. ACM SIGCHI Bulletin, 16(4):11–17, 1985. 1, 2, 4
1985
-
[28]
Norouzi, K
N. Norouzi, K. Kim, J. Hochreiter, M. Lee, S. Daher, G. Bruder, and G. Welch. A systematic survey of 15 years of user studies published in the intelligent virtual agents conference. In Proceedings of the 18th International Conference on Intelligent Virtual Agents , IV A ’18, 6...
2018
-
[29]
K. L. Nowak and F. Biocca. The effect of the agency and anthropomor- phism on users’ sense of telepresence, copresence, and social presence in virtual environments. Presence: Teleoperators and Virtual Environ- ments, 12(5):481–494, 10 2003. doi: 10.1162/105474603322761289 3
2003 doi
-
[30]
Schmidt, O
S. Schmidt, O. Ariza, and F. Steinicke. Intelligent blended agents: Reality–virtuality interaction with artificially intelligent embodied vir- tual humans. Multimodal Technologies and Interaction, 4(4):85, 2020. 1
2020
-
[31]
Schmidt, T
S. Schmidt, T. Rolff, H. V oigt, M. Offe, and F. Steinicke. Natural expression of a machine learning model’s uncertainty through verbal and non-verbal behavior of intelligent virtual agents. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technol...
2024
-
[32]
Shiwa, T
T. Shiwa, T. Kanda, M. Imai, H. Ishiguro, and N. Hagita. How quickly should communication robots respond? In Proceedings of the 3rd ACM/IEEE international conference on Human robot interaction , pp. 153–160, 2008. 2
2008
-
[33]
Sonlu, U
S. Sonlu, U. Güdükbay, and F. Durupinar. A conversational agent framework with multi-modal personality expression. ACM Transac- tions on Graphics (TOG), 40(1):1–16, 2021. 1
2021
-
[34]
Villar, M
A. Villar, M. Callegaro, and Y . Yang. Where am i? a meta-analysis of experiments on the effects of progress indicators for web surveys. Social Science Computer Review, 31(6):744–762, 2013. 8
2013
-
[35]
H. Wan, J. Zhang, A. A. Suria, B. Yao, D. Wang, Y . Coady, and M. Prpa. Building LLM-based AI agents in social virtual reality. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’24, article no. 65, 7 pages, 2024. doi: 10.1145/ 3613905.3651026 2
2024
-
[36]
Wang and J
I. Wang and J. Ruiz. Examining the use of nonverbal communication in virtual agents. International Journal of Human–Computer Interaction , 37(17):1648–1673, 2021. 2
2021
-
[37]
Yang and M
E. Yang and M. C. Dorneich. The effect of time delay on emotion, arousal, and satisfaction in human-robot interaction. In Proceedings of the human factors and ergonomics society annual meeting , vol. 59, pp. 443–447. SAGE Publications Sage CA: Los Angeles, CA, 2015. 2
2015
-
[38]
F.-C. Yang, K. Duque, and C. Mousas. The effects of depth of knowl- edge of a virtual agent. IEEE Transactions on Visualization and Com- puter Graphics, 2024. 1
2024
-
[39]
J. Zhu, R. Kumaran, C. Xu, and T. Höllerer. Free-form conversation with human and symbolic avatars in mixed reality. In 2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR) , pp. 751–760. IEEE, 2023. 1 9
2023
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.