{"id":"97711377-0d07-4638-90e1-8aa654a3a76e","arxiv_id":"2508.11781","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Behavioral filler animations made virtual agents' response delays feel more appropriate and natural, outperforming symbolic progress indicators and idle motion in an immersive VR study.","lead":"In a VR job-interview experiment with 24 people, the authors tested four ways to cover up a conversational agent's thinking delay: behavioral filler animations, idle motion, and two progress-bar indicators. The behavioral filler felt more natural, was preferred by two-thirds of participants, and improved perceived response time and presence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Filler type is perfectly confounded with character identity: each condition always used the same MetaHuman, so the behavioral filler's advantages could be due to the specific character rather than the filler itself.","rationale":"The paper is transparent and the within-subject design, counterbalanced condition order, and open data are strengths. The main statistical findings are plausible, and the consistency across multiple correlated dependent variables makes multiple testing a secondary concern. However, the character–filler confound is load-bearing: it threatens the causal attribution that is the paper's central claim. The reader's 'weakest_assumption' identifies exactly this issue, and I agree that it warrants a CONDITIONAL verdict. The proposed concrete test—rotating characters across filler types—would directly determine whether the effect is due to the filler or the character. Until such a test is done, the claim that behavioral fillers are superior should be treated as conditional, not established.","tokens_in":16046,"tokens_out":5068,"duration_ms":63153,"concrete_test":"Run a follow-up experiment using the same four MetaHuman characters and four filler types in a Latin-square design: each participant experiences each filler type with a different character, and across participants each character appears equally often with each filler. If BEHAVIORAL still shows significantly higher scores on the primary DVs (appropriateness, parasocial interaction, engagement, social realism, humanlikeness, naturalness) and is still preferred by a majority, the character confound is not responsible. If the pattern changes or disappears, the original claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The study claims that the BEHAVIORAL filler (thinking animations plus 'Hhm') caused significant improvements in perceived response-time appropriateness, parasocial interaction, engagement, social realism, humanlikeness, naturalness, and preference. However, the experimental design perfectly confounds filler type with character identity: Section 3.3 states that each condition always used a specific MetaHuman (two male, two female), and Section 5 acknowledges 'each filler type was always displayed with the same character... we can not exclude that character appearance might have influenced our results.' The dependent measures are all susceptible to a character's physical appearance, voice, and perceived attractiveness — especially humanlikeness, naturalness, and hiring preference. If the character assigned to BEHAVIORAL was simply more likable or humanlike, it could produce exactly the observed pattern: significant improvements across many correlated scales and a 16/24 preference. Counterbalancing the order of conditions does not break this confound because it only randomizes sequence, not the character–filler mapping. Thus the central causal claim that the filler itself is responsible is not secure; a character effect alone could explain the results.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a within-subject VR experiment (N=24) in which participants interviewed four MetaHuman ECAs while the agent displayed one of four delay-mitigation strategies during a 2/4/8-s thinking pause: BASE idle animation, BEHAVIORAL thinking animations with non-lexical \"Hhm\" fillers, EMBEDDED progress bar in a badge, and EXTERNAL progress bar in a thinking bubble. The authors measure perceived response-time appropriateness/expectedness, presence subscales, social presence, impression-of-agent scales (humanlikeness, intelligence, likeability, willingness, competence, naturalness, assuredness), gaze on face, and final hiring preference. They report significant benefits of BEHAVIORAL over all other conditions on appropriateness, parasocial interaction, engagement, humanlikeness, and naturalness, and over the two symbolic conditions on social realism and face gaze; 16/24 participants preferred BEHAVIORAL. H2 (symbolic indicators improve expectedness) and H3 (embedded better than external) were not confirmed.","tokens_in":16325,"tokens_out":7975,"duration_ms":94133,"significance":"If the causal interpretation were valid, the result would be practically useful for designers of LLM-based ECAs in VR, showing that naturalistic behavioral fillers can mask response delays more effectively than progress bars. The work extends prior screen-based findings to immersive VR and directly compares behavioral and symbolic strategies. The manuscript is generally clear, uses appropriate nonparametric tests with post-hoc corrections, and makes data and code available, which aids reproducibility. However, the central causal claim is weakened by the full confounding of filler type with character identity, so the study currently supports a conditional rather than a general causal conclusion.","major_comments":[{"comment":"The four filler conditions are perfectly confounded with character identity: Section 3.3 states that two male and two female MetaHumans were created, and Section 5 acknowledges that 'each filler type was always displayed with the same character.' Since the dependent variables include humanlikeness, naturalness, parasocial interaction, and hiring preference, the significant BEHAVIORAL advantages reported in Tables 2–3 and the 16/24 preference could in principle be caused by the specific character's appearance, voice, or perceived attractiveness rather than by the filler. Counterbalancing the presentation order does not break this confound because it does not randomize the character–filler mapping. The authors' acknowledgment that appearance 'might have influenced our results' understates the problem: with the current design, the two explanations are indistinguishable. I recommend either a","section":"§3.3, §5 (Limitations)"},{"comment":"The manuscript reports 13 subjective dependent variables and several eye-gaze measures without any correction for multiple testing across the family of measures. Some of the reported significant effects (e.g., parasocial interaction p=0.014, engagement p=0.016, naturalness p=0.013) would not survive a very conservative Bonferroni correction, although they would survive FDR. The consistency of the pattern mitigates the concern, but the number of 'significant' outcomes should be interpreted with caution. An explicit multiple-comparison strategy or pre-registered analysis plan would strengthen the claims; at minimum, report adjusted p-values or justify the unadjusted family-wise error rate.","section":"§3.1, Table 2"},{"comment":"The self-reported expectedness measure was not significant (p=0.066), but the authors use the gaze-at-response-onset results to argue that 'participants expected the response more naturally with the BEHAVIORAL fillers.' This is a post hoc interpretation, not a planned test. The gaze measure was not validated as an expectedness measure, and the claim goes beyond the data. Please label this as exploratory and remove it from the summary of hypothesis support.","section":"§4.1, §5"}],"minor_comments":[{"comment":"The preference counts are inconsistent: 16 (66.7%) + 3 (12.5%) + 1 (4.2%) + 2 (8.3%) = 22, not 24. Two participants are unaccounted for. Please clarify whether some participants did not state a preference or whether the counts/percentages contain an error.","section":"§4.5"},{"comment":"Minor typos: 'Humalikeness' in Figure 4 and 'd f1 = d fe f f ect' in Table headers should read 'Humanlikeness' and 'df1 = df_effect'.","section":"Figure 4, Tables 2/4/5"},{"comment":"Replication would benefit from more detail about the ECAs' voices (e.g., TTS engine, language/accent, whether all agents used the same voice) and about the animation state machine controlling the transition from filler to response. These details are also relevant to the character-confound concern.","section":"§3.3"},{"comment":"The relative gaze time on the face is reported as a proportion in Table 3 (e.g., 0.66) but as a percentage in Figure 3(g). Please use a consistent scale. Also clarify whether the 'cumulative average' smoothing affects the RGTF values reported.","section":"§4.4"},{"comment":"No sample size justification or power analysis is provided. Given the number of dependent variables, reporting confidence intervals for the main effect sizes would aid interpretation.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well organized and the empirical work is reported in a mostly transparent manner, with data/code availability being a clear strength. The main obstacle is the character–filler confound: it is not a superficial limitation but a structural feature of the design that prevents unambiguous causal attribution of the headline results. I would not support acceptance in the current form. A deconfounded follow-up study or a substantial rewriting that reduces the claims to condition-level observations would be needed. The missing two preference responses and the absence of a pre-registered multiple-comparison strategy are secondary but should also be fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, honest empirical comparison of delay fillers for embodied conversational agents in VR, with a clear preference for behavioral fillers over symbolic progress bars. The new bit is the direct VR comparison of an embedded badge, an external thinking bubble, and a multimodal behavioral filler, with eye-tracking. The study is carefully reported, the statistics fit the design, and the data and code are on OSF/GitLab. The main finding is plausible and consistent with Kum and Lee and Elfleet and Chollet.\n\nThe soft spots are real but not hidden. The structural confound: each filler condition always used the same MetaHuman, so character appearance or voice could drive effects on humanlikeness, naturalness, and presence. The authors flag this in Section 5 but cannot exclude it. That limits the causal claim that the filler itself is responsible. Given the effect sizes are large and the pattern is consistent across many scales, I suspect the effect is mostly real, but the confound prevents a clean causal read. A follow-up crossing characters with fillers would settle it.\n\nThey also test many dependent variables without cross-measure correction, so some significant results might be inflated. The strongest effect (response-time appropriateness) would survive most corrections, but the more marginal ones (e.g., social realism post-hocs) need care. The gaze-based 'expectedness' interpretation is post-hoc and treated as exploratory, which is the right call.\n\nThe paper is transparent about its limitations and the discussion is measured. It's a good contribution for HCI/IVR researchers, especially for designers of LLM-powered VR agents. I'd send it to peer review: the design and reporting are sound enough for reviewers to work with, and the confound is a clear agenda item rather than a fatal flaw. I'd cite it as evidence that behavioral fillers are promising, with the caveat about character identity.","headline":"Useful VR comparison of delay fillers with a real confound between filler type and character identity that the authors acknowledge but cannot fully rule out.","tokens_in":16747,"tokens_out":1565,"would_cite":true,"duration_ms":19085,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a VR job-interview study, a thinking animation with a vocal \"Hhm\" made a delayed conversational agent feel more human and natural than progress-bar indicators.","keywords":["Virtual Reality","Embodied Conversational Agent","Delay Mitigation","Behavioral Filler","Progress Indicator","Presence","Humanlikeness","Eye Tracking"],"falsifier":"Run the same study with each filler type paired with every character in a counterbalanced design while keeping delays identical; if the significant advantages of the behavioral filler disappear or shift across character pairings, the effect is character-driven, not filler-driven.","tokens_in":16010,"feed_emoji":"🤖","tokens_out":3802,"duration_ms":39447,"temperature":0.7,"pith_summary":"The paper asks whether an embodied conversational agent that needs several seconds to answer—say, because a large language model is computing—should fill the silence with human-like \"thinking\" behavior or with explicit progress indicators. It argues, on the basis of a 24-person within-subject VR job-interview study, that a multimodal behavioral filler (thinking animations plus a non-lexical \"Hhm\" vocalization) outperforms both a badge-mounted progress bar and a floating thinking-bubble progress bar. The behavioral filler significantly improved perceived response-time appropriateness, three presence subscales (parasocial interaction, engagement, social realism), humanlikeness, and naturalness, and was preferred by two-thirds of participants. The symbolic progress indicators did not improve expectedness or presence and drew gaze away from the agent's face. If the result generalizes, designers of LLM-powered VR agents should treat behavioral delay mitigation as the default, reserving progress indicators for settings where explicit time information matters more than naturalness.","feed_headline":"Thinking animations beat progress bars for delayed VR agents","feed_subtitle":"Behavioral fillers improved perceived response time, presence, humanlikeness, and naturalness versus progress bars.","key_machinery":"The central object is the multimodal behavioral filler: a set of nine \"thinking\" animations (gaze aversion, filler gestures) combined with a non-lexical verbal filler, \"Hhm,\" spoken with six pitches and occurring every three questions. This filler runs during the delay between the user's question and the agent's response. The comparison conditions are a base idle animation and two symbolic progress indicators (a badge-embedded bar and a floating thinking-bubble bar). The experimental machinery is a within-subject VR job interview with MetaHuman agents, pseudo-random delays of 2, 4, and 8 seconds, questionnaire measures of perceived response time, presence, and agent impression, plus continuo","core_discovery":"The paper reports a within-subject VR study with 24 participants who held simulated job interviews with four embodied conversational agents, one per condition. During scripted response delays of 2, 4, or 8 seconds, the agent either played idle animations (BASE), performed nine thinking animations plus a non-lexical 'Hhm' vocalization with six pitches (BEHAVIORAL), showed a green progress bar embedded in a visitor badge (EMBEDDED), or showed a floating thinking bubble with a progress bar (EXTERNAL). The central claim is that the BEHAVIORAL filler made the perceived response time significantly more appropriate, raised parasocial interaction, engagement, and social realism, and improved humanli","pith_inferences":["The effect may be partly driven by the specific MetaHuman character paired with the behavioral filler, since filler type and character were not crossed in the design; a replication with rotated character-filler pairings would isolate the filler itself.","If behavioral fillers work because they signal imminent speech, agents could deliberately time a finishing gesture to the arrival of the response, making even unpredictable delays feel natural.","In high-stakes or information-retrieval tasks, symbolic indicators might be more appropriate than the job-interview scenario suggests; the paper itself speculates along these lines, but future work would need to test it directly."],"forward_implications":["If the behavioral filler's effect holds, VR agents backed by LLMs can mask multi-second computation delays with low-cost animations and a brief vocal filler rather than loading UI.","Symbolic progress indicators, even embedded ones, did not help users predict when the response would arrive; the gaze data suggests they redirect attention away from the agent's face, which may make response onset feel abrupt.","The absence of significant differences between embedded and external progress bars implies that the placement of a symbolic indicator matters less than whether a symbolic indicator is used at all.","For real deployments with unpredictable response times, behavioral fillers need an animation state machine that can extend or finish naturally, a design problem the paper explicitly identifies.","For longer delays, behavioral fillers may lose their advantage because they provide no estimate of remaining wait time, so symbolic indicators or hybrids may be needed."],"supporting_citations":[{"why":"Showed that gestural fillers reduce user-perceived latency and improve naturalness and competence in screen-based agent interaction; provides the direct prior evidence for the behavioral filler.","marker":"[18]"},{"why":"Found that multimodal feedback improves presence and immersion with LLM-powered ECAs in VR; closest prior work and a baseline for the present study's design.","marker":"[9]"},{"why":"Established that percent-done progress indicators improve user satisfaction with waiting; motivates the symbolic progress-bar conditions.","marker":"[27]"},{"why":"Showed that conversational fillers make delayed web-based agent response time more acceptable; supplies the appropriateness measure and comparison.","marker":"[4]"},{"why":"Demonstrated that typing indicators increase social presence in chatbots; supports the hypothesis that symbolic indicators could help with ECAs.","marker":"[12]"},{"why":"Provides the social presence questionnaire items used to measure perceived copresence and realism.","marker":"[29]"}],"fun_headline_variants":["VR agents: thinking animations beat progress bars for delays","Behavioral fillers boost VR agent presence, not progress bars","VR delay fillers: Animations outshine progress indicators","Thinking animations improve VR agent experience over progress bars","VR study: Behavioral fillers win over symbolic progress bars"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The four filler types were always paired with the same MetaHuman character each, so differences in character appearance or voice, rather than the filler itself, could drive the reported perceptual differences.","fun_headline_variants_meta":{"raw":{"variants":["VR agents: thinking animations beat progress bars for delays","Behavioral fillers boost VR agent presence, not progress bars","VR delay fillers: Animations outshine progress indicators","Thinking animations improve VR agent experience over progress bars","VR study: Behavioral fillers win over symbolic progress bars"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000158,"raw_usage":{"total_tokens":1071,"prompt_tokens":762,"completion_tokens":309,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":230}},"tokens_in":506,"tokens_out":309,"duration_ms":3936,"temperature":1.0,"reasoning_tokens":230,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:46:14.009194+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same study with each filler type paired with every character in a counterbalanced design while keeping delays identical; if the significant advantages of the behavioral filler disappear or shift across character pairings, the effect is character-driven, not filler-driven.","supporting_citations":[{"cited_title":"Kum and M","cited_arxiv_id":null,"evidence_quote":"Showed that gestural fillers reduce user-perceived latency and improve naturalness and competence in screen-based agent interaction; provides the direct prior evidence for the behavioral filler."},{"cited_title":"Elfleet and M","cited_arxiv_id":null,"evidence_quote":"Found that multimodal feedback improves presence and immersion with LLM-powered ECAs in VR; closest prior work and a baseline for the present study's design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Established that percent-done progress indicators improve user satisfaction with waiting; motivates the symbolic progress-bar conditions."},{"cited_title":"Boukaram, M","cited_arxiv_id":null,"evidence_quote":"Showed that conversational fillers make delayed web-based agent response time more acceptable; supplies the appropriateness measure and comparison."},{"cited_title":"The chatbot is typing","cited_arxiv_id":null,"evidence_quote":"Demonstrated that typing indicators increase social presence in chatbots; supports the hypothesis that symbolic indicators could help with ECAs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the social presence questionnaire items used to measure perceived copresence and realism."}],"review_version":1}