{"id":"232aaac9-2b46-400b-8a02-1022788f521e","arxiv_id":"2507.22352","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Natural conversational fillers improve perceived response time for VR agents when LLM responses are delayed above four seconds, while artificial wait indicators do not.","lead":"This paper tests whether conversational fillers, such as thinking gestures and filler phrases, can mask response delays from LLM-powered virtual agents in VR. It finds that delays above four seconds hurt user experience, and that natural fillers improve perceived response speed, while artificial loading icons do not.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Q1's 'responding meaningfully' does not prevent the Natural filler's spoken utterance from itself counting as a meaningful response, so the H2a result may measure earlier audio onset rather than perceived latency.","rationale":"The reader's weakest assumption is exactly the load-bearing soft spot. The paper is otherwise well designed: explicit hypotheses, G*Power-based sample size, counterbalanced Latin square, ART repeated-measures ANOVA with Holm-Bonferroni-corrected post-hoc tests, and an open-source pipeline. The latency main effect (H1a) is robust because Q1 differences between delay levels occur within the same filler condition, where audio onset is constant. The filler result (H2a), however, compares conditions that differ in the objective timing of the first audible speech: Natural fillers are spoken at the start of the delay, so the agent literally begins talking earlier. Q1's 'meaningfully' is the only defense, and it is weak because 'Hmm, let me think about that...' is a meaningful conversational turn in human dialogue. Without a manipulation check, the significant Q1 advantage for Natural over None and over Artificial at High delay cannot distinguish perceived-latency mitigation from response-onset redefinition. The abstract's 'above 4 seconds' threshold is a secondary overgeneralization, but the Q1 construct issue is more central because it bears directly on the paper's novel positive claim. A post-condition manipulation check or a gesture-only Natural condition would settle the question; until then, the reader's CONDITIONAL verdict is appropriate and should remain unchanged.","tokens_in":33874,"tokens_out":4966,"duration_ms":68252,"concrete_test":"Run the same 3×3 within-subjects design with one added post-condition item: 'I counted the agent's filler speech as the start of its response' (Likert scale). If Natural-condition agreement is significantly above Neutral at Medium and High delay, the Q1 advantage is confounded. Alternatively, replace Natural voice lines with gesture-only fillers; if the Q1 advantage over None disappears, the effect is driven by audible speech onset rather than by perceived delay mitigation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main H2a result rests entirely on Q1 (§3.2.1): 'From the moment I stopped talking, the agent was quick to start responding meaningfully.' In the Natural condition, the agent speaks a filler line—e.g., 'Hmm, let me think about that...'—immediately after the user stops talking. That utterance is not meaningless: it acknowledges the user and signals deliberation. A participant who treats it as the start of a meaningful response will rate Q1 high regardless of perceived wait, because the objective time to first meaningful speech is shorter in Natural than in None or Artificial. The word 'meaningfully' does not remove the confound; it may even push participants to exclude only clearly non-semantic sounds, while natural filler lines remain semantically interpretable. No manipulation check (§3.2, §4.2) asks whether participants counted filler speech as the beginning of the agent's response. Thus the Medium/High Q1 advantage for Natural fillers (p < 0.01 and p < 0.0001, §5.2) is consistent with a definitional artifact rather than evidence of perceived-latency mitigation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a within-subjects VR experiment (N=54) in which participants held free-form, task-guided conversations with nine LLM-powered embodied agents under three response-latency levels (1.5s, 4.0s, 6.5s) and three filler conditions (None, Artificial wait indicator, Natural conversational filler). The authors use ART ANOVA with Holm-Bonferroni-corrected ART-C contrasts and find that latency significantly degrades perceived response time and several broader perception ratings, that Natural fillers significantly improve the single-item perceived-response-time rating at Medium and High latency (H2a), and that Artificial wait indicators do not produce significant effects (H3a/H3b unsupported). The paper interprets these results as showing that response delays above 4 seconds degrade quality of experience, that natural fillers mitigate perceived delay, and it contributes an open-source Unity/ASR/LLM/TTS pipeline for deploying conversational agents in VR.","tokens_in":34076,"tokens_out":9689,"duration_ms":121177,"significance":"The study is a well-powered full-factorial experiment in a realistic LLM-based conversational pipeline, and the open-source release is a practical contribution to VR conversational-agent research. Strengths include a power analysis, balanced-Latin-square counterbalancing, appropriate nonparametric ART ANOVA with Holm-Bonferroni correction, and a credible interactive system with a measured 1.5s SRT. If the H2a finding survives a cleaner measure of perceived latency, it would usefully extend prior WoZ and pre-recorded filler results to real-time LLM-driven free-form conversation. However, the central interpretation depends on a single questionnaire item whose wording does not fully separate filler speech from response speech, and the 'above 4 seconds' threshold claim goes beyond the sampled latency grid. The paper also makes summary claims in the Discussion and Conclusion that are broader than the significant results.","major_comments":[{"comment":"The main positive result (H2a) is measured by a single item, Q1: 'From the moment I stopped talking, the agent was quick to start responding meaningfully.' In the Natural condition the agent utters a semantically interpretable filler line ('Hmm, let me think about that...') and performs a thinking gesture immediately after the user stops talking, so the objective time to the first audible speech is shorter in Natural than in None or Artificial. The word 'meaningfully' was intended to prevent participants from counting filler speech, but no manipulation check in §3.2 or §4.2 verifies that they did so; PSQ6 asks whether fillers were noticed but is not used to test this interpretation. As a result, the Medium/High Q1 advantage of Natural over None (p < 0.01 and p < 0.0001 in §5.2) is also consistent with participants treating the filler utterance as the start of the response, making the effect partly definitional rather than a genuine improvement in perceived latency. Please add a manipulation check (for example, a post-condition item that distinguishes 'the agent said something while thinking' from 'the agent answered my question') or re-analyze with a measure that isolates the final response onset, and adjust the H2a claim accordingly.","section":"§3.2.1, §3.1.2, §5.2"},{"comment":"The abstract, §5.3, and the design recommendations state that response delays 'above 4 seconds' degrade quality of experience, but the experiment samples only 1.5s, 4.0s, and 6.5s. Q1 shows all three pairwise delay differences, but the broader perception metrics (Q2-Q6) show significant degradation only between Low and High, with Q2 the only broader metric distinguishing Medium from High and no broader metric distinguishing Low from Medium. The data therefore establish that 6.5s is worse than 1.5s and that 4.0s and 6.5s are perceived as slower than 1.5s on Q1, but they do not locate a threshold 'above 4 seconds.' Please soften the threshold claim to 'the two longer delays tested' or add a latency grid that brackets the boundary.","section":"§5.3, Figure 5"},{"comment":"The Discussion and Conclusion go beyond the supported effects. The only significant filler effects are on Q1 at Medium and High latency; no filler effect reaches significance on Q2-Q6, and the paper appropriately reports that H2b, H3a, and H3b are unsupported. Yet §7 concludes that Natural fillers 'enhance VR user experience' and 'reduce latency's negative repercussions,' and §5.2 states that fillers 'improve tolerance for delayed responses.' Please align these summary claims with the measured single-item Q1 effect and the broader null results, and avoid implying that the study demonstrates improvement in global user experience.","section":"§5.2, §7"},{"comment":"The claim that conversational fillers increase willingness to use the slowest system rests on a comparison of PSQ5 and PSQ10, which are repeated measures from the same 54 participants. The reported chi-square test treats the two sets of responses as independent; a paired analysis (e.g., a McNemar test on collapsed agree/disagree categories or a paired ordinal model) is needed because the same participants answer both questions. Please re-analyze these data and adjust §5.2 if the conclusion changes.","section":"§4.2.5"}],"minor_comments":[{"comment":"The phrase 'free-from conversation' appears to be a typo for 'free-form conversation'; please correct it.","section":"§1"},{"comment":"The set of 36 Holm-Bonferroni-adjusted contrasts should be defined explicitly in the text or supplementary material; currently the reader cannot tell which family of comparisons the correction controls.","section":"§4.1"},{"comment":"The custom questionnaires are appropriately acknowledged in Section 6 as unvalidated, but the single-item Q1 and the aggregated RoSAS-based Q4/Q5 would benefit from a brief psychometric justification or a pilot validation, especially because Q1 carries the main result.","section":"§3.2, §6"},{"comment":"In the copy under review, the text inside Figures 4-8 appears as long unreadable token sequences (e.g., '/uni000000...' strings); please ensure the final PDF contains clean figure graphics with readable labels.","section":"Figures 4-8"},{"comment":"The inductive analysis of text justifications is described as preliminary, but no inter-rater reliability or coding agreement is reported; a brief statement on the coding process would help readers assess the thematic counts.","section":"§4.2.6"},{"comment":"The observed filler effect sizes are small (η²p = 0.017-0.081), whereas the power analysis targeted a medium-to-large effect; the null results for Q2-Q6 should be explicitly framed as 'not supported' rather than as evidence of no effect, which the text mostly does but could state more consistently.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable fit for CUI and the open-source pipeline is a genuine artifact. My main concern is that the headline H2a result may reflect a measurement artifact in Q1; if the authors can validate the item or reframe the claim, I would support publication. I see no indication of problematic citation patterns or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading if you care about latency in embodied conversational agents. The genuinely new thing is the test bed: a real ASR-Llama-TTS pipeline in VR with free-form task conversations, with both delay and filler type manipulated. Prior work mostly used Wizard-of-Oz, scripted exchanges, or pre-recorded clips; this study actually varies latency in a working system, and the authors release the pipeline. The experiment is also carefully run: power analysis, 54 participants, 18-order Latin square, ART ANOVA, Holm-Bonferroni-corrected post hocs. I believe the main effect of delay is real, and the null result for artificial wait indicators is a useful practical finding.\n\nThe abstract's claim that \"latency above 4 seconds degrades quality of experience\" is broader than the pairwise evidence. The None-condition data show reliable differences for Low vs Medium and Low vs High on Q1, but Medium vs High is significant only for Q1 and Q2; for most broader questionnaire items the pairwise gap is Low vs High only. So \"above 4 seconds\" is an inference, not a direct result.\n\nMore important is Q1. The wording—\"quick to start responding meaningfully\"—was chosen to prevent the Natural filler's speech from being counted as a response, but no manipulation check shows that participants made that distinction. Fillers like \"Hmm, let me think about that...\" are themselves audible, meaningful acknowledgment; a participant could reasonably rate response onset as earlier in Natural simply because speech began earlier. That means H2a might partly measure audio onset, not perceived wait. The authors flag the wording risk in section 3.2.1 but do not resolve it. I still think the result is directionally consistent with prior Wizard-of-Oz filler work, and Natural also beats Artificial at High, so I would not call this a dealbreaker—but the paper needs a manipulation check or a revised measure.\n\nOther issues are minor: the custom questionnaires are unvalidated (acknowledged by the authors), RoSAS is compressed into two items, and no raw data or analysis scripts are included. There are no fitted parameters and no real circularity beyond the measurement-level issue.\n\nThis paper is for practitioners building LLM-driven VR agents and for researchers studying response-time mitigation. It deserves a serious referee. My recommendation: engage with it, and ask the authors to address the Q1 confound and rephrase the 4-second claim to match the statistical comparisons.","headline":"A well-run, fully interactive VR study of latency and fillers whose headline claims are a bit broader than the pairwise data will bear—still worth a serious referee.","tokens_in":34648,"tokens_out":2691,"would_cite":true,"duration_ms":34489,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that in VR conversations with LLM-powered agents, response delays above four seconds degrade users' perceived response time and experience, while natural conversational fillers—thinking gestures plus filler…","keywords":["conversational fillers","response latency","LLM-powered agents","virtual reality","perceived response time","quality of experience","user study"],"falsifier":"Run the same study with a natural filler that uses only the thinking gesture and no voice line; if the perceived-response-time advantage disappears, the effect came from participants counting the filler phrase as the start of the answer rather than from a genuine reduction in perceived waiting.","tokens_in":33681,"feed_emoji":"🗣️","tokens_out":4806,"duration_ms":53253,"temperature":0.7,"pith_summary":"This paper asks whether conversational fillers can soften the experience of waiting for an LLM-powered virtual agent to reply in free-form conversation, and whether conventional loading indicators do the same. The authors report that delays above four seconds degrade perceived response time and broader user experience, that natural fillers (a thinking gesture plus a filler phrase) significantly improve perceived response time at four and 6.5 second delays, and that artificial wait indicators do not. If correct, the findings give interaction designers a concrete, low-cost mitigation for the latency that remains unavoidable when speech recognition, generation, and synthesis run in sequence.","feed_headline":"Fillers ease perceived AI delays at 4-plus seconds","feed_subtitle":"Natural gesture-plus-voice fillers beat loading icons in VR chats with LLM agents.","key_machinery":"The load-bearing intervention is the natural conversational filler: a multimodal delay-mitigation cue in which the agent performs a \"thinking\" gesture (head turn, chin touch, subtle breathing) and speaks a filler voice line during the response delay. It works by occupying the wait with social cues that mimic human deliberation, and the paper measures its effect through a six-item post-condition survey whose first item asks whether the agent \"was quick to start responding meaningfully.\"","core_discovery":"This paper claims that response delay is a first-class usability problem for free-form, LLM-powered embodied conversational agents, and that the right interface-level remedy is a natural conversational filler rather than a conventional loading indicator. In a within-subjects VR study with 54 participants conversing with nine agents across three scenarios, delay degraded perceived response time and broader perception metrics, with effects becoming pronounced at 4.0s and 6.5s. Natural fillers—a thinking gesture plus a filler phrase such as \"Hmm, let's see...\"—significantly improved perceived response time at both of those delay levels compared with no filler, supporting H2a; artificial wait indicators did not produce significant improvement, and neither filler type rescued the broader experience metrics.","pith_inferences":["The gesture-versus-voice split in participant preferences suggests a personalization opportunity: a single filler design may not serve all users, and offering selectable filler modality could strengthen mitigation.","The null result for artificial wait indicators may be specific to generic spinners; more communicative indicators that signal content (for example, showing that the agent is checking a record or that a response stage is underway) might behave differently, and the paper's recommendation to explore such indicators is a natural next test.","The four-second threshold is measured in a leisurely task-guided VR setting; time-pressure or high-stakes contexts such as emergency or medical conversation may compress tolerance, so the threshold is likely an upper bound rather than a universal constant.","In non-embodied channels such as phone voice assistants or text chatbots, the natural filler's gestural component is absent, so the voice-line alone may be the transferable part; the paper's embodied results do not directly establish that transfer."],"forward_implications":["Designers of LLM-powered conversational agents should target system response times under four seconds, because longer delays significantly degrade perceived response time, engagement, impression, competence, and willingness to interact again.","When delays of four to 6.5 seconds are unavoidable, natural conversational fillers (gesture plus filler phrase) can recover perceived response time, though they do not restore the broader perception metrics.","Artificial loading icons and processing sounds should not be relied on as a latency-mitigation strategy, since the study found no significant benefit over no filler.","Studies using LLM-based free-form agents should report response latency, because high delays bias user perceptions and limit comparability across results.","The open-source deployment pipeline makes the system reproducible for future VR agent studies."],"supporting_citations":[{"why":"Prior scripted-agent study showing conversational fillers mitigate the effects of delayed virtual agent response time; the effect this paper extends to free-form LLM agents.","marker":"[12]"},{"why":"Shows gestural fillers reduce user-perceived latency in conversations with digital humans; source of the Natural filler design.","marker":"[53]"},{"why":"Demonstrates conversational fillers improve human-robot interaction without harming perceived intelligence; basis for the filler set and expectations.","marker":"[101]"},{"why":"Tests strategies for bridging time-to-content in spoken dialogue systems and finds fillers rated more appropriate than silence; a direct comparison point.","marker":"[64]"},{"why":"Provides delay thresholds and habituation effects for robot response speed; informs the latency levels chosen.","marker":"[83]"},{"why":"The Robotic Social Attributes Scale (RoSAS) from which the discomfort and competence survey items (Q4, Q5) were adapted.","marker":"[14]"},{"why":"Aligned Rank Transform procedure used for the 3x3 repeated-measures ANOVA on Likert-scale responses.","marker":"[102]"},{"why":"ART-C contrast tests used for post-hoc pairwise comparisons with Holm-Bonferroni correction.","marker":"[27]"}],"fun_headline_variants":["Natural fillers mask AI lag in VR conversations","4-second LLM delay? Use a thinking gesture","Fillers outperform loading icons for VR agent lag","Gesture plus filler trims perceived AI wait time","Conversational fillers help at 4+ second delays"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main measurement assumption is that the survey question about the agent being \"quick to start responding meaningfully\" measures perceived latency and not the fact that, in the Natural condition, the agent audibly speaks during the wait, so participants may treat the filler phrase itself as the start of the reply.","fun_headline_variants_meta":{"raw":{"variants":["Natural fillers mask AI lag in VR conversations","4-second LLM delay? Use a thinking gesture","Fillers outperform loading icons for VR agent lag","Gesture plus filler trims perceived AI wait time","Conversational fillers help at 4+ second delays"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00012,"raw_usage":{"total_tokens":1017,"prompt_tokens":804,"completion_tokens":213,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":420,"completion_tokens_details":{"reasoning_tokens":139}},"tokens_in":420,"tokens_out":213,"duration_ms":3243,"temperature":1.0,"reasoning_tokens":139,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:46:49.895245+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same study with a natural filler that uses only the thinking gesture and no voice line; if the perceived-response-time advantage disappears, the effect came from participants counting the filler phrase as the start of the answer rather than from a genuine reduction in perceived waiting.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows gestural fillers reduce user-perceived latency in conversations with digital humans; source of the Natural filler design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Tests strategies for bridging time-to-content in spoken dialogue systems and finds fillers rated more appropriate than silence; a direct comparison point."}],"review_version":1}