{"id":"f27760af-9f15-4c9b-9dad-5a6e6a68c68b","arxiv_id":"2505.20082","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Visible avatar laughter cues in VR co-viewing shift engagement from personal immersion toward socially coordinated, norm-governed participation.","lead":"This study compared people watching comedy in VR with only voice chat versus with avatars that visibly laugh, using 24 participants in groups of three. It found that visible laughter cues made viewers more socially engaged and emotionally contagious, but also more self-conscious about reacting.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The VLC vs VO contrast changes avatar presence and laughter cues simultaneously; without an idling-avatar control, the causal claim that embodied laughter cues drive the effects is unidentified.","rationale":"The reader's weakest_assumption identified exactly the same load-bearing concern: the VLC condition adds avatar presence and visible expressions simultaneously, so the observed effects cannot be attributed to laughter cues rather than to having an embodied avatar at all. This is the most serious threat to the paper's central claim because the abstract and RQs are phrased causally ('embodied laughter cues shifted...'), while the experimental contrast does not isolate that variable. The authors' explicit decision in Section 3.1 to omit an idling-avatar condition demonstrates awareness of the confound but does not resolve it. A neutral-avatar control is feasible in principle and would directly test whether the laughter animations add anything beyond embodiment. Some secondary issues exist, such as multiple-comparison inflation on the engagement subscales and the chained-laughter measure being defined by button presses tied to visible animations, but these are less central than the identification problem. The qualitative analysis is thoughtful and the chained-laughter effect is suggestive, so I would not reject the paper; I would require either the additional control or a narrowed claim, which is what a conditional verdict already communicates. Therefore the reader's verdict should remain unchanged.","tokens_in":21135,"tokens_out":4135,"duration_ms":47418,"concrete_test":"Add a third within-subjects condition, Avatar-Idle: participants are embodied in the same avatars and see co-viewers in the same living-room and mirror setup, but the laughter trigger is either disabled or produces no animation while voice laughter remains. Compare Avatar-Idle against VLC on the three engagement subscales that reached significance (Endurability, Novelty, Involvement), on chained-laughter counts, and on the interview themes of self-monitoring and expressive norms. If Avatar-Idle produces the same increases over VO as VLC does, the causal role of the laughter animations specifically is falsified and the paper should be reframed around avatar embodiment. If Avatar-Idle resembles VO, the current attribution survives this check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The design cannot identify 'embodied laughter cues' as the causal ingredient. Section 3.1 contrasts VO (voice chat only, no visible avatars) with VLC (voice chat plus full-body avatars, mirror, and laughter animations). The manipulation therefore changes at least three factors at once: having an avatar at all, seeing co-viewers' avatars, and seeing laughter animations. Any difference in engagement, self-awareness, or norms could be caused by avatar presence or embodiment alone. Participants' own quotes support this: P01 attributes togetherness to 'Having avatars,' P20 to 'being embodied in an avatar,' and P05 to 'With avatars present.' The Section 3.1 rationale for omitting an idling-avatar condition (audiovisual mismatch) explains a design constraint but does not provide a control that isolates laughter cues. Moreover, Section 3.1 states that in both conditions participants could manually trigger laughter expressions using the same interface, but in VO there is no visible avatar; this makes the trigger in VO either meaningless or an invisible action, so the conditions also differ in whether the trigger button affords visible expression. Without an idling-avatar control, the abstract's causal claim that embodied laughter cues shifted engagement and produced expressive norms is not supported by this comparison; a weaker claim about 'embodied expressive avatars' could be supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a within-subjects VR experiment (N=24 in eight triads) comparing voice-only co-viewing with a condition in which participants were embodied in expressive avatars with manually triggered laughter animations. Through questionnaires, system logs of chained laughter, and post-hoc interviews, the authors argue that visible embodied laughter cues shifted engagement from individual immersion to socially coordinated participation, heightened self-awareness and emotional contagion, and fostered expressive norms. The paper also derives design implications for balancing personal immersion with interpersonal emotional accommodation when co-viewing in social VR.","tokens_in":21393,"tokens_out":4569,"duration_ms":55522,"significance":"The topic is timely and the mixed-method design is well suited to generating rich, ecologically grounded observations about avatar-mediated co-viewing. The authors include an a priori power analysis, counterbalancing, deception to reduce demand characteristics, and a thematic analysis with saturation checks, and they describe the implemented VRChat world in enough detail for replication. The qualitative materials and interview quotes are valuable for hypothesis generation. However, the central causal claim is not identified by the experimental contrast, and the quantitative engagement results are weak under multiple-testing correction. The paper's contribution would be stronger as an exploratory account of avatar-mediated co-viewing than as a demonstration that laughter cues specifically cause the observed shifts.","major_comments":[{"comment":"The VO condition has no avatars at all, while the VLC condition adds full-body avatars, a mirror, and laughter animations simultaneously. The contrast therefore cannot identify 'embodied laughter cues' as the active ingredient; any difference could be due to having an avatar, seeing co-viewers' avatars, or seeing one's own avatar. The authors' stated rationale for omitting an idling-avatar condition explains a design constraint but does not provide the needed control. Participants' own quotes (P01, P05, P20) attribute togetherness to 'Having avatars', 'being embodied in an avatar', and 'With avatars present', not specifically to laughter animations. In addition, in VO the laughter trigger produced no visible feedback, so the two conditions also differ in whether the trigger affords visible expression. Please either add an idling-avatar control or reframe the central claim from 'embodied laughter cues' to 'embodied expressive avatars' and qualify the causal attribution throughout the paper.","section":"Section 3.1 and abstract"},{"comment":"Three of five engagement subscales are nominally significant (Endurability p=.014, Novelty p=.021, Involvement p=.011) with small effects (r=.18 to .32). With five related subscales, a Bonferroni correction sets the threshold at .01, and none of these p-values survives; the Involvement result (p=.011) is close but still above the corrected threshold. The paper should report adjusted p-values or a multivariate/omnibus test, and the text should not describe these results as 'significant' without qualification. The qualitative findings can stand independently, but the quantitative support for RQ1 is thinner than the current wording implies.","section":"Section 4.1.4 and Table 2"},{"comment":"The chained-laughter measure relies on an ad-hoc 5-second window and assumes that a cue following within that window is caused by the preceding cue. Simultaneous or video-triggered laughter could produce the same pattern, and no sensitivity analysis is reported for the window choice. The p=.006 result (V=0 over eight triads) is driven by a strong within-group difference, but the operational definition should be validated (e.g., with shorter/longer windows or a permutation baseline) before being used as the main quantitative evidence for emotional contagion. Moreover, in VO participants could press the same trigger but with no visible avatar, so the functional meaning of a 'chained laughter' event in VO is unclear and not comparable to VLC.","section":"Section 3.5.3 and Section 4.2.2"}],"minor_comments":[{"comment":"The compensation text contains a typo: 'Amaazon' should be 'Amazon'.","section":"Section 3.3"},{"comment":"In the P01 quote, 'press [the the laughter] button' contains a duplicated 'the' and should be corrected.","section":"Section 4.3.2"},{"comment":"The word 'accomodative' in Section 5.3.1 and 'Accomodation' in reference [81] should be spelled 'accommodative' and 'Accommodation', respectively.","section":"Section 5.3.1 and reference [81]"},{"comment":"The table would benefit from confidence intervals for the effect sizes, and the text should clarify whether the chained-laughter counts are aggregated at the group level or the individual level.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The experimental confound between avatar presence and laughter cues is substantial, and I considered recommending rejection. I settled on major revision because the qualitative findings and design implications remain informative if the claims are reframed as a comparison of voice-only versus avatar-mediated co-viewing, rather than as evidence specifically about laughter cues. The authors should be given the opportunity to revise the claims and analysis, but the central causal language in the abstract and discussion must be changed or supported by an additional control condition."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth reading. It's a controlled within-subjects comparison of voice-only versus voice plus embodied expressive avatars with manually triggered laughter animations, in triads of familiar participants watching comedy in VR. That's a new empirical setup: prior work used system-controlled avatars or 2D danmaku, and this one lets users drive their own avatar laughter. The qualitative analysis is genuinely good—the themes around self-monitoring, emotional contagion, and accommodation are well supported by the quotes, and the mapping to Communication Accommodation Theory is sensible. The chained-laughter measure (a follow within 5 seconds) shows a large effect (p=.006, r=.87) and aligns with the qualitative accounts. That's the strongest piece of evidence.\n\nThe soft spot is the design. The VLC condition adds avatar presence, visible co-viewer avatars, a mirror, and laughter animations all at once relative to VO. The abstract attributes the shift to 'embodied laughter cues,' but the comparison cannot isolate laughter cues from mere avatar presence. Several participant quotes ('Having avatars,' 'being embodied in an avatar') point to avatar presence generally. The authors' stated reason for skipping an idling-avatar control (audiovisual mismatch) is a legitimate design constraint, but it leaves the central attribution unidentified. A revised version that either adds a neutral-avatar condition or narrows the claim to 'embodied expressive avatars' would be much stronger.\n\nThe quantitative engagement results are also thinner than the abstract suggests. Three of five subscales are nominally significant with small effects (r=0.18–0.32), and none survive a Bonferroni correction for the five subscales. The chained-laughter result survives its own test but the window definition is self-authored and the multiple-comparison situation across all tests is not addressed. The data are not shared, which is a missed opportunity for a field where replication is hard.\n\nOverall, this is a solid paper with a plausible finding and a clear confound. It deserves a serious referee. I'd recommend conditional acceptance: ask for a neutral-avatar control or a re-framed claim, corrected p-values, and ideally the logs or interview data. The qualitative contribution alone justifies the revision cycle.","headline":"Solid mixed-methods study of VR co-viewing with a real confound between avatar presence and laughter cues; the causal claim about laughter cues needs either a neutral-avatar control or a narrower framing.","tokens_in":21904,"tokens_out":2835,"would_cite":false,"duration_ms":30022,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that visible, user-triggered avatar laughter in VR changes co-viewing comedy from solo immersion into socially coordinated participation, amplifying emotional contagion and prompting shared norms about when to laugh.","keywords":["co-viewing","virtual reality","embodied avatars","laughter cues","emotional contagion","engagement","expressive norms","communication accommodation"],"falsifier":"Run the same study with three conditions, voice only, an idling avatar that is present but never animates laughter, and the expressive avatar, and compare chained laughter, engagement scores, and interview reports of social pressure. If the idling-avatar condition behaves like voice-only, the laughter animation is the causal ingredient; if it behaves like the expressive condition, avatar presence itself is the cause.","tokens_in":20928,"feed_emoji":"😂","tokens_out":13502,"duration_ms":133919,"temperature":0.7,"pith_summary":"The paper tries to establish that in virtual reality, seeing co-viewers' embodied laughter, not just hearing their voices, changes how people watch comedy together. In a within-subjects experiment with eight acquainted triads (N=24), the authors compared voice-only co-viewing with co-viewing in which participants were embodied as expressive avatars that displayed full-body and facial laughter cues the user could trigger. They found that visible laughter cues raised Involvement, Endurability, and Novelty scores and increased chained laughter events, while interview data showed attention shifting from the screen toward socially coordinated watching, with participants monitoring and adjusting their own laughter to fit the group. The authors interpret this as emotional contagion through visible mimicry and the emergence of expressive norms, with a real trade-off: participants sometimes felt social pressure and missed the freedom of laughing unobserved. If the claim holds, designers of remote co-viewing systems should treat visible embodied reactions as a way to foster shared emotional experiences, but must give users control over how visible those reactions are.","feed_headline":"Visible avatar laughter turns voice-only viewing into a group event","feed_subtitle":"In eight triads, visible laughter cues boosted involvement and emotional contagion over voice-only co-viewing.","key_machinery":"The central object is the user-triggered embodied laughter cue: a button-pressed avatar animation combining whole-body seated-laughter motion with Facial Action Coding System (FACS) facial expressions, including cheek raising, lip-corner pulling, and lips parting, displayed on each participant's avatar and visible to the whole triad through a mirror above the shared screen. The mirror lets each viewer see both their own avatar's laughter and co-viewers' reactions, creating a self-observation and mutual-observation loop. This machinery carries the argument by being the only systematic difference between the two conditions, and it operationalizes emotional contagion as chained laughter events, a follow-up laughter cue triggered within five seconds of another participant's cue.","core_discovery":"The paper's central claim is that visible, user-triggered avatar laughter, not the voice channel alone, is what shifts triadic VR co-viewing of comedy from individual immersion to socially coordinated participation. In the comparison of Voice Only versus Voice + Laughter Cues, the expressive condition significantly increased self-reported Involvement, Endurability, and Novelty and significantly increased the rate of chained laughter events, with no significant change in focused attention or perceived usability. The qualitative analysis argues that these cues heighten self-awareness of one's own emotional expression, create emotional contagion through observable motor mimicry, and lead groups to develop normative expectations about when and how to laugh. The authors present the downside as integral to the finding: with visible expressive avatars, participants reported social pressure to conform, hesitation about initiating laughter, and occasional deliberate withholding of reactions to preserve personal boundaries.","pith_inferences":["Editorial inference: the same visible-cue mechanism should transfer to other emotions and genres if the avatar animation matches the target emotion; a horror or concert setting with user-triggered fearful or awe expressions would be a direct test.","Editorial inference: the five-second chained-laughter count likely undercounts contagion in the voice-only condition, where smiles and silent laughs can follow audio cues without a logged button press; adding head-tracking or vocal-pitch logs would sharpen the comparison.","Editorial inference: the tension between connection and conformity suggests a tunable emotional-visibility control as a design direction, with the prediction that users will choose different visibility settings for close ties versus strangers and for comedy versus sensitive content."],"forward_implications":["If the central claim is right, adding user-triggered embodied laughter cues to remote co-viewing will raise involvement, endurability, and novelty relative to voice-only co-viewing, without necessarily sharpening focused attention on the content.","Designers should expect visible expressive cues to generate social norms: viewers monitor co-viewers and accommodate the timing of their laughter, so systems need per-user and per-content control over how visible emotional expressions are.","Seeing one's own avatar laugh can intensify felt amusement, implying that avatar animation is not only a signal to others but also a way to shape the viewer's own emotional experience.","For content a viewer wants to enjoy privately or react to honestly, an invisible or low-expressivity mode is likely preferable, so co-viewing systems should support both visible and hidden expressive modes."],"supporting_citations":[{"why":"Measures triad closeness with the Inclusion of Other in the Self scale, supporting the claim that the groups were acquainted as intended.","marker":"[3]"},{"why":"Documents how mismatched auditory and visual cues can undermine social presence, which the authors cite to justify not including an idling-avatar control.","marker":"[5]"},{"why":"Provides the thematic-analysis procedure used to derive the interview themes that carry the qualitative claims.","marker":"[11]"},{"why":"Defines the Facial Action Coding System action units (AU6, AU12, AU25) used to build the avatars' facial laughter expressions.","marker":"[22]"},{"why":"Supplies the theory of communication accommodation used to interpret participants' convergence and divergence in laughter expression.","marker":"[28]"},{"why":"Defines emotional contagion and its link to visible motor mimicry, the core mechanism the study attributes to avatar laughter cues.","marker":"[38]"},{"why":"Provides the standardized engagement questionnaire whose subscales operationalize the engagement research question.","marker":"[61]"},{"why":"Supplies the prior laughter-animation design and VR comedy-viewing setup that this study extends to mutual, user-triggered co-viewer expressions.","marker":"[62]"}],"fun_headline_variants":["Visible avatar laughter turns VR co-viewing into a social event","Seeing avatars laugh makes VR comedy co-viewing more social","Avatar laughter cues shift VR co-viewing from solo to social","Embodied laughter in VR boosts group emotional contagion","Avatar laughter cues boost engagement but bring social pressure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the visible laughter animation, not simply having an avatar present, causes the observed shifts, because the expressive condition adds embodiment and visible expression together while the voice-only baseline has neither, and the authors explicitly chose not to include an idling-avatar control because mismatched audio and visual cues can undermine social presence.","fun_headline_variants_meta":{"raw":{"variants":["Visible avatar laughter turns VR co-viewing into a social event","Seeing avatars laugh makes VR comedy co-viewing more social","Avatar laughter cues shift VR co-viewing from solo to social","Embodied laughter in VR boosts group emotional contagion","Avatar laughter cues boost engagement but bring social pressure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000757,"raw_usage":{"total_tokens":3362,"prompt_tokens":943,"completion_tokens":2419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":2338}},"tokens_in":559,"tokens_out":2419,"duration_ms":20025,"temperature":1.0,"reasoning_tokens":2338,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:59:21.056027+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same study with three conditions, voice only, an idling avatar that is present but never animates laughter, and the expressive avatar, and compare chained laughter, engagement scores, and interview reports of social pressure. If the idling-avatar condition behaves like voice-only, the laughter animation is the causal ingredient; if it behaves like the expressive condition, avatar presence itself is the cause.","supporting_citations":[{"cited_title":"Bailenson, Kim Swinth, Crystal Hoyt, Susan Persky, Alex Dimov, and Jim Blascovich","cited_arxiv_id":null,"evidence_quote":"Documents how mismatched auditory and visual cues can undermine social presence, which the authors cite to justify not including an idling-avatar control."},{"cited_title":"2015.Communication Accommodation Theory","cited_arxiv_id":null,"evidence_quote":"Supplies the theory of communication accommodation used to interpret participants' convergence and divergence in laughter expression."},{"cited_title":"O’Brien and Elaine G","cited_arxiv_id":null,"evidence_quote":"Provides the standardized engagement questionnaire whose subscales operationalize the engagement research question."}],"review_version":1}