{"id":"252a5a27-2c8a-4abd-b753-7f54e0869605","arxiv_id":"2507.22241","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"People initiating professional conversations in social VR move through three steps (recognizing availability, capturing attention, breaking the ice), and success depends on how faithfully VR preserves the social meaning of non-verbal cues, a quality the authors call verisimilitude.","lead":"This study followed 16 professionals through 23 social VR events and interviewed them afterward, mapping how they start informal conversations in virtual professional settings. It introduces the idea of verisimilitude, or how faithfully virtual non-verbal cues keep their real-world social meaning, as the central factor shaping these interactions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'progressively higher verisimilitude' claim is underdetermined: the data show different challenges at each step, but no evidence that participants themselves experienced a monotonic escalation in verisimilitude requirements.","rationale":"The reader identified retrospective self-reports and researcher observation as the weakest assumption. I agree that the data source is a concern, but the more specific and load-bearing issue is the leap from those self-reports to the monotonic 'progressively higher verisimilitude' thesis. The paper's own evidence suggests that later steps are harder and that participants increasingly rely on non-verisimilar workarounds, which is compatible with a qualitative difference across steps but does not demonstrate an increasing requirement on a single verisimilitude dimension. This is not a challenge to the study's qualitative richness or to the usefulness of the three-step framing; those stand on their own. The concern is about the strength of the central theoretical claim as stated. A re-coding of the transcripts would settle whether the ordering is present in participant accounts or is an analyst-level interpretation. I recommend keeping the CONDITIONAL verdict, but making the condition explicit: the authors should either present transcript evidence of cross-step verisimilitude comparisons or soften the claim to reflect that different steps present different verisimilitude challenges. The paper is transparent, well-written, and offers valuable design implications, so rejection is not warranted; nevertheless, the headline contribution should not overstate what the data can support.","tokens_in":25783,"tokens_out":3831,"duration_ms":52153,"concrete_test":"Re-code the interview transcripts for explicit cross-step comparisons: identify every utterance in which a participant compares the trustworthiness, fidelity, or social meaning of cues across two or more of the three phases (availability recognition, attention capture, ice-breaking). If fewer than one third of the 16 participants make at least one such cross-phase comparison, the 'progressively higher verisimilitude' claim should be downgraded to 'different verisimilitude challenges at each step.' A stronger check: have two independent coders, blind to the paper's thesis, rate each participant's described need for verisimilitude in each phase on a 3-point scale, and report the proportion of participants whose within-participant pattern is monotonically increasing. If the proportion is not substantial, the central ordering claim fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim (Section 5, Figure 1) is that initiating opportunistic interactions in social VR comprises three successive steps, each requiring a progressively higher degree of verisimilitude than the preceding one(s). This monotonic ordering is load-bearing: it is the headline contribution and the basis for the design implications. However, the presented evidence does not establish it. Section 3.4 describes an inductive coding process that converged on 'verisimilitude' as a focal concept, but the paper offers no participant statement in which someone compares the required fidelity of cues across the three steps. The quotes in Sections 4.1-4.3 are about difficulties within each step, not about a perceived ordering of verisimilitude requirements. Moreover, the ice-breaking findings in Section 4.3 undercut the monotonic claim: participants succeeded through avatar customization and host-mediated warm-up activities, which are mechanisms that do not depend on high-verisimilitude non-verbal cues, and Section 4.3 states that 'the range of non-verbal cues, as well as the verisimilitude of those cues, are often insufficient.' Thus the data are equally or better described as revealing different kinds of verisimilitude challenges at each step, rather than a progressively higher requirement. Because 'degree of verisimilitude' is never operationalized across different cue types (gaze, handclap, emoji, proximity, avatar appearance), the monotonic claim risks being an analytic imposition rather than an empirical finding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative study of how people initiate opportunistic interactions at professional events in social VR. The authors shadowed 16 participants across 23 events on four platforms and conducted post-event interviews. From an inductive analysis, they propose that initiation comprises three successive steps—availability recognition, attention capture, and ice-breaking—and that each step requires a progressively higher degree of what they call verisimilitude, defined as the perceived preservation of real-world social meanings of non-verbal cues. They also describe strategies participants use at each step and propose design implications for platforms and hosts.","tokens_in":26064,"tokens_out":4310,"duration_ms":47164,"significance":"If the central model holds, the paper makes a valuable contribution by connecting theories of non-verbal communication to the understudied context of professional networking in social VR. Its strengths include rich empirical material, triangulation of observations with interviews, an auditing step in coding, and specific technical mechanisms (gaze rules, emoji lag, speech icons, handclapping paths) that ground the analysis. However, the load-bearing claim of a monotonic increase in verisimilitude across the three steps is not supported by the presented evidence, so the headline contribution needs revision before publication.","major_comments":[{"comment":"The claim that each step 'requires a progressively higher degree of verisimilitude' is not established by the data presented in Sections 4.1–4.3. The participant quotes describe difficulties within each step, but none of them compares the required fidelity of cues across steps, and Section 4.3.2 shows that ice-breaking succeeded through avatar customization and host-mediated warm-ups that do not depend on high-verisimilitude non-verbal cues. Because this monotonic ordering is the paper's central claim, it should either be supported with comparative evidence (e.g., explicit participant statements or cross-step observational comparisons) or softened to 'different kinds of verisimilitude challenges at each step.'","section":"Section 5 and Figure 1"},{"comment":"The paper never operationalizes 'degree of verisimilitude.' Verisimilitude is defined as the perceived preservation of real-world meanings, but no criterion is given for judging a higher versus lower degree, and the concept spans heterogeneous cue types (gaze, handclap, emoji, proximity, avatar appearance). Without such an analytic criterion, the monotonic-ordering claim is difficult to verify or falsify. The authors should specify how they compared verisimilitude across steps and cue types.","section":"Section 3.4 and throughout"}],"minor_comments":[{"comment":"The quotation introduced with 'As explained by P13 and several others' is attributed to P14 in the following line; the attribution and the introductory clause should be reconciled.","section":"Section 4.3.1"},{"comment":"The sentence 'one straightforward means, for example, to is to provide' contains a typo; it should read 'one straightforward means, for example, is to provide.'","section":"Section 5.1"},{"comment":"In the description of P6's 'give and take protocol,' the pronoun shifts from 'he' to 'she' within the same sentence; the referent should be corrected.","section":"Section 4.2.3"},{"comment":"The Gather.town study is cited as 'Sanchez et al.' in the text, but the reference list entry is 'Palos-Sanchez et al.'; the citation and reference list should be aligned.","section":"Section 2.3"},{"comment":"The text references 'IEEE VR 2021' but reference [2] is titled 'IEEEVR2020'; the year in the text should be checked against the cited source.","section":"Section 2.3"},{"comment":"The column headers in Table 1 appear garbled, with repeated entries such as 'Role Platform Role Platform'; the table formatting should be corrected.","section":"Table 1"},{"comment":"In the Limitations section, 'all participants in our study located in the United States' should be 'all participants in our study were located in the United States.'","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The empirical material is rich and the three-step structure is plausible, but the monotonic verisimilitude claim is the paper's main theoretical contribution and it is currently under-supported. I recommend asking the authors to either provide direct evidence of the escalation or reframe the contribution as identifying distinct verisimilitude-related challenges at different stages. The paper's fit with CSCW is good, and the ethics reflection is a notable strength, though it could be condensed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid CSCW paper with a genuinely useful empirical core, and it deserves a serious referee. The one piece I'd push back on is the headline monotonic claim, which the stress-test note nails.\n\nWhat's new: the paper is the first to look specifically at initiation of opportunistic interactions at professional events in social VR, not casual hangouts. The data are rich—23 observed events across 4 platforms, 16 participants, post-event interviews, triangulation, an audit step. The concrete technical mechanisms are valuable: platform-specific gaze rules (Altspace looks at sounds, Engage blinks on its own), emoji lag, speech icons as availability signals, Venu's handclap shortcuts, host-driven warm-ups like collective clapping and avatar cloning. Those grounded details are the contribution, and they support the three-step model (availability recognition, attention capture, ice-breaking) convincingly.\n\nThe soft spot is real. The data show different verisimilitude challenges at each step, but I don't see participants comparing required fidelity across steps, and the paper doesn't operationalize 'degree of verisimilitude' across cue types. So the 'progressively higher' wording in Section 5 and Figure 1 is an analytic gloss. Worse, the ice-breaking section undercuts it: avatar customization and host warm-ups work precisely because they don't depend on high-verisimilitude non-verbal cues, and Section 4.3 states cues are often insufficient. The framework survives if you recast it as 'different kinds of verisimilitude demands' rather than a monotonic ladder. That's a revision, not a rejection.\n\nOther concerns are minor: the US-only, self-selected sample is acknowledged, and the proxy-consent ethics for bystanders is a genuine grey area, but the authors flag it honestly in Section 6 and don't overclaim. The ethics reflection is better than most.\n\nWho's it for: CSCW/HCI people studying social VR, remote professional events, or non-verbal cues in virtual environments. It will be cited. Send it to review, with a request that the authors either find evidence for the escalation claim or soften it.","headline":"A well-executed qualitative study of how professionals initiate conversations in social VR; the three-step framework is useful, but the 'progressively higher verisimilitude' claim is the one piece that outruns the data.","tokens_in":26576,"tokens_out":2329,"would_cite":true,"duration_ms":27235,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Professionals initiating unplanned chats in social VR pass through three successive steps—availability recognition, attention capture, and ice-breaking—and each step demands a higher degree of verisimilitude from non-verbal cues.","keywords":["social VR","opportunistic interactions","professional events","non-verbal cues","verisimilitude","avatar-mediated communication","qualitative study","CSCW"],"falsifier":"Compare two otherwise identical VR networking events, one where avatar gaze follows the user's real gaze and one where gaze is machine-generated; the model predicts that only the machine-gaze version produces systematic false 'she looked at me' approach attempts. A larger behavioral trace should also show failed initiations clustering at distinct stages rather than spreading evenly across all moments.","tokens_in":25601,"feed_emoji":"🥽","tokens_out":7334,"duration_ms":83932,"temperature":0.7,"pith_summary":"This paper asks how professionals manage the unplanned conversations that matter for networking when the event happens in virtual reality. It argues that initiating such an interaction is not a single act, as it is face-to-face, but three successive steps: recognizing that someone is available, capturing their attention, and breaking the ice. Each step requires a progressively higher degree of verisimilitude, defined as the user's perceived extent to which social VR preserves the real-world social meaning of non-verbal cues. The authors base this account on shadowing 16 experienced users across 23 professional VR events and interviewing them afterward. The claim matters because current platforms replace the social machinery behind cues such as gaze, proximity, and timing with technical substitutes that quietly change what those cues mean, and designers and hosts currently lack a systematic map of where the breakdowns occur.","feed_headline":"VR networking chat fails in three stages, each needing more realism","feed_subtitle":"Professionals read gaze, claps, and avatar style before speaking; platforms quietly change what those cues mean.","key_machinery":"Verisimilitude is the paper's central lens: the extent to which a social VR environment preserves, in the user's perception, the real-world social meaning of a non-verbal cue. It is not objective fidelity but perceived meaning-preservation, so the same cue can carry different social weight for different users and platforms. The three-step model—availability recognition, attention capture, and ice-breaking—is the other load-bearing structure, and verisimilitude explains why progression is hard: each step relies on cues whose platform-level preconditions, such as user control of gaze, timing of gestures, multimodal sensing, and shared appearance, are progressively harder to preserve.","core_discovery":"On its own terms, the paper's central discovery is that people initiate opportunistic interactions in professional social VR through a three-step sequence: availability recognition, attention capture, and ice-breaking. At each step users lean on non-verbal cues—gaze and wandering for availability, claps, emojis, and proximity for attention, avatar appearance for ice-breaking—and judge whether those cues carry the same social meanings they would in physical settings. The platform's technical substitutions, such as emoji menus replacing facial expressions or system-controlled gaze replacing intentional gaze, disturb those meanings, and the disturbance is progressive: what is good enough to infer availability is not good enough to capture attention, and what suffices for attention is not enough to start a conversation. The paper concludes that verisimilitude is both a resource and a source of uncertainty, and that hosts, artificial status indicators, and embedded tutorials are the current and near-future workarounds.","pith_inferences":["The progressive-verisimilitude claim yields a sharp untested prediction: interventions that raise the realism of one cue should help most at the step where that cue is the primary tool, and have diminishing returns at later steps.","Because the participants were US-based and drawn from art and design, technology, and education communities, the model may fit those professional norms better than others; in more relationally relaxed or culturally different professional settings, the step boundaries could blur.","Prior work on casual social VR shows users initiating encounters through deliberately playful and sometimes wild gestures, which suggests the escalating-verisimilitude pattern may be specific to professional contexts rather than a general property of social VR.","The hosts' success at ice-breaking through collective rituals, such as synchronized clapping and avatar cloning, hints that shared playful action may achieve the social outcome of verisimilitude without realism, a route the paper leaves under-explored."],"forward_implications":["If the three-step model is right, platform designers cannot treat starting a chat as a single feature; each step needs its own cue support and its own failure diagnostics.","Because participants increasingly confirmed availability through system indicators such as speech icons, the design of those indicators directly shapes who gets approached and when.","Hosts already carry much of the ice-breaking burden, so supporting hosts with real-time awareness of quiet clusters or likely conversation hotspots would shift part of the social load onto the system.","A unified cross-platform policy for how non-verbal cues behave, analogous to standardized emoji, would reduce the trial-and-error cost users pay when moving between platforms.","Situated tutorials embedded in live events, rather than standalone manuals, are a plausible path to building users' skill at estimating verisimilitude."],"supporting_citations":[{"why":"Provides the prior evidence that social VR users initiate encounters through non-verbal cues; the paper carries this finding into professional events and adds the verisimilitude concern.","marker":"[43]"},{"why":"Sets up the face-to-face baseline in which initiating an interaction is a single, brief process, the contrast the three-step model is built against.","marker":"[59]"},{"why":"Documents that social VR is primarily used for meeting new people, framing the shift to professional networking contexts.","marker":"[22]"},{"why":"Shows casual social VR users creatively repurposing cues, the recreational baseline from which professional caution diverges.","marker":"[40]"},{"why":"Reports that 2D spatial videoconferencing leaves users short of interruptibility cues, motivating the search for richer non-verbal information in VR.","marker":"[56]"},{"why":"Supplies the iterative sampling and coding procedure through which the three steps and verisimilitude were derived.","marker":"[26]"}],"fun_headline_variants":["VR event networking: realism fails at three distinct stages","Three-step VR small talk hinges on realism of non-verbal cues","Gaze, claps, avatars: where VR social cues break down","Verisimilitude decides who gets spoken to at VR conferences","Why VR ice-breaking stalls even when availability works"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis rests on the premise that participants' post-event recollections and the researcher's observations faithfully capture what actually drove their decisions, including private judgments about whether a cue's social meaning survived the move into VR.","fun_headline_variants_meta":{"raw":{"variants":["VR event networking: realism fails at three distinct stages","Three-step VR small talk hinges on realism of non-verbal cues","Gaze, claps, avatars: where VR social cues break down","Verisimilitude decides who gets spoken to at VR conferences","Why VR ice-breaking stalls even when availability works"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000158,"raw_usage":{"total_tokens":1262,"prompt_tokens":1017,"completion_tokens":245,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":161}},"tokens_in":633,"tokens_out":245,"duration_ms":3732,"temperature":1.0,"reasoning_tokens":161,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:54:04.003297+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare two otherwise identical VR networking events, one where avatar gaze follows the user's real gaze and one where gaze is machine-generated; the model predicts that only the machine-gaze version produces systematic false 'she looked at me' approach attempts. A larger behavioral trace should also show failed initiations clustering at distinct stages rather than spreading evenly across all moments.","supporting_citations":[{"cited_title":"Talking without a Voice","cited_arxiv_id":null,"evidence_quote":"Provides the prior evidence that social VR users initiate encounters through non-verbal cues; the paper carries this finding into professional events and adds the verisimilitude concern."},{"cited_title":"Pillet-Shore","cited_arxiv_id":null,"evidence_quote":"Sets up the face-to-face baseline in which initiating an interaction is a single, brief process, the contrast the three-step model is built against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents that social VR is primarily used for meeting new people, framing the shift to professional networking contexts."},{"cited_title":"Palos-Sanchez, Pedro Baena-Luna, and Daniel Silva-O’Connor","cited_arxiv_id":null,"evidence_quote":"Reports that 2D spatial videoconferencing leaves users short of interruptibility cues, motivating the search for richer non-verbal information in VR."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the iterative sampling and coding procedure through which the three steps and verisimilitude were derived."}],"review_version":1}