{"id":"cd55cdd9-b8f3-479a-87ff-7089358e7497","arxiv_id":"2508.17676","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A VR meeting system with an embodied agent standing in for an absent participant, who can later rewatch from the agent's viewpoint, increases perceived inclusion and access to absentees' input.","lead":"This paper introduces SEAM, a virtual reality meeting format in which absent participants are represented by embodied stand-in agents that answer questions during the meeting, and later review the recording from the agent's viewpoint. Two studies with 45 participants suggest the format helps attendees factor in absentees' views and helps absentees feel included.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'enhance' claim has no no-stand-in baseline, and the task instructed participants to include Lee; SEAM's benefit over current asynchronous practice is untested even under Wizard-of-Oz conditions.","rationale":"I agree with the reader that the WoZ/AI mismatch is a genuine limitation, but I do not think it is the most load-bearing point: the Wizard-of-Oz stand-in is the best-case stand-in, and if the data cannot show an enhancement even for that best case, the AI fidelity question is moot. The paper's own framing as a technology probe (Section 5) makes a no-baseline design defensible for exploration, but the abstract and conclusion make a comparative claim ('enhance meetings,' 'benefiting both present and absent collaborators') that the design cannot support. The task instruction requiring participants to take Lee's preferences into account directly produces the 14/15 decision-inclusion result, and the self-report inclusion measures in Study 1 show no clear quantitative effect (Table 3: -0.33 [-0.89, 0.22]). Study 2 provides qualitative evidence that absentees felt included, but participants watched recordings in which attendees had been instructed to engage the stand-in, again without a no-stand-in control. The missing baseline is therefore not a minor add-on; it is the reference point against which 'enhancement' is defined. A controlled comparison (concrete_test) would settle this. Because the paper is explicitly exploratory and the concept may still be valuable, the reader's CONDITIONAL verdict remains appropriate; my concern does not change the verdict, but it should be written into the conditions.","tokens_in":31446,"tokens_out":5524,"duration_ms":63976,"concrete_test":"Run a between-subjects follow-up of Study 1 with three arms: (A) stand-in as in the paper, (B) no stand-in, with Lee's preferences provided as a written note to both attendees (current practice), and (C) Lee present synchronously. Keep the same weekend-planning task and same instructions except remove the explicit instruction that all three preferences must be considered, or hold it constant across arms if the comparison is about presence rather than information access. Pre-register the same cumulative-probit analyses as Tables 1-4, plus a measure of whether the group's final choice matches Lee's preferences.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is comparative: the stand-in 'can enhance meetings' (abstract, Section 11). Yet neither study includes a condition in which an absentee is absent without a stand-in, or in which attendees catch up via notes or recordings, which is the baseline SEAM is meant to improve upon. Section 5 states the authors deliberately did not use a baseline condition, and the task instructions (Section 5.2) told participants that by the end of the meeting they 'should have agreed ... taking all three participants' preferences into account.' Consequently, the headline result that 14/15 groups incorporated Lee's preferences is partly a compliance effect and cannot be attributed to the stand-in. The quantitative social-presence comparison (Table 1) only contrasts the stand-in and recorded avatars against a live human attendee, showing a gap of -0.75 [-0.97, -0.61] SD for the stand-in; it never shows whether a stand-in raises perceived presence above an absentee who is merely mentioned in notes. Thus, even if the Wizard-of-Oz stand-in behaved perfectly, the paper's data would not establish that SEAM enhances meetings relative to current asynchronous practice. The WoZ/AI gap (Section 10) is real but secondary: it concerns fidelity of the simulated stand-in, whereas the missing baseline concerns the validity of the enhancement claim itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SEAM (Stand-in Enhanced Asynchronous Meetings), a vision for asynchronous VR meetings in which an absent participant is represented by an embodied stand-in agent. Attendees interact with the stand-in during the meeting; later, the absentee watches the recording from the stand-in's first-person perspective and can add responses. The manuscript reports two mixed-method studies with 45 participants using a Wizard-of-Oz stand-in: Study 1 has pairs of attendees conducting a decision-making task with a stand-in and then watching the recording from the stand-in's viewpoint, while Study 2 has new participants watch the Study 1 recordings as true absentees. The qualitative analysis identifies themes about interaction strategies, social presence, decision-making, inclusion, and exclusion; the quantitative analysis uses Bayesian ordinal models to compare social presence ratings across conditions. The paper concludes that the stand-in 'can enhance meetings,' presents design implications and ethical/trust considerations, and describes a follow-on personalized LLM-powered stand-in system. The central stated contribution is a proof-of-concept user-experience exploration of embodied asynchronous meetings rather than a controlled comparative evaluation.","tokens_in":31592,"tokens_out":5652,"duration_ms":59643,"significance":"If the enhancement claim were supported, the paper would make a useful contribution to CSCW and HCI: it opens a design space for asynchronous embodied meetings, provides a working VR prototype, and offers concrete design considerations for stand-in representation, playback, and trust. The paper has genuine strengths: two complementary studies, a systematic inductive qualitative analysis, Bayesian ordinal models with compatibility intervals and explicit caveats, clear reporting of limitations, and a detailed system description. The largest weakness is that the headline comparative claim is not backed by a no-stand-in baseline condition, and the stand-in behavior was manually enacted by a researcher, so the results characterize a simulated stand-in rather than a real AI-powered one. With appropriate reframing or an additional baseline condition, the work could be a solid exploratory contribution; in its current form, the abstract and conclusion overstate the evidence.","major_comments":[{"comment":"The paper's headline claim that 'the stand-in can enhance meetings' (abstract; Section 11) is comparative, but the studies include no baseline condition in which an absentee is absent without a stand-in or in which attendees catch up via notes or recordings. Section 5 explicitly states no baseline was used, and Section 5.2 instructed participants that by the end of the meeting they 'should have agreed ... taking all three participants' preferences into account,' so the observation in Section 7.2.2 that 14/15 groups incorporated Lee's preferences is partly a compliance effect rather than evidence attributable to the stand-in. Table 1 only compares the stand-in and the recorded avatar against a live human attendee (-0.75 [-0.97, -0.61] SD for the stand-in on the perception-of-self subscale); it never shows whether the stand-in raises perceived social presence above a no-stand-in asynchronous practice. Consequently, the current data support an exploratory account of how attendees and absentees experience a WoZ stand-in, but they do not establish that SEAM enhances meetings relative to current asynchronous practice. I recommend either adding a baseline condition or reframing the contribution as an exploration of user experience with the enhancement claim explicitly qualified.","section":"Section 5, 5.2, 7.2.2, 11"},{"comment":"The stand-in in both user studies was Wizard-of-Oz controlled: Section 4.4 states the researcher manually played back recorded responses, and Section 10 concedes that the approach 'does not reflect the constraints, unpredictabilities, and response delays inherent in real-world AI systems.' Section 9 then presents a personalized LLM stand-in system, but that system was not evaluated with participants in the reported studies; its described behaviors (e.g., 'manage the discussion' and 'contribute input if topics of interest are mentioned') are claims about the prototype, not about measured user experience. The central user-experience findings therefore hold only for the manually enacted stand-in, and the paper should not imply that the LLM-powered version necessarily produces the same effects. Please either add an evaluation of the Section 9 system or clearly delimit all results to the technology probe.","section":"Section 4.4, 9, 10"},{"comment":"The model specification for the social-presence analysis is inconsistent: the text gives 'Social presence ~ Attendance + (1|Participant) + (1|Questions)', Table 1's caption includes an additional '(1|Meetings)' term, and Table 5 reports the model without the Meeting random intercept. Because the reported compatibility intervals depend on which random effects are included, please align the formula in the text, tables, and appendix and confirm which random effects were used for the estimates in Table 1.","section":"Section 7.1, Table 1, Table 5"}],"minor_comments":[{"comment":"The sentence 'Present attendees can easily access information that drives decision-making in the meeting perceive high social presence of absentees' is missing a verb or punctuation; it should be rewritten for clarity.","section":"Abstract"},{"comment":"Several subsection headings render as garbled placeholder characters (e.g., the heading following 'Non-verbal behaviours are crucial for face-to-face collaboration'), making related-work content difficult to read; the production/PDF encoding should be fixed.","section":"Section 2"},{"comment":"The instruction 'taking all three participants' preferences into account' conflates the stand-in with Lee as a participant; since the paper later uses the 14/15 result as evidence, clarify how Lee was described to participants and whether they understood Lee as a person they should include.","section":"Section 5.2"},{"comment":"The caption of Figure 12 says the intervals are 95% compatibility intervals, while the text says they are 'not 95% confidence intervals'; reword this to avoid confusing Bayesian compatibility intervals with frequentist confidence intervals.","section":"Section 7.1"},{"comment":"The statement about the authors' own two months of use over eleven asynchronous meetings is anecdotal; it should not be offered alongside the user-study findings without an explicit label as informal first-hand experience.","section":"Section 8.1"},{"comment":"The co-location caveat (participants performed the Study 1 task in the same room) should be stated where the quantitative social-presence results are first reported in Section 7.1, not only in Limitations.","section":"Section 10"}],"recommendation":"major_revision","confidential_remarks":"The missing baseline is the paper's most consequential issue, and I agree with the skeptic's framing that even a perfect Wizard-of-Oz stand-in does not, by itself, validate the comparative 'enhance' claim. The paper is nonetheless a reasonable exploratory design contribution, and the qualitative data are rich enough that the manuscript could become acceptable with a substantial reframing or an additional no-stand-in condition. I do not see circular reasoning or unsupported novelty claims beyond the baseline problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read SEAM as an exploratory design study, not a comparative evaluation, so I was less bothered by the Wizard-of-Oz setup. The novel combination here is real: an embodied stand-in that attendees can converse with, plus later first-person replay from that stand-in's viewpoint for delayed participation. The closest prior work, Ditto, stayed in 2D and never put the absent user back in the scene; Wang et al. transformed recordings but did not have a live stand-in. The paper also gives a fairly complete description of an LLM-powered prototype, and the qualitative results are genuinely interesting—attendees alternated between treating the stand-in as human and as a limited bot, and absentees felt included largely because of eye contact and listening behaviors. The authors report their Bayesian ordinal models carefully, with compatibility intervals and an explicit warning against significance claims. They also acknowledge in Section 10 that the WoZ stand-in diverges from real AI constraints, and that participants were mostly students in the same room. So the paper is honest about its scope.\n\nThe soft spot is not the WoZ gap; it is that the central claim of 'enhancing' meetings is untested. Study 1 had no condition without a stand-in, and the task instructions explicitly told participants to take all three preferences into account. That makes the 14/15 groups outcome a compliance finding, not evidence that the stand-in drove inclusion. The quantitative comparisons show only that a stand-in yields lower social presence than a live human, not that it beats taking notes or watching a Zoom recording. So the abstract's 'can enhance meetings' goes beyond the data. This is a common exploratory-paper problem, and it is fixable: the authors could reframe the claim as 'a proof of concept with promising qualitative indicators' or add a modest comparison to an equivalent video-based condition.\n\nI would send this to review. It deserves referee time because the design concept is timely and the implementation detail is useful, but I'd ask for the claims to be scaled back or the comparison added. If I were doing asynchronous-collaboration work, I'd cite it for the concept and the prototype description, not for the empirical 'enhancement' result.","headline":"A worthwhile proof-of-concept for a novel asynchronous VR meeting format, but the headline 'enhancement' claim is not supported because there is no baseline condition.","tokens_in":32176,"tokens_out":1532,"would_cite":true,"duration_ms":19107,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Embodied stand-ins let absent colleagues stay in VR meetings—and catch up in first person.","keywords":["asynchronous collaboration","virtual reality","embodied conversational agent","stand-in","social presence","meeting recordings","Wizard-of-Oz","first-person perspective"],"falsifier":"Deploy the LLM-powered stand-in system the paper describes in a real multi-iteration meeting and count how often attendees address it after it makes an incorrect or delayed response; if addressing rates fall to near zero or absentees report that the stand-in's errors made them feel misrepresented, the claimed enhancement does not survive the move from Wizard-of-Oz to autonomous AI.","tokens_in":31203,"feed_emoji":"🤖","tokens_out":5210,"duration_ms":55286,"temperature":0.7,"pith_summary":"This paper introduces SEAM—Stand-in Enhanced Asynchronous Meetings—in which a colleague who cannot attend a VR meeting sends an embodied virtual agent, configured with their responses to the agenda, to take their place. Attendees address the stand-in directly; later, the absent colleague watches the recording from the stand-in's first-person perspective and can pause it to record their own response. Through two Wizard-of-Oz studies with 45 participants, the authors aim to establish that this arrangement benefits both sides: present attendees gain access to preferences that push the decision forward, and absentees feel included by the social attention their stand-in attracted. The paper is a proof of concept for asynchronous meetings that feel more like synchronous conversations.","feed_headline":"Stand-in agents bring absent colleagues into VR meetings","feed_subtitle":"Two studies with 45 users find stand-ins let decisions move forward and absentees feel included.","key_machinery":"The load-bearing mechanism is the stand-in itself: an embodied conversational agent that turns its gaze toward the active speaker, nods, and delivers pre-recorded responses to agenda items while shrugging, pointing, or gesturing to reinforce meaning. Because these behaviours are triggered in real time during the meeting, the stand-in is not a recording; it adapts to the conversation and gives attendees something to address. The second mechanism is the first-person playback: the absentee re-experiences the meeting from the stand-in's exact position and can pause playback to record a spoken, embodied reply, which is timestamped for inclusion in later iterations. Together they convert an absence into a deferred conversational presence.","core_discovery":"The central discovery is that a stand-in can make an absent participant conversationally available without being synchronously present. In the studies, attendees addressed the stand-in as if it were the absentee, used its preconfigured answers as genuine input in deciding on a shared plan, and many reported perceiving three people in the room rather than two. Absentees who later watched the recording from the stand-in's viewpoint reported feeling included, because attendees had looked at, asked, and waited for their stand-in; the embodied attention, eye contact, and listening behaviours carried that feeling. The same data shows a measurable gap: the stand-in's social presence was rated lower than a live attendee's (about -0.75 standard units on the Networked Minds scale), and even a recording of real people retained some of that gap, so the paper's claim is that stand-ins enhance—not fully replicate—presence.","pith_inferences":["Inference: the same design could extend to partial participation—a stand-in that escalates questions to the absentee over mobile messaging mid-meeting—moving SEAM from all-or-nothing absence to granular availability.","Inference: because participants said they wanted stand-ins to learn from past meetings and to sound like the absentee, the concept's long-term value may lie less in the avatar body and more in the memory and voice models that make responses feel attributable to a specific person.","Inference: the study's 'trigger words' finding suggests that if real AI stand-ins require explicit address (like 'Hey Lee') to respond, attendees may come to treat them as voice assistants; a testable design fix is to give stand-ins ambient, attention-based triggers instead."],"forward_implications":["If the central claim holds, meeting software can let a missing stakeholder's preferences shape decisions during the live discussion instead of after the fact.","Absentees can catch up on meetings as an embodied participant rather than by reading notes or scrubbing a flat video, preserving the reasoning, affect, and attention that drove the decision.","Attendees will treat a well-behaved stand-in as a delayed participant, adapting their communication style as they discover what the stand-in can and cannot answer.","The measured social-presence gap implies that stand-ins will need better response intelligence and more human-like behaviour before they can stand in for someone in high-stakes meetings."],"supporting_citations":[{"why":"Supplies the Networked Minds social-presence inventory used for the quantitative comparison across attendee and absentee conditions.","marker":"[5]"},{"why":"The closest prior concept: Ditto, a personalized embodied agent for absent participants in 2D meetings; SEAM positions itself against its 2D, synchronous-only scope.","marker":"[34]"},{"why":"Establishes causality-preserving asynchronous replay in mixed reality, the basis for recording and later re-experiencing spatial meetings.","marker":"[17]"},{"why":"Shows how nonverbal transformations of recorded VR interactions affect social presence, motivating the first-person playback design.","marker":"[60]"},{"why":"Just-in-time information retrieval agents, cited as evidence that timely context-aware information in meetings improves decision-making.","marker":"[48]"},{"why":"The general inductive approach used to analyze the qualitative interview and observation data.","marker":"[56]"},{"why":"The brms Bayesian mixed ordinal regression software used to estimate the social-presence effects.","marker":"[8]"}],"fun_headline_variants":["VR stand-ins let absentees attend meetings by proxy","Agent stand-ins keep meetings moving when you're away","SEAM: your VR agent takes your place in meetings","Stand-in agents give absentees a meeting presence","Absentee? Send your agent to stay in the VR loop"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The stand-in's behaviour was scripted and triggered by a researcher rather than produced by a real AI, so the user experiences reported here depend on the assumption that an autonomous stand-in would respond with comparable timing, accuracy, and naturalness.","fun_headline_variants_meta":{"raw":{"variants":["VR stand-ins let absentees attend meetings by proxy","Agent stand-ins keep meetings moving when you're away","SEAM: your VR agent takes your place in meetings","Stand-in agents give absentees a meeting presence","Absentee? Send your agent to stay in the VR loop"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1492,"prompt_tokens":881,"completion_tokens":611,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":532}},"tokens_in":497,"tokens_out":611,"duration_ms":6994,"temperature":1.0,"reasoning_tokens":532,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:01:09.285393+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deploy the LLM-powered stand-in system the paper describes in a real multi-iteration meeting and count how often attendees address it after it makes an incorrect or delayed response; if addressing rates fall to near zero or absentees report that the stand-in's errors made them feel misrepresented, the claimed enhancement does not survive the move from Wizard-of-Oz to autonomous AI.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Networked Minds social-presence inventory used for the quantitative comparison across attendee and absentee conditions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows how nonverbal transformations of recorded VR interactions affect social presence, motivating the first-person playback design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Just-in-time information retrieval agents, cited as evidence that timely context-aware information in meetings improves decision-making."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The general inductive approach used to analyze the qualitative interview and observation data."}],"review_version":2}