{"id":"b6caa6f9-e871-43f7-9477-9c5acf8a1d79","arxiv_id":"2502.00880","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Viewers of a recorded VR inspection tour remember more and enjoy it more when they can control their viewpoint and perform guided interactions, compared with passive video playback.","lead":"The paper tests whether letting viewers interact with a recorded VR tour helps them remember it. It finds that independent viewpoint control improves experience and that some interactivity can aid recall, but the evidence is modest and partly confounded.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'additional interactivity improves recall' claim rests on comparisons against a video-playback condition that confounds interactivity with embodiment; INTERACTIVE never significantly beats RAILS or NAVIGATION on recall.","rationale":"The reader correctly identifies the PASSIVE-condition confound between interactivity and immersion, and flags the abstract's overstatement. My stress-test goes one step further: the paper's own pairwise results already contain the decisive null result. INTERACTIVE does not significantly outperform RAILS or NAVIGATION on auditory or spatial recall, and RAILS itself significantly outperforms PASSIVE on spatial recall. This means the recall advantages attributed to 'additional interactivity' are not separable from the advantages of embodied viewpoint control. The paper's discussion in Section 6.1 partially acknowledges the viewpoint interpretation, and Section 6.2 explicitly attributes the experience differences to the distinct design of PASSIVE, which is honest and helpful. The central direction of the findings—that embodied viewpoint control improves user experience and spatial recall relative to passive video—is defensible, but the abstract's phrasing overreaches. The paper deserves a CONDITIONAL verdict with a required revision of the abstract and a request for effect sizes or confidence intervals, rather than rejection, because the data do support the weaker and still meaningful claims about embodiment and viewpoint control.","tokens_in":17865,"tokens_out":3492,"duration_ms":37705,"concrete_test":"Reanalyze the existing dataset to compute pairwise effect sizes and 95% confidence intervals for INTERACTIVE vs RAILS and INTERACTIVE vs NAVIGATION on auditory and spatial recall, using the same ART-C/Tukey procedure as Sections 5.5 and 5.6. If the confidence intervals include zero and the comparisons remain non-significant, the abstract's claim that additional interactivity improves recall should be revised. To fully settle the causal question, run a new condition in which participants are embodied in the VE but watch the guide's first-person video on a floating display, matching embodiment while removing independent viewpoint control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's strongest claim is that 'additional interactivity can improve auditory and spatial recall of key information conveyed during the tour.' In the reported analyses, however, this claim is supported only by pairwise comparisons that use the PASSIVE condition as the baseline. In Section 5.5, INTERACTIVE differs from PASSIVE on auditory recall, and in Section 5.6 both INTERACTIVE and RAILS differ from PASSIVE on spatial recall. Critically, INTERACTIVE is not reported as significantly better than RAILS or NAVIGATION on either recall measure. Because the PASSIVE condition (Section 4.3) is not merely a non-interactive version of the other conditions: it is a first-person video shown on a virtual display 1 m away in an empty VE, with no embodied viewpoint control, no spatial presence, and no ability to move through the environment. Thus every PASSIVE-vs-other difference conflates the effect of added interactivity with the effect of embodied viewpoint control and immersion. The correct control for isolating 'additional interactivity' is RAILS, which provides embodied viewpoint control without playback interactions; the data show no significant recall advantage for INTERACTIVE over RAILS. At most, the data support the weaker conclusion that embodied viewpoint control improves experience and spatial recall relative to video playback, and that INTERACTIVE improves auditory recall relative to video playback, but there is no statistically supported claim that adding interactivity beyond viewpoint control improves recall.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a VR guided-tour system for asynchronous collaboration in additive manufacturing inspection, in which a recorded guiding avatar narrates an inspection and pauses playback for the observer to perform mimicry interactions. A between-subjects study (N=40, n=10 per condition) compared four conditions: PASSIVE (first-person video on a virtual display), RAILS (embodied viewpoint without playback interactions), NAVIGATION (multiscale navigation interactions), and INTERACTIVE (navigation plus perspective-sharing interactions). Dependent measures were custom verbal and spatial recall tests, SSQ, UES-SF, NASA-TLX, a 1-10 revisit rating, and an interview. The authors report that INTERACTIVE significantly outperformed PASSIVE on verbal recall, that INTERACTIVE and RAILS significantly outperformed PASSIVE on spatial recall, and that PASSIVE was rated lower on several experience dimensions. They conclude that viewpoint control improves experience and that additional interactivity can improve recall.","tokens_in":18077,"tokens_out":4609,"duration_ms":44487,"significance":"The work addresses a genuinely under-studied area: asynchronous collaboration in CVEs via guided tours, and it contributes a concrete interaction design with four incremental levels plus a mixed-method evaluation. The statistical reporting is generally careful: nonparametric tests, Holm-Bonferroni correction for questionnaires, ART ANOVA with Tukey-corrected contrasts for recall data, and explicit hypotheses stated before the experiment. If the recall findings were supported, they would justify design guidance for interactive playback in asynchronous VR tours. However, the strongest claimed result, that extra interactivity beyond viewpoint control improves recall, is not supported by the reported pairwise contrasts, because the only significant recall differences are against PASSIVE, a condition that lacks both interactivity and embodied viewpoint control. The study still provides useful exploratory evidence and qualitative insight, but the conclusions must be substantially reframed.","major_comments":[{"comment":"The abstract and conclusion claim that 'additional interactivity can improve auditory and spatial recall,' but no significant difference between INTERACTIVE and RAILS or between INTERACTIVE and NAVIGATION was found on either recall measure. RAILS has no playback interactions but does have embodied viewpoint control; therefore the only comparison that isolates added interactivity from viewpoint control, INTERACTIVE vs. RAILS, is non-significant. The sentence in §6.1 ('The evidence, therefore, suggests that additional interactivity beyond simple viewpoint control can help increase observers' recall') is not statistically supported by the reported contrasts. Please reframe the central claim as 'embodied viewpoint control improves recall relative to passive video playback' and present any benefit of additional playback interactions as exploratory.","section":"§5.5, §5.6, §6.1"},{"comment":"The PASSIVE condition is confounded: it is a first-person video displayed on a virtual screen one meter away in an empty VE, whereas all other conditions place the participant inside the environment with head-tracked viewpoint control. Consequently, every PASSIVE-versus-other comparison in §§5.2, 5.4, 5.5, and 5.6 conflates playback interactivity with immersion and embodiment. This is not merely a wording issue; it undermines the attribution of recall and experience differences to interactivity. The correct no-interactivity control is RAILS, which provides embodied viewpoint control without playback interactions. The manuscript should present RAILS as the primary baseline for isolating interactivity and explicitly state that PASSIVE differences cannot be attributed to interactivity alone.","section":"§4.3"},{"comment":"The recall instruments are short (three verbal and three spatial items per test), the verbal items are binary-scored with no partial credit, and no pilot validation or reliability evidence is reported. With n=10 per condition, the null comparisons (e.g., INTERACTIVE vs. RAILS) have low power and cannot be interpreted as evidence of no difference. The paper should either report sensitivity or power analyses with confidence intervals for the key contrasts, or explicitly label the recall-related conclusions as preliminary and hypothesis-generating.","section":"§4.2, Table 1, §5.5"},{"comment":"The discussion describes the PASSIVE condition as a 'third-person perspective,' but §4.3 defines it as a first-person video shown on a virtual display one meter away. This inconsistency reflects the same confound and should be reconciled. In addition, the abstract's phrase 'fully passive playback' should be qualified: the PASSIVE condition is not only passive in interaction but also lacks embodied viewpoint control and physical immersion, so the phrase overstates what the comparison can show.","section":"§6.2"}],"minor_comments":[{"comment":"The name 'Gaylean' should be 'Galyean' when referring to the guided-navigation author.","section":"§2.1"},{"comment":"The tests are called 'Test A' and 'Test B' in Table 1 but 'Test 1' and 'Test 2' in the procedure section; please use consistent labels.","section":"Table 1 and §4.5"},{"comment":"The phrase 'a significant pairing with the I NTERACTION condition' should read 'the INTERACTIVE condition.'","section":"§6.2"},{"comment":"The figures use asterisks for significance but do not report effect sizes or confidence intervals; adding these, or at least reporting the mean and standard error in the text for each significant pair, would strengthen the presentation.","section":"Figures 3 and 4"},{"comment":"The name 'RAILS' is never explained; a one-sentence definition of the term would help readers understand the condition name.","section":"§4.3"},{"comment":"Some references contain encoding artifacts (e.g., 'Fr´econ & N¨ou'); please clean the LaTeX source so author names render correctly.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The PASSIVE-condition confound is the central issue. The actual study and system are valuable: the interaction designs are well specified, the qualitative data are rich, and the nonparametric analysis is careful. A revision that reframes the abstract and conclusions around embodied viewpoint control rather than playback interactivity, and that uses RAILS as the key baseline, could make the paper publishable without new data collection. The authors must remove unsupported causal claims about 'additional interactivity' and report the INTERACTIVE-versus-RAILS contrast transparently. The novelty is modest but appropriate for IEEE VR."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this one for its design space, not for its abstract. The paper compares four levels of playback interactivity in a VR guided tour for asynchronous collaboration, and that systematic ordering is genuinely new—prior work mostly looked at annotations or synchronous collaboration. The mimicry interactions (navigate, point, align heads, dock a cutting plane) are thoughtfully designed, and the study itself is clean by HCI standards: hypotheses pre-registered in the text, appropriate nonparametric tests with corrections, and transparent reporting of significant pairs.\n\nThe results that do hold up are about passive vs. embodied viewing. The PASSIVE condition—a first-person video on a virtual display in an empty VE—is worse on engagement, usability, reward, frustration, and spatial recall than the other conditions. That is a useful, defensible design guideline: if you record a tour, let the viewer stand in the space and control their viewpoint.\n\nThe problem is the abstract's claim that \"additional interactivity can improve auditory and spatial recall.\" That is supported only by pairwise comparisons against PASSIVE. INTERACTIVE does not significantly beat RAILS on either recall measure, and RAILS is the proper control for isolating interactivity beyond viewpoint control. So the incremental-interactivity claim is not statistically supported. The authors come close to admitting this in Section 6.1, but the abstract still overstates.\n\nOther soft spots are minor in comparison: n=10 per cell, custom short recall tests without validation, and no effect sizes or confidence intervals. The authors acknowledge the sample size in limitations. I would have liked the data or analysis scripts to be released, but their absence is not disqualifying.\n\nThis paper deserves a serious referee. The contribution is practical, the system is real, and the empirical work is honest. I would ask for revisions: temper the abstract, reframe the analysis around the RAILS comparison as the baseline for interactivity, and ideally add effect sizes. With those changes, it is a solid IEEE VR paper. For anyone studying asynchronous collaboration or VR guided tours, the condition ordering and qualitative feedback are worth citing.\n\nRoss","headline":"A competent VR user study whose headline claim about interactivity improving recall overreaches: the only significant recall gains are against a confounded passive-video baseline, not against the viewpoint-control condition.","tokens_in":18612,"tokens_out":2094,"would_cite":true,"duration_ms":23662,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding required interactions to a recorded VR tour—pausing playback until the observer mimics the guide's navigation and pointing—improves what the observer remembers and how engaged they feel, relative to passive playback.","keywords":["virtual reality","asynchronous collaboration","guided tours","playback interactivity","information recall","user experience","additive manufacturing","collaborative virtual environments"],"falsifier":"Run a follow-up experiment with a fifth condition: the same first-person guide video played in the 3D environment with full head-tracked viewpoint control but no movement and no required interactions. If that condition's spatial and auditory recall match the RAILS (free-viewpoint, no-interaction) condition, then the benefits come from viewpoint control rather than from the playback interactions; if it matches PASSIVE, the benefits come from being physically inside the environment rather than watching a flat screen.","tokens_in":17669,"feed_emoji":"🥽","tokens_out":6830,"duration_ms":64274,"temperature":0.7,"pith_summary":"This paper claims that adding required interactions to a recorded VR guided tour—pausing playback until the observer mimics the guide's navigation and gaze—improves what the observer remembers and how engaged they feel, relative to passive playback. The evidence comes from a between-subjects study of 40 participants in four conditions: fully passive video, free viewpoint with no interactions, viewpoint plus navigation interactions, and full interactivity. Interactive viewing significantly improved auditory recall over passive viewing, and both full interactivity and mere free viewpoint improved spatial recall over passive viewing. Passive viewing was also rated significantly worse on engagement, usability, and reward, and produced more simulator sickness. If right, the work gives concrete guidance for designing asynchronous VR collaboration systems: build in playback interactions and let observers control their own viewpoint.","feed_headline":"Interactive VR tours sharpen recall and engagement","feed_subtitle":"Guided tour study shows letting observers mimic the guide beats video playback.","key_machinery":"The central object is the 'guide mimicry' playback system: the tour pauses at key moments (scale changes, pointing, viewpoint sharing, and plane placement) and resumes only when the observer reproduces the guide's action—navigating down or up, pointing at the guide's controller, aligning their head in the guide's head (WYSIWIS), or docking the cutting plane. This mechanism forces the observer's motor and attentional system to re-enact the guide's exploration, which the paper links to embodied cognition accounts of perception. It also gives the experiment its independent variable: the amount of mimicry required.","core_discovery":"The central discovery is that the extent of interactivity in guided tour playback matters measurably for recall and experience. Participants who could see the tour from their own viewpoint in the environment (RAILS) and those who additionally had to perform the guide's actions (INTERACTIVE) both located defects significantly more accurately than participants who watched a first-person video on a flat virtual display (PASSIVE). Those who performed the full set of mimicking interactions also recalled significantly more of the guide's spoken content than the passive group. The paper interprets this as evidence that embodied, active exploration—not just viewing—supports information retention in asynchronous VR collaboration, and it concludes that independent viewpoint control is the main driver of user experience, with additional interactivity adding recall benefits.","pith_inferences":["The study's own null result—viewpoint plus navigation interactions did not beat passive viewing on spatial recall—suggests the spatial benefit may come from viewpoint control rather than from the interactions themselves; a follow-up that adds head-tracked viewpoint to passive video playback would separate these factors.","If the mimicry mechanism works through embodied attention, then the number and timing of pauses likely matter: too many required actions could interrupt narrative flow, and the paper notes some participants navigated ahead of the guide, so adaptive or time-limited interaction prompts are worth testing.","The results may generalize beyond manufacturing to any narrated spatial task—medical anatomy tours, safety walkthroughs, or historical reconstructions—where later observers must remember both what was said and where things are located.","Because the same tour content was used across all conditions, the measured differences are attributable to delivery mode; applying the same interaction design to user-generated recordings rather than confederate recordings could test whether real experts' tours benefit the same way."],"forward_implications":["Asynchronous expert tours in VR—such as inspection walkthroughs—should let the later observer control their own viewpoint, not just play back a first-person video.","Adding required mimicry interactions to playback, pausing until the user performs the guide's navigation or pointing, can measurably raise recall of the tour's spoken content.","Spatial recall benefits from at least free viewpoint in the environment; the strongest spatial recall in this study came with full interactivity, though not significantly better than free viewpoint alone.","Passive video-in-VR playback risks lower engagement, higher frustration, and more simulator sickness; designers should avoid it as a baseline.","Recall benefits appeared both immediately and after a buffer activity, suggesting the improvements are not merely short-lived.","No significant differences were found between conditions on simulator sickness increases overall, though the passive group showed significant pre-post increases on all subscales.","The interactive condition produced the highest average verbal and spatial recall scores, even though only the comparison with passive viewing reached significance."],"supporting_citations":[{"why":"Supplies the author-controlled navigation concept that the guided tour playback extends.","marker":"[14]"},{"why":"Provides the user-preference evidence for control in immersive experiences, motivating the interaction design.","marker":"[40]"},{"why":"Direct comparison of viewpoint control versus video training, a baseline for the spatial recall results.","marker":"[33]"},{"why":"Supplies the embodied cognition framework used to explain why mimicking actions improves recall.","marker":"[46]"},{"why":"Closest prior system with a presenting avatar for asynchronous assembly tasks, which the guided tour builds on.","marker":"[36]"},{"why":"Identifies the gap in asynchronous collaboration research that this work addresses.","marker":"[24]"},{"why":"Provides the named WYSIWIS perspective-sharing concept used in the interactive condition.","marker":"[48]"}],"fun_headline_variants":["Interactive VR playback boosts recall and experience","VR tour interactivity improves recall and satisfaction","Active control in VR tours enhances memory and engagement","Mimicking guide in VR tour boosts recall and enjoyment","Independent viewpoint in VR tours aids recall and experience"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study treats the four conditions as differing only in how much interaction they allow, but the PASSIVE condition also removes the observer's own embodied viewpoint in the environment, so the recall and engagement gains attributed to interactivity could instead come from that difference in viewpoint control.","fun_headline_variants_meta":{"raw":{"variants":["Interactive VR playback boosts recall and experience","VR tour interactivity improves recall and satisfaction","Active control in VR tours enhances memory and engagement","Mimicking guide in VR tour boosts recall and enjoyment","Independent viewpoint in VR tours aids recall and experience"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1307,"prompt_tokens":825,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":412}},"tokens_in":441,"tokens_out":482,"duration_ms":4958,"temperature":1.0,"reasoning_tokens":412,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T17:21:03.584898+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a follow-up experiment with a fifth condition: the same first-person guide video played in the 3D environment with full head-tracked viewpoint control but no movement and no required interactions. If that condition's spatial and auditory recall match the RAILS (free-viewpoint, no-interaction) condition, then the benefits come from viewpoint control rather than from the playback interactions; if it matches PASSIVE, the benefits come from being physically inside the environment rather than watching a flat screen.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the author-controlled navigation concept that the guided tour playback extends."},{"cited_title":"Pausch et al","cited_arxiv_id":null,"evidence_quote":"Provides the user-preference evidence for control in immersive experiences, motivating the interaction design."},{"cited_title":"Lovreglio et al","cited_arxiv_id":null,"evidence_quote":"Direct comparison of viewpoint control versus video training, a baseline for the spatial recall results."},{"cited_title":"Shapiro et al","cited_arxiv_id":null,"evidence_quote":"Supplies the embodied cognition framework used to explain why mimicking actions improves recall."},{"cited_title":"Mayer et al","cited_arxiv_id":null,"evidence_quote":"Closest prior system with a presenting avatar for asynchronous assembly tasks, which the guided tour builds on."},{"cited_title":"Irlitti et al","cited_arxiv_id":null,"evidence_quote":"Identifies the gap in asynchronous collaboration research that this work addresses."},{"cited_title":"Stefik et al","cited_arxiv_id":null,"evidence_quote":"Provides the named WYSIWIS perspective-sharing concept used in the interactive condition."}],"review_version":1}