{"id":"0c0e2576-e5e4-454a-8356-852aa20e80c0","arxiv_id":"2411.15583","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A systematic review and meta-analysis of 45 CVR studies finds inconsistent effect sizes across viewing modalities and widespread terminological and questionnaire problems in user-experience evaluation.","lead":"This paper systematically reviewed 45 studies of cinematic virtual reality viewing modalities and meta-analyzed 13 of them, finding that effect sizes vary widely across studies and that terms like presence, immersion, and narrative engagement are used inconsistently. It is worth reading because it makes clear why CVR user-experience results are hard to compare and argues for standardized evaluation methods.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pooling SMDs for presence, immersion, and narrative engagement as one outcome may create the cross-study inconsistency; construct equivalence is the load-bearing assumption.","rationale":"I agree with the reader's weakest assumption: construct mixing is the most load-bearing vulnerability. The meta-analytic result that motivates the paper's headline—unexpected inconsistency in effect sizes even within the same modality—is only meaningful if the effect sizes estimate a common outcome. The paper's own evidence that the field uses terms interchangeably weakens rather than strengthens the case for pooling, because interchangeable terminology does not make the instruments psychometrically equivalent; it may reflect confusion. The descriptive contribution (questionnaire irregularity, self-developed items, incomplete reporting) is well supported and would survive even if the meta-analysis were discarded, so the verdict should remain conditional rather than reject. I would not change the reader's CONDITIONAL verdict; the required fix is to make the outcome definition and effect-size computation transparent and to re-estimate with construct-specific models. The paper's Limitations section already concedes small effect-size counts and restricted moderator analysis, which reinforces the need for this check.","tokens_in":29541,"tokens_out":4376,"duration_ms":42687,"concrete_test":"Recode the 26 effect sizes by outcome construct (presence vs. immersion vs. narrative engagement) as identified in Table 4, and rerun the CHE-RVE model separately for each construct, using effect sizes computed from the original papers' reported means and standard deviations with the correct within-subject variance formula for repeated-measures designs. If within-construct I^2 drops substantially or the construct-specific estimates diverge, the cross-study inconsistency is an artifact of pooling distinct outcomes; if heterogeneity remains high within each construct, the inconsistency claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim—inconsistent effect sizes across studies using the same viewing modality—rests on pooling standardized mean differences computed from instruments that measure different constructs. Section 3.5 defines the outcome as 'presence, immersion and narrative engagement'; the actual instruments (iPQ, PQ, SUS, IEQ, MNEQ, NTS, and self-developed items; Tables 2 and 4) are validated for distinct constructs. Section 4.4 shows overlapping questionnaire items and interchangeable terminology, but item overlap does not establish construct equivalence, especially when studies adapt or select subsets of items. If these instruments are not measuring the same latent outcome, the pooled SMDs, the I^2 values, and the 'inconsistencies' reported in Section 4.2 and Figure 3 are partly artifacts of mixing outcome constructs rather than genuine modality effects. The paper's own moderator analysis (Table 5, 'Questionnaire') cannot resolve this because it groups 'valid' questionnaires together and does not include construct as a moderator. Additional reporting problems—Table 5 omits the implicit diegetic guidance category discussed in Section 4.2.1, and Section 4.2.2 labels the limited rotation coefficient as 'with an agency'—further weaken the numerical meta-analytic claims. The descriptive findings about questionnaire misuse, terminological confusion, and incomplete reporting are independently supported by the coding and do not depend on the pooled effect sizes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic review (45 studies) and meta-analysis (13 studies, 26 effect sizes) of viewing modalities in cinematic virtual reality (CVR), covering guidance cues, intervened/forced rotation, perspective shifting, and avatar assistance. The meta-analysis pools standardized mean differences (SMDs) on presence, immersion, and narrative engagement between a swivel-chair control condition and experimental conditions, using random-effects models and robust variance estimation under several assumed sampling correlations. The review finds no significant overall modality effects, large cross-study heterogeneity within modalities, and documents measurement problems: interchangeable use of presence/immersion/narrative engagement, frequent use of self-developed or adapted questionnaires, and incomplete reporting of statistics. The paper concludes that evaluation practice, not any intrinsic modality effect, limits the comparability of CVR user-experience studies.","tokens_in":29691,"tokens_out":7247,"duration_ms":62462,"significance":"The descriptive, corpus-based findings are a useful contribution. The screening is PRISMA-based, the coding book is specified, inter-rater agreement is reported, and the robust variance estimation is accompanied by a sensitivity analysis over the sampling correlation. If the descriptive claims hold, the paper offers an evidence-based argument for standardizing CVR terminology and measurement, which is of clear value to HCI and CVR researchers. The meta-analytic part, however, is too small and too construct-heterogeneous to support the headline 'inconsistency' claim on its own; its value is mainly as a quantitative illustration of the field's reporting problems rather than as a reliable estimate of modality effects.","major_comments":[{"comment":"The meta-analysis outcome is defined as 'presence, immersion and narrative engagement,' but the SMDs are pooled across instruments designed for different constructs (IPQ/PQ/SUS for presence, IEQ for immersion, MNEQ/NTS for narrative engagement; Table 4). Section 4.4 shows overlapping questionnaire items, but item overlap does not establish construct equivalence, especially because some studies use self-developed items or single items. Consequently, the pooled SMDs and the I² statistics that drive the 'inconsistency' claim in Section 4.2 and Figure 3 are partly artifacts of outcome-construct mixing. The moderator analysis in Table 5 groups only by questionnaire type (valid/adapted/selected/self-developed), not by construct; a moderator for construct (presence vs. immersion vs. narrative engagement), or restricting the pooling to a single construct, is needed before cross-study inconsistency can be attributed to modality effects or study design.","section":"§3.5, §4.2, §4.4, Table 5"},{"comment":"The reported RVE results contain internal inconsistencies that prevent interpretation. The text says the estimated effects range from 0.961 (SE = 0.294) for 'viewing modality with an agency' to 0.068 for explicit diegetic guidance, but Table 5 shows 0.961 under 'Limited rotation' and 0.309 (at ρ = 0.6) under 'With agency.' Moreover, Table 5 has no row for 'With implicit diegetic guidance,' even though Modality 2 (implicit diegetic guidance) is a central category in Section 4.2.1 and appears in the forest plot; the Wald test comparing modalities therefore does not compare all six modalities described. The positive coefficients for limited and forced rotation also contradict the statement in Section 4.2.1 that Modality 6 'generally shows negative results.' Finally, the 'Studies' column in the Questionnaire block sums to 15 studies, conflicting with the stated 14 studies in the meta-analysis. These issues must be corrected before the RVE estimates and p-values can be used.","section":"§4.2.2, Table 5"},{"comment":"The effect-size extraction does not account for the design of the primary studies. Most experiments (64.7%) used within-subjects designs, yet the coding book says Hedge's g was computed from 'means and standard deviations from groups' or from F/t/p values. For within-subjects data, computing an SMD from the two marginal distributions without the paired-difference correlation yields different and generally more variable effect sizes than a proper within-subject SMD. The sensitivity analysis over ρ addresses dependence among multiple effect sizes within a study, but it does not address the missing within-pair correlation in the SMD itself. Please specify how within-subject SMDs were computed, or exclude within-subject studies in a sensitivity analysis.","section":"§3.5, §4.1.3"}],"minor_comments":[{"comment":"Reference [65] is used for both Hong and Kim's rotational gain work and for Aitamurto et al.'s half-sphere experiment; the latter should cite [3].","section":"References / §2.2"},{"comment":"There are several typos in the author affiliations, including 'Chian' for China and 'Netherland' for Netherlands.","section":"Title page"},{"comment":"The inter-rater agreement formula '87.5% [i.e., 8−1/8]' is ambiguous; it should read (8−1)/8, and the authors should clarify what the 8 coded items were.","section":"§3.4"},{"comment":"The appendix table in Figure 5 is visually cluttered; the mapping from effect-size IDs to study IDs and conditions is difficult to follow and should be reformatted for readability.","section":"Appendix Figure 5"},{"comment":"The MNEQ row lists citation [73] (Kennedy et al., SSQ) instead of the correct source [22] (Busselle & Bilandzic).","section":"Table 2"},{"comment":"The phrase 'with an agency' is confusing and inconsistent with the coding book's 'with agency' and the more natural 'with an avatar'; please standardize the terminology.","section":"§4.2.1"}],"recommendation":"major_revision","confidential_remarks":"The systematic-review portion is the paper's real strength and should be foregrounded. The meta-analysis, as currently presented, is not reliable enough to support the paper's central claim about cross-study inconsistency because of construct mixing and the internal inconsistencies in Table 5. If the authors cannot reanalyze with a construct moderator or otherwise resolve these issues, they should explicitly reframe the meta-analysis as exploratory and make the descriptive measurement critique the primary contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Start with the punchline: the systematic review is the real contribution; the meta-analysis is a small, fragile add-on that makes the paper look more formal than it can support. The descriptive findings—how CVR studies actually measure experience, the questionnaire improvisation, the terminology swapping, the missing summary statistics—are useful, concrete, and supported by the coded corpus. The meta-analytic numbers are not the draw.\n\nWhat's new: none of the cited CVR reviews attempt a quantitative synthesis of viewing-modality effect sizes. This paper does, and it also provides a systematic coding of questionnaire and terminology practice that is absent from earlier reviews. The authors are appropriately hedged about 'potential' impacts, and the descriptive conclusions do not depend on the pooled effects.\n\nThe soft spots are real but not lethal. The meta-analysis pools standardized mean differences from instruments validated for presence, immersion, and narrative engagement as if they measured one outcome. The paper itself documents that these terms are used interchangeably, which cuts both ways: it proves the measurement mess, but it also means the pooled SMDs and I² values are partly artifacts of construct mixing rather than clean modality effects. With 13 papers, 26 effect sizes, and 2–6 per modality, there is not enough power for moderator analysis anyway. There are also internal reporting problems: Section 4.2.2 labels the limited rotation coefficient as 'with an agency' when Table 5 puts it under limited rotation; Table 5 drops the implicit diegetic guidance category that Section 4.2.1 discusses; the one-effect-size-per-study-per-modality retention rule appears only in a figure footnote. No data or code are shipped, so the extraction cannot be audited.\n\nCredit where it is due: the screening follows PRISMA, the coding has inter-rater checking, and the narrative about evaluation practice is honest and well-grounded. The terminological-confusion argument is not an artifact of the meta-analysis; it is visible in the corpus directly.\n\nWho this is for: people designing CVR user studies will get a practical map of which questionnaires are being used and where the field lacks standards. The meta-analysis itself should not drive anyone's decisions.\n\nRecommendation: yes, send it to referees. Expect major revision. A referee should require sensitivity analyses by construct, a corrected Table 5, and a transparent effect-size table with inclusion rules. If those are fixed, the descriptive core deserves publication.","headline":"The descriptive review of CVR evaluation practice is solid and worth publishing; the meta-analysis is too small and too construct-mixed to carry the headline numbers.","tokens_in":30314,"tokens_out":4871,"would_cite":true,"duration_ms":42208,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A meta-analysis of 13 CVR experiments finds no viewing modality with a reliable effect on user experience; the paper attributes the inconsistency to measurement practice.","keywords":["cinematic virtual reality","viewing modality","user experience","systematic review","meta-analysis","presence","immersion","narrative engagement"],"falsifier":"A concrete check: if a re-analysis of the 26 reported effect sizes, grouped by the exact questionnaire used instead of by modality, showed little reduction in heterogeneity, the paper's claim that measurement practice drives the inconsistency would be weakened. Conversely, a well-powered study using one validated instrument across several CVR contents and several modalities that still finds large, modality-dependent effect-size swings would point to content variability rather than instrument noise.","tokens_in":29258,"feed_emoji":"🎬","tokens_out":7677,"duration_ms":67752,"temperature":0.7,"pith_summary":"Cinematic Virtual Reality (CVR)—narrative 360-degree video watched on a head-mounted display—promises viewers an unscripted field of view, and researchers have tried to shape that field with guidance cues, intervened rotation, perspective shifts, and avatars. This paper asks whether these viewing modalities actually change the viewer's experience, pooling 45 experiments and carrying 13 of them into a meta-analysis. The pooled evidence is inconclusive: no modality shows a statistically reliable effect, and effect sizes swing from $g=-1.13$ to $g=2.10$ even within the same modality. The paper's diagnosis is that the field's measurement practices are the main problem: presence, immersion, and narrative engagement are used interchangeably; validated questionnaires are mixed with self-designed items; and many studies do not report the statistics meta-analysis needs. If this diagnosis is right, the next bottleneck in CVR research is methodological standardization rather than another round of content experiments.","feed_headline":"Meta-analysis finds no reliable CVR modality effect","feed_subtitle":"Across 45 studies, inconsistent effect sizes point to measurement practice as the real barrier.","key_machinery":"The engine of the paper is a six-way coding of viewing modalities—explicit or implicit, diegetic or non-diegetic guidance cues, agency (avatar assistance), and limited or forced rotation—combined with a random-effects meta-analysis that expresses every result as a standardized mean difference (Hedges' $g$) between the experiment and a \"Swivel-Chair\" CVR baseline. Because studies contribute multiple dependent effect sizes, the authors pool them with a correlated hierarchical effects working model and robust variance estimation, testing sampling correlations of $\\rho=0$, $0.3$, $0.6$, and $0.9$. The same coding book records questionnaire type, terminology mixing, sample size, design, and data-reporting completeness, which lets the paper pair its quantitative null result with a qualitative account of why the numbers do not line up.","core_discovery":"The central claim is that the existing evidence does not yet support any conclusion about which CVR viewing modalities improve or harm user experience. Across 26 effect sizes from 14 studies, the authors found no statistically significant pooled effect for any of the six modality categories, and a robust Wald test could not reject the hypothesis that all modalities have the same average effect ($p\\approx 0.74$). What stood out was not the size of any single effect but the inconsistency: effect sizes ranged from $g=-1.13$ to $g=2.10$, with substantial heterogeneity in several subgroups. The paper attributes this inconsistency to evaluation practice. It documents that \"presence,\" \"immersion,\" and \"narrative engagement\" are frequently treated as synonyms, that several standardized questionnaires share identical items, that 42.4% of the questionnaire-based studies rely at least partly on self-developed instruments, and that many studies report incomplete summary statistics. The conclusion is that unrigorous and nonstandard evaluation, more than any specific modality, is what makes the field's quantitative findings unreliable.","pith_inferences":["If construct mixing were the main source of inconsistency, re-coding the 26 effect sizes by the exact questionnaire used (rather than by modality) should reduce heterogeneity; the paper's tables hint that questionnaire type alone may not explain the variance, but the cell sizes are too small to tell.","A testable prediction follows: a single well-powered study using one validated instrument across several content types and several modalities would either tighten the pooled estimates or show that content variability, not measurement noise, drives the spread.","The same terminological and questionnaire migration likely troubles adjacent areas such as social viewing and collaborative virtual reality, where the same instruments are used; those fields could inherit the same unreliability.","Because 32 of 45 reviewed papers could not enter the meta-analysis for lack of usable statistics, recovering their data or encouraging raw-data sharing would be the quickest way to test whether the null result is an artifact of selective reporting."],"forward_implications":["No viewing modality can currently be declared beneficial or harmful on the basis of pooled evidence; the meta-analyzed studies are consistent with zero average effect.","Comparisons between CVR experiments are unreliable until the field agrees on what \"presence,\" \"immersion,\" and \"narrative engagement\" mean and adopts shared instruments.","Research on attention-driven modalities needs manipulation checks: most modalities are assumed to change attention, but the surveyed studies rarely verify that assumption.","Future meta-analyses will need fuller reporting of means and standard deviations; incomplete reporting is a major reason only 13 of 45 reviewed papers contributed effect sizes.","A validated CVR-specific questionnaire, built from the non-overlapping content of existing instruments, would directly address the terminological confusion the paper documents."],"supporting_citations":[{"why":"Reporting guideline that structures the screening flow and defines the 45-paper corpus.","marker":"[147]"},{"why":"Supplies the correlated hierarchical effects working model with robust variance estimation used to pool dependent effect sizes.","marker":"[109]"},{"why":"Provides the guidance-cue taxonomy the authors use to categorize modalities and interpret directional patterns.","marker":"[121]"},{"why":"First classifies guidance cues into diegetic and non-diegetic dimensions, the basis for the guidance modalities.","marker":"[103]"},{"why":"Provides the motorized swivel-chair baseline and the counter-example where forced rotation can be comfortable.","marker":"[58]"},{"why":"Supplies the Immersion Experience Questionnaire, one of the instruments entangled in the terminological overlap.","marker":"[71]"},{"why":"Supplies the narrative-engagement questionnaire whose items overlap with presence and immersion scales.","marker":"[22]"},{"why":"Supplies the Simulator Sickness Questionnaire the paper criticizes as over-used and not VR-specific.","marker":"[73]"},{"why":"Supplies the iGroup Presence Questionnaire, the most frequently used presence measure in the reviewed studies.","marker":"[68]"},{"why":"Supplies the classic Presence Questionnaire used for presence measurement in the corpus.","marker":"[156]"}],"fun_headline_variants":["No reliable CVR modality effect after meta-analysis","CVR studies fail to show any winning viewing mode","Why CVR research can't agree: measurement issues","Meta-analysis reveals no proven CVR modality benefit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The meta-analysis assumes that scores from different questionnaires—IPQ, IEQ, MNEQ, SUS, and self-designed items—can be pooled as one shared measure of \"user experience\"; if those instruments measure distinct constructs, the pooled estimates and the observed inconsistency partly reflect construct mixing rather than true modality effects.","fun_headline_variants_meta":{"raw":{"variants":["No reliable CVR modality effect after meta-analysis","CVR studies fail to show any winning viewing mode","Why CVR research can't agree: measurement issues","Meta-analysis reveals no proven CVR modality benefit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1316,"prompt_tokens":1037,"completion_tokens":279,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":218}},"tokens_in":653,"tokens_out":279,"duration_ms":3671,"temperature":1.0,"reasoning_tokens":218,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:08:45.209794+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: if a re-analysis of the 26 reported effect sizes, grouped by the exact questionnaire used instead of by modality, showed little reduction in heterogeneity, the paper's claim that measurement practice drives the inconsistency would be weakened. Conversely, a well-powered study using one validated instrument across several CVR contents and several modalities that still finds large, modality-dependent effect-size swings would point to content variability rather than instrument noise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Immersion Experience Questionnaire, one of the instruments entangled in the terminological overlap."}],"review_version":1}