{"id":"1876cef0-6dba-40ff-8fb2-f5710735b3c9","arxiv_id":"2607.26183","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Embodied VR enactment of a Hakka dish raises engagement and transmission intentions but suppresses concurrent cultural narration (action masking), unlike a matched VR video.","lead":"A VR cooking game that lets players use hand-tracking to prepare a Hakka stuffed bitter melon dish raises engagement and the wish to spread the tradition more than watching a VR video of the same steps—but players miss much of the chef's cultural narration while their hands are busy. The result maps out a trade-off that matters for digital heritage and for any immersive training where doing and listening compete.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'gamified VR' arm differs from the control not only in enactment but also in game elements (progress bar, badge, audio, hints, error/retry); the central attribution to embodied enactment is confounded and needs a 2×2 factorial check.","rationale":"The reader's weakest assumption identifies the enactment–gamification confound as the primary threat; my reading confirms this as the most load-bearing concern. The paper's central contribution is that embodied enactment of constraints, rather than passive observation, raises engagement and cultural awareness and produces a mode-dependent trade-off. But the operational comparison holds 'medium' constant at VR while varying both enactment and a bundle of game elements. Because these elements are themselves designed to increase motivation and affect (§3.3.3), any observed difference is over-determined. The paper's own design rationale (DG3) says game elements support engagement, so attributing the engagement gains to enactment is circular without an un-gamified enactment arm. The action-masking finding is similarly ambiguous: the gamified arm's feedback and progress cues could capture attention independently of manual load. Thus the central claim—that embodied VR transmits a complementary layer—is not yet identified. The paper has genuine strengths: it is transparent about statistics, provides rich qualitative data, and honestly reports non-significant post-Holm effects. But the confound is structural, not a minor limitation. My proposed 2×2 factorial test directly isolates the two factors and would settle the attribution. Because the reader already marked the verdict CONDITIONAL for exactly this reason, my recommendation is UNCHANGED.","tokens_in":35168,"tokens_out":4477,"duration_ms":46289,"concrete_test":"Run a 2×2 between-subjects factorial follow-up: (A) interactive+gamified (current design), (B) interactive+non-gamified (same hand-tracking and physics, no progress bar, badge, or optional hints; retain only immediate consequence feedback), (C) video+gamified (current video with progress bar, badge, and audio cues overlaid), (D) video+non-gamified (current control). With n≈20 per cell, measure the same GEQ/IMI/awareness outcomes. If the interaction main effect (A,B > C,D) holds across both game-framing levels, the enactment interpretation is supported; if the game-framing main effect (A,C > B,D) is significant instead, the reported advantages owe to gamification, not enactment. Also compare action-masking: if self-reported masking appears in B as well as A, sensorimotor load is the cause; if only A, game feedback is the cause.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 asserts the two arms 'differ only in whether participants physically enacted or passively observed,' but this is unsupported. The enactment arm includes a progress bar, achievement badge, step-completion audio, error-and-retry states, hints, and a physics-based interactive environment (§3.3.2–3.3.3); the video control (§3.4.1) has none of these. These are not inert: progress and achievement feedback are known drivers of engagement and positive affect; error–retry creates intermittent high-load events and interactive affordances themselves can raise sensory/imaginative immersion. Therefore the significant increases in IMI Interest/Enjoyment (δ=0.81), GEQ Sensory & Imaginative (δ=0.62), and Transmission (δ=0.70) may be caused by gamification/interactivity rather than by isomorphic bodily enactment—the paper's central independent variable. The action-masking argument is likewise confounded: game feedback and progress tracking may draw attention away from narration regardless of sensorimotor load. The per-item quiz dissociation (85% vs 100% on the narrated-symbolism item; 70% vs 50% on blanching rationale) is non-significant and supports the masking claim only weakly. The central claim that 'embodied VR transmits a different, complementary layer of cultural knowledge' requires that enactment, not game framing, be the causal factor.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Hakka Kitchen, a gamified VR experience for the Hakka dish stuffed bitter melon, grounded in chef interviews and a procedural dictionary. In a between-subjects study (N=40), participants either enacted the recipe via hand-tracked VR or watched a matched video in VR. The gamified condition reported higher Interest/Enjoyment, Sensory & Imaginative Immersion, Positive Affect, and heritage-awareness Transmission/Vitality, with no difference on a six-item procedural quiz. Qualitative analysis reports that enactment creates somatic memory anchors and that concurrent cultural narration is partially masked by action. The authors conclude that embodied VR transmits a different, complementary layer of cultural knowledge than passive observation.","tokens_in":35477,"tokens_out":5088,"duration_ms":54171,"significance":"The study has clear strengths: the statistical reporting is transparent and appropriate (Mann–Whitney U with Cliff's delta and bootstrapped CIs, Holm-corrected flags), the video-in-VR control is better than the no-control or desktop baselines common in prior VR-ICH work, and the procedural-dictionary design process is a transferable methodological contribution. If the central claims could be causally isolated, the paper would usefully complicate the assumption that embodied VR is uniformly additive for cultural transmission. However, the central independent variable is confounded with gamification/interactivity, and the action-masking evidence is largely qualitative and underpowered. The contribution is therefore currently suggestive rather than demonstrated.","major_comments":[{"comment":"The assertion that the two arms 'differ only in whether participants physically enacted or passively observed' is not supported. The enactment condition includes a progress bar, achievement badge, step-completion audio, error-and-retry states, hints, and a physics-based interactive environment; the video control has none of these. These elements are known engagement drivers. The observed increases in IMI Interest/Enjoyment (δ=0.81), GEQ Sensory & Imaginative (δ=0.62), and Transmission (δ=0.70) may therefore be due to game/interactivity features rather than embodied enactment per se. To support the central causal claim, the design needs a factorial manipulation (enactment × game feedback) or an additional control that includes equivalent progress/feedback elements without requiring physical enactment. At minimum, the conclusions should be reframed as effects of the interactive gamified ex","section":"§4.1, §3.3.2–3.3.3, §3.4.1"},{"comment":"The action-masking conclusion is not adequately supported. The per-item quiz dissociation (gamified 17/20 vs video 20/20 on the narrated-symbolism item; 14/20 vs 10/20 on the blanching-rationale item) is non-significant (Fisher's exact p>.12), and both groups are near ceiling on the six-item quiz (5.15 vs 5.05 out of 6), so item-level tests have low sensitivity. The qualitative reports of 'not noticing the chef' could reflect attention capture by game elements (progress bar, badge, error cues) rather than cognitive load from sensorimotor coordination. No direct measure of cognitive load or attention is provided. The central claim that enactment masks narration while observation transmits it requires stronger evidence, e.g., an experimental manipulation of narration timing, a process-tracing measure, or an improved quiz without ceiling effects.","section":"§5.2.3, §5.1.4"},{"comment":"The interpretation that the bottleneck is 'amodal' (central working-memory competition) rather than modality-specific is an overreach. In both conditions, narration is presented auditorily during the procedure, so the design cannot distinguish between competition for central amodal resources and modality-specific interference. The paper honestly acknowledges in §6.4.1 that step-boundary triggering did not prevent masking, but that does not resolve the interpretive issue. This is secondary to the main mode-dependence claim, but the amodal conclusion should be flagged as speculative or tested with a design that varies narration modality/timing.","section":"§6.2, §6.4.4"}],"minor_comments":[{"comment":"The gamified-VR group is older on average (26.4 vs 22.7) and no test or control for age is reported. Given the age range and possible effects on VR familiarity and engagement, a sensitivity analysis or covariate check would strengthen the between-group comparison.","section":"Table 2"},{"comment":"The caption states Fisher's exact p>.05 while Section 5.1.4 reports all per-item p>.12; please make the reported thresholds consistent.","section":"Figure 7"},{"comment":"Section 6.4.4 is admirably honest about the control video's limitations (error-free demonstration, viewpoint mismatch), but this should be prominently disclosed in Section 4.1 where the 'differ only in' claim is made, rather than deferred to the limitations.","section":"§4.1, §6.4.4"},{"comment":"The qualitative themes are plausible, but the codebook table consists largely of example quotes. Reporting inter-rater reliability or code frequencies would help confirm that the 'nine of twenty' action-masking theme is not an artifact of selective quotation.","section":"Appendix A.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely publishable after revision, but the main confound is serious. I would not reject, because the empirical material, design process, and statistical transparency are valuable and the authors are explicit about many limitations. However, the current framing overclaims causal isolation of embodiment. A revised version that either runs the factorial follow-up or carefully limits the claims to the interactive gamified experience would be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a genuinely useful empirical paper with an honest, nuanced result, but the headline interpretation is overreach. The authors compare a gamified VR cooking experience against a content-matched VR video and find gains in engagement, positive affect, and cultural transmission for the interactive condition. That part is solid. The interesting bit is the action-masking pattern: participants who enacted the cooking steps missed parts of the cultural narration, while the video group absorbed it. That complicates the usual story that embodiment is uniformly good for cultural heritage, and it gives designers a concrete lever (schedule narration in low-load windows). The qualitative data support the somatic account in a way that is not just post hoc branding.\n\nThe stats are handled transparently: Mann-Whitney U with Cliff's delta, bootstrapped CIs, Holm correction, and non-significant results reported as non-significant. That is refreshing. The procedural dictionary from chef interviews is a reasonable, reusable process contribution.\n\nWhere it gets wobbly is the causal claim. Section 4.1 says the two conditions 'differ only in whether participants physically enacted or passively observed.' They do not. The gamified condition also has a progress bar, badge, step-completion audio, hints, and error/retry states. Those are known engagement drivers. So the quantitative differences could be from gamefulness rather than from isomorphic bodily enactment. The authors frame the contribution around 'enactment of constraints,' but the design doesn't isolate enactment from gamification. A 2x2 factorial (enactment × game elements) is the obvious next step. This is not a fatal flaw—the qualitative responses do gesture at felt physical learning—but it is a serious gap between the data and the interpretation.\n\nThe action-masking evidence is also thinner than the narrative suggests. The quiz item dissociation (17/20 vs 20/20 and 14/20 vs 10/20) is non-significant, and the main support is self-report. The authors are appropriately cautious, but I would not let that slide in peer review.\n\nAlso: no data or code, N=40, single session, post-only measures. The limitations section acknowledges most of this honestly.\n\nWho is this for? Researchers in VR-ICH and embodied learning. It deserves a serious referee—the question is real and the study is well-run. But the revision needs to either add a control condition that separates game elements from enactment, or reframe the claims to be about a gamified interactive experience broadly, with enactment as one hypothesized component.\n\nI'd send it to peer review.","headline":"A solid empirical VR-ICH study with a genuinely interesting action-masking finding, but the central causal claim about embodied enactment is confounded with gamification elements.","tokens_in":35966,"tokens_out":2260,"would_cite":true,"duration_ms":23836,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hakka Kitchen claims that embodying a traditional dish's constraints in VR transmits a complementary layer of cultural heritage, at the cost of absorbing concurrent narration.","keywords":["virtual reality","intangible cultural heritage","culinary heritage","embodied cognition","gamification","action masking","kinesthetic empathy","VR video comparison"],"falsifier":"Run a three-arm between-subjects design: interactive enactment, passive VR video, and passive VR video with identical game elements overlaid (progress bar, badge, error audio, hints) but no hand interaction. If the third arm matches the engagement and transmission gains, the effect is game feedback, not enactment. Alternatively, log narration recall split by high- versus low-load steps; action masking predicts recall drops precisely during manual phases.","tokens_in":35027,"feed_emoji":"🥘","tokens_out":4118,"duration_ms":42559,"temperature":0.7,"pith_summary":"The paper argues that process-based culinary heritage, like making stuffed bitter melon, cannot be fully transmitted by passive video because its meaning lives in physical constraints: slice thickness, pith removal, blanching timing, and filling tightness. It builds a gamified VR experience in which players enact those constraints with hand-tracked gestures, then compares them against a matched VR video of the same procedure in a between-subjects study with 40 participants. Enactment raised sensory and imaginative immersion, positive affect, and the desire to transmit awareness of the heritage. But heritage recognition is mode-dependent: players absorbed the chef's cultural narration less while enacting, yet gained felt bodily knowledge of the constraints and a kinesthetic empathy for the artisan's labor. The two modes transmit complementary layers; neither transmits both at once.","feed_headline":"Doing beats watching for cultural heritage in VR cooking","feed_subtitle":"Players who cooked felt the dish's constraints in their hands, but missed more of the chef's story—both are needed.","key_machinery":"The procedural dictionary: a table translating chef interviews into process stages, somatic constraints, game mechanics, and cultural connotations. The load-bearing identity is enactment of constraints—culturally meaningful tolerances (thickness, pith removal, timing, stuffing limits) become the game mechanics themselves, using isomorphic hand-tracking so gestures structurally resemble real cooking actions. The mechanism includes a withheld-parameter timer, reversible failure states, and stage-gated narration. The paper's explanatory device is action masking: the motor demands of enactment occupy the attentional bandwidth needed to encode concurrent narration, an amodal bottleneck that narra","core_discovery":"The paper's central claim is that representing interactive procedures instead of static content can empower cultural awareness in virtual ICH practice. In Hakka Kitchen, expert-elicited tacit checkpoints become game mechanics: slices outside the 1–3 cm band wobble, unremoved pith blocks blanching, wrong timer settings trigger error feedback, and overfilled rings burst. Compared with a VR video containing the same five-stage content and narration, enactment raised interest/enjoyment, sensory and imaginative immersion, positive affect, and the heritage-awareness dimensions tied to transmission and vitality, with no difference on a six-item procedural quiz. The central finding is the trade-off:","pith_inferences":["The study cannot yet separate enactment from gamification: a control video that also carries a progress bar, badge, and error-feedback audio would test whether the engagement gains come from doing or from game feedback.","Action masking should be sharpest in threshold-bound, outcome-bearing crafts and weakest in expressive heritage where movement itself carries meaning; this predicts a domain-specific design rule rather than a universal VR-heritage effect.","A longitudinal follow-up with delayed quizzes and real-kitchen transfer could show whether somatic anchors produce durable retention even though immediate quiz scores were equal.","Spatial or environmental cultural cues, rather than concurrent voiceover, might deliver heritage meaning during high-load steps without competing for attention."],"forward_implications":["For culinary ICH, embodied VR can transmit felt, threshold-bound knowledge—what 1–3 cm, fully scraped pith, and 'not too full' feel like—that matched video leaves as abstract instruction.","Enactment increases interest, sensory/imaginative immersion, positive affect, and the perceived transmission value of the heritage, with the strongest robust effect on transmission awareness.","Cultural narration is reliably absorbed during passive observation but partially masked during manual enactment; enacted constraints are retained better than narrated-only meaning.","Immediate procedural quiz performance and perceived competence do not differ between modes, so the enacted benefit is affective and valuing, not declarative recall.","Design should keep culturally meaningful constraints as the mechanic, make feedback externalize the tacit judgment, and schedule narration during low-load windows."],"fun_headline_variants":["VR cooking game: feel more, recall less about the culture","Immersive cooking VR boosts sensory engagement, not story uptake","Hands-on VR heritage cooking: more empathy, less narration","In VR, cooking a dish beats watching a video for cultural feel","VR gameplay in Hakka Kitchen: richer experience, muted narration"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central comparison assumes the gamified VR condition differs from the control only in embodied enactment, but it also includes game elements—progress bar, achievement badge, step-completion audio, error-and-retry states, hints, and an interactive environment—so the measured advantages may come from interactivity or game feedback rather than enactment per se.","fun_headline_variants_meta":{"raw":{"variants":["VR cooking game: feel more, recall less about the culture","Immersive cooking VR boosts sensory engagement, not story uptake","Hands-on VR heritage cooking: more empathy, less narration","In VR, cooking a dish beats watching a video for cultural feel","VR gameplay in Hakka Kitchen: richer experience, muted narration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000251,"raw_usage":{"total_tokens":1388,"prompt_tokens":734,"completion_tokens":654,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":567}},"tokens_in":478,"tokens_out":654,"duration_ms":7079,"temperature":1.0,"reasoning_tokens":567,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:32:09.720043+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a three-arm between-subjects design: interactive enactment, passive VR video, and passive VR video with identical game elements overlaid (progress bar, badge, error audio, hints) but no hand interaction. If the third arm matches the engagement and transmission gains, the effect is game feedback, not enactment. Alternatively, log narration recall split by high- versus low-load steps; action masking predicts recall drops precisely during manual phases.","supporting_citations":[],"review_version":1}