{"id":"7ebf9493-6e66-4048-bb26-4e023d186b23","arxiv_id":"2501.15456","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A proof-of-concept VR workflow that combines speech-based prompt refinement, egocentric focal adjustment, and segment-wise iteration to co-create 360-degree panoramic videos with AI.","lead":"Imagine360 is a prototype VR system in which people co-create 360-degree panoramic videos by speaking prompts, letting an AI agent refine them, and choosing the focal direction for each new video segment. An eight-person pilot found that users preferred this iterative human-AI workflow over one-shot prompt generation, though it increased mental workload.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The pilot's Wizard-of-Oz setup means the evaluation tests a human stand-in for the AI agent, not the automated Imagine360 system, so the headline claim of demonstrated effectiveness is unsupported.","rationale":"The reader's weakest_assumption correctly identifies the Wizard-of-Oz simulation as the main threat to the central claim. My stress-test converges on the same point but sharpens it: the WoZ setup does not merely add noise; it replaces the AI agent with a human experimenter, so the study cannot attribute any observed effect to the automated system described in Section 3. The paper's own disclosure in Section 4.1 confirms that the core integration (API with Unity, automated angle adjustment) was not functional during the evaluation. Therefore, the abstract's verb 'demonstrates' is not supported by the evidence. I considered whether the larger problem is the small sample size or non-significant primary comparison, but those are secondary: even with n=80, a WoZ evaluation would not establish the effectiveness of an untested automated pipeline. The appropriate verdict remains CONDITIONAL: the work is a reasonable design-concept pilot with honest limitations, but its conclusions must be re-scoped to the WoZ-simulated interaction, not to the deployed system. The concrete test I propose would settle whether the WoZ condition is a valid proxy; until such a test is run, the central claim should be treated as unverified.","tokens_in":7966,"tokens_out":2874,"duration_ms":29939,"concrete_test":"Run a fidelity check for the WoZ simulation: in a follow-up within-subjects study, compare the original WoZ condition against a fully integrated Imagine360 pipeline (or a high-fidelity automated simulation) on the same tasks, with objective logs of agent actions and timing. If user experience ratings or performance scores differ significantly between conditions, the original WoZ results are not a valid test of the automated system. Alternatively, if historical WoZ session logs exist, have independent raters code each experimenter action against the Section 3 workflow; any deviation in prompt refinement, suggestion frequency, or recentering beyond user instructions would demonstrate that human facilitation, not the AI agent, drove the outcomes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1, Task 3 discloses that the study used a Wizard-of-Oz approach: the experimenter manually imported generated videos and adjusted the user's visual center. The paper's central claim—that Imagine360's co-creative approach 'effectively integrates temporal and spatial creative controls'—requires that this WoZ simulation faithfully reproduces the automated pipeline described in Section 3, including speech-to-text, prompt optimization, suggestion behavior, segment assembly, and egocentric recentering. No fidelity check is reported: there are no logs of experimenter interventions, no comparison with what the Section 3 system would have produced, and no independent verification that the experimenter followed the specified protocol. If the experimenter's judgment, responsiveness, or implicit guidance—rather than the co-creation design—drove participant preference and performance ratings, then the results do not validate the automated system. Moreover, the experimenter was not blind to the hypotheses and had a stake in the system's success, introducing a demand-characteristic confound. The concern is load-bearing because every positive result in the paper (performance ratings, interview preferences) is attributed to the co-creation framework, yet the 'AI agent' in the study was, in effect, a human. The paper even notes that only 4 of 24 trials involved perspective changes, so the spatial-control component was rarely exercised; combined with the WoZ simulation, the evidence for the system's claimed capabilities is thin.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Imagine360 is a proof-of-concept VR system that lets users generate 360° panoramic videos through iterative speech-based prompts, AI prompt refinement, and egocentric focal recentering. The paper describes the system architecture (video generation via Runway Gen-3 Alpha Turbo, equirectangular projection, voice input via Whisper-1, prompt optimization via GPT) and reports an eight-participant pilot study comparing three conditions: non-AI pre-recorded videos, linear AI-driven generation, and human-agent co-creation. Quantitative measures (NASA-TLX, Boden creativity ratings) and semi-structured interviews are used to argue that the co-creative workflow better integrates temporal and spatial creative controls. The evaluation used a Wizard-of-Oz protocol in which the experimenter manually imported generated videos and adjusted the user's visual center.","tokens_in":8226,"tokens_out":3698,"duration_ms":32572,"significance":"If the effectiveness claims were fully supported, the paper's contribution would be a workflow design rather than a new generative model: it demonstrates that segment-wise iterative speech refinement and egocentric recentering are usable and preferred over one-shot prompting in immersive VR. The paper is transparent about its proof-of-concept status and about the Wizard-of-Oz simulation, and the qualitative interview data provide useful design insights. However, the central claim is not as strong as the abstract states: the principal performance metric is not statistically significant, the sample size is eight, and the automated system was not actually evaluated. The paper's value at this stage is as a pilot that motivates future work, not as a demonstration of the effectiveness of the Imagine360 agent.","major_comments":[{"comment":"The performance advantage claimed for co-creation is not statistically supported: the mean performance ratings (8.875 vs. 8.75 vs. 8.5) are reported without any test statistic, while the only significant effects reported are higher mental load and effort for co-creation. The abstract's phrase 'demonstrates that Imagine360's co-creative approach effectively integrates temporal and spatial creative controls' therefore overstates the evidence; please report the relevant test (with effect size) or reframe the claim as a qualitative/pilot finding.","section":"Abstract and §4.2"},{"comment":"The evaluation used a Wizard-of-Oz protocol: the experimenter manually imported generated videos and adjusted the user's visual center. Because no fidelity check, intervention log, or inter-rater reliability measure is reported, and because the experimenter was not blind to the hypotheses, the results cannot be attributed to the automated Imagine360 system described in Section 3. This is load-bearing because all positive outcomes are attributed to the co-creation framework. Please either evaluate the actual automated pipeline (even in a limited pilot) or explicitly scope the conclusions to the WoZ-simulated workflow and add a fidelity assessment.","section":"§4.1 Task 3"},{"comment":"Spatial control was rarely exercised: only 4 of 24 Task 3 trials (17%) involved perspective changes, yet the central claim is that the system integrates temporal and spatial creative controls. The paper should report what happened in those four trials and discuss why participants preferred composition-level adjustments; without this, the spatial-control component of the claim is supported only by interview preferences, not by usage data.","section":"§4.2"},{"comment":"The statistical reporting is incomplete. The Friedman test is written as χ2(3,N) without giving N or degrees of freedom; with three conditions the df should be 2, not 3. Please specify the test details, the correction used for post-hoc Wilcoxon tests, and effect sizes, and clarify what 'significant relevance and value' refers to (no test is reported for the Boden measures).","section":"§4.1 and §4.2"}],"minor_comments":[{"comment":"The word 'demonstrates' should be softened to 'suggests' or 'provides initial evidence from a pilot study,' given the small sample and the Wizard-of-Oz simulation.","section":"Abstract"},{"comment":"The text in Section 3 says the prompt optimization uses GPT-3.5 Turbo, while Figure 2 and other parts of the paper say GPT-4o; please reconcile this inconsistency.","section":"§3 and Figure 2"},{"comment":"The caption for Figure 1 is very long and mixes philosophical framing with system description; consider shortening it and moving the philosophical context to the introduction.","section":"Figure 1"},{"comment":"The procedure for the linear AI-driven condition is underspecified: please state how the three prompts were converted into videos and whether the experimenter also manually imported and adjusted those videos, so that the comparison with Task 3 is clear.","section":"§4.1 Task 2"},{"comment":"The sentence 'Our system shows significant relevance and value in outcomes and participants' creativity' does not report a test statistic or a p-value; please specify which comparison and which measure this refers to.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"This paper reads more like a workshop or late-breaking-results contribution than a full archival paper at its current stage. The design idea is interesting and the qualitative insights are worth preserving, but the abstract and conclusion need to be scaled back to match the evidence, and the Wizard-of-Oz limitation needs to be addressed head-on. If the authors can either provide a fidelity check for the WoZ simulation or reframe the claims as being about a simulated workflow, the contribution could become publishable after major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a legitimate design-concept paper describing a workflow that, as a combination, I haven't seen before—speech-refined prompts, egocentric focal recentering, and segment-wise iterative generation for 360° panoramic video in VR. The system description is clear and the pilot is reported honestly, including the Wizard-of-Oz disclosure. But the abstract's claim that it 'demonstrates' effective integration of temporal and spatial controls outruns the evidence.\n\nThe good parts: the integration is genuinely new relative to 360DVD, Interact360, and ImmerseSketch, all cited. The paper is transparent that the evaluation used a WoZ simulation because API-Unity integration was unstable. The qualitative interviews give plausible, useful signals for co-creation in VR.\n\nThe soft spots are concentrated in the evaluation. Eight participants, a non-significant primary performance comparison, and a WoZ stand-in for the core agent. The stress test is right that this is load-bearing: every positive result is attributed to the co-creation framework, yet the 'AI agent' in Task 3 was the experimenter manually importing videos and adjusting the visual center. The experimenter wasn't blind and no fidelity check was reported, so demand characteristics are a real confound. Also, only 4 of 24 trials actually used perspective changes, so the spatial-control component was rarely exercised. That makes the 'spatial creative controls' claim thinner than the title suggests.\n\nNone of this is fatal if the paper is re-scoped. The authors already acknowledge the WoZ and outline future work. The right fix is to present this as a design-concept evaluation, integrate the pipeline or at least add a fidelity check and larger or external-judge data, and soften the conclusions. The prototype and workflow are worth building on.\n\nWho should read this: HCI people working on co-creativity, immersive tools, or generative video UIs. It's a useful design exploration even though the empirical backing is modest.\n\nRecommendation: send it to peer review. A serious referee can separate the prototype value from the overclaimed evaluation and push for the needed revision or re-scope.\n\nBest,","headline":"A clearly described co-creation workflow for 360° video in VR, but the pilot study is a Wizard-of-Oz simulation with n=8, so the headline claims outrun the evidence.","tokens_in":8729,"tokens_out":2075,"would_cite":false,"duration_ms":19404,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a VR workflow where users co-create 360-degree video with an AI agent through iterative speech prompts and egocentric focal re-centering outperforms both non-AI and one-shot AI workflows on user-valued outcomes.","keywords":["360-degree video generation","panoramic video","human-AI co-creation","virtual reality","speech-based prompting","egocentric view adjustment","iterative generation","Wizard-of-Oz pilot"],"falsifier":"A fully automated end-to-end run of the real system, with no experimenter mediating video imports or visual-center changes, compared against the pilot's co-creation condition; if preference, performance, and creativity scores do not reproduce, the Wizard-of-Oz mediation rather than the design was doing the work. A second decisive observation is the frequency of egocentric focal changes in a larger sample: if usage stays near the pilot's 4 of 24 trials even with onboarding, the spatial-control component of the claimed integration is unsupported.","tokens_in":7765,"feed_emoji":"🥽","tokens_out":9104,"duration_ms":82547,"temperature":0.7,"pith_summary":"Imagine360 is a proof-of-concept VR system that turns 360-degree video generation into an iterative dialogue between user and AI agent. Instead of issuing one prompt and waiting for a finished panorama, the user speaks an intention, the agent refines it into a text prompt, a short video segment is generated from the last frame of the previous segment, and the user can rotate to choose the focal center of the next generation. The paper's central claim is that this co-creative loop integrates temporal control (what happens next) with spatial control (where the user is looking), and the eight-participant pilot comparing it with pre-recorded non-AI video and one-shot AI workflows supports that claim on user-valued outcomes. The contribution is the interaction design and pilot evidence, not a new generative model.","feed_headline":"Voice-driven co-creation beats one-shot AI prompts in VR test","feed_subtitle":"A pilot with eight users found iterative, speech-refined generation more valued than passive prompting in a VR headset.","key_machinery":"The load-bearing mechanism is the agent-mediated feedback loop: each new video segment is generated from the last frame of the segment the user just watched, re-centered to the user's current egocentric angle, and driven by a speech prompt that the agent refines and offers suggestions for. The raw generated video is then post-processed into an equirectangular projection—a 2:1 flat map of a sphere with blended edges, a blurred background, and reduced foreground height—so it can be viewed continuously in the VR headset. This loop connects seeing (the frame still in view) to imagining (the next segment's prompt) and is what the paper claims integrates temporal and spatial controls.","core_discovery":"On the paper's own terms, the discovery is that people can co-create 360-degree video in VR when generation is broken into segments and each segment starts from the visual state the user actually saw. Users issue speech commands, have them refined or suggested by an AI agent, and can change the focal angle of the next image prompt by physically rotating; the system maps image length to angular degrees and re-centers the panorama accordingly. The pilot's performance ratings were highest for this co-creation condition (mean 8.875 vs 8.75 for non-AI and 8.5 for linear AI), and participants reported strong relevance and value in the outcomes, with a correlation of $r = 0.81$ between relevance and value. The paper also reports that the co-creation condition imposed higher mental load, and that only 17% of trials used a perspective change, because users preferred adjusting overall composition.","pith_inferences":["Editorial inference: because the evaluation used a Wizard-of-Oz stand-in, the pilot validates the co-creation concept more than the implemented automation; an end-to-end automated system could produce different preference results.","Editorial inference: the low rate of focal-angle changes (4 of 24 trials) suggests egocentric spatial control may be secondary for this user group, and a tutorial or more salient affordance could reveal whether the low usage is discoverability or preference.","Editorial inference: the post-processing choices that make 2D output feel panoramic (blur, height reduction, edge blending) trade visual fidelity for seamlessness, so future studies could separate workflow satisfaction from output-quality judgments.","Editorial inference: the last-frame continuation loop transfers naturally to non-panoramic creative tools such as storyboard or animation previz, where each revision starts from the last rendered frame."],"forward_implications":["If the pilot result holds, VR experiences can be authored in real time from the user's speech and gaze instead of being limited to pre-designed scenes.","Segment-wise generation from the last frame turns temporal continuity into a user control: every 10-second segment is a direct revision of what was just seen.","Egocentric focal re-centering gives users a spatial control that one-shot text-to-video prompting does not offer.","The higher measured mental load and the low 17% perspective-change rate imply that the next iteration should reduce co-creation overhead or make spatial control easier to discover.","The refined prompt suggestions allow users with little prompting experience to steer generation by confirming or choosing from agent-provided options."],"supporting_citations":[{"why":"Establishes the controllable 360-degree panoramic video generation pipeline that this paper adopts as the generation foundation.","marker":"[30]"},{"why":"Supplies the creativity framework (novelty, value, surprise, relevance) used to score participant outcomes.","marker":"[5]"},{"why":"Supplies the six-dimension workload index used to measure mental load and effort in the pilot.","marker":"[17]"},{"why":"Provides the machine-in-the-loop model of human control with AI support that motivates the co-creation design.","marker":"[11]"},{"why":"Defines alternating and task-divided modes of human-AI collaboration that the iterative workflow instantiates.","marker":"[21]"},{"why":"Demonstrates a prior VR system that blends generated content into panoramic scenes, positioning the new workflow relative to earlier immersive generative tools.","marker":"[6]"},{"why":"Provides the text-to-video generation approach whose motion dynamics the panorama post-processing adapts.","marker":"[16]"},{"why":"Documents equirectangular constraints (2:1 aspect, edge continuity, curved motion) that guide the prompt descriptors and post-processing.","marker":"[29]"}],"fun_headline_variants":["VR co-creation beats one-shot AI for 360° video","Speech-refined VR video generation wins in pilot","Co-creative VR video: users prefer iterative prompts","Voiced co-creation outranks passive AI in VR","Users steer 360° VR video with speech and rotation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pilot's co-creation task was run by a human experimenter who manually imported each generated video and adjusted the user's visual center, so the central claim assumes that this Wizard-of-Oz stand-in faithfully reproduces the automated agent's behavior and that participant preference reflects the design rather than the experimenter's responsiveness.","fun_headline_variants_meta":{"raw":{"variants":["VR co-creation beats one-shot AI for 360° video","Speech-refined VR video generation wins in pilot","Co-creative VR video: users prefer iterative prompts","Voiced co-creation outranks passive AI in VR","Users steer 360° VR video with speech and rotation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1551,"prompt_tokens":879,"completion_tokens":672,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":591}},"tokens_in":495,"tokens_out":672,"duration_ms":6566,"temperature":1.0,"reasoning_tokens":591,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:15:55.603936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A fully automated end-to-end run of the real system, with no experimenter mediating video imports or visual-center changes, compared against the pilot's co-creation condition; if preference, performance, and creativity scores do not reproduce, the Wizard-of-Oz mediation rather than the design was doing the work. A second decisive observation is the frequency of egocentric focal changes in a larger sample: if usage stays near the pilot's 4 of 24 trials even with onboarding, the spatial-control component of the claimed integration is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the creativity framework (novelty, value, surprise, relevance) used to score participant outcomes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the six-dimension workload index used to measure mental load and effort in the pilot."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines alternating and task-divided modes of human-AI collaboration that the iterative workflow instantiates."}],"review_version":1}