Pith. sign in

REVIEW 4 major objections 5 minor 36 references

"See What I Imagine, Imagine What I See": Human-AI Co-Creation System for 360$^\circ$ Panoramic Video Generation in VR

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that a VR workflow where users co-create 360-degree video with an AI agent through iterative speech prompts and egocentric focal re-centering outperforms both non-AI and one-shot AI workflows on user-valued outcomes.

desk verdict A clearly described co-creation workflow for 360° video in VR, but the pilot study is a Wizard-of-Oz simulation with n=8, so the headline claims outrun the evidence. read the letter →

arxiv 2501.15456 v1 pith:NTBYCIO6 submitted 2025-01-26 cs.HC

classification cs.HC
keywords 360-degreevideogenerationpanoramichuman-AIco-creationvirtualrealityspeech-basedpromptingegocentricviewadjustmentiterativeWizard-of-Ozpilot
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Imagine360 is a proof-of-concept VR system that turns 360-degree video generation into an iterative dialogue between user and AI agent. Instead of issuing one prompt and waiting for a finished panorama, the user speaks an intention, the agent refines it into a text prompt, a short video segment is generated from the last frame of the previous segment, and the user can rotate to choose the focal center of the next generation. The paper's central claim is that this co-creative loop integrates temporal control (what happens next) with spatial control (where the user is looking), and the eight-participant pilot comparing it with pre-recorded non-AI video and one-shot AI workflows supports that claim on user-valued outcomes. The contribution is the interaction design and pilot evidence, not a new generative model.

What carries the argument

The load-bearing mechanism is the agent-mediated feedback loop: each new video segment is generated from the last frame of the segment the user just watched, re-centered to the user's current egocentric angle, and driven by a speech prompt that the agent refines and offers suggestions for. The raw generated video is then post-processed into an equirectangular projection—a 2:1 flat map of a sphere with blended edges, a blurred background, and reduced foreground height—so it can be viewed continuously in the VR headset. This loop connects seeing (the frame still in view) to imagining (the next segment's prompt) and is what the paper claims integrates temporal and spatial controls.

What would settle it

A fully automated end-to-end run of the real system, with no experimenter mediating video imports or visual-center changes, compared against the pilot's co-creation condition; if preference, performance, and creativity scores do not reproduce, the Wizard-of-Oz mediation rather than the design was doing the work. A second decisive observation is the frequency of egocentric focal changes in a larger sample: if usage stays near the pilot's 4 of 24 trials even with onboarding, the spatial-control component of the claimed integration is unsupported.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that people can co-create 360-degree video in VR when generation is broken into segments and each segment starts from the visual state the user actually saw. Users issue speech commands, have them refined or suggested by an AI agent, and can change the focal angle of the next image prompt by physically rotating; the system maps image length to angular degrees and re-centers the panorama accordingly. The pilot's performance ratings were highest for this co-creation condition (mean 8.875 vs 8.75 for non-AI and 8.5 for linear AI), and participants reported strong relevance and value in the outcomes, with a correlation of $r = 0.81$ between relevance and value. The paper also reports that the co-creation condition imposed higher mental load, and that only 17% of trials used a perspective change, because users preferred adjusting overall composition.

Load-bearing premise

The pilot's co-creation task was run by a human experimenter who manually imported each generated video and adjusted the user's visual center, so the central claim assumes that this Wizard-of-Oz stand-in faithfully reproduces the automated agent's behavior and that participant preference reflects the design rather than the experimenter's responsiveness.

Editorial extensions

If this is right

  • If the pilot result holds, VR experiences can be authored in real time from the user's speech and gaze instead of being limited to pre-designed scenes.
  • Segment-wise generation from the last frame turns temporal continuity into a user control: every 10-second segment is a direct revision of what was just seen.
  • Egocentric focal re-centering gives users a spatial control that one-shot text-to-video prompting does not offer.
  • The higher measured mental load and the low 17% perspective-change rate imply that the next iteration should reduce co-creation overhead or make spatial control easier to discover.
  • The refined prompt suggestions allow users with little prompting experience to steer generation by confirming or choosing from agent-provided options.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: because the evaluation used a Wizard-of-Oz stand-in, the pilot validates the co-creation concept more than the implemented automation; an end-to-end automated system could produce different preference results.
  • Editorial inference: the low rate of focal-angle changes (4 of 24 trials) suggests egocentric spatial control may be secondary for this user group, and a tutorial or more salient affordance could reveal whether the low usage is discoverability or preference.
  • Editorial inference: the post-processing choices that make 2D output feel panoramic (blur, height reduction, edge blending) trade visual fidelity for seamlessness, so future studies could separate workflow satisfaction from output-quality judgments.
  • Editorial inference: the last-frame continuation loop transfers naturally to non-panoramic creative tools such as storyboard or animation previz, where each revision starts from the last rendered frame.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. Imagine360 is a proof-of-concept VR system that lets users generate 360° panoramic videos through iterative speech-based prompts, AI prompt refinement, and egocentric focal recentering. The paper describes the system architecture (video generation via Runway Gen-3 Alpha Turbo, equirectangular projection, voice input via Whisper-1, prompt optimization via GPT) and reports an eight-participant pilot study comparing three conditions: non-AI pre-recorded videos, linear AI-driven generation, and human-agent co-creation. Quantitative measures (NASA-TLX, Boden creativity ratings) and semi-structured interviews are used to argue that the co-creative workflow better integrates temporal and spatial creative controls. The evaluation used a Wizard-of-Oz protocol in which the experimenter manually imported generated videos and adjusted the user's visual center.

Significance. If the effectiveness claims were fully supported, the paper's contribution would be a workflow design rather than a new generative model: it demonstrates that segment-wise iterative speech refinement and egocentric recentering are usable and preferred over one-shot prompting in immersive VR. The paper is transparent about its proof-of-concept status and about the Wizard-of-Oz simulation, and the qualitative interview data provide useful design insights. However, the central claim is not as strong as the abstract states: the principal performance metric is not statistically significant, the sample size is eight, and the automated system was not actually evaluated. The paper's value at this stage is as a pilot that motivates future work, not as a demonstration of the effectiveness of the Imagine360 agent.

major comments (4)
  1. [Abstract and §4.2] The performance advantage claimed for co-creation is not statistically supported: the mean performance ratings (8.875 vs. 8.75 vs. 8.5) are reported without any test statistic, while the only significant effects reported are higher mental load and effort for co-creation. The abstract's phrase 'demonstrates that Imagine360's co-creative approach effectively integrates temporal and spatial creative controls' therefore overstates the evidence; please report the relevant test (with effect size) or reframe the claim as a qualitative/pilot finding.
  2. [§4.1 Task 3] The evaluation used a Wizard-of-Oz protocol: the experimenter manually imported generated videos and adjusted the user's visual center. Because no fidelity check, intervention log, or inter-rater reliability measure is reported, and because the experimenter was not blind to the hypotheses, the results cannot be attributed to the automated Imagine360 system described in Section 3. This is load-bearing because all positive outcomes are attributed to the co-creation framework. Please either evaluate the actual automated pipeline (even in a limited pilot) or explicitly scope the conclusions to the WoZ-simulated workflow and add a fidelity assessment.
  3. [§4.2] Spatial control was rarely exercised: only 4 of 24 Task 3 trials (17%) involved perspective changes, yet the central claim is that the system integrates temporal and spatial creative controls. The paper should report what happened in those four trials and discuss why participants preferred composition-level adjustments; without this, the spatial-control component of the claim is supported only by interview preferences, not by usage data.
  4. [§4.1 and §4.2] The statistical reporting is incomplete. The Friedman test is written as χ2(3,N) without giving N or degrees of freedom; with three conditions the df should be 2, not 3. Please specify the test details, the correction used for post-hoc Wilcoxon tests, and effect sizes, and clarify what 'significant relevance and value' refers to (no test is reported for the Boden measures).
minor comments (5)
  1. [Abstract] The word 'demonstrates' should be softened to 'suggests' or 'provides initial evidence from a pilot study,' given the small sample and the Wizard-of-Oz simulation.
  2. [§3 and Figure 2] The text in Section 3 says the prompt optimization uses GPT-3.5 Turbo, while Figure 2 and other parts of the paper say GPT-4o; please reconcile this inconsistency.
  3. [Figure 1] The caption for Figure 1 is very long and mixes philosophical framing with system description; consider shortening it and moving the philosophical context to the introduction.
  4. [§4.1 Task 2] The procedure for the linear AI-driven condition is underspecified: please state how the three prompts were converted into videos and whether the experimenter also manually imported and adjusted those videos, so that the comparison with Task 3 is clear.
  5. [§4.2] The sentence 'Our system shows significant relevance and value in outcomes and participants' creativity' does not report a test statistic or a p-value; please specify which comparison and which measure this refers to.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the pilot is an empirical usability comparison, not a derivation, and its Wizard-of-Oz limitation is a validity threat rather than a circular step.

full rationale

This paper does not contain a derivation chain of the kind the circularity pass targets. Imagine360 is a proof-of-concept prototype that composes third-party models (Whisper-1, GPT-3.5/4o, Gen-3 Alpha Turbo) with a panorama-projection post-process, and the evaluation is an eight-participant within-subjects pilot comparing three workflows. There are no fitted parameters, no predictive equations, and no "uniqueness" claims; consequently none of the enumerated circular patterns (self-definitional, fitted-input-called-prediction, self-citation load-bearing, imported uniqueness, ansatz-by-citation, renaming) is present. The single-authored paper contains no self-citations that carry the argument. The most important limitation is disclosed in Section 4.1, Task 3: "In Task 3, we employed a Wizard of Oz (WoZ) approach, manually importing the generated videos into the VR headset and adjust user's visual center." This means the study validates a human-simulated agent rather than the automated pipeline of Section 3, and the paper reports no fidelity check between the experimenter's behavior and the intended agent behavior. This is a real threat to the external validity of the headline claim, but it is not circularity under the rubric: the study does not assume the conclusion to prove it; it empirically compares conditions with a confounded implementation. Similarly, the possibility that participants rated the more effortful co-creation condition favorably (IKEA effect), or that only 4 of 24 trials exercised perspective changes, weakens the evidence base without making the claim equivalent to its inputs. These concerns belong in a correctness/validity assessment, not in the circularity score. Verdict: no significant circularity, score 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The claims rest on domain assumptions about self-report validity, the faithfulness of the Wizard-of-Oz substitution, and the adequacy of the 2D-to-equirectangular post-processing; plus a small-sample statistical assumption. No numerical model parameters were fitted.

assumptions (4)
  • domain assumption Self-reported ratings on NASA-TLX and Boden scales capture creative outcome quality.
    Effectiveness and creativity are measured only through participant self-report (Section 4.2); no external judges, behavioral logs, or objective output metrics are used.
  • ad hoc to paper The Wizard-of-Oz protocol faithfully emulates the automated agent.
    Section 4.1 Task 3 substitutes a human experimenter for the system's AI agent due to integration failures; preferences measured in this condition are attributed to the system in the abstract.
  • domain assumption Post-processing (2:1 aspect ratio, 50% Gaussian blur, 75% foreground height) produces a valid 360-degree immersive experience.
    Section 3 'Panorama Projection' describes these heuristics without geometric validation; participants reported blurriness from the headset, but no checks on projection fidelity are reported.
  • standard math Repeated-measures nonparametric statistics are valid with eight participants.
    Section 4.2 applies Friedman and post-hoc Wilcoxon tests to small ordinal Likert data; no power analysis or effect sizes are reported.
invented entities (1)
  • Imagine360 AI co-creation agent
    purpose: Refines speech prompts, suggests edits, recenters panorama, and drives iterative segment generation
    Described as a working component in Section 3, but the pilot study used a human wizard in its place (Section 4.1), so no end-to-end implementation was demonstrated for evaluation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "See What I Imagine, Imagine What I See": Human-AI Co-Creation System for 360$^\circ$ Panoramic Video Generation in VR." pith.science (2026). https://pith.science/paper/NTBYCIO6

@misc{pith2026250115456,
  author       = {Pith},
  title        = {Pith review of: "See What I Imagine, Imagine What I See": Human-AI Co-Creation System for 360$^\circ$ Panoramic Video Generation in VR},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NTBYCIO6}},
  note         = {Machine review of arXiv:2501.15456}
}
read the original abstract

The emerging field of panoramic video generation from text and image prompts unlocks new creative possibilities in virtual reality (VR), addressing the limitations of current immersive experiences, which are constrained by pre-designed environments that restrict user creativity. To advance this frontier, we present Imagine360, a proof-of-concept prototype that integrates co-creation principles with AI agents. This system enables refined speech-based text prompts, egocentric perspective adjustments, and real-time customization of virtual surroundings based on user perception and intent. An eight-participant pilot study comparing non-AI and linear AI-driven workflows demonstrates that Imagine360's co-creative approach effectively integrates temporal and spatial creative controls. This introduces a transformative VR paradigm, allowing users to seamlessly transition between 'seeing' and 'imagining,' thereby shaping virtual reality through the creations of their minds.

Figures

Figures reproduced from arXiv: 2501.15456 by the authors.

Figure 1
Figure 1. The boundary between experience and representation has long been debated in philosophical inquiry. Building on the [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Co-Creation Workflow of Panorama Video Generation in VR. The system enables users to (1) generate panoramic [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Qualitative Comparisons of Participant Genera [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 18 canonical work pages

  1. [1]

    The Matrix

    1999. The Matrix. IMDb: tt0133093

  2. [2]

    Inception

    2010. Inception. IMDb: tt1375666

  3. [3]

    Naofumi Akimoto, Yuhi Matsuo, and Yoshimitsu Aoki. 2022. Diverse plausible 360-degree image outpainting for efficient 3DCG background creation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11441–11450

  4. [4]

    Jie An, Songyang Zhang, Harry Yang, Sonal Gupta, Jia-Bin Huang, Jiebo Luo, and Xi Yin. 2023. Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation. arXiv preprint arXiv:2304.08477 (2023)

  5. [5]

    Margaret A. Boden. 2004. The Creative Mind: Myths and Mechanisms. Psychology Press

  6. [6]

    Zeyu Cai, Zhelong Huang, Xu Zheng, Yexin Liu, Chao Liu, Zeyu Wang, and Lin Wang. 2024. Interact360: Interactive Identity-driven Text to 360° Panorama Generation. In 2024 IEEE Conference on Artificial Intelligence (CAI) . 728–736. doi:10.1109/CAI59869.2024.00141

  7. [7]

    Kai-Yuan Cheng. 2014. Self and the Dream of the Butterfly in the Zhuangzi. Philosophy East and West 64 (07 2014), 563–597. doi:10.1353/pew.2014.0051

  8. [8]

    Xinhua Cheng, Nan Zhang, Jiwen Yu, Yinhuai Wang, Ge Li, and Jian Zhang. 2023. Null-space diffusion sampling for zero-shot point cloud completion. In Proceed- ings of the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI)

Show all 36 references
  1. [9]

    Lydia B Chilton, Ecenaz Jen Ozmen, Sam H Ross, and Vivian Liu. 2021. VisiFit: Structuring Iterative Improvement for Novice Designers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21)(Yokohama, Japan). Association for Computing Machinery...

  2. [10]

    John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: Visual Sketching of Story Generation with Pretrained Language Models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22...

  3. [11]

    Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, and Noah A. Smith. 2018. Creative Writing with a Machine in the Loop: Case Studies on Slogans and Stories. In Proceedings of the 23rd International Conference on Intelli- gent User Interfaces (IUI ’18) (Tokyo, Japan...

  4. [12]

    Robbersmyr, and Kristian Muri Knausgård

    Anurag Dalal, Daniel Hagen, Kjell G. Robbersmyr, and Kristian Muri Knausgård

  5. [13]

    Katy Ilonka Gero, Vivian Liu, and Lydia Chilton. 2022. Sparks: Inspiration for Science Writing Using Language Models. In Proceedings of the 2022 ACM Designing Interactive Systems Conference (DIS ’22) (Virtual Event, Australia). Association for Computing Machinery, New York, NY...

  6. [14]

    William Gibson. 1984. Neuromancer. Ace Books, New York. ISBN: 978- 0441569595

  7. [15]

    Frederic Gmeiner, Humphrey Yang, Lining Yao, Kenneth Holstein, and Nikolas Martelaro. 2023. Exploring Challenges and Opportunities to Support Designers in Learning to Co-Create with AI-Based Manufacturing Design Tools. InProceedings of the 2023 CHI Conference on Human Factors ...

  8. [16]

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2023. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

  9. [17]

    Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Human mental workload. Elsevier, 139–183

  10. [18]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33 (2020), 6840–6851

  11. [19]

    Francisco Ibarrola, Tomas Lawton, and Kazjon Grace. 2023. A Collaborative, Interactive and Context-Aware Drawing Agent for Co-Creative Design. IEEE Transactions on Visualization and Computer Graphics (2023), 1–13. doi:10.1109/ TVCG.2023.3293853

  12. [20]

    Immanuel Kant. 1790. Critique of Judgment. Penguin Classics, London. Original work published in 1790

  13. [21]

    Anna Kantosalo and Hannu Toivonen. 2016. Modes for creative human-computer collaboration: Alternating and task-divided co-creativity. In Proceedings of the Seventh International Conference on Computational Creativity . 77–84

  14. [22]

    Alfred Lan, Tai-Chen Tsai, Chih-Chuan Huang, Pu Ching, Tse-Yu Pan, and Min- Chun Hu. 2024. ImmerseSketch: Transforming Creative Prompts into Vivid 3D Environments in VR. In ACM SIGGRAPH 2024 Posters (Denver, CO, USA) (SIGGRAPH ’24). Association for Computing Machinery, New Yor...

  15. [23]

    Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. 2024. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence. 4296–4304

  16. [24]

    Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. GLIDE: Towards photore- alistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741 (2021)

  17. [25]

    Arthur Schopenhauer. 1818. The World as Will and Representation . Dover Publications, New York. Original work published in 1818

  18. [26]

    Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. 2022. Make-a-video: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792 (2022)

  19. [27]

    Neal Stephenson. 1992. Snow Crash. Bantam Books, New York. ISBN: 978- 0553380958

  20. [28]

    Dmitry Tochilkin, David Pankratz, Zexiang Liu, Zixuan Huang, Adam Letts, Yangguang Li, Ding Liang, Christian Laforte, Varun Jampani, and Yan-Pei Cao

  21. [29]

    Hai Wang, Xiaoyu Xiang, Yuchen Fan, and Jing-Hao Xue. 2024. Customizing 360- Degree Panoramas through Text-to-Image Diffusion Models . In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). IEEE Computer Society, Los Alamitos, CA, USA, 4921–4931. doi:10...

  22. [30]

    TripoSR: Fast 3D Object Reconstruction from a Single Image.arXiv preprint arXiv:2403.02151 (2024)

  23. [31]

    Santiago Negrete Yankelevich and Nora Angelica Morales Zaragoza. 2014. The apprentice framework: planning and assessing creativity. In Proceedings of the International Conference on Computational Creativity (Jönköping, Sweden). As- sociation for Computational Creativity, 280–283

  24. [32]

    Qian Wang, Weiqi Li, Chong Mou, Xinhua Cheng, and Jian Zhang. 2024. 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model. arXiv preprint arXiv:2401.06578 (2024). Yunge Wen

  25. [33]

    Rosina Yuan, Antony Tang, Qianyuan Zou, Masoumeh Hesam Mahmoudinezhad, Yuewei Zhang, and Iain Anderson. 2024. Finger Painting in VR: Multi-Dynamic Gestural Input for VR Painting. In SIGGRAPH Asia 2024 XR (SA ’24) . Association for Computing Machinery, New York, NY, USA, Articl...

  26. [34]

    Emilie Yu, Fanny Chevalier, Karan Singh, and Adrien Bousseau. 2024. 3D-Layers: Bringing Layer-Based Color Editing to VR Painting. 43, 4, Article 101 (July 2024), 15 pages. doi:10.1145/3658183

  27. [36]

    Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. 2022. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018 (2022)

  28. [2024]

    arXiv preprint arXiv:2405.03417 (2024)

    Gaussian Splatting: 3D Reconstruction and Novel View Synthesis, a Review. arXiv preprint arXiv:2405.03417 (2024)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.