{"id":"4c75c169-4e31-4132-96d9-5763e7b4f5af","arxiv_id":"2505.15973","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper maps how storytellers prefer to use AI-generated text, audio, images, videos, and 3D content to augment AR stories, based on a 223-video analysis and two user studies with 30 participants.","lead":"This paper studies whether generative AI can help storytellers create multi-modal content for augmented reality storytelling. It analyzes 223 AR storytelling videos, builds a testbed with text-to-image, text-to-video, text-to-music, and text-to-3D tools, and runs two studies with 30 storytellers to learn which content types they prefer.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The modality-to-element preference mapping in Fig. 5 is confounded with per-modality AIGC quality: poor video generation may artificially depress video preference for Character, Background, and Sentiment.","rationale":"The Reader's verdict is CONDITIONAL and identifies the non-systematic YouTube corpus as the weakest assumption. That is a legitimate concern, but I find a more load-bearing internal confound: the study's central deliverable, the modality-element preference mapping, is measured on outputs of specific generative models whose quality varies by modality. The paper's own data show that video AIGC is rated much lower than image and text AIGC, and its own qualitative analysis attributes video's limited use to low quality. Under these conditions, preference percentages in Fig. 5 conflate 'this modality suits this story element' with 'this modality's generated output is acceptable for this element.' The authors' rebuttal in Section 5.5.2, that poor video quality is an algorithm limitation rather than a modality-suitability limitation, is plausible but cannot be validated without a controlled comparison that holds content quality constant across modalities. This does not invalidate the paper as an exploratory study; the testbed, interaction findings, and design considerations remain useful. However, the strongest claim of a reusable modality-to-element mapping is weaker than stated. Since the Reader already assigned CONDITIONAL, my concern reinforces that verdict without requiring a change. agreement is partial because the Reader's weakest assumption (corpus representativeness) is real but distinct from the quality confound I emphasize; both point in the same direction of limiting generalizability, but the confound is internal to the study's design and can be tested on existing data.","tokens_in":24134,"tokens_out":2599,"duration_ms":26653,"concrete_test":"Re-analyze the existing Study 1 and Study 2 data with a mixed-effects logistic regression predicting each participant's modality choice for each element from element type, the participant's quality rating for that modality (from Fig. 6 or per-element ratings in Study 2), and their interaction, with random intercepts for participants and stories. If modality quality rating is a significant predictor of choice, or if the element-by-quality interaction is significant, then the Fig. 5 modality-element mapping is confounded with AIGC quality. A stronger follow-up would repeat Study 1 with the same five stories but with human-created stock content matched per element; if the video preference share for Background or Sentiment rises materially relative to Fig. 5, the original mapping reflects model limitations rather than modality suitability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the study establishes a reusable mapping of modality suitability to atomic story elements (image for Character/Background/Sentiment, video for Development). That mapping is load-bearing because it drives the design-space taxonomy, the testbed, and the design recommendations. However, preference data in Study 1 were collected after participants previewed content produced by specific generative models: Stable Diffusion for images, MDM plus Text2Video-Zero for video, MusicGen for audio, and DreamFusion for 3D. These models differ sharply in output quality, and the paper reports exactly that: video quality received the lowest rating (AVG=2.8, SD=1.13 in Figure 6), while text and images were rated highest. Section 5.5.1 even states that participants 'only used video at a big scale in the development element where details are safe to ignore' and that 'the quality of the videos was not good enough.' Section 5.5.2 argues that poor video quality is 'due to limitations in algorithmic development rather than the suitability of the video as a modality.' That argument is untestable with the current data because preference and quality are measured on the same generated artifacts. If participants avoid video for Character, Background, and Sentiment partly because the generated video is low-quality, then the Fig. 5 mapping does not establish that video is inherently less suitable for those elements; it only shows that current video AIGC is less acceptable. Thus the reusable design-space mapping is not yet separated from the implementation quality of the specific generative models. This is an internal confound, not merely a disagreement with consensus, and it weakens the strongest claim even if the 223-video corpus were perfectly representative.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory study of multi-modal generative AI (GenAI) for AR storytelling. The authors analyze 223 YouTube videos to derive a design space with two dimensions: five modalities (text, audio, image, video, 3D) and four atomic elements (Character, Background, Sentiment, Development). They implement a testbed that generates content in these modalities from textual narratives and displays it through an AR interface, then conduct two studies with 30 experienced storytellers and presenters (split into two groups of 15). The reported findings are a modality-to-element preference mapping (image preferred for Character, Background, and Sentiment; video preferred for Development), quality ratings of generated content, Likert-scale evaluations of co-creation interactions and suitability, and qualitative insights about alignment, selective augmentation, and context awareness. The paper concludes with design considerations for future AR storytelling systems with GenAI.","tokens_in":24417,"tokens_out":4735,"duration_ms":42085,"significance":"If the claims are properly qualified, the paper makes a useful exploratory contribution: it provides a taxonomy of modalities and story elements, a working testbed for studying GenAI-based AR authoring, and an empirical snapshot of author preferences. Strengths include the explicit acknowledgement of corpus-selection limitations, the reporting of inter-rater agreement for video filtering (κ=0.76), and the inclusion of the full story stimuli in an appendix. However, the central preference mapping is coupled to the specific generative models used in the testbed, and the claimed 'empirical comparison' with human-generated content is not supported by the data. Because the load-bearing conclusions overreach beyond what the design can establish, the paper needs revision before its claims can be accepted as stated.","major_comments":[{"comment":"The contribution list claims an 'empirical comparison of AIGC with human-generated content,' but the manuscript reports no human-generated baseline and no systematic comparison. Section 5.5.2 describes only participants' subjective impressions ('Some of the participants felt the content was as good as human-generated content') and the authors' argument that video quality reflects algorithmic limitations rather than modality suitability. This is not an empirical comparison; the claim should be removed or replaced with a condition in which participants rate matched human-generated content.","section":"§1, Contribution bullet 3; §5.5.2"},{"comment":"The central modality-to-element preference mapping is confounded with per-modality AIGC quality. Preferences in Figure 5 were expressed after previewing outputs from specific models (Stable Diffusion for images, MDM plus Text2Video-Zero for video, MusicGen for audio, DreamFusion for 3D), and Figure 6 reports that video quality was rated lowest (AVG=2.8, SD=1.13). Section 5.5.1 states that 'the quality of the video was not good enough' and that participants used video only where 'details are safe to ignore.' The rebuttal in §5.5.2, that this is 'due to limitations in algorithmic development rather than the suitability of the video as a modality,' is not testable from the current data because each modality is instantiated by a single model. The paper should either frame the conclusions as preferences under current AIGC quality or include a controlled comparison with quality matched across modalities.","section":"§5.1, Fig. 5; §5.5.1; Fig. 6"},{"comment":"The design-space taxonomy is load-bearing for the study: it determines the elements and modalities used in Study 1's pre-highlighted elements, the testbed capabilities, and the interpretation of participant preferences. However, the corpus was assembled through a manual, non-systematic search (stated explicitly in §3.1.1), and inter-rater reliability is reported only for the filtering step (κ=0.76), not for the open coding of modalities/elements or for the final placement of videos in the design space. The study therefore provides no independent check that the four atomic elements are complete or reliably identifiable, which limits the generality of the subsequent preference findings. Please report coding reliability for the design-space dimensions or validate the taxonomy on an independent sample.","section":"§3.1.1, §3.1.3"}],"minor_comments":[{"comment":"The abstract reports 'N=30' without clarifying that this number was split into two studies of 15 participants each; please state this split explicitly for accuracy.","section":"Abstract; §4.2.1"},{"comment":"The x-axis label contains a typo: 'Accepable' should be 'Acceptable.'","section":"Fig. 6"},{"comment":"The heading 'Empirical comparison to human-generated content' overstates what is presented, because no human-generated baseline is included; consider renaming the subsection to 'Participants' perceptions of AIGC versus human-generated content.'","section":"§5.5.2"},{"comment":"The first sentence reads 'Mutli-modal Gen-AI'; this should be corrected to 'Multi-modal Gen-AI.'","section":"§8 Conclusion"},{"comment":"The Mann-Whitney U tests are used to support the statement that there is 'no significant difference' between the two study conditions; however, with N=15 per group, non-significance does not establish equivalence, and no effect sizes or confidence intervals are reported. Please temper these interpretations.","section":"§5.3, §5.4"},{"comment":"The caption states 'All images are generated by ChatGPT,' but ChatGPT alone is not an image-generation model; please specify the image model used (e.g., DALL-E through ChatGPT) for reproducibility.","section":"§6.1.1, Fig. 11"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable exploratory study for an HCI audience, but the public claims exceed the evidence in two places: the stated human-comparison contribution and the generality of the modality-to-element preference mapping. Both are fixable through reframing and additional analysis rather than requiring a fundamentally new study, so I recommend major revision rather than rejection. The authors should also consider reporting the coding reliability for the design-space dimensions, since that taxonomy is the pivot on which the testbed and preference findings rest."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper actually ships something new: an empirically derived design space (5 modalities x 4 atomic elements) from 223 AR storytelling videos, plus a preference map from 30 experienced storytellers saying which modality they would use for which element. That is a useful starting point for anyone building AR storytelling authoring tools. Second, the headline map is partly an artifact of the generative models they chose. Video got the lowest quality ratings by a clear margin, and participants explicitly said they only used video for Development because the videos were too poor for details. So the preference for image over video for Character, Background, and Sentiment is confounded with the uneven quality of Text2Video-Zero/MDM. The authors argue this is an algorithmic limitation rather than a modality suitability issue—that may be right, but it is untestable with their data.\n\nWhat the paper does well: the corpus analysis is reasonably rigorous (Fleiss kappa 0.76, open coding with three authors), the testbed is integrative, and the two studies are sensibly designed—Study 1 assigned elements, Study 2 let people choose freely, and they report both preference and quality plus interaction experiences. The limitations section is honest about desktop AR, hand interaction, and the focus on modalities. The design considerations (selective augmentation, context-awareness, alignment across modalities) are sensible and grounded in participant quotes.\n\nSoft spots, in proportion. The biggest is the confound I mentioned. It does not kill the paper—as an exploratory mapping it is still informative—but it does mean Fig. 5 should be read as 'what these particular AIGC models support' rather than 'which modalities suit which elements.' The contribution list also overclaims an 'empirical comparison of AIGC with human-generated content'—Section 5.5.2 is just participants' impressions; there is no human-generated baseline. That should be fixed in revision. The corpus is admittedly non-systematic, which is a minor concern for an exploratory study. And no code or data is released, which limits reproducibility. Circularity is structural but not disqualifying: they derived the taxonomy from the corpus, built the testbed from it, then used it to elicit preferences—so the study cannot independently validate the taxonomy, but that is not what an exploratory study claims.\n\nWho this is for: HCI researchers and system builders in AR storytelling and GenAI authoring. It is a reasonable empirical anchor for design decisions, not a definitive answer to a long-open question. It deserves a serious referee—I would send it to review, asking for the claims to be scaled back to what the data supports, and ideally for a follow-up that varies generative models or adds a human baseline.\n\nRecommendation: accept for peer review, with major revision.","headline":"Useful exploratory mapping of modality preferences for AIGC in AR storytelling, but the headline map is tangled up with the uneven quality of the specific generative models used.","tokens_in":24973,"tokens_out":2539,"would_cite":true,"duration_ms":21474,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A study of 223 AR videos maps which AI-generated media fits which story element, and 30 storytellers confirm the mapping in practice.","keywords":["augmented reality","storytelling","generative AI","multi-modal content generation","human-AI interaction","design space","authoring tools","user study"],"falsifier":"A systematically sampled corpus of AR storytelling videos that reveals a common modality outside the five (for example haptic feedback) or a common atomic element outside the four (for example audience interaction) would falsify the taxonomy's completeness; likewise, a replication with a state-of-the-art text-to-video model that erases the current video-for-development-only preference would falsify the modality-element mapping as a stable property of AR storytelling.","tokens_in":23926,"feed_emoji":"🖼️","tokens_out":4822,"duration_ms":40444,"temperature":0.7,"pith_summary":"The paper asks whether AI-generated content can carry the multimodal load of augmented-reality storytelling, and answers with a qualified yes backed by a design space and two user studies. Analyzing 223 AR storytelling videos on YouTube, the authors identify five content modalities—text, audio, image, video, and 3D—and four atomic story elements: character, background, sentiment, and development. In studies with 30 experienced storytellers and presenters using a GenAI-powered authoring testbed, participants preferred images for characters, backgrounds, and sentiment, and video for plot development, while rating generated text highest in quality and generated video lowest. The result matters because it gives future AR authoring systems a reusable mapping of which AI modality to offer for which narrative element, and it pinpoints the open problems: aligning outputs across modalities, making prompts controllable, and deciding what to augment at all.","feed_headline":"Image is the go-to AI modality for AR story characters and mood","feed_subtitle":"A 223-video corpus and 30 storytellers point to images for characters, backgrounds, and sentiment, and video for plot development.","key_machinery":"The load-bearing object is a two-dimensional design space: five Modalities (Text, Audio, Image, Video, 3D) crossed with four atomic Elements (Character, Background, Sentiment, Development). The authors derive it by open-coding 223 YouTube AR storytelling videos, then use it to structure a testbed in which a storyteller selects a sentence, chooses a modality, and receives generator output from Stable Diffusion (image), MusicGen (audio), a motion-diffusion plus Text2Video-Zero pipeline (video), and a text-to-3D model; the AR interface then triggers the saved content by speech while the narrator gestures with hand-tracked interaction. The design space organizes both the corpus analysis and the study tasks, so the preferences reported are preferences over cells of this five-by-four grid.","core_discovery":"The central claim is that multi-modal AIGC is suitable for AR storytelling, provided the modality is matched to the story element. From the 223-video analysis the authors derive a design space of five modalities and four atomic elements (Character, Background, Sentiment, Development). Their two studies, each with 15 participants, show that images are strongly preferred for characters (51%), backgrounds (47%), and sentiment (47%); video is preferred for development (40%), with text close behind (30%); and 3D is a secondary choice for characters. They also find that participants rate generated text highest in quality (4.43/5) and video lowest (2.8/5), and that while co-creating with AI feels fast and enjoyable, guiding the generation to match intention is the hardest part. The authors further report that participants could mostly tell AIGC apart from human-made content but said it did not hurt the storytelling, and that cross-modal inconsistencies and literal misinterpretation of metaphors are the main blockers to wider use.","pith_inferences":["The modality-element preferences are partly confounded by current model quality: participants avoided video for small details because outputs were poor, so a replication with a stronger text-to-video model might shift the mapping.","The four-element taxonomy may be reusable beyond AR, as a general vocabulary for choosing generative media in slideware, virtual reality, or interactive fiction.","The non-systematic YouTube sampling means the taxonomy should be treated as a starting hypothesis; a systematic corpus study could add elements such as interactivity or user choice that this work deliberately excludes.","If cross-modal alignment is solved, the same testbed pattern could support live 'when to augment' decisions during presentations rather than pre-authored augmentation only."],"forward_implications":["Future AR storytelling authoring tools can present images as the default augmentation for characters, backgrounds, and emotional tone, and reserve video for temporal development.","Because participants found text the clearest and most reliable output, tools should keep text as a fallback or complement even when visuals are preferred.","The low video-quality ratings imply that improvements in text-to-video generation could enlarge the role of video beyond development into other elements.","The repeated difficulty in prompting suggests authoring systems need example-based, iterative, or in-AR prompting beyond plain text.","The observed cross-modal misalignment argues for generating each story entity once and sharing its representation across modalities."],"supporting_citations":[{"why":"Supplies the speech-driven augmented presentation paradigm that this testbed extends with GenAI content.","marker":"[56]"},{"why":"Provides the prior example of AI-generated visuals augmenting live verbal storytelling.","marker":"[57]"},{"why":"Establishes the GenAI visual storytelling approach this study generalizes to multiple modalities.","marker":"[8]"},{"why":"Supplies the expressiveness, immersion, exploration, and alignment factors used in the questionnaires.","marker":"[54]"},{"why":"Generates the character motion that anchors the testbed's video pipeline.","marker":"[102]"},{"why":"Turns the generated motion into the video modality in the testbed.","marker":"[47]"},{"why":"Produces the image modality in the content generator.","marker":"[87]"},{"why":"Produces the audio modality in the content generator.","marker":"[19]"},{"why":"Produces the 3D modality in the content generator.","marker":"[81]"}],"fun_headline_variants":["Images lead for AR story characters, backgrounds, sentiment","Video drives AR plot development, text scores highest quality","GenAI in AR storytelling: images for mood, video for plot","Cross-modal gaps and literal AI block AR story immersion","AR storytellers pick images for character, video for plot"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings stand on the assumption that the 223 manually collected YouTube videos represent the space of AR storytelling well enough for the five-by-four taxonomy to be the right frame for the testbed and the study tasks; the authors say plainly that the corpus was not collected by systematic search.","fun_headline_variants_meta":{"raw":{"variants":["Images lead for AR story characters, backgrounds, sentiment","Video drives AR plot development, text scores highest quality","GenAI in AR storytelling: images for mood, video for plot","Cross-modal gaps and literal AI block AR story immersion","AR storytellers pick images for character, video for plot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1365,"prompt_tokens":938,"completion_tokens":427,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":347}},"tokens_in":554,"tokens_out":427,"duration_ms":4066,"temperature":1.0,"reasoning_tokens":347,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:08:57.405043+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematically sampled corpus of AR storytelling videos that reveals a common modality outside the five (for example haptic feedback) or a common atomic element outside the four (for example audience interaction) would falsify the taxonomy's completeness; likewise, a replication with a state-of-the-art text-to-video model that erases the current video-for-development-only preference would falsify the modality-element mapping as a stable property of AR storytelling.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the speech-driven augmented presentation paradigm that this testbed extends with GenAI content."},{"cited_title":"Bruce\" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang","cited_arxiv_id":null,"evidence_quote":"Provides the prior example of AI-generated visuals augmenting live verbal storytelling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the expressiveness, immersion, exploration, and alignment factors used in the questionnaires."},{"cited_title":"Helping Nemo!","cited_arxiv_id":null,"evidence_quote":"Generates the character motion that anchors the testbed's video pipeline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Produces the image modality in the content generator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Produces the 3D modality in the content generator."}],"review_version":1}