{"id":"ade59244-19ae-4267-806b-9401d7f928a6","arxiv_id":"2504.13119","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"VLM-generated metaphorical narratives can be anchored to real-world objects in AR and modestly improve user engagement, but the evidence is limited by small samples and missing baselines.","lead":"This paper presents an AR storytelling system that uses a vision language model to read everyday objects in a room as symbolic story elements, then wires the generated metaphors into Unity through a JSON interface. The authors report small user studies suggesting the approach makes people view their surroundings with new narrative meaning.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 70% statistic is unauditable: the flagship evidence for 'stories from the environment itself' does not correspond to any reported measurement.","rationale":"The reader correctly identified missing baselines, incomplete metrics, and small samples, and their verdict of CONDITIONAL is reasonable. My concern is more specific: the paper's flagship quantitative claim—the 70% reinterpretation rate—does not appear in any table or method description, and the Discussion introduces several other unreported percentages (57%, 23%, 42%, 78%). This is not a matter of interpretation; it is a traceability failure in the central evidence. I therefore agree with the conditional verdict, but the acceptance conditions should explicitly require reporting the raw survey data or replacing the unauditable percentages with numbers that can be recomputed. I would not reject the paper: the three-phase evaluation and the JSON-anchor architecture are plausible, and the user studies, though small, provide some directional support. The concern is fixable through transparent reporting and, ideally, a control condition in the Section 6 deployment. Hence UNCHANGED (still CONDITIONAL), with the condition broadened to include auditability of the headline statistics.","tokens_in":13209,"tokens_out":5059,"duration_ms":46116,"concrete_test":"Locate the raw post-experiment survey responses or item definitions behind '70% of participants reported seeing real-world objects differently' in Section 6.2 / Table 6. Determine whether 70% corresponds to the share of participants rating Recognition above a pre-specified threshold on the 1–7 scale; if possible, recompute this from per-participant logs. If the value cannot be produced, revise the Abstract and Section 6.2 to state the actual reported statistic (e.g., mean Recognition = 4.80/7) and either add the underlying item or drop the percentage claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract and conclusion—'70% of participants reported seeing real-world objects differently'—is not supported by any reported result in the body. The only related number is Recognition mean = 4.80/7 in Table 6 (17-participant integration study), but a 7-point Likert mean is not a percentage, and no survey item, response threshold, or raw count is given to justify 70%. The same pattern is repeated elsewhere: Discussion cites '57% reduction in 3D coordinate errors,' 'user ratings 23% higher than baseline,' '42% reduction in user disorientation,' and '78% of non-expert users struggled,' none of which can be recomputed from Tables 2, 5, or 6 or from the method text. Because this 70% figure is the only direct empirical support for the paper's central claim that narratives are generated from environmental symbolism rather than superimposed on it, the absence of traceability is load-bearing. The Section 6.1 study also has no control or baseline condition, so even the reported subjective ratings (spatial fit 5.31/7, recognition 4.80/7) cannot isolate the contribution of metaphor-grounded object detection from generic VLM story generation. The narrowness of the deployment (one office scene, 3 key objects, 10 branch items, 17 participants) is a secondary issue; the primary issue is that the headline number cannot be audited.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a scenario-metaphor framework for AR storytelling that integrates Vision Language Models (VLMs) with a structured JSON bridge to map metaphorical narratives onto AR anchors. The framework decomposes object semantics into physical, functional, and metaphorical layers, and is evaluated in three phases: a STAM capability benchmark for VLM spatial/narrative reasoning (Table 2), a cognitive-alignment user study (n=26; Tables 3-5), and a system integration study in an office scene (n=17; Table 6). The paper claims that 70% of participants reported seeing real-world objects differently when narratives were grounded in environmental symbolism, and that the system achieves spatial fit of 5.31/7, among other results.","tokens_in":13479,"tokens_out":4418,"duration_ms":40712,"significance":"If fully substantiated, the framework addresses a real gap: moving AR storytelling beyond object-label semantics toward relation- and metaphor-aware narrative generation. The proposed JSON intermediary and the three-phase evaluation protocol are potentially reusable contributions, and the user studies target a meaningful question about whether VLM-generated metaphors can enhance spatial engagement. However, the current manuscript reports several headline statistics that cannot be recomputed from the tables or method text, and the key studies lack baselines, variance measures, and full data. The significance of the contribution is therefore conditional on a substantial reporting overhaul.","major_comments":[{"comment":"The claim that '70% of participants reported seeing real-world objects differently' is load-bearing for the paper's central conclusion, but it is not traceable to any reported measurement. The only related result is Recognition = 4.80/7 in Table 6, which is a 7-point Likert mean, not a percentage; no survey item, response distribution, or threshold is provided that would yield 70%. This claim must be either backed by the exact item, counts, and threshold, or removed from the abstract and conclusion.","section":"Abstract / Section 9 / Table 6"},{"comment":"Several quantitative claims appear without any supporting computation: 'user ratings are 23% higher than baseline,' 'reducing 3D coordinate errors by 57%,' 'reduced user disorientation by 42%,' '78% of non-expert users struggled,' and a '118%' error surge in dense scenes. None of these numbers can be derived from Tables 2, 5, or 6 or from the method text, and Section 6.1 describes no baseline or control condition against which such percentages could be computed. The authors must either provide the precise measurements, definitions, and statistical tests for each claim, or delete the unsupported figures.","section":"Sections 7 and 9"},{"comment":"Table 2 contains empty cells for CE and DT, yet Section 4.5 discusses these as key limitations; moreover, every column lacks variance, number of trials, and per-condition sample sizes. Without these values, the STAM-based conclusions about VLM capability, including the claimed weakness in coordinate estimation and dynamic tolerance, are not auditable. Please report the actual CE and DT values, along with standard deviations and trial counts, or state explicitly that they were not measurable and explain why.","section":"Table 2 / Section 4"},{"comment":"The text states that 'Story 2 outperformed Story 1 across all dimensions except for understanding,' but Table 4 shows that the differences are not statistically significant for most dimensions (e.g., Reasonable 1 vs 2 p=0.1306; Suitable p=0.8658). Additionally, the authors report that nearly all participants selected whichever story they read second as more engaging, which is a severe order confound for the 'Interesting' dimension (p=0.0494 for 1 vs 2). This confound is acknowledged but not controlled, so the claim of strong support for RQ2 is overstated; a within-subject counterbalancing analysis or an explicit regression on presentation order is needed.","section":"Section 5.3 / Table 4"},{"comment":"The system integration study has no baseline or control condition and tests only a single office scene with 3 key objects, 10 branch items, and 17 participants. Consequently, the favorable ratings (e.g., Spatial Fit = 5.31/7, Motivation = 5.19/7) cannot be attributed to the metaphor-grounded object detection or the JSON bridge rather than to generic VLM story generation, novelty effects, or the specific scene. A comparison condition (e.g., stories generated without the metaphorical layer) or a clear statement that the study is a feasibility demonstration rather than an efficacy test is required before the claimed causal contribution is made.","section":"Section 6.1"}],"minor_comments":[{"comment":"The framework is first called 'STEAM' and later 'STAM'; please use a single consistent acronym throughout.","section":"Section 3.1"},{"comment":"The text contains an unresolved table reference: 'which is shown in Table ??'—this must be fixed.","section":"Section 4"},{"comment":"The heading '2.3 System Implementation' appears in the middle of the Related Work section, and the subsequent paragraph returns to interactive narratives; the section ordering appears accidentally misplaced.","section":"Section 2.2 / 2.3"},{"comment":"References [24] and [25] are duplicates of the same Retargetable AR paper; one should be removed and citations updated.","section":"References"},{"comment":"Figure 1 is not cited in the text; please add a citation where the pipeline is first described.","section":"Figures"},{"comment":"The table reports only mean ratings; standard deviations and per-item participant counts should be added, and the number of participants per survey item should be stated.","section":"Table 6"},{"comment":"The phrase 'exceeded user expectations by about 18.9% on average compared to the median value' is undefined; please clarify the baseline and computation method.","section":"Section 5.2"},{"comment":"The scenario names 'Living Area', 'Work Area', 'Special Environment', and 'Lab (Macro)' are not defined in the text, and the caption does not indicate how many images or trials per scenario were used.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's technical idea is potentially interesting, but the gap between the narrative claims (70%, 23%, 57%, 42%, 78%) and the reported tables is unusually wide, and the primary 70% statistic has no corresponding measurement anywhere in the body. I would ask the editor to require the authors to provide a data appendix or supplementary material with raw response counts, survey instruments, and exact computations for every percentage and percentage reduction claimed, or to delete the unsupported figures. Without this, the empirical core of the paper cannot be verified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a read, but the headline number doesn't exist. The 70% claim in the abstract and conclusion is presented as the key evidence for \"stories from the environment itself,\" yet no survey item, response threshold, or count in the body produces that percentage. The closest is Recognition mean 4.80/7 in a 17-person study, and a Likert mean isn't a percentage. Same problem with the \"23% higher than baseline,\" \"57% reduction,\" \"42% reduction,\" and \"78% of non-experts\" — none can be recomputed from Tables 2, 5, or 6. That's load-bearing because the central claim rests on it.\n\nWhat is genuinely new is the design: decomposing object semantics into physical, functional, and metaphorical layers, and using a bidirectional JSON layer to map VLM metaphor output to AR anchors. That's a concrete answer to an actual integration problem, and the three-phase validation structure (STAM capability benchmark, user metaphor study, in-the-wild AR deployment) is a sensible shape, even if execution is rough. I also credit them for openly reporting the bad results, like immersion at 2.81/7 and the comprehension tradeoff for metaphor-heavy stories.\n\nThe empirical softness is real but typical of a systems paper at this stage: n=26 and n=17, one office scene with 3 key objects, empty cells in Table 2 for CE and DT, no baseline in the integration study, and no error bars on the user ratings. The lack of a baseline matters most — the 5.31/7 spatial fit can't be attributed to the metaphor layer rather than to generic VLM story generation. The stress-test note holds up; I don't think it's wrong.\n\nOn balance, the framework is a plausible contribution to AR/HCI, not a broad scientific advance. The writing is a bit rough (the \"Conference'17\" template is a giveaway), and the consistency problems between abstract and body should have been caught. But the authors know their limitations and are transparent about the tradeoffs. This deserves peer review — a serious referee could help them cut the unsupported percentages and add one clean baseline condition. If they replaced the invented statistics with the actual 4.80/7 and 5.31/7 scores, the paper would be honest and still useful.\n\nI'd bring it to a reading group as an example of how LLM+AR evaluation gets overclaimed. I wouldn't cite it in my own work until the numbers are cleaned up.","headline":"A genuinely new VLM-to-AR pipeline with a solid design contribution, but the flagship 70% statistic doesn't exist in the data and the evaluation needs major cleanup before the claims can be trusted.","tokens_in":13969,"tokens_out":2082,"would_cite":false,"duration_ms":18674,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A scene-driven AR storytelling framework uses a vision-language model to turn real objects into metaphor anchors, and reports that 70% of users re-see their environment.","keywords":["Augmented Reality","Storytelling","Vision Language Model","Metaphor","Scene semantics","Spatial narrative","Spatial anchoring","STAM evaluation"],"falsifier":"Run the same pipeline in a room with more than ten visually similar objects and measure coordinate error and story coherence; the paper itself reports that coordinate errors surge by 118% in dense scenes, so a controlled replication would reveal whether the 5.31/7 spatial fit survives realistic clutter, or whether the JSON bridge only works in sparse, curated settings.","tokens_in":13023,"feed_emoji":"📖","tokens_out":8011,"duration_ms":70446,"temperature":0.7,"pith_summary":"This paper tries to establish that augmented-reality stories can be generated from the visual environment itself rather than pasted on top of it, provided the environment is read through a three-layer semantic lens and the story is passed through a structured bridge into the AR runtime. The lens splits object meaning into physical, functional, and metaphorical layers, so a wedding ring next to a bottle of pills can carry loss or crisis instead of registering as two labelled objects. The bridge is a JSON file whose fragments carry topic, core object, interaction mode, symbolic meaning, and trigger conditions, which lets a vision-language model's metaphors become anchored, explorable AR content while keeping spatial coherence. In user tests on a single office scene, 70% of participants reported seeing real-world objects differently, spatial fit scored 5.31/7, and exploration motivation scored 5.19/7, which the paper reads as evidence for a new object-driven narrative paradigm.","feed_headline":"AR stories emerge from a room's objects, not from scripts","feed_subtitle":"A metaphor layer plus a JSON bridge made 70% of users see their real objects differently.","key_machinery":"The load-bearing mechanism is the scenario-metaphor layer paired with a bidirectional JSON interface. The metaphor layer decomposes each object into physical, functional, and metaphorical meanings so the vision-language model can tell a kitchen knife from a hidden bedroom knife and attach different narrative roles to visually similar objects. The JSON interface then structures the model's output as object, mainstory, and fragments, with each fragment specifying the core object, interaction mode, symbolic meaning, narrative content, and trigger condition; this gives the AR runtime deterministic handles for anchoring prefabs and launching branches while preserving the model's metaphorical language. The STAM evaluation framework—scoring Spatial, Temporal, Adaptive, and Metaphorical dimensions—is the measurement instrument that connects these components to the empirical claims.","core_discovery":"At the core of the paper is a scene-driven narrative pipeline: a vision-language model takes spatial images or video as input, identifies the objects that carry the strongest metaphorical charge, and generates both a linear main story and branching fragments around them. Each fragment is written into a structured JSON schema and includes the object name, trigger condition, interactive agent, interaction mode, symbolic meaning, and displayed narrative text, allowing the AR runtime to anchor the content to real tracked objects and to support user-triggered story branches. The authors claim this arrangement resolves the tension between VLM creativity and physical plausibility: it reduced 3D coordinate errors by 57% compared with their baseline pipelines and earned a spatial consistency rating of 5.31/7 from 17 participants. The experiments also show the trade-off: metaphor-rich stories were rated more interesting but less understandable, immersion fell to 2.81/7, and dense scenes with more than ten objects raised coordinate errors by 118%, so the paper's claim is for a working paradigm with known limits, not a complete solution.","pith_inferences":["The same object-as-metaphor bridge could generalize to non-narrative AR content—museum annotations, learning prompts, or therapeutic cues—wherever the symbolism of a physical arrangement carries meaning.","The observed story-order effect suggests a two-pass presentation design—a literal grounded story first, then a metaphor-rich retelling—that the paper did not test directly but which could resolve the comprehensibility trade-off.","Because metaphor appropriateness and layout understanding varied with user expertise in the study, a personalization layer that adapts metaphor abstraction to the user's narrative familiarity is a plausible next step.","A direct cross-cultural test would be to run the Generation Evaluation Phase with participant groups from different cultural backgrounds and compare metaphor appropriateness scores, addressing the paper's own limitation statement."],"forward_implications":["AR storytelling can shift from hand-authored branching scripts to real-time generation from one scene image, since the pipeline turns VLM output into anchorable fragments in a single pass.","Metaphor-grounded narratives increase engagement: 85% of participants scanned more objects than required and motivation scored 5.19/7, so object-driven stories appear to promote exploration.","Users re-see their physical environment: 70% reported object reinterpretation, implying that the narrative changes perception of the real room rather than adding decorations to it.","Focusing on one or two key metaphorical objects per scene, instead of metaphorizing everything, improves comprehension and cuts user disorientation by 42% in the reported experiments.","VLM 3D coordinate inference remains the bottleneck, so practical deployments should keep anchor counts small and supplement the VLM with deterministic localization."],"supporting_citations":[{"why":"Supplies the generative engine: the GPT-4o vision-language model parses scene images, selects metaphorical objects, and writes the narrative fragments.","marker":"[19]"},{"why":"GPT4Scene is the comparison model for understanding 3D scenes from video, used in the STAM capability benchmark.","marker":"[20]"},{"why":"Defines the prior approach that treats spatial sampling and script generation as separate processes, the baseline the new pipeline unifies.","marker":"[13]"},{"why":"The location-aware AR narrative system this work extends by adding metaphorical semantics and a structured VLM-to-AR bridge.","marker":"[14]"},{"why":"Establishes that vision-language models possess fine-grained spatial understanding, the premise on which the method builds.","marker":"[11]"},{"why":"Supplies the spatial-reasoning evaluation style that STAM extends toward state-based and metaphorical reasoning.","marker":"[28]"}],"fun_headline_variants":["AR stories emerge from object metaphors, not scripts","VLM-driven AR: rooms become active storytellers","Object symbolism powers AR narratives, study finds","Scene-driven AR: from objects to branching tales","Metaphor layer cuts AR coordinate errors by 57%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The spatial-coherence claim rests on a single office deployment with 3 key objects, 10 branching items, and 17 participants, so the VLM-plus-JSON pipeline has not been shown to anchor metaphors reliably in dense or varied real environments.","fun_headline_variants_meta":{"raw":{"variants":["AR stories emerge from object metaphors, not scripts","VLM-driven AR: rooms become active storytellers","Object symbolism powers AR narratives, study finds","Scene-driven AR: from objects to branching tales","Metaphor layer cuts AR coordinate errors by 57%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000835,"raw_usage":{"total_tokens":3696,"prompt_tokens":1053,"completion_tokens":2643,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":2569}},"tokens_in":669,"tokens_out":2643,"duration_ms":19746,"temperature":1.0,"reasoning_tokens":2569,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:13:34.172279+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same pipeline in a room with more than ten visually similar objects and measure coordinate error and story coherence; the paper itself reports that coordinate errors surge by 118% in dense scenes, so a controlled replication would reveal whether the 5.31/7 spatial fit survives realistic clutter, or whether the JSON bridge only works in sparse, curated settings.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the generative engine: the GPT-4o vision-language model parses scene images, selects metaphorical objects, and writes the narrative fragments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that vision-language models possess fine-grained spatial understanding, the premise on which the method builds."}],"review_version":1}