{"id":"62788e15-f164-44fa-b81e-289aecc9641b","arxiv_id":"2411.17541","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A qualitative study of 29 failed AR/VR startups produces the Metaverse Innovation Canvas, but its evaluation does not demonstrate that the canvas improves startup viability.","lead":"This paper analyzes 29 failed AR/VR startups, identifies common failure factors, and proposes the Metaverse Innovation Canvas, a business planning tool tailored to XR ventures. It is of interest because it offers a practical framework for founders, though its effectiveness claim rests on a small expert evaluation of the same cases used to design the tool.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central viability claim rests on a circular, uncontrolled evaluation: the same failed startups used to design the MIC were used to test it, with no Lean Canvas baseline and only expert opinion as outcome.","rationale":"The reader's weakest assumption is exactly the vulnerable point: the evaluation design cannot support the causal claim. I do not see an internal inconsistency in the MIC's block logic, and the failure-factor list is plausible as a design artifact. However, the claim of 'enhancing viability' is a causal claim about founders' behavior and startup outcomes, and Section 5 provides no counterfactual. The sample size of three consultants is a secondary issue; even fifty consultants would not fix the absence of a baseline and hold-out cases. The manuscript's own limitations section admits the qualitative-only evaluation. Therefore the reader's REJECT verdict is appropriate, though confidence is low because the artifact is a plausible design suggestion. My read does not change the verdict: UNCHANGED.","tokens_in":11835,"tokens_out":3738,"duration_ms":35131,"concrete_test":"Conduct a preregistered controlled evaluation with held-out cases: assemble 10 XR startup descriptions (5 failed, 5 successful) not among the 29 used in Stage 1; randomly assign consultants or founders to fill either the MIC or the Lean Canvas for each case; have blind assessors count the number of pre-registered risk categories (XR value proposition, motion load, scalability, problem definition) surfaced and rate the actionability of proposed mitigations. If MIC does not outperform Lean Canvas on these pre-registered measures, the 'effectiveness in surfacing issues' claim is not supported; a prospective cohort tracking actual venture pivots or survival would be needed for the 'enhancing viability' claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that the MIC 'enhances viability' and is 'effective in surfacing overlooked usability issues and technology constraints' depends on Section 5. That section reports three startup consultants completing the MIC for five failed AR/VR startups drawn from the same 29-case failure set whose analysis directly informed the canvas design in Sections 2.1 through 2.3. Because the canvas blocks are essentially prompts built from those failure factors, consultants applying the canvas to those same cases are led to rediscover the categories that generated the tool. There is no comparison with the Lean Canvas, no hold-out set of unseen startups, no successful-startup control, and no longitudinal or quantitative outcome measure linking canvas use to venture outcomes. The consultants' open-ended comments about usability blocks being 'crucial' are opinions about the tool, not measurements of its effect. The paper's own Limitations section concedes the sole reliance on qualitative consultant feedback. The most load-bearing flaw is therefore not the small sample size by itself but the absence of any counterfactual: the study cannot distinguish 'the MIC surfaces risks' from 'experts can label obvious problems in dead startups when prompted by a checklist derived from those very problems.' Without a baseline, the causal claim in the abstract is unsubstantiated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 29 failed AR/VR startups from 2016-2022, identifies five failure-factor categories, and uses these factors to design a one-page business-modeling tool, the Metaverse Innovation Canvas (MIC). The canvas adds XR-specific blocks for problem framing, XR-exclusive value propositions, motion-based interaction load, usability scenarios, social/virtual-economy opportunities, and future scalability, on top of Lean Canvas viability elements. The evaluation consists of three startup consultants completing the MIC for five failed AR/VR startups drawn from the same dataset, followed by semi-structured interviews. Based on this feedback, the abstract and conclusion claim that the MIC is 'effective in surfacing overlooked usability issues and technology constraints upfront, enhancing the viability of future metaverse startups.'","tokens_in":12020,"tokens_out":5372,"duration_ms":52862,"significance":"The failure-factor corpus assembled in Section 3 is a potentially useful descriptive baseline for an understudied domain, and the MIC is a well-motivated design artifact: its blocks respond directly to documented failure modes and to known limitations of generalized canvas tools. The paper is also honest in Section 6 about relying on qualitative consultant feedback. However, the empirical claim that the canvas 'enhances viability' is not supported by the presented evaluation, which is circular, uncontrolled, and lacks any outcome measure. With the causal claims tempered and the evaluation reframed as a formative expert feedback study, the artifact and failure analysis could make a modest contribution to XR entrepreneurship tooling. The paper does not include reproducible code or machine-checked proofs, but the canvas itself is presented in sufficient detail to be independently applied and tested.","major_comments":[{"comment":"The evaluation in Section 5 cannot support the abstract's claim that the MIC 'enhanc[es] the viability of future metaverse startups.' The five cases in Table 1 are drawn from the same 29-case failure set whose analysis in Section 3 directly generated the canvas blocks described in Section 4. Since the consultants completed the MIC for those same failure cases, their identification of missing XR value propositions, usability concerns, and scalability issues is at least partly an artifact of the canvas prompting them to look for exactly those failure categories. There is no control tool (e.g., the Lean Canvas), no hold-out set of unseen startups, and no before/after measure. The authors should either add an independent evaluation, such as a blinded within-subjects comparison on unseen cases, or remove the causal viability wording and present Section 5 as a formative usability test of the artifact.","section":"Section 5; Abstract"},{"comment":"The evidence reported in Section 5 is expert opinion, not a measurement of viability. The results consist of open-ended consultant statements about the importance of value propositions, usability blocks, and scalability, with no quantitative score, no pre/post comparison, no longitudinal follow-up, and no objective indicator linking canvas completion to venture outcomes. The authors themselves acknowledge in Section 6 that the study relies 'solely' on qualitative data from startup consultants. The conclusion that the canvas 'enhanc[es] viability' therefore exceeds what the data can establish; the conclusion should be limited to a claim that three experts found the canvas useful for structuring problem analysis.","section":"Section 5; Table 1; Section 6"},{"comment":"The coding procedure behind the failure-factor counts (e.g., '22 of the 29 startups struggled with scalability issues,' '18 out of the 29 failed startups') is underdocumented. The manuscript does not report the number of coders, the codebook, inter-coder agreement, or an audit trail, so the categories in Section 3 cannot be independently verified or reproduced. Because these factors are the empirical foundation for every block in the MIC, the authors should provide the coding scheme and a trace of how each startup was coded to a failure factor, either in the main text or as an appendix.","section":"Section 2.1; Section 3"}],"minor_comments":[{"comment":"The number of focus group sessions is inconsistent: Section 2 says 'Four focus group sessions were conducted,' while Section 2.1 says 'There were five group sessions in total.'","section":"Section 2; Section 2.1"},{"comment":"The citation for the MIC figure is unresolved: 'Figure[?]' appears in the sentence introducing the canvas, and no figure is actually included in the manuscript.","section":"Section 4.1"},{"comment":"The recruitment description is confusing: the paragraph says startups were selected by leveraging the authors' extensive network, while the immediately preceding sentences say participants were recruited through LinkedIn and 'not from the first connections of the authors.' Please clarify the distinction between startup selection and participant recruitment.","section":"Section 2.1"},{"comment":"There are several typographical and grammatical errors, including 'The entreprenurs have demonstratedtwo common mistakes in desiging startups,' 'Bringing an application to XRis avalue proposition. This iswrong,' and 'highlighted po tential pitfalls.' A careful proofreading pass is needed.","section":"General"},{"comment":"The 'Measurable Metrics' subsection introduces usability-related metrics (expected usable session, daily engagement, revenue per user), but these metrics are never used in the evaluation in Section 5. Either connect them to the evaluation or state explicitly that they are a design proposal for future testing.","section":"Section 4.7"}],"recommendation":"major_revision","confidential_remarks":"I considered rejection because the causal claim in the abstract is not supported by the circular evaluation. I settled on major revision because the core artifact and failure-factor analysis are salvageable: the authors can reframe the contribution as a design proposal with formative expert feedback, add a clear statement that viability effects are not measured, and document the coding process. If the authors decline to temper the causal claims, I would recommend rejection on the grounds that the current evidence cannot validate the stated central contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plain English. Candid. The paper's real contribution is the MIC itself: a Lean Canvas variant with blocks that force XR-specific thinking—motion-based interaction load, XR-unique value propositions, AR/VR UX opportunities, and future scalability constraints. Those blocks are sensibly derived from the 29-case failure analysis, and the failure factors (scalability, usability, value proposition, interaction load, problem definition) are consistent with the ergonomics and entrepreneurship literature. Credit where due: the authors are transparent about their autoethnographic method, and the canvas prompts address genuine gaps in general-purpose business modeling tools for XR. The 29-startup dataset is new, albeit small and network-sourced, and the coding process is not documented.\n\nThe soft spot is the evaluation in Section 5. The consultants were asked to fill out the canvas for five failed startups drawn from the same 29 used to derive the canvas blocks. Since the canvas is essentially a checklist built from those failure factors, the consultants' positive feedback mostly confirms that the checklist can label obvious problems in already-failed cases. There is no baseline (no Lean Canvas comparison), no hold-out set, and no longitudinal or quantitative outcome measure. The paper's own Limitations section concedes the sole reliance on qualitative consultant feedback. So the abstract's claim that the MIC 'enhances the viability of future metaverse startups' is unsubstantiated. That said, the circularity is a weakness in the evidence, not a fatal flaw in the artifact itself. The canvas is a plausible design tool; the authors just can't claim efficacy from this evaluation.\n\nFor a reader, this is a useful practical proposal for startup incubators and XR design educators. The failure-factor analysis is a reasonable starting point for future work, even if it doesn't break new scientific ground. I'd send it to peer review because the artifact is concrete, the domain is underserved, and the authors have been honest about limitations. The referee should push for a revised evaluation (e.g., blind comparison with Lean Canvas on unseen cases, or at least a clearly framed feasibility study rather than efficacy claim). Serious thinker: yes—the thinking is coherent and the limitations are acknowledged, even if the central claim overreaches.","headline":"A reasonable XR-specific canvas built on a modest failure analysis, but the evaluation can't support the abstract's causal claim—still worth a referee's time.","tokens_in":614,"tokens_out":1510,"would_cite":false,"duration_ms":22267,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Metaverse Innovation Canvas is a one-page business tool built from the failures of 29 AR/VR startups to make founders confront usability, XR-only value, and scalability early.","keywords":["Lean Canvas","Mixed Reality","Augmented Reality","Virtual Reality","Innovation","Entrepreneurship","Metaverse","Startup failure"],"falsifier":"Take a set of XR startup ideas, randomly assign people to ideate with the MIC or the Lean Canvas, and have evaluators blind to the tool count the usability and technology constraints each person surfaces; if the MIC group surfaces no more, the central claim fails. A stronger version would follow two cohorts of real founders, MIC users and non-users, and compare survival or pivot rates after a fixed period.","tokens_in":11600,"feed_emoji":"🥽","tokens_out":7831,"duration_ms":65398,"temperature":0.7,"pith_summary":"This paper sets out to explain why AR/VR startups fail and to give founders a practical tool that prevents the most common mistakes. By analyzing 29 failed ventures from 2016 to 2022, the authors identified recurring pitfalls: short-term thinking, scalability blocked by hardware limits, poor usability, unjustified motion-based interaction load, weak value propositions next to non-XR alternatives, and no clear user problem. From those findings they built the Metaverse Innovation Canvas (MIC), a one-page ideation template whose colored blocks force founders to answer XR-specific questions, from 'what is the user actually trying to do?' to 'does this interaction's physical burden earn its value?' Three startup consultants who applied the canvas to five failed ventures said it surfaced usability issues and technology constraints they would otherwise have missed.","feed_headline":"Map why AR/VR startups fail on one page","feed_subtitle":"Built from 29 dead ventures, the Metaverse Innovation Canvas forces founders to confront usability and scalability early.","key_machinery":"The Metaverse Innovation Canvas (MIC) is a one-page business ideation template with five color-coded groups of blocks — problem (red), solution (green), usability (blue), viability (yellow), and future scalability (purple). Its load-bearing mechanism is a set of targeted prompts that translate each identified failure factor into a question a founder must answer, including separate blocks for existing XR, non-XR, and non-digital alternatives, the XR unique value proposition, motion-based interaction load, AR/VR UX opportunities, social and virtual economy opportunities, interoperability features, a minimum viable experience scenario, and explicit scalability limitations and future threats. The same prompts are supported by measurable metrics (expected usable session, daily engagement, revenue per user) that turn usability from a vague concern into a planning quantity.","core_discovery":"The central claim is that the Metaverse Innovation Canvas is an effective ideation tool for extended reality ventures because it embeds XR-specific failure factors directly into business-model thinking. The canvas was derived from an analysis of 29 failed AR/VR startups, which yielded six failure factors: lack of long-term planning, scalability blocked by hardware limitations, poor usability, unjustified motion-based interaction load, unclear value propositions relative to non-XR alternatives, and failure to address a clear problem. The MIC responds with specialized blocks: separate alternatives for XR, non-XR, and non-digital solutions; an 'XR unique value proposition' block; usability prompts including motion-based interaction load, AR and VR UX opportunities, social and virtual economy opportunities, interoperability, and a minimum viable experience scenario; plus future scalability and measurable metrics. The paper's evaluation — three startup consultants completing the canvas for five of the failed startups — found that the tool helped surface overlooked usability issues and technology constraints, and the authors take this as evidence that the canvas enhances the viability of future metaverse startups.","pith_inferences":["In our reading, the strongest untested promise is comparative: a head-to-head study of MIC versus Lean Canvas on the same XR ideas, with blinded counts of surfaced constraints, would tell whether the specialized blocks do more than generic prompts.","The 'separate alternatives' and 'motion-based interaction load' ideas could plausibly transfer to other physically constrained technologies, such as brain-computer interfaces or wearable robotics, where interaction cost also needs to be justified.","Because the expert evaluation used the same startups that inspired the canvas, the tool's real-world impact on survival remains unmeasured; a longitudinal deployment study would be the natural next step.","The measurable metrics (session length, daily engagement, revenue per user) could eventually be validated against telemetry from operating XR products, turning the canvas from a heuristic into a forecast tool."],"forward_implications":["Founders who work through the MIC will name the user's problem and the non-XR alternatives before investing in development, which should cut wasted engineering.","Startup accelerators and investors can use the filled canvas as a structured artifact to compare XR ventures on usability and scalability thinking, not just pitch deck language.","The six failure factors give XR entrepreneurship researchers a typology for coding and comparing new cases.","The usability blocks connect UX design work to business planning, so usability issues become visible before a minimum viable product is built."],"supporting_citations":[{"why":"The Business Model Canvas, the base framework the MIC extends and critiques.","marker":"[20]"},{"why":"The Lean Startup, source of the Lean Canvas the paper adapts.","marker":"[22]"},{"why":"Running Lean, which introduced the Lean Canvas one-page business plan.","marker":"[13]"},{"why":"Narrative review of VR ergonomics and risks that documents the usability burdens the MIC addresses.","marker":"[27]"},{"why":"The Extended by Design toolkit, an XR design resource the MIC builds beyond.","marker":"[8]"},{"why":"The Blitz Canvas, a specialized business model canvas for software startups that MIC positions itself against.","marker":"[25]"},{"why":"Work on monopoly and survivorship bias that motivates the study of failed startups rather than success stories.","marker":"[28]"},{"why":"Customer-engagement strategy framework that supplies the user engagement goal checkboxes.","marker":"[10]"}],"fun_headline_variants":["Why AR/VR startups fail, mapped on one canvas","A canvas built from 29 failed startups to fix XR","New tool turns startup failures into XR success guide","Metaverse Innovation Canvas: learn from dead ventures","One-page canvas to avoid AR/VR startup pitfalls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument leans on the feedback of three startup consultants who filled out the canvas for the same five failed startups used to design it, with no control group and no evidence that canvas use changes what happens to a real venture.","fun_headline_variants_meta":{"raw":{"variants":["Why AR/VR startups fail, mapped on one canvas","A canvas built from 29 failed startups to fix XR","New tool turns startup failures into XR success guide","Metaverse Innovation Canvas: learn from dead ventures","One-page canvas to avoid AR/VR startup pitfalls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000405,"raw_usage":{"total_tokens":2101,"prompt_tokens":934,"completion_tokens":1167,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":1089}},"tokens_in":550,"tokens_out":1167,"duration_ms":9514,"temperature":1.0,"reasoning_tokens":1089,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:59:04.035660+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of XR startup ideas, randomly assign people to ideate with the MIC or the Lean Canvas, and have evaluators blind to the tool count the usability and technology constraints each person surfaces; if the MIC group surfaces no more, the central claim fails. A stronger version would follow two cohorts of real founders, MIC users and non-users, and compare survival or pivot rates after a fixed period.","supporting_citations":[],"review_version":1}