{"id":"0a8a590e-851c-4a8f-b431-f95e868ad3cf","arxiv_id":"2508.13074","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A mixed-reality museum-fire game makes players choose between a cat and a cultural artifact, and its preliminary evaluation reportedly deepens empathy and moral reflection.","lead":"This paper presents Ashes or Breath, a mixed-reality game in which players must choose between saving a living cat or a priceless artifact in a simulated museum fire. The authors report that this immersive format intensifies empathy and moral reflection, but only the abstract could be verified for this review.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central empirical claim is unverifiable from the supplied source: the abstract reports no baseline, sample, or instruments, and the full evaluation text is corrupted, so the causal attribution to MR-HMDs cannot be assessed.","rationale":"The reader's weakest_assumption is exactly the load-bearing concern: the empirical claim requires a non-immersive control to attribute effects to MR-HMDs, and the abstract provides no evidence of such a control. The full text being corrupted means the concern cannot be resolved without additional source access. I agree with the reader's verdict of UNVERDICTED rather than ACCEPT or REJECT, because the visible evidence floor is too low to assess correctness. My stress-test does not change that verdict; it reinforces it by articulating why the missing control is not a minor methodological preference but a logical requirement for the causal claim. The only route to a stronger verdict is to inspect the full evaluation, which is why the concrete test is to obtain a readable copy and check for the presence of a non-immersive comparison condition.","tokens_in":9143,"tokens_out":1176,"duration_ms":14533,"concrete_test":"Obtain a readable copy of the full text (e.g., the arXiv source PDF rather than the corrupted extraction) and locate the evaluation section. Check whether the study includes a control group or within-subjects condition where participants experience the identical museum-fire dilemma in a non-immersive format (e.g., text or 2D video). If no such comparison exists, the causal claim about MR-HMDs should be weakened to a descriptive demonstration. If a control exists, verify that the reported differences on empathy/introspection measures are statistically tested and that effect sizes and instruments are reported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline contribution is the final abstract sentence: 'Preliminary evaluations suggest that embedding moral dilemmas into everyday environments via MR-HMDs intensifies empathy, deepens introspection, and encourages users to reconsider their moral assumptions.' This is a causal claim about the MR-HMD medium. For it to hold, the evaluation must isolate the medium from the dilemma content. That requires a control condition presenting the same dilemma non-immersively (e.g., text, video, or desktop VR), validated measures of empathy and introspection, and a sample with sufficient statistical power. The abstract provides none of these. The supplied full text is mojibake (e.g., '����� �� ������'), making the evaluation section unreadable, so it is impossible to confirm whether such a control exists. If no such baseline exists, the claim reduces to 'a dramatic moral dilemma evokes emotional responses,' which is neither new nor specific to MR-HMDs. The paper's motivation explicitly critiques 'abstract, disembodied scenarios,' yet it never states in the abstract that it compared against one. This is not an accusation of misconduct; it is an honest boundary condition: the central claim's evidentiary base is currently inaccessible and, at the abstract level, incomplete.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'Ashes or Breath,' a Mixed Reality head-mounted-display (MR-HMD) game that places players in a museum-fire moral dilemma: save a living cat or a priceless cultural artifact. The experience is designed through an iterative, values-centered process and includes embodied interaction, spatial immersion, irreversible choices, narrative consequences, and a reflective room. The abstract's central empirical claim is that preliminary evaluations suggest this MR-HMD embedding 'intensifies empathy, deepens introspection, and encourages users to reconsider their moral assumptions.' The supplied full text is corrupted (mojibake) and unreadable, so the evaluation design, data, and analysis cannot be inspected. The review is therefore based primarily on the abstract and what can be inferred from the readable fragments.","tokens_in":9333,"tokens_out":2736,"duration_ms":35886,"significance":"If the central claim were adequately supported, the paper would make a useful contribution to ethics-based experiential learning in HCI, showing that MR-HMDs can function not just as interaction media but as 'stages for ethical encounter.' The design concept is creative and addresses a real gap in abstract, text-based moral-dilemma teaching. However, the manuscript currently does not provide the evidence needed to support the modality-specific causal claim. The paper's value is therefore contingent on a substantially more detailed and rigorous empirical report. Strengths visible at this stage include the concrete artifact, the integration of spatial and embodied interaction with narrative consequences, and a design process grounded in values. These are promising, but the current submission does not yet establish the claimed outcomes.","major_comments":[{"comment":"The abstract claims that embedding moral dilemmas into everyday environments via MR-HMDs 'intensifies empathy, deepens introspection, and encourages users to reconsider their moral assumptions.' This is a causal claim about the MR-HMD medium. No sample size, comparison condition, measurement instruments, statistical results, or qualitative analysis are reported in the abstract, and the full-text evaluation section is unreadable in the supplied source due to character corruption. If the evaluation did not include a non-immersive control condition presenting the same dilemma (e.g., text, video, or desktop VR), then the claim conflates the effect of the dilemma content with the effect of the MR-HMD medium. The authors must either provide a readable evaluation section with a clear baseline/comparison design or substantially weaken the claim to describe user experience without medium-specific","section":"Abstract, final sentence"},{"comment":"Because the full text is mojibake, I cannot verify whether the evaluation included any validated measures of empathy, introspection, or moral reflection. If the 'preliminary evaluations' are anecdotal or based on in-house questionnaires without established psychometric properties, the central claim remains unsupported. At minimum, the revised manuscript must report participant demographics, sample size, experimental design, procedure, instruments, and results (including effect sizes and confidence intervals where appropriate). Without these details, the phrase 'preliminary evaluations suggest' does not meet the evidentiary standard for a journal publication.","section":"Evaluation section (unreadable in supplied full text)"},{"comment":"The game was designed through an iterative, values-centered process and then evaluated by the same team for the value-targeted outcomes. This creates a structural risk of confirmatory bias. The manuscript should state whether the evaluation used pre-registered hypotheses, independent raters, blinded analysis, or other safeguards. If no such safeguards exist, the evaluation should be presented as a formative usability study and the abstract's causal-sounding conclusion should be revised accordingly.","section":"Values-centered design and evaluation"},{"comment":"If the evaluation used only a small convenience sample (as 'preliminary' suggests), the claim that users 'reconsider their moral assumptions' is not generalizable. Moreover, participants in an MR moral-dilemma study may give socially desirable responses, especially when the game explicitly thematizes empathy and reflection. The manuscript needs to address demand characteristics, for example through behavioral measures, open-ended reflections with independent coding, or a between-subjects design with a comparison condition. This is not a demand for a specific methodology, but the current report gives no indication that any of these concerns were addressed.","section":"Generalizability and demand characteristics"}],"minor_comments":[{"comment":"The phrase 'everyday environments' is odd for a museum fire, which is not an everyday setting. Consider 'real-world relevant environments' or similar.","section":"Abstract"},{"comment":"The term MR-HMD is used without expansion; spell out 'Mixed Reality head-mounted display' at first occurrence.","section":"Abstract"},{"comment":"The choices are described as 'irreversible,' but in a game context they are not ethically irreversible. Clarify that irreversibility is in-game only.","section":"Abstract"},{"comment":"The supplied full text is not readable due to character-encoding corruption. The authors should ensure that the submitted PDF/source renders correctly; this is a precondition for review.","section":"Full text"}],"recommendation":"major_revision","confidential_remarks":"The apparent encoding corruption is likely an artifact of the arXiv source, but as submitted the paper cannot be properly reviewed. I am recommending major revision rather than reject because the central problem is verifiability: the authors can fix it by providing a readable full text and a detailed evaluation section with sample, measures, comparison condition, and results. If, after reading the actual evaluation, it turns out there is no baseline condition at all, the paper may need to be reframed as a design case study rather than a claim about MR-HMD effects. I would also flag that the fit of this paper depends on the journal's interest in design contributions versus empirical validation; without the missing evidence, it reads as a work-in-progress report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The full text we were given is mojibake, so the only assessable content is the abstract. On that evidence, this is a design contribution with an unverified empirical claim, and the review has to stop there.\n\nWhat is actually new: the specific artifact—a mixed-reality HMD museum-fire dilemma, save a cat or a priceless artifact, with an irreversible embodied choice and a reflective room—is a real design. The values-centered iterative process and the setting (everyday environment rather than lab) are sensible for experiential ethics education. That part reads well and is worth a look for anyone designing moral-dilemma experiences.\n\nSoft spots: the abstract's final sentence claims that embedding moral dilemmas via MR-HMDs 'intensifies empathy, deepens introspection, and encourages users to reconsider their moral assumptions.' That is a causal claim about the medium, and the abstract gives no sample, no comparison condition, no measures, no results. We literally cannot tell whether they compared against a non-immersive version of the same dilemma, or whether they just showed that a dramatic dilemma is engaging. The evaluation section is unreadable in the supplied source, so the central claim is currently unverifiable. This is not an accusation; it is the boundary of what can be fairly inferred. If the full manuscript has a proper controlled evaluation, this is a solid HCI contribution. If not, the claim collapses to something like 'scenarios with vivid content evoke emotion,' which is not new.\n\nThe paper deserves a serious referee if the real full text contains a real evaluation; the design and the question justify referee time. A desk reject would be wrong for a novel artifact. But if you are thinking of citing the empirical claim, wait until you can see the data.\n\nFor peer review: send it out, and ask the referee to check the evaluation design and specifically whether the claimed modality effect is actually isolated from the dilemma content.","headline":"A novel MR ethics-dilemma game with an unverifiable central claim — the design deserves review, but the empirical evidence needs proof.","tokens_in":9856,"tokens_out":1923,"would_cite":false,"duration_ms":22631,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that staging a moral dilemma in a physical space via a head-mounted mixed-reality display—choosing between a cat and a cultural artifact in a burning museum—intensifies empathy, deepens introspection, and makes users recon","keywords":["mixed reality","moral dilemmas","empathy","introspection","ethics education","head-mounted display","embodied interaction","cultural heritage"],"falsifier":"Run the same museum-fire dilemma and reflective aftermath with one group experiencing the MR-HMD version and another group experiencing the identical dilemma as text, video, or desktop simulation. If validated measures of empathy, introspection, and moral-reasoning change show roughly equal gains in both groups, the specific claim that MR-HMD embedding drives the effect collapses.","tokens_in":8977,"feed_emoji":"🕹️","tokens_out":4061,"duration_ms":48070,"temperature":0.7,"pith_summary":"This paper develops and tests a mixed-reality game, Ashes or Breath, in which players wearing a head-mounted display must choose whether to save a living cat or a priceless cultural artifact from a burning museum. The aim is to show that putting moral dilemmas directly into physical, everyday spaces through mixed reality makes ethical conflict feel immediate and personally consequential. The paper argues that embodied interaction, spatial immersion, an irreversible choice, and a reflective aftermath deepen empathy and introspection more than traditional abstract, text-based scenarios. Preliminary evaluations are reported as supporting this view, positioning augmented reality as a stage for ethical encounter rather than merely an interaction medium.","feed_headline":"Mixed reality makes moral dilemmas more visceral, early tests suggest","feed_subtitle":"A headset game stages an irreversible choice between a cat and an artifact to deepen empathy.","key_machinery":"The key mechanism is the MR-HMD staging itself: embodied interaction and spatial immersion anchor an abstract moral dilemma in a concrete, felt, everyday environment. The design adds three reinforcing elements: an irreversible choice that cannot be undone, emotionally charged in-game consequences, and a reflective room afterwards that invites players to examine their decision from multiple perspectives. Together these turn the moral dilemma from a proposition to be weighed into an experience to be undergone.","core_discovery":"The central claim is that the modality of presentation is part of the moral lesson: a dilemma encountered spatially, with the body, inside a real environment engages ethical reasoning differently from one read as a thought experiment. Ashes or Breath operationalizes this by staging an irreversible choice between a cat and a cultural artifact during a museum fire, followed by narrative consequences in a reflective room. Players are encouraged to confront both the immediate emotional cost and the societal value of cultural legacy. The paper reports that preliminary evaluations with this mixed-reality headset experience suggest intensified empathy, deeper introspection, and a reconsideration of","pith_inferences":["The strongest unstated prediction is that the same museum-fire dilemma presented as text, video, or desktop simulation would produce measurably weaker empathy and reflection than the MR-HMD version; this comparison is directly testable.","If spatial presence is the active ingredient, the effect should grow with the physical fidelity of the environment: walking through a real museum gallery should matter more than a headset-only virtual room.","The cat-versus-artifact conflict is a reusable template for other value clashes—personal safety, cultural memory, animal life, institutional duty—suggesting a family of situated moral-dilemma experiences.","Because the choices are irreversible and emotionally charged, designers may need to build in psychological safeguards and debriefing procedures; that ethical-care implication is left implicit in the paper."],"forward_implications":["If the effect holds, ethics instruction can be redesigned around felt, situated choices rather than abstract thought experiments.","The MR-HMD format itself, not only the dilemma content, becomes the object of evaluation for empathy and moral reflection.","The irreversible-decision-plus-reflection-room structure provides a reusable design pattern for other moral-education experiences.","The paper positions augmented reality as an ethical stage, so HCI evaluation criteria may need to include moral outcomes such as empathy and introspection alongside usability measures."],"supporting_citations":[],"fun_headline_variants":["MR game forces a cat-or-artifact choice, boosts empathy","Headset game: save a cat or an artifact, feel the burn","Moral dilemmas get visceral in mixed reality test","Mixed reality turns ethics into a body-level choice"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the reported gains in empathy, introspection, and moral reconsideration come specifically from placing the dilemma in a physical space via the MR headset, which requires comparing it against the same dilemma presented non-immersively.","fun_headline_variants_meta":{"raw":{"variants":["MR game forces a cat-or-artifact choice, boosts empathy","Headset game: save a cat or an artifact, feel the burn","Moral dilemmas get visceral in mixed reality test","Mixed reality turns ethics into a body-level choice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00053,"raw_usage":{"total_tokens":2350,"prompt_tokens":664,"completion_tokens":1686,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":1619}},"tokens_in":408,"tokens_out":1686,"duration_ms":12750,"temperature":1.0,"reasoning_tokens":1619,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:09:35.916272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same museum-fire dilemma and reflective aftermath with one group experiencing the MR-HMD version and another group experiencing the identical dilemma as text, video, or desktop simulation. If validated measures of empathy, introspection, and moral-reasoning change show roughly equal gains in both groups, the specific claim that MR-HMD embedding drives the effect collapses.","supporting_citations":[],"review_version":1}