{"id":"261bd004-b707-48f2-9e06-7f5b1d119630","arxiv_id":"2504.18410","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-phase VR scenario with LLM role-play can prompt reflection and mixed, history-dependent emotions in adults with a history of parental verbal abuse, based on interviews with 12 participants.","lead":"This paper built a virtual reality experience in which people who lived through parental verbal abuse first play the hurtful parent, then watch an AI mother rephrase their abusive lines warmly. In interviews, 12 participants said the role switch prompted reflection on their past, though reactions varied with personal history. It is a small prototype study, not a clinical trial.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No control or neutral-prime comparison, plus a TV-drama memory prime immediately before the VR session and VR-directed interview prompts, leaves the causal attribution of reported reflections to the VR-LLM system untested.","rationale":"I read the paper in good faith. It is an honest qualitative report of a novel VR-LLM prototype and a 12-participant interview study. The descriptive results are plausible and the Discussion and Limitation sections are appropriately hedged. However, the abstract and conclusion make a causal-sounding claim: the experience 'encourages reflection' and 'fosters supportive emotions.' The weakest point in the evidentiary chain is not the coding scheme or the sample size; it is the provenance of the reported recollections. The procedure deliberately primes autobiographical recall with a TV-drama clip, then conducts semi-structured interviews that explicitly invite participants to connect the VR content to their past experiences. Without a control or neutral-prime condition, or without pre-/post-measures of reflection, the study cannot separate the VR system's contribution from the priming clip, the interview prompts, or demand characteristics. The reader's weakest_assumption identifies exactly this issue, and I agree with it. My concrete test targets the single most diagnostic check: locating the first unprompted memory mention in the interview timeline. If the paper's authors can show that participants spontaneously referenced specific past events during Scene 1, before any interviewer prompt, the causal claim would be considerably strengthened. If not, the paper should be reworded to make clear that it reports participant perceptions of a primed experience, not evidence that the VR-LLM system itself elicited the reflection. This does not change the CONDITIONAL verdict; it supports it. A REJECT would be too harsh because the prototype and descriptive findings have value, and the requested controls are feasible in follow-up work.","tokens_in":10728,"tokens_out":3105,"duration_ms":34539,"concrete_test":"Obtain the interview transcripts and code, for each participant, the temporal location of the first unprompted mention of a specific past parental-abuse memory relative to (a) the TV drama clip, (b) the first VR scene, and (c) the interviewer's first VR-related question. If the first specific memories appear only after the clip or only after the interviewer asks about the VR experience, the attribution to the VR system is unsupported. A stronger confirmatory version: run a separate arm with 12 new participants using the same VR session but a neutral or no priming clip and an interview that begins with an open-ended 'tell me about your experience' question before any VR-specific prompts; if recollection rates remain comparable, the prime is not the driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the dual-phase VR-LLM experience 'encourages reflection on their past experiences and fosters supportive emotions.' The strongest defensible reading is descriptive: participants reported these reflections in post-hoc interviews. The causal attribution to the VR system, however, is load-bearing and untested. In Methods, Procedure, the authors state: 'We played a segment from a TV drama depicting a mother verbally abusing her daughter, sourced from popular clips on the widely known Chinese social media platform Xiaohongshu to help participants recall and bring into the scenes of familial verbal abuse.' This prime occurred immediately before the VR session. The interview protocol then explicitly asked participants how their role-playing 'related to their past experiences' and whether the LLM responses 'evoked memories of their past selves.' With no control condition, no neutral-prime comparison, and no pre/post reflection measure, the reported recollections could originate from the TV clip, from the interview prompts, or from demand characteristics, rather than from the designed VR-LLM interaction. The qualitative finding that 11 of 12 participants said the first scene reproduced home situations is consistent with this confound, because participants were recruited for having experienced parental verbal abuse, were shown an abuse scene, and were then asked about their past. This is not an accusation of misreporting; it is a correctness risk in the causal claim. The descriptive finding stands, but the paper's stronger wording in the abstract and conclusion goes beyond what the procedure can support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes a dual-phase VR-LLM experience designed to prompt reflection on parental verbal abuse. Participants with self-reported histories of such abuse first role-play an abusive parent speaking scripted hurtful lines to an LLM-driven child, then observe an LLM-driven mother rephrase the same lines into warm, supportive language. The authors report a qualitative study with 12 Chinese adults (18–34) using semi-structured interviews and thematic analysis. The findings are organized into three themes: the first scene evoked recollection of past events and reflection; the second scene fostered supportive emotions; participants viewed LLMs as promising but in need of personalization. The abstract concludes that the experience 'encourages reflection on their past experiences and fosters supportive emotions,' while the Discussion frames this as a contribution to AI-driven emotional support design.","tokens_in":10918,"tokens_out":3939,"duration_ms":37876,"significance":"If taken as a descriptive qualitative study, the paper offers a clearly described dual-phase role-reversal design, a detailed system implementation (Unity, Oculus Quest 2, GPT-4, GPT-SoVITS, Xunfei API), and a rich set of participant quotes that ground the reported themes. The design rationale draws on psychodrama theory, and the Limitations section is unusually candid about demographic and ethical risks. The main value is as a formative design exploration: it shows that a VR-LLM role-switching experience can elicit vivid recollections and mixed emotions in this population, and it identifies personalization as a key challenge. However, the paper's central causal claim that the experience 'prompts' reflection and supportive feelings is not supported by the present study design, which lacks controls and includes a memory-priming stimulus. The strongest defensible claim is descriptive: participants reported these effects in post-hoc interviews.","major_comments":[{"comment":"The causal attribution in the abstract and findings is not supported by the design. The procedure states that 'We played a segment from a TV drama depicting a mother verbally abusing her daughter, sourced from popular clips on the widely known Chinese social media platform Xiaohongshu to help participants recall and bring into the scenes of familial verbal abuse' immediately before the VR session. With no control condition, no neutral-prime comparison, and no pre/post measures, participants' recollections and reported reflections could be elicited by the priming clip, by the interview questions (which explicitly asked how the role-play related to past experiences and whether LLM responses evoked memories), or by demand characteristics. Because the abstract claims the experience 'encourages reflection' and the findings section states it 'prompts reflective and supportive feelings,' this confound is load-bearing. I recommend reframing the conclusions as descriptive self-report findings, or adding a baseline/control arm and pre-post reflection measures to support causal claims.","section":"Methods, Procedure"},{"comment":"The thematic categories ('recollection of past events,' 'emotional impact,' 'LLM response effectiveness,' 'environmental design influence') closely mirror the interview-guide items (i)-(x) listed in Methods, Procedure. The Analysis section describes open coding and thematic analysis but provides no coding procedure, codebook, inter-rater reliability, or member-checking information. As a result, the finding that 'All participants apart from P4 stated that in the first scene, they reproduced some or all of the scenes that occurred at home' may largely reflect the interview prompts rather than an emergent property of the experience. Please report the coding process in more detail and discuss how the themes were distinguished from the interview structure.","section":"Analysis"},{"comment":"The counts of participants associated with each claim are sometimes ambiguous or inconsistent. For example, the paper states 'Most participants perceived that in Scene 2, the LLM had effectively translated... (P1, P2, P3, P4, P6, P7, P8, P9, P10, P11, P12)', which is 11 of 12, while 'this rephrased language felt supportive... (P1, P2, P3, P5, P6, P9, P10, P11)' is only 8. Some of these parenthetical lists appear to include participants who made related but distinct comments, and no participant-level matrix is provided. Please clarify whether these counts represent explicit endorsements of the stated claim and consider a summary table mapping each theme to participant IDs.","section":"Qualitative Findings"}],"minor_comments":[{"comment":"Figure 4 contains a stray Chinese phrase ('LLM回复的图（要英文）') and an ambiguous annotation 'if true?'; the figure should be cleaned and the switch condition explained in the caption.","section":"Design, Figure 4"},{"comment":"The phrase 'conflicting” role' appears to be a formatting artifact; it should read 'conflicting role.'","section":"Discussion"},{"comment":"The paper uses 'GPT-4' and 'GPT-SoVITS Fast API' without version numbers or a citation for GPT-SoVITS; please add specific model identifiers and a reference.","section":"System Implementation"},{"comment":"Participant demographics are sparse (only age range and sex); report mean age/SD and how 'experience of parental verbal abuse' was verified (e.g., screening questions or a validated scale).","section":"Participants"},{"comment":"No analytic software or number of coders is reported; adding these details would improve reproducibility.","section":"Analysis"},{"comment":"The footnote states 'These authors contributed equllly to the work'; 'equllly' should be 'equally.'","section":"Author footnote"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the TV-drama memory prime is well founded and is the main barrier to accepting the paper's central claim as stated. The study is better positioned as a design exploration and pilot qualitative investigation; with a re-scoped framing or additional controlled data, it could be publishable. I would not reject, as the design and qualitative data have clear value for the HCI community. I also note that the related-work coverage is somewhat thin in areas such as VR exposure therapy and LLM safety, which the authors may be asked to expand."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core here is the design: first-person role-play as an abusive parent speaking to an LLM child, then third-person observation of an LLM mother rephrasing the participant's own stored lines into warm, supportive speech. I have not seen that exact combination in the cited LLM intervention work, and it is a plausible mechanism for reflection that deserves to be tested further. The qualitative data are quoted generously and the findings section is mostly careful to describe what participants reported rather than overclaiming. The paper also discloses its ethical screening (SDS exclusion) and the potentially distressing nature of the task, which is good practice for this population.\n\nThe soft spots are real but not fatal. The stress-test note is fair: a TV-drama clip depicting maternal verbal abuse was shown immediately before the VR session to prime recall, and the interview guide explicitly asked about VR elements and past memories. With no control or neutral-prime comparison, the abstract's wording that the experience 'encourages reflection' goes beyond what the procedure can support. The strongest defensible claim is descriptive - participants reported reflection and mixed emotions in interviews - and the paper mostly sticks to that in the findings. I would soften the abstract and conclusion and either drop the priming clip or add a comparison condition. The lack of shared interview protocol, coding scheme, transcripts, or system code also makes the qualitative analysis hard to verify; that is an addressable reproducibility gap, not a reason to reject.\n\nMinor notes: the sample is small and culturally homogeneous, which the paper acknowledges; the claim that 11 of 12 participants reproduced home situations is exactly what you would expect given recruitment and priming, so it does not validate the VR system specifically. The central descriptive finding survives, but the causal attribution to the VR-LLM interaction is untested.\n\nWho is this for? Researchers working on LLM-based emotional support, VR interventions for trauma reflection, and psychodrama-inspired HCI. It is a solid workshop-or-full-paper candidate with revision, not a desk reject. I would send it to peer review and let reviewers push for a controlled follow-up. I would not cite it in my own work yet, but I would bring it to a reading group as an example of an innovative prototype with honest limitations.\n\nRecommendation: serious referee, conditional acceptance path.","headline":"A genuinely novel dual-phase VR-LLM prototype for reflecting on parental verbal abuse, with honest qualitative data, but the causal claim is undercut by the priming clip and no control condition.","tokens_in":11506,"tokens_out":939,"would_cite":false,"duration_ms":11499,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a two-phase VR experience, in which adults first role-play a verbally abusive parent against an LLM child and then watch an LLM mother rephrase those same words, prompts self-reported reflection on past parental…","keywords":["virtual reality","large language models","parental verbal abuse","emotional reflection","role reversal","psychodrama","LLM-driven dialogue","qualitative study"],"falsifier":"Run the same dual-phase role-switch with one group that receives the pre-session TV clip and one that does not, and also compare VR against a plain text-LLM interface. If recollection and reflection scores do not drop when the clip is removed or when VR is replaced by text, the central claim that the immersive VR-LLM experience drives the effect is falsified.","tokens_in":10496,"feed_emoji":"🎭","tokens_out":5680,"duration_ms":59830,"temperature":0.7,"pith_summary":"This paper reports a descriptive, qualitative claim: a virtual-reality experience built on large language models can lead adults with a history of parental verbal abuse to recall past family scenes and feel a degree of emotional relief. The experience is deliberately two-phase. In the first phase, the participant speaks as the abusive parent to an LLM playing a child; in the second phase, they become an observer while an LLM mother transforms the participant's own recorded abusive lines into warm, supportive language. Interviews with 12 Chinese adults aged 18 to 34 indicate that most participants recognized the LLM child's responses as realistic, recalled specific home situations, and found the rephrased dialogue soothing, though several felt the idealized mother was too distant from their real parents. The paper's contribution is not a controlled clinical outcome but a demonstration that a role-reversal and mirroring structure in VR can elicit reflective and supportive feelings, with personal history shaping how strongly those feelings land.","feed_headline":"12 adults report reflection after VR role-swap on parental abuse","feed_subtitle":"First they play the abusive parent, then an LLM mother rewrites those same lines; interviews show it lands unevenly.","key_machinery":"The carrying mechanism is the perspective shift between two VR scenes. In Scene 1, the user speaks as the verbally abusive parent to an LLM-driven child, with voice input converted to text and the child's escalating emotional responses returned as synthesized speech; in Scene 2, the same stored abusive utterances are replayed one by one to an LLM-driven mother, who rephrases them into warm, constructive language while the user watches from a third-person observer position. The two phases operationalize role reversal (adopting the position of the person one is in conflict with) and mirroring (seeing one's own behavior from the outside), both borrowed from psychodrama. The scene transition is triggered by a structured signal in the LLM's response, and visual elements—cold rising water and dissipating particles in Scene 1, warm light and floating particles in Scene 2—are designed to amplify the emotional contrast between the two vantage points.","core_discovery":"The central claim is that a dual-phase VR-LLM interaction, grounded in psychodrama's role reversal and mirroring, elicits reflective recall of parental verbal abuse and generates supportive feelings, while the strength and valence of these effects depend on each participant's personal history. The authors summarize their finding as follows: the experience 'prompts reflective and supportive feelings, yet evokes varied emotional responses shaped by personal histories.' The evidence is qualitative: in interviews, most participants reported reproducing scenes from home while voicing the parent's role, substituting themselves into the LLM child, and finding the Scene 2 reframing soothing; a subset instead reacted to it as idealized or textbook-like, creating emotional distance. The paper does not claim a controlled clinical outcome; it claims that these reported reflections and emotions arose in connection with the designed experience.","pith_inferences":["The pre-session TV clip showing mother-daughter verbal abuse is a plausible alternative source of the reported recollections; a version of the study without the clip, or with a neutral priming video, would test whether the VR scenes themselves carry the reflective effect.","The 'idealized mother' reaction suggests a design frontier: the rephrasing should be calibrated to the participant's actual parental register, not a uniformly warm tone, or it may widen the emotional distance it aims to close.","The role-swap mechanism is not obviously specific to parent-child abuse; the same dual-phase structure (speaker first, witness second) could be tried for other self-blame or conflict memories, such as bullying or romantic conflict, and the paper's qualitative method would transfer directly.","If the effect is confirmed, the next measurable step is to ask whether one session changes communication self-efficacy or mood in the days afterward, since the present study only records immediate interview reactions."],"forward_implications":["If the finding holds, a two-phase role switch can produce recollections of parental verbal abuse in adults without a therapist in the room, which is the precondition for a self-guided reflective tool.","The same mechanism can turn a user's own harsh words into a visible model of supportive parenting; several participants explicitly said they wished their parents had spoken to them that way.","Because reactions varied with personal history, a fixed LLM persona will leave some users emotionally distant (the 'textbook' or 'idealized' reaction), so personalization is a necessary next step rather than an optional polish.","The working pipeline—voice input, speech-to-text, LLM generation, synthesized voice, and VR staging—is reusable for other emotionally difficult conversations, so the paper functions as a design template as well as a study."],"supporting_citations":[{"why":"Supplies evidence that verbal abuse leaves lasting emotional scars, motivating the target population and the need for reflective intervention.","marker":"[7, 25]"},{"why":"Provides the psychodrama role-reversal technique that grounds the first-person abusive-parent phase of the experience.","marker":"[16]"},{"why":"Provides the psychodrama mirroring technique that grounds the third-person observer phase in Scene 2.","marker":"[4]"},{"why":"Supports the claim that adjusting virtual environment parameters can elicit emotion, justifying the cold-to-warm scene design.","marker":"[8]"},{"why":"Justifies choosing the mother as the warm rephraser in Scene 2 by citing mothers' generally more supportive parenting style.","marker":"[37]"},{"why":"Defines the self-rating depression scale used to screen participants and exclude those with severe depression for ethical reasons.","marker":"[40]"}],"fun_headline_variants":["VR role-swap with LLM parent spurs reflection on abuse","LLM rewrites abusive lines in VR; reflection varies by history","Play the abuser, hear an LLM mother reword it: mixed impact","Dual-phase VR with LLM evokes reflection, but unevenly"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported reflections are attributed to the VR-LLM experience, but participants watched a TV clip depicting mother-daughter verbal abuse immediately before the VR session, so the study never separates memories triggered by that priming clip from memories triggered by the VR scenes.","fun_headline_variants_meta":{"raw":{"variants":["VR role-swap with LLM parent spurs reflection on abuse","LLM rewrites abusive lines in VR; reflection varies by history","Play the abuser, hear an LLM mother reword it: mixed impact","Dual-phase VR with LLM evokes reflection, but unevenly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000283,"raw_usage":{"total_tokens":1625,"prompt_tokens":854,"completion_tokens":771,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":693}},"tokens_in":470,"tokens_out":771,"duration_ms":7262,"temperature":1.0,"reasoning_tokens":693,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:16:43.149197+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same dual-phase role-switch with one group that receives the pre-session TV clip and one that does not, and also compare VR against a plain text-LLM interface. If recollection and reflection scores do not drop when the clip is removed or when VR is replaced by text, the central claim that the immersive VR-LLM experience drives the effect is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the psychodrama role-reversal technique that grounds the first-person abusive-parent phase of the experience."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the psychodrama mirroring technique that grounds the third-person observer phase in Scene 2."},{"cited_title":"D.; Schmidt, M.; Hein- zle, A.-K.; Beutl, L.; Hlavacs, H.; and Kryspin-Exner, I","cited_arxiv_id":null,"evidence_quote":"Supports the claim that adjusting virtual environment parameters can elicit emotion, justifying the cold-to-warm scene design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies choosing the mother as the warm rephraser in Scene 2 by citing mothers' generally more supportive parenting style."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the self-rating depression scale used to screen participants and exclude those with severe depression for ethical reasons."}],"review_version":1}