REVIEW 3 major objections 6 minor 46 references
Can Code Outlove Blood? An LLM-based VR Experience to Prompt Reflection on Parental Verbal Abuse
T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper claims that a two-phase VR experience, in which adults first role-play a verbally abusive parent against an LLM child and then watch an LLM mother rephrase those same words, prompts self-reported reflection on past parental…
desk verdict A genuinely novel dual-phase VR-LLM prototype for reflecting on parental verbal abuse, with honest qualitative data, but the causal claim is undercut by the priming clip and no control condition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the perspective shift between two VR scenes. In Scene 1, the user speaks as the verbally abusive parent to an LLM-driven child, with voice input converted to text and the child's escalating emotional responses returned as synthesized speech; in Scene 2, the same stored abusive utterances are replayed one by one to an LLM-driven mother, who rephrases them into warm, constructive language while the user watches from a third-person observer position. The two phases operationalize role reversal (adopting the position of the person one is in conflict with) and mirroring (seeing one's own behavior from the outside), both borrowed from psychodrama. The scene transition is triggered by a structured signal in the LLM's response, and visual elements—cold rising water and dissipating particles in Scene 1, warm light and floating particles in Scene 2—are designed to amplify the emotional contrast between the two vantage points.
What would settle it
Run the same dual-phase role-switch with one group that receives the pre-session TV clip and one that does not, and also compare VR against a plain text-LLM interface. If recollection and reflection scores do not drop when the clip is removed or when VR is replaced by text, the central claim that the immersive VR-LLM experience drives the effect is falsified.
Extended reading notes
Core claim
The central claim is that a dual-phase VR-LLM interaction, grounded in psychodrama's role reversal and mirroring, elicits reflective recall of parental verbal abuse and generates supportive feelings, while the strength and valence of these effects depend on each participant's personal history. The authors summarize their finding as follows: the experience 'prompts reflective and supportive feelings, yet evokes varied emotional responses shaped by personal histories.' The evidence is qualitative: in interviews, most participants reported reproducing scenes from home while voicing the parent's role, substituting themselves into the LLM child, and finding the Scene 2 reframing soothing; a subset instead reacted to it as idealized or textbook-like, creating emotional distance. The paper does not claim a controlled clinical outcome; it claims that these reported reflections and emotions arose in connection with the designed experience.
Load-bearing premise
The reported reflections are attributed to the VR-LLM experience, but participants watched a TV clip depicting mother-daughter verbal abuse immediately before the VR session, so the study never separates memories triggered by that priming clip from memories triggered by the VR scenes.
Editorial extensions
If this is right
- If the finding holds, a two-phase role switch can produce recollections of parental verbal abuse in adults without a therapist in the room, which is the precondition for a self-guided reflective tool.
- The same mechanism can turn a user's own harsh words into a visible model of supportive parenting; several participants explicitly said they wished their parents had spoken to them that way.
- Because reactions varied with personal history, a fixed LLM persona will leave some users emotionally distant (the 'textbook' or 'idealized' reaction), so personalization is a necessary next step rather than an optional polish.
- The working pipeline—voice input, speech-to-text, LLM generation, synthesized voice, and VR staging—is reusable for other emotionally difficult conversations, so the paper functions as a design template as well as a study.
Reading between the lines
- The pre-session TV clip showing mother-daughter verbal abuse is a plausible alternative source of the reported recollections; a version of the study without the clip, or with a neutral priming video, would test whether the VR scenes themselves carry the reflective effect.
- The 'idealized mother' reaction suggests a design frontier: the rephrasing should be calibrated to the participant's actual parental register, not a uniformly warm tone, or it may widen the emotional distance it aims to close.
- The role-swap mechanism is not obviously specific to parent-child abuse; the same dual-phase structure (speaker first, witness second) could be tried for other self-blame or conflict memories, such as bullying or romantic conflict, and the paper's qualitative method would transfer directly.
- If the effect is confirmed, the next measurable step is to ask whether one session changes communication self-efficacy or mood in the days afterward, since the present study only records immediate interview reactions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript describes a dual-phase VR-LLM experience designed to prompt reflection on parental verbal abuse. Participants with self-reported histories of such abuse first role-play an abusive parent speaking scripted hurtful lines to an LLM-driven child, then observe an LLM-driven mother rephrase the same lines into warm, supportive language. The authors report a qualitative study with 12 Chinese adults (18–34) using semi-structured interviews and thematic analysis. The findings are organized into three themes: the first scene evoked recollection of past events and reflection; the second scene fostered supportive emotions; participants viewed LLMs as promising but in need of personalization. The abstract concludes that the experience 'encourages reflection on their past experiences and fosters supportive emotions,' while the Discussion frames this as a contribution to AI-driven emotional support design.
Significance. If taken as a descriptive qualitative study, the paper offers a clearly described dual-phase role-reversal design, a detailed system implementation (Unity, Oculus Quest 2, GPT-4, GPT-SoVITS, Xunfei API), and a rich set of participant quotes that ground the reported themes. The design rationale draws on psychodrama theory, and the Limitations section is unusually candid about demographic and ethical risks. The main value is as a formative design exploration: it shows that a VR-LLM role-switching experience can elicit vivid recollections and mixed emotions in this population, and it identifies personalization as a key challenge. However, the paper's central causal claim that the experience 'prompts' reflection and supportive feelings is not supported by the present study design, which lacks controls and includes a memory-priming stimulus. The strongest defensible claim is descriptive: participants reported these effects in post-hoc interviews.
major comments (3)
- [Methods, Procedure] The causal attribution in the abstract and findings is not supported by the design. The procedure states that 'We played a segment from a TV drama depicting a mother verbally abusing her daughter, sourced from popular clips on the widely known Chinese social media platform Xiaohongshu to help participants recall and bring into the scenes of familial verbal abuse' immediately before the VR session. With no control condition, no neutral-prime comparison, and no pre/post measures, participants' recollections and reported reflections could be elicited by the priming clip, by the interview questions (which explicitly asked how the role-play related to past experiences and whether LLM responses evoked memories), or by demand characteristics. Because the abstract claims the experience 'encourages reflection' and the findings section states it 'prompts reflective and supportive feelings,' this confound is load-bearing. I recommend reframing the conclusions as descriptive self-report findings, or adding a baseline/control arm and pre-post reflection measures to support causal claims.
- [Analysis] The thematic categories ('recollection of past events,' 'emotional impact,' 'LLM response effectiveness,' 'environmental design influence') closely mirror the interview-guide items (i)-(x) listed in Methods, Procedure. The Analysis section describes open coding and thematic analysis but provides no coding procedure, codebook, inter-rater reliability, or member-checking information. As a result, the finding that 'All participants apart from P4 stated that in the first scene, they reproduced some or all of the scenes that occurred at home' may largely reflect the interview prompts rather than an emergent property of the experience. Please report the coding process in more detail and discuss how the themes were distinguished from the interview structure.
- [Qualitative Findings] The counts of participants associated with each claim are sometimes ambiguous or inconsistent. For example, the paper states 'Most participants perceived that in Scene 2, the LLM had effectively translated... (P1, P2, P3, P4, P6, P7, P8, P9, P10, P11, P12)', which is 11 of 12, while 'this rephrased language felt supportive... (P1, P2, P3, P5, P6, P9, P10, P11)' is only 8. Some of these parenthetical lists appear to include participants who made related but distinct comments, and no participant-level matrix is provided. Please clarify whether these counts represent explicit endorsements of the stated claim and consider a summary table mapping each theme to participant IDs.
minor comments (6)
- [Design, Figure 4] Figure 4 contains a stray Chinese phrase ('LLM回复的图(要英文)') and an ambiguous annotation 'if true?'; the figure should be cleaned and the switch condition explained in the caption.
- [Discussion] The phrase 'conflicting” role' appears to be a formatting artifact; it should read 'conflicting role.'
- [System Implementation] The paper uses 'GPT-4' and 'GPT-SoVITS Fast API' without version numbers or a citation for GPT-SoVITS; please add specific model identifiers and a reference.
- [Participants] Participant demographics are sparse (only age range and sex); report mean age/SD and how 'experience of parental verbal abuse' was verified (e.g., screening questions or a validated scale).
- [Analysis] No analytic software or number of coders is reported; adding these details would improve reproducibility.
- [Author footnote] The footnote states 'These authors contributed equllly to the work'; 'equllly' should be 'equally.'
Circularity Check
No circular derivation: qualitative design study whose findings are self-reported interview data, not consequences of fitted inputs or self-cited theorems.
full rationale
This paper is a qualitative HCI/VR design study and does not contain a formal derivation chain. It fits no parameters, defines no equations, and makes no prediction that is mathematically forced by its inputs. The core findings are participants' self-reported reflections and emotions elicited in semi-structured interviews; these are empirical observations, not quantities computed from the system design. The coding themes (e.g., 'recollection of past events', 'emotional impact', 'environmental design influence') are descriptive labels applied to interview excerpts, not constructs that are definitionally identical to the interview prompts. For example, asking participants whether LLM responses 'evoked memories of their past selves' does not by construction produce the finding that the VR-LLM experience encourages reflection; participants could have answered negatively, and some did report divergence (e.g., P8, P9). No load-bearing self-citation appears: the design is justified by established psychodrama literature (Kellermann, Chesner) and general VR emotion research, and the authors do not invoke a prior uniqueness theorem or their own prior work to rule out alternatives. The LLM's rephrasing behavior in Scene 2 is a designed intervention, and participants' perceptions of it are evaluated outcomes, not used to prove the system works by definition. The paper's own limitations section identifies demographic and ethical concerns, and the reader's worry about the TV-drama prime and interview prompts is a methodological threat to causal attribution, not circularity: the recollections are not logically entailed by the prime or the interview guide. Under the stated rules, this warrants a score of 0.
Assumptions & free parameters
assumptions (6)
- domain assumption Role reversal in psychodrama promotes emotional reflection.
- domain assumption Observing one's own scene from a third-person perspective (mirroring) yields a broader, more authentic perspective.
- domain assumption Adjustable virtual environment parameters (lighting, weather, particles) reliably induce target emotions.
- domain assumption Choosing the mother as the 'language translator' aligns with audience expectations because mothers are more accepting and supportive than fathers.
- domain assumption Participant self-report in a semi-structured interview after the experience is a valid indicator of reflection and emotional response.
- domain assumption The SDS cutoff (Y > 0.6) adequately protects participants from psychological risk in an emotionally triggering VR task.
Cite this review
Pith. "Pith review of Can Code Outlove Blood? An LLM-based VR Experience to Prompt Reflection on Parental Verbal Abuse." pith.science (2026). https://pith.science/paper/R3FIDUDI
@misc{pith2026250418410,
author = {Pith},
title = {Pith review of: Can Code Outlove Blood? An LLM-based VR Experience to Prompt Reflection on Parental Verbal Abuse},
year = {2026},
howpublished = {\url{https://pith.science/paper/R3FIDUDI}},
note = {Machine review of arXiv:2504.18410}
}
read the original abstract
Parental verbal abuse leaves lasting emotional impacts, yet current therapeutic approaches often lack immersive self-reflection opportunities. To address this, we developed a VR experience powered by LLMs to foster reflection on parental verbal abuse. Participants with relevant experiences engage in a dual-phase VR experience: first assuming the role of a verbally abusive parent, interacting with an LLM portraying a child, then observing the LLM reframing abusive dialogue into warm, supportive expressions as a nurturing parent. A qualitative study with 12 participants showed that the experience encourages reflection on their past experiences and fosters supportive emotions. However, these effects vary with participants' personal histories, emphasizing the need for greater personalization in AI-driven emotional support. This study explores the use of LLMs in immersive environment to promote emotional reflection, offering insights into the design of AI-driven emotional support systems.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Afifah, W. 2016. Keep on guard against violent commu- nication (verbal-abuse) within early childhood education. In Annual Conference on Islamic Early Childhood Educa- tion (ACIECE), volume 1, 75–86
work page 2016
-
[2]
Alessa, A., and Al-Khalifa, H. 2023. Towards design- ing a chatgpt conversational companion for elderly peo- ple. In Proceedings of the 16th International Conference on PErvasive Technologies Related to Assistive Environ- ments, PETRA ’23, 667–674. New York, NY , USA: Asso- ciation for Computing Machinery
work page 2023
-
[3]
R.; Dalgleish, T.; and Joseph, S
Brewin, C. R.; Dalgleish, T.; and Joseph, S. 1996. A dual representation theory of posttraumatic stress disorder. Psychological review 103(4):670
work page 1996
-
[4]
Chesner, A. 2019. One-to-one psychodrama psychother- apy: Applications and technique. Routledge
work page 2019
-
[5]
Chirico, A.; Ferrise, F.; Cordella, L.; and Gaggioli, A
-
[6]
Cohen, J. A., and Mannarino, A. P. 2015. Trauma- focused cognitive behavioral therapy for traumatized chil- dren and families. Child and adolescent psychiatric clinics of North America 24(3):557
work page 2015
-
[7]
Danese, A.; Moffitt, T. E.; Harrington, H.; Milne, B. J.; Polanczyk, G.; Pariante, C. M.; Poulton, R.; and Caspi, A. 2009. Adverse childhood experiences and adult risk factors for age-related disease: depression, inflammation, and clustering of metabolic risk markers. Archives of pe- diatrics & adolescent medicine 163(12):1135–1143
work page 2009
-
[8]
D.; Schmidt, M.; Hein- zle, A.-K.; Beutl, L.; Hlavacs, H.; and Kryspin-Exner, I
Felnhofer, A.; Kothgassner, O. D.; Schmidt, M.; Hein- zle, A.-K.; Beutl, L.; Hlavacs, H.; and Kryspin-Exner, I
Show all 46 references
-
[9]
Ferrari, A. M. 2002. The impact of culture upon child rearing practices and definitions of maltreatment. Child abuse & neglect 26(8):793–813
2002
-
[10]
M.; Boxer, G.; Ambeau, A
Fletcher, C. M.; Boxer, G.; Ambeau, A. V .; and De- Marais, S. 2018. More to the story: Synthesizing narrative therapy with the adaptive information processing model. Journal of Counselor Practice 9(2)
2018
-
[11]
Gibb, B. E. 2002. Childhood maltreatment and nega- tive cognitive styles: A quantitative and qualitative review. Clinical Psychology Review 22(2):223–246
2002
-
[12]
T.; Hubbard, E
Ho, H.-R.; White, N. T.; Hubbard, E. M.; and Mutlu, B
-
[13]
Hu, Z.; Hou, H.; and Ni, S. 2024. Grow with your ai buddy: Designing an llms-based conversational agent for the measurement and cultivation of children’s mental resilience. In Proceedings of the 23rd Annual ACM Inter- action Design and Children Conference, 811–817
2024
-
[14]
G.; Cohen, P.; Smailes, E
Johnson, J. G.; Cohen, P.; Smailes, E. M.; Skodol, A. E.; Brown, J.; and Oldham, J. M. 2001. Childhood ver- bal abuse and risk for personality disorders during ado- lescence and early adulthood. Comprehensive psychiatry 42(1):16–23
2001
-
[15]
Kellermann, P. 1992. Focus on psychodrama: The ther- apeutic aspects of psychodrama. Jessica Kingsley
1992
-
[16]
Kellermann, P. F. 1994. Role reversal in psychodrama. Psychodrama since Moreno: Innovations in theory and practice 263–279
1994
-
[17]
Kellermann, P. F. 1999. Ethical concerns in psy- chodrama. Journal of the British Psychodrama Associa- tion 14:3–19
1999
-
[18]
Kumar, H.; Yu, K.; Chung, A.; Shi, J.; and Williams, J. J. 2023. Exploring the potential of chatbots to provide mental well-being support for computer science students. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V . 2, SIGCSE 2023, 1339. N...
2023
-
[19]
M., and Quinn, J
Lawson, D. M., and Quinn, J. 2013. Complex trauma in children and adolescents: Evidence-based practice in clinical settings. Journal of clinical psychology69(5):497– 509
2013
-
[20]
B.; and Reichart, R
Lissak, S.; Calderon, N.; Shenkman, G.; Ophir, Y .; Fruchter, E.; Klomek, A. B.; and Reichart, R. 2024. The colorful future of llms: Evaluating and improving llms as emotional supporters for queer youth. arXiv preprint arXiv:2402.11886
2024 arXiv
-
[21]
Liu, D.; Zhou, H.; and An, P. 2024. ”when he feels cold, he goes to the seahorse”—blending generative ai into mul- timaterial storymaking for family expressive arts therapy. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24. New York, NY ...
2024
-
[22]
Luken, A.; Nair, R.; and Fix, R. L. 2021. On racial disparities in child abuse reports: Exploratory mapping the 2018 ncands. Child maltreatment 26(3):267–281
2021
-
[23]
D.; Anderson, S
Luxton, D. D.; Anderson, S. L.; and Anderson, M
-
[24]
Malchiodi, C. A. 2011. Handbook of art therapy. Guil- ford Press
2011
-
[25]
Ney, P. G. 1987. Does verbal abuse leave deeper scars: A study of children and parents. The Canadian Journal of Psychiatry 32(5):371–378
1987
-
[26]
Ntoutsi, E.; Fafalios, P.; Gadiraju, U.; Iosifidis, V .; Ne- jdl, W.; Vidal, M.-E.; Ruggieri, S.; Turini, F.; Papadopou- los, S.; Krasanakis, E.; et al. 2020. Bias in data-driven arti- ficial intelligence systems—an introductory survey. Wiley Interdisciplinary Reviews: Data Mi...
2020
-
[27]
Pease, B., and Pease, A. 2008. The definitive book of body language: The hidden meaning behind people’s ges- tures and expressions. Bantam
2008
-
[28]
C.; Seinfeld, S.; Aglioti, S
Peck, T. C.; Seinfeld, S.; Aglioti, S. M.; and Slater, M
-
[29]
Perella-Holfeld, F.; Sallam, S.; Petrie, J.; Gomez, R.; Irani, P.; and Sakamoto, Y . 2024. Parent and educator con- cerns on the pedagogical use of ai-equipped social robots. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 8(3)
2024
-
[30]
Pinilla, A.; Garcia, J.; Raffe, W.; V oigt-Antons, J.-N.; and M¨oller, S. 2021. Visual representation of emotions in virtual reality. PsyArXiv
2021
-
[31]
Polcari, A.; Rabi, K.; Bolger, E.; and Teicher, M. H
-
[32]
Slater, M., and Sanchez-Vives, M. V . 2016. Enhanc- ing our lives with immersive virtual reality. Frontiers in Robotics and AI 3:74
2016
-
[33]
Tang, Y .; Chen, L.; Chen, Z.; Chen, W.; Cai, Y .; Du, Y .; Yang, F.; and Sun, L. 2024. Emoeden: Applying generative artificial intelligence to emotional learning for children with high-function autism. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Sy...
2024
-
[34]
H.; Samson, J
Teicher, M. H.; Samson, J. A.; Polcari, A.; and Mc- Greenery, C. E. 2006. Sticks, stones, and hurtful words: relative effects of various forms of childhood maltreat- ment. American journal of psychiatry 163(6):993–1000
2006
-
[35]
Wang, Y ., and Tang, T. Y . 2024. Position paper: A per- sonalized large language model (llm)-based chat compan- ion for autistic children early intervention. In Compan- ion of the 2024 on ACM International Joint Conference on Pervasive and Ubiquitous Computing , UbiComp ’24, ...
2024
-
[36]
Wang, Y .; Tang, M.; He, Y .; and Tang, T. Y . 2024. In- teractive design with autistic children using llm and iot for personalized training: The good, the bad and the challeng- ing. In Companion of the 2024 on ACM International Joint Conference on Pervasive and Ubiquitous Com...
2024
-
[37]
Yaffe, Y . 2023. Systematic review of the differences between mothers and fathers in parenting styles and prac- tices. Current psychology 42(19):16011–16024
2023
-
[38]
funny how?
Zargham, N.; Avanesi, V .; Reicherts, L.; Scott, A. E.; Rogers, Y .; and Malaka, R. 2023. “funny how?” a serious look at humor in conversational agents. In Proceedings of the 5th International Conference on Conversational User Interfaces, CUI ’23. New York, NY , USA: Associati...
2023
-
[39]
Zayfert, C., and Becker, C. B. 2019. Cognitive- behavioral therapy for PTSD: A case formulation ap- proach. Guilford Publications
2019
-
[40]
Zung, W. W. 1965. A self-rating depression scale. Arch Gen Psychiatry 12(1):63–70. PMID: 14221692
1965
-
[2013]
Consciousness and cognition 22(3):779–787
Putting yourself in the skin of a black avatar re- duces implicit racial bias. Consciousness and cognition 22(3):779–787
-
[2014]
Child Abuse & Neglect 38(1):91–102
Parental verbal affection and verbal aggression in childhood differentially influence psychiatric symptoms and wellbeing in young adulthood. Child Abuse & Neglect 38(1):91–102
-
[2015]
Is virtual reality emotionally arousing? investigating five emotion inducing virtual park scenarios.International journal of human-computer studies 82:48–56
-
[2016]
In Artificial in- telligence in behavioral and mental health care
Ethical issues and artificial intelligence technolo- gies in behavioral and mental health care. In Artificial in- telligence in behavioral and mental health care . Elsevier. 255–276
-
[2018]
Frontiers in psychology8:2351
Designing awe in virtual reality: An experimental study. Frontiers in psychology8:2351
-
[2023]
In Proceedings of the 22nd Annual ACM Interaction Design and Children Conference, IDC ’23, 355–366
Designing parent-child-robot interactions to facili- tate in-home parental math talk with young children. In Proceedings of the 22nd Annual ACM Interaction Design and Children Conference, IDC ’23, 355–366. New York, NY , USA: Association for Computing Machinery
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.