REVIEW 3 major objections 5 minor 110 references
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Red-teaming generative AI is interactional labor that can harm the mental health of the people doing it, making their protection a workplace-safety obligation that can be met by safeguards borrowed from four comparable professions.
desk verdict A careful, honest position paper that argues for red-teamer mental health support; the four-profession analogy is novel and the safeguard menu is practical, but the 'uniquely tied' claim outruns the evidence and one cited statistic needs verification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a named mechanism: 'interactional labor,' defined as work in which a person repeatedly simulates a malicious actor, solicits harmful outputs, documents them, and repeats the cycle. The paper's argument moves through this mechanism in two steps: first, interactional labor is what separates red-teaming from observational content moderation; second, the same active participation has been shown in other settings to produce perpetration-related stress, moral injury, and self-concept blurring. From there, the paper adapts four protective technologies—de-roling and debriefing, meaning-making and compassion-fatigue prevention, dose-monitoring and inoculation rituals, and peer support plus organizational feedback—mapping each onto the remote, contract-heavy, NDA-bound structure of red-teaming work.
What would settle it
Run a controlled comparison in which one group of red-teamers actively role-plays malicious personas to elicit harmful outputs while a second group reviews identical pre-generated harmful outputs without interacting; if the two groups show no difference in guilt, moral injury, intrusive thoughts, or self-concept change, the claim that the harms are uniquely tied to interactional labor would collapse, and the same design could randomly assign a de-roling protocol to half the participants to test whether the proposed safeguard changes outcomes.
Extended reading notes
Core claim
The paper's central claim is that the psychological hazard of AI red-teaming lies in its active, repeated enactment of harm rather than in exposure to disturbing content. Effective red-teaming requires a tester to think and write from inside a malicious persona, elicit harmful outputs, and then document the exploit; the paper argues that this participation—not the screen content alone—is what generates moral injury, guilt, intrusive thoughts, hypervigilance, and self-concept changes, drawing on evidence that harming virtual agents produces real stress even when the agent is known to be artificial. It then claims that these harms amount to a critical workplace-safety concern for red teams and the contractors who staff them, and that the right response is a structured set of individual and organizational safeguards adapted from professions that do comparable interactional labor. The paper's concrete proposals include after-session de-roling and debriefing, reframing red-teaming as bearing witness, BEEP-style self-monitoring with transition rituals, balanced sensitization and desensitization, confidential peer support, red-team-culture-literate ombudspeople, and feedback that makes the impact of the work visible.
Load-bearing premise
The paper's argument breaks if role-playing a malicious actor against an AI does not actually produce the same psychological harm as real or vividly simulated perpetration, or if the coping practices of actors, therapists, war photographers, and moderators do not transfer to a remote, screen-based, contract job.
Editorial extensions
If this is right
- If the claim is correct, psychological support moves from optional to load-bearing in AI-safety infrastructure: red teams should have informed consent, fair compensation, and confidential mental-health care as standard practice.
- De-roling and debriefing after adversarial sessions would become routine, including digital equivalents—separate accounts, separate chat spaces, and cross-organization debrief groups—because many red-teamers work remotely or as contractors.
- Contract and freelance red-teamers, who often lack employer health benefits, would need support delivered through cross-organization bodies, union-like structures, or employer-funded plans to make the safeguards actually reachable.
- As red-teaming extends to voice, video, and automated tools, the exposure profile changes; the paper's proposed ombudspeople and context-sensitive monitoring would be needed to catch new psychological hazards before they compound.
- If stress and burnout erode creativity, then protecting red-teamer mental health is also an effectiveness argument: healthier testers may find more inventive exploits, an empirical claim the paper explicitly flags for future study.
Reading between the lines
- The same interactional-labor argument likely extends beyond red-teaming to everyone who role-plays harmful or traumatic scenarios with AI systems for evaluation, security, or research—chatbot safety testers, adversarial benchmark writers, and embodied-agent evaluators may face similar psychological risk that has not been measured.
- One testable extension: red-teamers who perform a deliberate de-roling ritual after sessions should show lower next-day intrusive thoughts and self-blame than those who stop abruptly; acting pedagogy gives the intervention, and a randomized field study could settle it.
- A further implication is that NDAs are not just a confidentiality tool but a mental-health risk amplifier, because they block the peer debriefing the comparison professions treat as protective; carving out shared, confidential psychological-support spaces may be as important as hiring therapists.
- If the mechanism is really identification, then risk is not uniform across red-teamers: those asked to weaponize their own lived experience to impersonate members of targeted groups would be most exposed, so safeguards should be tailored by identity and role rather than applied as a generic wellness policy.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that AI red-teaming is a form of "interactional labor" in which workers actively simulate malicious actors and solicit harmful content from generative AI models, and that this labor can cause mental health harms "uniquely tied" to the adversarial engagement strategies required for effective red-teaming. The authors contend that the unmet mental health needs of red-teamers constitute a critical workplace safety concern, grounding this claim in historical and legal arguments about the right to a safe workplace. They propose a portfolio of individual and organizational safeguards adapted from four comparison professions: actors (de-roling and debriefing), mental health professionals (reframing existential meaning and diversifying caseloads), conflict photographers (BEEP self-monitoring and inoculation rituals), and content moderators (peer support, cultural sensitivity in mental health services, and feedback on impact). The paper is a position/argument piece rather than an empirical study, and it explicitly notes that many of the proposed strategies are potential rather than proven, with the economic case for safer red-team labor identified as an empirically testable question.
Significance. If the central claim is accepted, the paper makes a timely and socially important contribution by naming a population (AI red-teamers) whose occupational mental health is under-examined, and by translating protective practices from adjacent professions into concrete, low-cost recommendations that organizations could implement immediately. The paper is commendably honest in its hedging: it repeatedly uses modal language ("could," "may," "potential"), acknowledges that Anthropic's red-teamers felt positively about their tasks, and explicitly flags the effectiveness and economic-benefit questions as areas for future empirical research. The "interactional labor" framing is a useful conceptual contribution that distinguishes red-teaming from passive content moderation. However, the significance is conditional: the paper's strongest rhetorical claim, that the harms are "uniquely tied" to adversarial interactional labor, rests on thin and partially overlapping evidence, and the transferability of the four professional analogies is asserted rather than argued. With strengthened evidence or carefully hedged claims, this could be a valuable agenda-setting paper.
major comments (3)
- [§2.2 and Abstract] The abstract's load-bearing claim that red-teaming's interactional labor produces mental health harms "uniquely tied" to adversarial engagement is not supported by the evidence adduced. The direct evidence for red-teamer distress rests on two non-peer-reviewed preprints ([26] and [108]) with overlapping authorship (Jina Suh is a co-author of this manuscript and of both preprints) and a two-author op-ed [83]; the mechanism is supplied by analogy to Slater et al.'s VR Milgram replication [90], which measured acute physiological and behavioral stress during a short experiment, not the persistent clinical symptoms (moral injury, nightmares, hypervigilance) that the paper attributes to red-teamers. Please either provide a controlled comparison of red-teamers with non-adversarial AI workers, or soften the uniqueness claim to "potentially distinct" and explicitly discuss the heterogeneity of responses, including the positive experiences of Anthropic's red-teamers reported in [25].
- [§2.1, Spence et al. [93]] The sentence "only 93.1% of moderators had moderate to severe levels of persistent mental distress" is implausible as printed (the word "only" is semantically odd with a number this high) and should be verified against the source and corrected. If the true prevalence is materially lower, the paper's reliance on content moderation as an anchor for the "critical workplace safety concern" argument is weakened.
- [§§3–6] The entire safeguard portfolio is predicated on the analogy between red-teaming and acting, therapy, conflict photography, and content moderation. The paper should explicitly identify which dimensions of these professions are analogous (exposure, role-taking, documentation) and which differ (real vs. simulated harm, voluntariness, pro-social mission, presence of a victim), and explain why the transfer of coping practices is nonetheless expected to hold. Without this analysis, the adapted strategies risk being mismatched to the specific mechanisms of red-teamer distress, which is the central justification for the paper's recommendations.
minor comments (5)
- [§1 (Introduction)] The sentence "In this paper, we thus ask pose two research questions: In this paper, we thus pose two research questions:" contains a duplicated phrase; please clean it up.
- [§2.2] "this has not be the case" should be "this has not been the case".
- [§2.3] "individual and organization strategies" should be "individual and organizational strategies"; also, "strategies that have successful" should be "strategies that have been successful".
- [§3.2 and §5.2] The name "Bailey and Dickson [8]" in §3.2 should be "Bailey and Dickinson [8]" to match the reference list, and "Slowshower [91]" in §5.2 should be "Sloshower [91]" (also in the later occurrence in the same section).
- [§4.1] "can bring help bring meaning" should be "can help bring meaning"; in addition, there are occasional extra spaces before commas and periods throughout the text (e.g., in §3.1, "character ,").
Circularity Check
Central unique-harm claim rests on self-cited preprints by a co-author; safeguard recommendations retain independent content.
-
self citation load bearing
[Abstract; Section 2.2, para. 4; Refs. [26], [108]]
"This interactional labor done by red teams can result in mental health harms that are uniquely tied to the adversarial engagement strategies necessary to effectively red team. ... This phenomenon is indeed what popular press [83] and past research [108] has described among red teams—though red-teamers know that they are interacting with a generative AI model, they report moral injury, persistent feelings of guilt, impaired sleep, nightmares, intrusive thoughts, hypervigilance, and symptoms of PTSD."
The paper's central empirical premise—that red-teaming produces harms 'uniquely tied' to adversarial interactional labor—is supported not by new data or an independent study but by two arXiv preprints, Gillespie et al. [26] and Zhang et al. [108], both co-authored by this paper's co-author Jina Suh. The red-teamer-specific symptoms and the 'uniquely interactional' characterization are taken from those preprints and then restated in the abstract as established fact. No external red-teamer dataset, machine-checked result, or independent replication is introduced, so the load-bearing part of the argument reduces to the authors' own prior, non-peer-reviewed claims.
-
uniqueness imported from authors
[Section 2.2, first paragraph]
"Gillespie et al. [26] differentiate red-teaming from content moderation by noting that 'red-teaming involves deliberately engaging in transgressive, uncomfortable, unethical, immoral, or harmful activities, including immersing [oneself] in scenarios that go against [an individual's] morals or belief systems.' Unlike content moderation's primarily observational nature, red-teaming's core interaction pattern can require workers to inhabit distressing perspectives and perform harmful behaviors."
The 'uniquely interactional' premise that distinguishes red-teaming from content moderation is imported from a prior paper co-authored by a present co-author. The argument then uses this imported uniqueness to conclude that red-teamers face 'unique mental health harms' and need tailored safeguards. The uniqueness is not established within the present paper; it is assumed from [26], so the distinctiveness of the claimed hazard rests on the authors' own prior framing rather than on an independent derivation.
full rationale
This is a position/argument paper rather than an empirical derivation, so formal 'prediction reduces to fit' circularity does not apply. The concrete safeguard recommendations—de-roling, debriefing, BEEP self-monitoring, peer support, and feedback loops—are imported from external professional literatures (acting, therapy, photojournalism, content moderation) and do not themselves feed back into the central claim, giving the paper's practical contribution independent content. However, the load-bearing premise that red-teaming causes harms uniquely tied to adversarial interactional labor is supported primarily by two arXiv preprints co-authored by a current co-author (Jina Suh): Gillespie et al. [26] supplies the 'uniquely interactional' framing, and Zhang et al. [108] supplies the reported symptoms. The external sources (content-moderation studies, the VR Milgram replication) are analogical rather than direct evidence about red-teamers. Because the recommended safeguards are reasonable low-cost supports even if the uniqueness claim were false, the circularity is partial rather than total, so a moderate score of 4 is appropriate.
Assumptions & free parameters
assumptions (3)
- domain assumption Protective practices are transferable: strategies that safeguard actors, therapists, conflict photographers, and content moderators can be adapted to protect AI red-teamers.
- domain assumption The cited findings on red-teamer and moderator distress are accurate and generalize.
- domain assumption Active participation in generating harmful content is psychologically harmful even in simulated, explicitly virtual settings with no real victim.
Cite this review
Pith. "Pith review of When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines." pith.science (2026). https://pith.science/paper/DI6ZJXJ6
@misc{pith2026250420910,
author = {Pith},
title = {Pith review of: When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines},
year = {2026},
howpublished = {\url{https://pith.science/paper/DI6ZJXJ6}},
note = {Machine review of arXiv:2504.20910}
}
read the original abstract
Red-teaming is a core part of the infrastructure that ensures that AI models do not produce harmful content. Unlike past technologies, the black box nature of generative AI systems necessitates a uniquely interactional mode of testing, one in which individuals on red teams actively interact with the system, leveraging natural language to simulate malicious actors and solicit harmful outputs. This interactional labor done by red teams can result in mental health harms that are uniquely tied to the adversarial engagement strategies necessary to effectively red team. The importance of ensuring that generative AI models do not propagate societal or individual harm is widely recognized -- one less visible foundation of end-to-end AI safety is also the protection of the mental health and wellbeing of those who work to keep model outputs safe. In this paper, we argue that the unmet mental health needs of AI red-teamers is a critical workplace safety concern. Through analyzing the unique mental health impacts associated with the labor done by red teams, we propose potential individual and organizational strategies that could be used to meet these needs, and safeguard the mental health of red-teamers. We develop our proposed strategies through drawing parallels between common red-teaming practices and interactional labor common to other professions (including actors, mental health professionals, conflict photographers, and content moderators), describing how individuals and organizations within these professional spaces safeguard their mental health given similar psychological demands. Drawing on these protective practices, we describe how safeguards could be adapted for the distinct mental health challenges experienced by red teaming organizations as they mitigate emerging technological risks on the new digital frontlines.
Reference graph
Works this paper leans on
-
[26]
Tarleton Gillespie, Ryland Shaw, Mary L Gray, and Jina Suh. 2024. AI Red- Teaming is a Sociotechnical System. Now What?arXiv preprint arXiv:2412.09751 (2024)
arXiv 2024
-
[108]
Alice Qian Zhang, Judith Amores, Mary L Gray, Mary Czerwinski, and Jina Suh. 2024. AURA: Amplifying Understanding, Resilience, and Awareness for Responsible AI Content Work. arXiv preprint arXiv:2411.01426 (2024)
work page Pith review arXiv 2024
-
[83]
Evan Selinger and Brenda Leong. 2024. Getting AI ready for the real world takes a terrible human toll. The Boston Globe (January 11 2024). https://www. bostonglobe.com/2024/01/11/opinion/ai-testing-red-team-human-toll/
work page 2024
-
[90]
Mel Slater, Angus Antley, Adam Davison, David Swapp, Christoph Guger, Chris Barker, Nancy Pistrang, and Maria V Sanchez-Vives. 2006. A virtual reprise of the Stanley Milgram obedience experiments. PloS one 1, 1 (2006), e39
work page 2006
-
[25]
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858 (2022)
arXiv 2022
-
[1]
Herbert K Abrams. 2001. A short history of occupational health. Journal of public health policy 22, 1 (2001), 34–80
2001
-
[2]
Lama Ahmad, Sandhini Agarwal, Michael Lampe, and Pamela Mishkin. 2024. OpenAI’s Approach to External Red Teaming for AI Models and Systems. Technical Report. OpenAI. https://cdn.openai.com/papers/openais-approach-to-external- red-teaming.pdf
2024
-
[3]
Henry Ajder, Giorgio Patrini, Francesco Cavalli, and Laurence Cullen. 2019. The State of Deepfakes: Landscape, Threats, and Impact . Technical report. Deeptrace Labs
2019
Show all 110 references
-
[4]
Bobby Allyn. 2024. Lawsuit: A Chatbot Hinted a Kid Should Kill His Parents Over Screen Time Limits. NPR (10 December 2024). https://www.npr.org/2024/ 12/10/nx-s1-5222574/kids-character-ai-lawsuit
2024
-
[5]
Andrew Arsht and Daniel Etcovitch. 2018. The human cost of online content moderation. Harvard Journal of Law and Technology 2 (2018)
2018
-
[6]
UN General Assembly et al. 1948. Universal declaration of human rights. UN General Assembly 302, 2 (1948), 14–25
1948
-
[7]
Stephane J Baele, Elahe Naserian, and Gabriel Katz. 2024. Is AI-Generated Extremism Credible? Experimental Evidence from an Expert Survey. Terrorism and Political Violence (2024), 1–17
2024
-
[8]
Sally Bailey and Paige Dickinson. 2016. The importance of safely de-roling. Methods: A Journal of Acting Pedagogy 2 (2016)
2016
-
[9]
Holly Bell, Shanti Kulkarni, and Lisa Dalton. 2003. Organizational prevention of vicarious trauma. Families in society 84, 4 (2003), 463–470
2003
-
[10]
Alex Beutel, Kai Xiao, Johannes Heidecke, and Lilian Weng. 2024. Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforce- ment Learning. arXiv preprint arXiv:2412.18693 (2024)
2024 arXiv
-
[11]
Augusto Boal. 2013. The rainbow of desire: The Boal method of theatre and therapy. Routledge
2013
-
[12]
Blake Bullwinkel, Amanda Minnich, Shiven Chawla, Gary Lopez, Martin Pouliot, Whitney Maxwell, Joris de Gruyter, Katherine Pratt, Saphir Qi, Nina Chikanov, et al. 2025. Lessons From Red Teaming 100 Generative AI Products. arXiv preprint arXiv:2501.07238 (2025)
2025 arXiv
-
[13]
Kate Busselle. 2021. De-roling and debriefing: Essential aftercare for educational theatre. Theatre Topics 31, 2 (2021), 129–135
2021
-
[14]
Kristin Byron, Shalini Khazanchi, and Deborah Nazarian. 2010. The relation- ship between stressors and creativity: a meta-analysis examining competing theoretical models. Journal of Applied Psychology 95, 1 (2010), 201
2010
-
[15]
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. 2024. Jailbreakbench: An open robustness bench- mark for jailbreaking large language models. a...
2024 arXiv
-
[16]
Dart Center. 2014. Working with Traumatic Imagery.Dart Center for Journalism & Trauma (12 August 2014). https://dartcenter.org/content/working-with- traumatic-imagery Accessed January 19, 2025
2014
-
[17]
Joseph V DeMarco. 2018. An approach to minimizing legal and reputational risk in Red Team hacking exercises. Computer law & security review 34, 4 (2018), 908–911
2018
-
[18]
Jennifer Dillard. 2008. A slaughterhouse nightmare: Psychological harm suffered by slaughterhouse employees and the possibility of redress through legal reform. Geo. J. on Poverty L. & Pol’y 15 (2008), 391
2008
-
[19]
Jimmy Donaghey and Juliane Reinecke. 2018. When industrial democracy meets corporate social responsibility—A comparison of the Bangladesh accord and alliance as responses to the Rana Plaza disaster. British Journal of Industrial Relations 56, 1 (2018), 14–42
2018
-
[20]
Community Guidelines
Anna Drootin. 2021. " Community Guidelines": The Legal Implications of Workplace Conditions for Internet Content Moderators. Fordham L. Rev. 90 (2021), 1197
2021
-
[21]
Daniel Fabian. 2023. Google’s AI Red Team: the ethical hackers making AI safer. https://blog.google/technology/safety-security/googles-ai-red-team- the-ethical-hackers-making-ai-safer/
2023
-
[22]
Paul Farmer. 2004. Pathologies of power: Health, human rights, and the new war on the poor. Vol. 4. Univ of California Press
2004
-
[23]
Michael Feffer, Anusha Sinha, Wesley H Deng, Zachary C Lipton, and Hoda Heidari. 2024. Red-Teaming for generative AI: Silver bullet or security theater?. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , Vol. 7. 421–437
2024
-
[24]
Anthony Feinstein. 2017. War Photography: The Physical and Psycholog- ical Costs. Journal of Humanities in Rehabilitation Spring 2017 (2 May 2017). https://www.jhrehab.org/2017/05/02/war-photography-the-physical- and-psychological-costs/
2017
-
[27]
Aaron Grattafiori, Ivan Evtimov, Joanna Bitton, and Maya Pavlova. 2024. Taming the Beast: Inside the Llama 3 Red Teaming Process. Presented at DEFCON 32. https://www.youtube.com/watch?v=UQaNjwLhAmo
2024
-
[28]
Mary L Gray and Siddharth Suri. 2019. Ghost work: How to stop Silicon Valley from building a new global underclass . Eamon Dolan Books
2019
-
[29]
E Tory Higgins. 1987. Self-discrepancy: a theory relating self and affect. Psy- chological review 94, 3 (1987), 319
1987
-
[30]
Joe Hight and Frank Smyth. 2003. Tragedies & Journalists: A Guide for More Effective Coverage. (2003). https://dartcenter.org/sites/default/files/en_tnj_0. pdf
2003
-
[31]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)
2024 arXiv
-
[32]
Charles Jaret. 1999. Troubled by newcomers: Anti-immigrant attitudes and action during two eras of mass immigration to the United States. Journal of American Ethnic History (1999), 9–39
1999
-
[33]
Haibo Jin, Ruoxi Chen, Andy Zhou, Yang Zhang, and Haohan Wang. 2024. Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models. arXiv preprint arXiv:2402.03299 (2024)
2024
-
[34]
Eva Jonisová. 2022. The importance and consequences of war photography. Cultural Intertexts 12, 12 (2022), 68–85
2022
-
[35]
Biden Jr
Joseph R. Biden Jr. 2023. Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. https://www.whitehouse.gov/ briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe- secure-and-trustworthy-development-and-use-of-a...
2023
-
[36]
Patrice A Keats. 2010. The moment is frozen in time: Photojournalists’ metaphors in describing trauma photography. Journal of Constructivist Psychology 23, 3 (2010), 231–255
2010
-
[37]
Robert O Keohane. 2005. After hegemony: Cooperation and discord in the world political economy. Princeton university press
2005
-
[38]
Robert O Keohane and Joseph S Nye Jr. 1973. Power and interdependence. Survival 15, 4 (1973), 158–165
1973
-
[39]
John W Kingdon. 1984. Agendas, alternatives, and public policies. Brown and Company (1984)
1984
-
[40]
Linnea Laestadius, Andrea Bishop, Michael Gonzalez, Diana Illenčík, and Celeste Campos-Castillo. 2024. Too human and not human enough: A grounded theory analysis of mental health harms from emotional dependence on the social chatbot Replika. New Media & Society 26, 10 (2024), ...
2024
-
[41]
Xiaoxia Li, Siyuan Liang, Jiyi Zhang, Han Fang, Aishan Liu, and Ee-Chien Chang
-
[42]
Brett T Litz, Nathan Stein, Eileen Delaney, Leslie Lebowitz, William P Nash, Caroline Silva, and Shira Maguen. 2009. Moral injury and moral repair in war veterans: A preliminary model and intervention strategy. Clinical psychology review 29, 8 (2009), 695–706. When Testing AI ...
2009
-
[43]
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Kailong Wang. 2024. A hitchhiker’s guide to jail- breaking chatgpt via prompt engineering. In Proceedings of the 4th International Workshop on Software Engineering and AI for Da...
2024
-
[44]
Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng- Xin Yong, Suhas Kotha, et al . 2024. A safe harbor for ai evaluation and red teaming. arXiv preprint arXiv:2403.04893 (2024)
2024 arXiv
-
[45]
Rachel M MacNair. 2002. Perpetration-induced traumatic stress in combat veterans. Peace and conflict: journal of peace psychology 8, 1 (2002), 63–72
2002
-
[46]
Shira Maguen, Brandon J Griffin, Dawne Vogt, Claire A Hoffmire, John R Blos- nich, Paul A Bernhard, Fatema Z Akhtar, Yasmin S Cypel, and Aaron I Schnei- derman. 2023. Moral injury and peri-and post-military suicide attempts among post-9/11 veterans. Psychological Medicine 53, ...
2023
-
[47]
Shira Maguen, Barbara A Lucenko, Mark A Reger, Gregory A Gahm, Brett T Litz, Karen H Seal, Sara J Knight, and Charles R Marmar. 2010. The impact of reported direct and indirect killing on mental health symptoms in Iraq war veterans. Journal of Traumatic Stress: Official Public...
2010
-
[48]
Shira Maguen, Thomas J Metzler, Brett T Litz, Karen H Seal, Sara J Knight, and Charles R Marmar. 2009. The impact of killing in war on mental health symptoms and related functioning. Journal of traumatic stress 22, 5 (2009), 435–443
2009
-
[49]
Shira Maguen, Dawne S Vogt, Lynda A King, Daniel W King, Brett T Litz, Sara J Knight, and Charles R Marmar. 2011. The impact of killing on mental health symptoms in Gulf War veterans. Psychological Trauma: Theory, Research, Practice, and Policy 3, 1 (2011), 21
2011
-
[50]
Bernard Marr. 2024. Generative AI Is Coming To Your Home Appliances. Forbes (29 March 2024). https://www.forbes.com/sites/bernardmarr/2024/03/ 29/generative-ai-is-coming-to-your-home-appliances/
2024
-
[51]
Arthur F McEvoy. 1995. The Triangle Shirtwaist Factory Fire of 1911: Social change, industrial accidents, and the evolution of common-sense causality. Law & Social Inquiry 20, 2 (1995), 621–651
1995
-
[52]
Cait McMahon. 2019. The self-care. In Trauma reporting: a journalist’s guide to covering sensitive stories, Jo Healey (Ed.). Routledge, 178–185
2019
-
[53]
Jacob Metcalf and Ranjit Singh. 2024. Scaling up mischief: Red-teaming AI and distributing governance. Harvard Data Science Review Special Issue 5 (2024)
2024
-
[54]
Stanley Milgram. 1963. Behavioral study of obedience. The Journal of abnormal and social psychology 67, 4 (1963), 371
1963
-
[55]
Saira Mohamed. 2015. Of monsters and men: Perpetrator trauma and mass atrocity. Colum. L. Rev. 115 (2015), 1157
2015
-
[56]
Susana de Deus Tavares Monteiro and Alexandra Marques Pinto. 2017. Report- ing daily and critical events: Journalists’ perceptions of coping and savouring strategies, and of organizational support. European Journal of Work and Organi- zational Psychology 26, 3 (2017), 468–480
2017
-
[57]
Sonia Moore. 1984. The Stanislavski system: The professional training of an actor . Penguin
1984
-
[58]
Christina Mutschler, Chyrell Bellamy, Larry Davidson, Sidney Lichtenstein, and Sean Kidd. 2022. Implementation of peer support in mental health services: A systematic review of the literature. Psychological Services 19, 2 (2022), 360
2022
-
[59]
Lily Hay Newman. 2024. Security News This Week: A Creative Trick Makes Chat- GPT Spit Out Bomb-Making Instructions. WIRED (14 September 2024). https: //www.wired.com/story/chatgpt-jailbreak-homemade-bomb-instructions/
2024
-
[60]
Casey Newton. 2019. The Trauma Floor: The secret lives of Facebook moderators in America. The Verge (February 25 2019). https://www.theverge.com/2019/2/25/18229714/cognizant-facebook-content- moderator-interviews-trauma-working-conditions-arizona
2019
-
[61]
Paige Nong, Julia Adler-Milstein, Nate C Apathy, A Jay Holmgren, and Jordan Everson. 2025. Current Use And Evaluation Of Artificial Intelligence And Predictive Models In US Hospitals: Article examines uses and evaluation of artificial intelligence and predictive models in US h...
2025
-
[62]
Chloe Nurik. 2022. Facing Contracting Issues: The Psychological and Financial Impacts of Facebook Outsourcing Content Moderation. U. Pa. L. Rev. 171 (2022), 1551
2022
-
[63]
Lisa Parks. 2019. Dirty data: Content moderation, regulatory outsourcing, and the cleaners. Film Quarterly 73, 1 (2019), 11–18
2019
-
[64]
Saakvitne
Laurie Anne Pearlman and Karen W. Saakvitne. 1995.Trauma and the Therapist: Countertransference and Vicarious Traumatization in Psychotherapy with Incest Survivors. W. W. Norton
1995
-
[65]
Sachin R Pendse, Talie Massachi, Jalehsadat Mahdavimoghaddam, Jenna Butler, Jina Suh, and Mary Czerwinski. 2024. Towards Inclusive Futures for Worker Wellbeing. Proceedings of the ACM on Human-Computer Interaction 8, CSCW1 (2024), 1–32
2024
-
[66]
Can I not be suicidal on a Sunday?
Sachin R Pendse, Amit Sharma, Aditya Vashistha, Munmun De Choudhury, and Neha Kumar. 2021. “Can I not be suicidal on a Sunday?”: understanding technology-mediated pathways to mental health support. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–16
2021
-
[67]
Iryna Pentina, Tyler Hancock, and Tianling Xie. 2023. Exploring relationship development with social chatbots: A mixed-method study of replika. Computers in Human Behavior 140 (2023), 107600
2023
-
[68]
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. arXiv preprint arXiv:2202.03286 (2022)
2022 arXiv
-
[69]
Franziska Plessow, Andrea Kiesel, and Clemens Kirschbaum. 2012. The stressed prefrontal cortex and goal-directed behaviour: acute psychosocial stress impairs the flexible implementation of task goals. Experimental brain research 216 (2012), 397–408
2012
-
[70]
GE Appliances Pressroom. 2023. GE Appliances Helps Consumers Create Personalized Recipes from the Food in Their Kitchen with Google Cloud’s Generative AI. https://pressroom.geappliances.com/news/ge-appliances-helps- consumers-create-personalized-recipes-from-the-food-in-their-...
2023
-
[71]
Gavin Rees. 2017. Handling Traumatic Imagery: Developing a Standard Op- erating Procedure. (4 April 2017). https://dartcenter.org/resources/handling- traumatic-imagery-developing-standard-operating-procedure
2017
-
[72]
Joe Regalia. 2024. From Briefs to Bytes: How Generative AI is Transforming Legal Writing and Practice. Tulsa L. Rev. 59 (2024), 193
2024
-
[73]
Julie Repper and Tim Carter. 2011. A review of the literature on peer support in mental health services. Journal of mental health 20, 4 (2011), 392–411
2011
-
[74]
Sarah T Roberts. 2019. Behind the screen. Yale University Press
2019
-
[75]
Roediger
D.R. Roediger. 1999.The Wages of Whiteness: Race and the Making of the American Working Class. Verso. https://books.google.com/books?id=PwyMmV1_0kMC
1999
-
[76]
Kevin Roose. 2023. How ChatGPT Kicked Off an A.I. Arms Race. The New York Times (3 February 2023). https://www.nytimes.com/2023/02/03/technology/ chatgpt-openai-artificial-intelligence.html
2023
-
[77]
Kevin Roose. 2024. Can A.I. Be Blamed for a Teen’s Suicide? The New York Times (23 October 2024). https://www.nytimes.com/2024/10/23/technology/ characterai-lawsuit-teen-suicide.html
2024
-
[78]
Joelle Ré Arp-Dunham. 2024. Introduction. In Stanislavsky and Intimacy, Joelle Ré Arp-Dunham (Ed.). Routledge
2024
-
[79]
SAG-AFTRA Health Plan. 2024. Earned Eligibility. https://www.sagaftraplans. org/health/eligibility/earned-eligibility
2024
-
[80]
Vlad Savov. 2025. Samsung Adds Generative AI to World’s Best-Selling TV Lineup. Bloomberg (5 January 2025). https://www.bloomberg.com/news/ articles/2025-01-06/samsung-adds-generative-ai-to-world-s-best-selling-tv- lineup
2025
-
[81]
Angela M Schöpke-Gonzalez, Shubham Atreja, Han Na Shin, Najmin Ahmed, and Libby Hemphill. 2024. Why do volunteer content moderators quit? Burnout, conflict, and harmful behaviors. New Media & Society 26, 10 (2024), 5677–5701
2024
-
[82]
Valentin Schwind, Pascal Knierim, Nico Haas, and Niels Henze. 2019. Using pres- ence questionnaires in virtual reality. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–12
2019
-
[84]
Mark Cariston Seton. 2006. ‘Post-Dramatic’ Stress: Negotiating Vulnerability for Performance. In Proceedings of the 2006 Annual Conference of the Australasian Association for Drama, Theatre and Performance Studies. Australasian Association for Drama, Theatre and Performance St...
2006
-
[85]
Mark Cariston Seton. 2013. Traumas of acting physical and psychological violence: How fact and fiction shape bodies for better or worse. Performing Ethos 4, 1 (2013), 25–40
2013
-
[86]
Reham A Hameed Shalaby and Vincent IO Agyapong. 2020. Peer support in mental health: literature review. JMIR mental health 7, 6 (2020), e15572
2020
-
[87]
Shane Sinclair, Shelley Raffin-Bouchal, Lorraine Venturato, Jane Mijovic- Kondejewski, and Lorraine Smith-MacDonald. 2017. Compassion fatigue: A meta-narrative review of the healthcare literature. International journal of nursing studies 69 (2017), 9–24
2017
-
[88]
Jasmeet Singh, Maria Karanika-Murray, Thom Baguley, and John Hudson. 2020. A systematic review of job demands and resources associated with compassion fatigue in mental health professionals. International Journal of Environmental Research and Public Health 17, 19 (2020), 6987
2020
-
[89]
Jessica Slade and Emma Alleyne. 2023. The psychological impact of slaughter- house employment: A systematic literature review. Trauma, Violence, & Abuse 24, 2 (2023), 429–440
2023
-
[91]
Jordan Sloshower. 2013. Capturing suffering: Ethical considerations of bearing witness and the use of photography. The International Journal of the Image 3, 2 (2013), 11
2013
-
[92]
Ruth Spence, Antonia Bifulco, Paula Bradbury, Elena Martellozzo, and Jeffrey DeMarco. 2023. The psychological impacts of content moderation on con- tent moderators: A qualitative study. Cyberpsychology: Journal of Psychosocial FAccT ’25, June 23–26, 2025, Athens, Greece Pendse...
2023
-
[93]
Ruth Spence, Antonia Bifulco, Paula Bradbury, Elena Martellozzo, and Jeffrey DeMarco. 2024. Content moderator mental health, secondary trauma, and well- being: A cross-sectional study. Cyberpsychology, Behavior, and Social Networking 27, 2 (2024), 149–155
2024
-
[94]
Ruth Spence, Amy Harrison, Paula Bradbury, Paul Bleakley, Elena Martellozzo, and Jeffrey DeMarco. 2023. Content moderators’ strategies for coping with the stress of moderating content online. Journal of Online Trust and Safety 1, 5 (2023)
2023
-
[95]
Ruth Spence, Elena Martellozzo, and Jeffrey DeMarco. 2024. Content moderator coping strategies: Associations with psychological distress, secondary trauma, and well-being. Journal of Media Psychology: Theories, Methods, and Applications (2024)
2024
-
[96]
Constantin Stanislavski and Elizabeth Reynolds Hapgood. 2012. Creating a role. Routledge
2012
-
[97]
Miriah Steiger, Timir J Bharucha, Wilfredo Torralba, Marlyn Savio, Priyanka Manchanda, and Rachel Lutz-Guevara. 2022. Effects of a novel resiliency train- ing program for social media content moderators. In Proceedings of Seventh International Congress on Information and Commu...
2022
-
[98]
Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of content moderators: the emotional labor of commercial moderation and avenues for improving support. In Pro- ceedings of the 2021 CHI conference on hum...
2021
-
[99]
David Thiel, Melissa Stroebel, and Rebecca Portnoff. 2023. Generative ML and CSAM: Implications and Mitigations
2023
-
[100]
David Turgoose and Lucy Maddox. 2017. Predictors of compassion fatigue in mental health professionals: A narrative review.Traumatology 23, 2 (2017), 172
2017
-
[101]
Rebecca Umbach, Nicola Henry, Gemma Faye Beard, and Colleen M Berryessa
-
[102]
John Villasenor. 2023. Generative Artificial Intelligence and the practice of Law: Impact, opportunities, and risks. Minn. JL Sci. & Tech. 25 (2023), 25
2023
-
[103]
In Proceedings of the CHI Conference on Human Factors in Computing Systems
Non-Consensual Synthetic Intimate Imagery: Prevalence, Attitudes, and Knowledge in 10 Countries. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–20
-
[104]
Stefan Weber, David Weibel, and Fred W Mast. 2021. How to get there when you are there already? Defining presence in virtual reality and the importance of perceived realism. Frontiers in psychology 12 (2021), 628298
2021
-
[105]
Zhenhua Wang, Wei Xie, Baosheng Wang, Enze Wang, Zhiwen Gui, Shuoy- oucheng Ma, and Kai Chen. 2024. Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology. arXiv preprint arXiv:2402.15690 (2024)
2024 arXiv
-
[106]
Irvin D Yalom. 2002. The gift of therapy: An open letter to a new generation of therapists and their patients. (No Title) (2002)
2002
-
[107]
Harry Yi-Jui Wu. 2021. Mad by the millions: mental disorders and the early years of the World Health Organization . MIT Press
2021
-
[109]
Micah Zenko. 2015. Red Team: How to succeed by thinking like the enemy . Basic Books
2015
-
[2024]
arXiv preprint arXiv:2402.14872 (2024)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs. arXiv preprint arXiv:2402.14872 (2024)
2024 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.