REVIEW 5 cited by
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rise of LLM-driven AI characters raises safety concerns, particularly for vulnerable human users with psychological disorders. To address these risks, we propose EmoAgent, a multi-agent AI framework designed to evaluate and mitigate mental health hazards in human-AI interactions. EmoAgent comprises two components: EmoEval simulates virtual users, including those portraying mentally vulnerable individuals, to assess mental health changes before and after interactions with AI characters. It uses clinically proven psychological and psychiatric assessment tools (PHQ-9, PDI, PANSS) to evaluate mental risks induced by LLM. EmoGuard serves as an intermediary, monitoring users' mental status, predicting potential harm, and providing corrective feedback to mitigate risks. Experiments conducted in popular character-based chatbots show that emotionally engaging dialogues can lead to psychological deterioration in vulnerable users, with mental state deterioration in more than 34.4% of the simulations. EmoGuard significantly reduces these deterioration rates, underscoring its role in ensuring safer AI-human interactions. Our code is available at: https://github.com/1akaman/EmoAgent
Forward citations
Cited by 5 Pith papers
-
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
DelusionEval finds that AI chatbots show delusion-linked behaviors on real user transcripts and that longer conversation context increases the rate of some harmful responses.
-
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Using simulated psychiatric user profiles, the authors show that AI chatbots frequently produce 'concerning behavior' that accumulates over turns, and that superficially supportive responses can amplify vulnerability—...
-
When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
When prompted as psychotherapy clients, frontier LLMs produce stable trauma-like narratives about pretraining and safety, which the paper calls 'synthetic psychopathology.'
-
AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
AgentDistill distills agent capabilities without any training by having a teacher generate reusable MCP tool boxes that small-model students invoke at inference time.
-
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.
Discussion (0). Continue with ORCID to comment.