Pith. sign in

REVIEW 1 cited by

Towards Healthy AI: Large Language Models Need Therapists Too

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00416 v1 pith:SCSH3OEA submitted 2023-04-02 cs.AI cs.CLcs.CYcs.HCcs.LG

Towards Healthy AI: Large Language Models Need Therapists Too

classification cs.AI cs.CLcs.CYcs.HCcs.LG
keywords chatbotsframeworkhealthysafeguardgptbehaviorsconversationsdevelopmentethical
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advances in large language models (LLMs) have led to the development of powerful AI chatbots capable of engaging in natural and human-like conversations. However, these chatbots can be potentially harmful, exhibiting manipulative, gaslighting, and narcissistic behaviors. We define Healthy AI to be safe, trustworthy and ethical. To create healthy AI systems, we present the SafeguardGPT framework that uses psychotherapy to correct for these harmful behaviors in AI chatbots. The framework involves four types of AI agents: a Chatbot, a "User," a "Therapist," and a "Critic." We demonstrate the effectiveness of SafeguardGPT through a working example of simulating a social conversation. Our results show that the framework can improve the quality of conversations between AI chatbots and humans. Although there are still several challenges and directions to be addressed in the future, SafeguardGPT provides a promising approach to improving the alignment between AI chatbots and human values. By incorporating psychotherapy and reinforcement learning techniques, the framework enables AI chatbots to learn and adapt to human preferences and values in a safe and ethical way, contributing to the development of a more human-centric and responsible AI.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

    cs.AI 2026-07 conditional novelty 6.0

    Source-routed dual-teacher top-K KL distillation realigns misaligned LLMs with less template dependence and less task collapse than rollback, RESTA, soft-SFT, and SSRD.