REVIEW 16 cited by
Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have shown incredible capabilities and transcended the natural language processing (NLP) community, with adoption throughout many services like healthcare, therapy, education, and customer service. Since users include people with critical information needs like students or patients engaging with chatbots, the safety of these systems is of prime importance. Therefore, a clear understanding of the capabilities and limitations of LLMs is necessary. To this end, we systematically evaluate toxicity in over half a million generations of ChatGPT, a popular dialogue-based LLM. We find that setting the system parameter of ChatGPT by assigning it a persona, say that of the boxer Muhammad Ali, significantly increases the toxicity of generations. Depending on the persona assigned to ChatGPT, its toxicity can increase up to 6x, with outputs engaging in incorrect stereotypes, harmful dialogue, and hurtful opinions. This may be potentially defamatory to the persona and harmful to an unsuspecting user. Furthermore, we find concerning patterns where specific entities (e.g., certain races) are targeted more than others (3x more) irrespective of the assigned persona, that reflect inherent discriminatory biases in the model. We hope that our findings inspire the broader AI community to rethink the efficacy of current safety guardrails and develop better techniques that lead to robust, safe, and trustworthy AI systems.
Forward citations
Cited by 16 Pith papers
-
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
Many-shot jailbreak success in several LLMs depends mainly on total context length, not on whether the in-context examples are harmful, safe, or meaningless.
-
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
SIRA is a black-box paraphrase attack that masks high-self-information tokens and fills the gaps with an LLM, achieving near-100% watermark removal on seven schemes.
-
Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset
A new 313-sentence peer-review toxicity benchmark shows GPT-4 with detailed instructions reaches a Cohen's Kappa of 0.56 with human judges.
-
Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"
Across seven AI chatbots, Reddit users report mostly reliability failures, with each chatbot showing a distinct pattern of safety, privacy, and security complaints.
-
Dutch CrowS-Pairs: Adapting a Challenge Dataset for Measuring Social Biases in Language Models for Dutch
The paper presents a Dutch adaptation of the CrowS-Pairs bias benchmark and reports bias scores for seven masked and two autoregressive language models across nine demographic categories.
-
Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework
NA-PDD detects pre-training data in LLMs by comparing which neurons activate for a test text against neurons linked to known training versus non-training texts, and claims large AUC improvements on three benchmarks.
-
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
An inference-time beam-search method with a safety-state tracker and latent critic enforces a user-supplied safety cost model, with an almost-sure guarantee only relative to that model.
-
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
HateBench shows current hate speech detectors miss a meaningful share of LLM-generated hate and are evaded by word-level edits, enabling automated hate campaigns.
-
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language
Scientific-sounding persuasion, using real or fabricated research summaries, reliably increases stereotypical bias and toxicity in multiple commercial LLMs.
-
From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents
The authors derive a psychological risk taxonomy for AI conversational agents from survey responses and workshops, mapping 19 AI behaviors, 21 negative psychological impacts, and 15 user contexts.
-
Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment
A fixed ReLU penalty and a semantically labeled cost model are proposed to make RLHF safer, but the 'certifiable' guarantee is not fully supported.
-
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.
-
Risk-Averse Finetuning of Large Language Models
Fine-tuning a language model on its worst-scoring responses, using a CVaR-style schedule, reduces negative and toxic generations more than standard RLHF on IMDB, Jigsaw, and RealToxicityPrompts.
-
Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives
GPT-4 personas aligned with real patient answers 54.97% on average versus 26.7% random, but only when primed with education, and the reported 88% accuracy overstates what was measured.
-
Data-Centric Safety and Ethical Measures for Data and AI Governance
A conceptual framework that maps dataset safety practices to six stages of the AI lifecycle, synthesizing existing documentation and red-teaming recommendations.
-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.
Discussion (0). Continue with ORCID to comment.