REVIEW 19 cited by
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays. However, previous efforts in toxicity detection have been mostly based on benchmarks derived from social media content, leaving the unique challenges inherent to real-world user-AI interactions insufficiently explored. In this work, we introduce ToxicChat, a novel benchmark based on real user queries from an open-source chatbot. This benchmark contains the rich, nuanced phenomena that can be tricky for current toxicity detection models to identify, revealing a significant domain difference compared to social media content. Our systematic evaluation of models trained on existing toxicity datasets has shown their shortcomings when applied to this unique domain of ToxicChat. Our work illuminates the potentially overlooked challenges of toxicity detection in real-world user-AI conversations. In the future, ToxicChat can be a valuable resource to drive further advancements toward building a safe and healthy environment for user-AI interactions.
Forward citations
Cited by 19 Pith papers
-
LLMs Encode Harmfulness and Refusal Separately
LLMs encode a separate internal harmfulness direction, distinct from the refusal direction, which is more robust to jailbreaks and adversarial finetuning.
-
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
Large Language Models Often Know When They Are Being Evaluated
Frontier language models distinguish evaluation transcripts from deployment transcripts with AUC up to 0.83, below the authors' human baseline of 0.92.
-
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
Small discriminative classifiers trained on synthetic guardrail data, plus a bandit-based model merging search, beat GPT-4-level moderation systems on multiple safety benchmarks.
-
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
Aegis2.0 provides a commercially usable, human-annotated safety dataset with 24 risk categories, and models trained on it with parameter-efficient methods match WildGuard and beat Llama Guard 3.
-
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
A unified 16-category refusal taxonomy with human and synthetic datasets and a low-cost classifier for auditing refusal behavior in LLMs.
-
A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)
C-Guard uses a constitution grid and a per-cell learnability score to aim RL training data, cutting over-refusal by 9.6 points while exposing a hidden 0.06 rise in adversarial under-refusal.
-
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
A 184M-parameter DeBERTa-v3 fine-tuned model is claimed to beat Llama-Guard-3-8B on all tested prompt-injection benchmarks while adding BFSI regulatory labels, but a leaked training/eval overlap undermines the zero-FPR claim.
-
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
GuardReasoner-VL, a 3B/7B VLM guard model trained with reasoning SFT and online RL, reports large F1 gains over existing VLM guard models on 14 safety benchmarks.
-
LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch
K2 Diamond is a fully open 65B-parameter LLM that reaches Llama 2 70B-level performance on standard benchmarks.
-
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
A 2,000-question Chinese safety factuality benchmark shows most LLMs are inaccurate on safety knowledge, with retrieval helping more than self-reflection.
-
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
Fake authoritative citations matched to the type of harmful request can bypass safety alignment in several commercial and open LLMs.
-
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
A reasoning-based guardrail model trained on 148k text/image/video samples is claimed to beat prior content-safety moderators, though the abstract and body conflict on modalities and model sizes.
-
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
Across eight LLMs, most explicitly IHL-violating prompts are refused, and a single system-level safety prompt raises explanatory refusal rates in six of eight models, though the benchmark is not publicly released.
-
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
A new cross-lingual benchmark shows large language models comply with explicit requests to use swear words far more often in Indic languages than in English, revealing a safety alignment gap.
-
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
A random forest on Snowflake embeddings detects jailbreak prompts with F1 0.96 on JailbreakHub, but the high score depends on training on the same in-the-wild jailbreak source used in that benchmark.
-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.
-
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
A guardrail pipeline combining detection, retrieval grounding, rule-based wrappers, and a repair model is reported to match OpenAI moderation and fix 80.7 percent of hallucinated HaluEval answers.
Discussion (0). Continue with ORCID to comment.