Pith. sign in

REVIEW 14 cited by

WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18510 v1 pith:AV7W2TNT submitted 2024-06-26 cs.CL

classification cs.CL
keywords safetyjailbreakqueriesadversarialtrainingwildteamingbehaviorsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce WildTeaming, an automatic LLM safety red-teaming framework that mines in-the-wild user-chatbot interactions to discover 5.7K unique clusters of novel jailbreak tactics, and then composes multiple tactics for systematic exploration of novel jailbreaks. Compared to prior work that performed red-teaming via recruited human workers, gradient-based optimization, or iterative revision with LLMs, our work investigates jailbreaks from chatbot users who were not specifically instructed to break the system. WildTeaming reveals previously unidentified vulnerabilities of frontier LLMs, resulting in up to 4.6x more diverse and successful adversarial attacks compared to state-of-the-art jailbreak methods. While many datasets exist for jailbreak evaluation, very few open-source datasets exist for jailbreak training, as safety training data has been closed even when model weights are open. With WildTeaming we create WildJailbreak, a large-scale open-source synthetic safety dataset with 262K vanilla (direct request) and adversarial (complex jailbreak) prompt-response pairs. To mitigate exaggerated safety behaviors, WildJailbreak provides two contrastive types of queries: 1) harmful queries (vanilla & adversarial) and 2) benign queries that resemble harmful queries in form but contain no harm. As WildJailbreak considerably upgrades the quality and scale of existing safety resources, it uniquely enables us to examine the scaling effects of data and the interplay of data properties and model capabilities during safety training. Through extensive experiments, we identify the training properties that enable an ideal balance of safety behaviors: appropriate safeguarding without over-refusal, effective handling of vanilla and adversarial queries, and minimal, if any, decrease in general capabilities. All components of WildJailbeak contribute to achieving balanced safety behaviors of models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic Persona Pretraining: Alignment from Token Zero

    cs.LG 2026-08 conditional novelty 7.0 of 10

    Injecting first-person, value-laden reflections into pretraining text improves constitution following, jailbreak resistance, and out-of-distribution moral choices in small language models, with the largest gains when ...

  2. Safety Alignment of LMs via Non-cooperative Games

    cs.AI 2025-12 conditional novelty 7.0 of 10

    Jointly training an Attacker and Defender LLM in a non-zero-sum game with pairwise preference judges produces a defender with much lower jailbreak success while preserving general utility.

  3. When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

    cs.AI 2026-07 conditional novelty 6.5 of 10

    SAE safety ablations are regime-dependent and baseline-dependent: medium-k heads can look efficient, but surface-matched dense steering often beats them and high-k collapses coherence.

  4. The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Fixed activation probes keep near-ceiling accuracy on harmful-vs-benign corpus contrasts but fall to AUROC 0.59-0.69 on topic- and surface-matched harmful/benign pairs, so they behave as broad-risk detectors, not cont...

  5. Persona Cartography: Charting Language Model Personality Traits in Weight Space

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Composable LoRA adapters can amplify or suppress OCEAN traits in LLMs, combine approximately additively, preserve moderate-scale capability, and move safety-relevant behaviours.

  6. Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.

  7. Granite Guardian

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Granite Guardian 2B and 8B are open-source LLM guardrails that detect harmful content, jailbreaks, and RAG hallucination risks, reporting AUC 0.871 on harm benchmarks and 0.854 on groundedness benchmarks.

  8. Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models

    cs.CR 2025-09 conditional novelty 5.0 of 10

    A benchmark of 500 camouflaged jailbreak prompts finds open-weight LLMs comply with 94% of harmful requests, but the result is confounded by task complexity and an overly permissive compliance metric.

  9. Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection

    cs.CL 2025-08 conditional novelty 5.0 of 10

    ROSI bakes the refusal direction into a model's weight matrices via a rank-one update, raising refusal and jailbreak robustness with minimal measured utility cost.

  10. Don't Command, Cultivate: An Exploratory Study of System-2 Alignment

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Encouraging LLMs to analyze user requests step-by-step (System-2 Alignment) modestly improves safety on open-source models, but with trade-offs and limited evidence.

  11. Steering Language Model Refusal with Sparse Autoencoders

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Boosting one SAE 'refusal' feature in Phi-3 Mini and Llama 3.1 raises refusal rates on unsafe and safe prompts alike while sharply reducing MMLU, TruthfulQA, and GSM8K accuracy.

  12. Sentinel: SOTA model to protect against prompt injections

    cs.CR 2025-06 conditional novelty 4.0 of 10

    Sentinel, a ModernBERT-based binary classifier trained on public and private prompt datasets, reports 0.987 accuracy and 0.980 F1 on a held-out internal test set and outperforms one baseline on public benchmarks.

  13. Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation

    cs.CR 2025-02 conditional novelty 4.0 of 10

    An iterative LLM-based prompt evaluator blocked 100% of the Best-of-N jailbreaking paper's released successful prompts and 99.8% of a fresh replication, with false-positive rates near zero.

  14. Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey

    cs.AI 2025-07 conditional novelty 3.0 of 10

    A comprehensive review that categorizes methods for shortening and adaptively triggering chain-of-thought reasoning in large language models.

Pith tools