Pith. sign in

REVIEW 3 cited by

Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.06237 v3 pith:OK27BE6B submitted 2023-11-10 cs.CL cs.CRcs.HC

classification cs.CLcs.CRcs.HC
keywords llmsteamingactivityattackingpeopledefininggroundedhighly
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition of how and why people perform such attacks, defining LLM red-teaming based on extensive and diverse evidence. Using a formal qualitative methodology, we interviewed dozens of practitioners from a broad range of backgrounds, all contributors to this novel work of attempting to cause LLMs to fail. We focused on the research questions of defining LLM red teaming, uncovering the motivations and goals for performing the activity, and characterizing the strategies people use when attacking LLMs. Based on the data, LLM red teaming is defined as a limit-seeking, non-malicious, manual activity, which depends highly on a team-effort and an alchemist mindset. It is highly intrinsically motivated by curiosity, fun, and to some degrees by concerns for various harms of deploying LLMs. We identify a taxonomy of 12 strategies and 35 different techniques of attacking LLMs. These findings are presented as a comprehensive grounded theory of how and why people attack large language models: LLM red teaming.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

    cs.CR 2025-10 conditional novelty 6.0 of 10

    A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.

  2. What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

    cs.HC 2026-08 accept novelty 5.0 of 10

    A synthesis of 161 empirical studies finds that industry responsible AI practices have professionalized since 2019, yet persistent gaps in training, organizational support, and tailored interventions remain.

  3. Improved Large Language Model Jailbreak Detection via Pretrained Embeddings

    cs.CR 2024-12 reject novelty 4.0 of 10

    A random forest on Snowflake embeddings detects jailbreak prompts with F1 0.96 on JailbreakHub, but the high score depends on training on the same in-the-wild jailbreak source used in that benchmark.

Pith tools