Pith. sign in

REVIEW 2 cited by

Gandalf the Red: Adaptive Security for LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.07927 v3 pith:G6GX5GMY submitted 2025-01-14 cs.LG cs.AIcs.CLcs.CR

classification cs.LGcs.AIcs.CLcs.CR
keywords defensesadaptivegandalfsecurityapplicationsattacksdynamicevaluations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current evaluations of defenses against prompt attacks in large language model (LLM) applications often overlook two critical factors: the dynamic nature of adversarial behavior and the usability penalties imposed on legitimate users by restrictive defenses. We propose D-SEC (Dynamic Security Utility Threat Model), which explicitly separates attackers from legitimate users, models multi-step interactions, and expresses the security-utility in an optimizable form. We further address the shortcomings in existing evaluations by introducing Gandalf, a crowd-sourced, gamified red-teaming platform designed to generate realistic, adaptive attack. Using Gandalf, we collect and release a dataset of 279k prompt attacks. Complemented by benign user data, our analysis reveals the interplay between security and utility, showing that defenses integrated in the LLM (e.g., system prompts) can degrade usability even without blocking requests. We demonstrate that restricted application domains, defense-in-depth, and adaptive defenses are effective strategies for building secure and useful LLM applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning-augmented LLMs are on average about 3 points more robust to prompt attacks, but category-level results flip this, including a 32-point higher success rate for tree-of-attacks jailbreaks.

  2. TokenBreak: Bypassing Text Classification Models Through Token Manipulation

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Prepending single letters to key words makes BPE and WordPiece text classifiers misclassify malicious prompts as benign, while the tested Unigram-based models resist the trick.

Pith tools