Pith. sign in

REVIEW 10 cited by

Coercing LLMs to do and reveal (almost) anything

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.14020 v1 pith:DY4HZAFB submitted 2024-02-21 cs.LG cs.CLcs.CR

classification cs.LGcs.CLcs.CR
keywords attacksllmsadversarialattackmodelalmostanalyzeanything
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

It has recently been shown that adversarial attacks on large language models (LLMs) can "jailbreak" the model into making harmful statements. In this work, we argue that the spectrum of adversarial attacks on LLMs is much larger than merely jailbreaking. We provide a broad overview of possible attack surfaces and attack goals. Based on a series of concrete examples, we discuss, categorize and systematize attacks that coerce varied unintended behaviors, such as misdirection, model control, denial-of-service, or data extraction. We analyze these attacks in controlled experiments, and find that many of them stem from the practice of pre-training LLMs with coding capabilities, as well as the continued existence of strange "glitch" tokens in common LLM vocabularies that should be removed for security reasons.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

    cs.CR 2025-10 conditional novelty 6.0 of 10

    A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.

  2. BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profit

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A tunable 'overthinking' backdoor can be implanted into large reasoning models so that repeating a trigger word N times forces N extra reasoning steps, increasing token use several-fold without hurting accuracy.

  3. How much do language models memorize?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A compression-based measurement puts GPT-style model memorization capacity at roughly 3.6 bits per parameter, with membership inference success following a sigmoid in the dataset-to-capacity ratio.

  4. The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Fine-tuned LLMs separate roles via task-type and begin-of-text shortcuts; PFT, which inserts a gap into position IDs during fine-tuning, reduces those shortcuts and improves closed-domain attack robustness.

  5. Has My System Prompt Been Used? Large Language Model Prompt Membership Inference

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A permutation test on BERT embeddings of LLM outputs can detect, with statistical significance, when response distributions differ because a chat service uses a different system prompt than a candidate prompt.

  6. Towards Action Hijacking of Large Language Model-based Agent

    cs.CR 2024-12 conditional novelty 6.0 of 10

    A RAG-based LLM application can be induced to assemble harmful SQL, code, or medical action plans from knowledge already stored in its database, with the user prompt itself carrying no forbidden words.

  7. Obfuscated Activations Bypass LLM Latent-Space Defenses

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Obfuscation attacks that jointly optimize for target behavior and for low monitor scores bypass sparse autoencoders, probes, and OOD detectors on LLMs, while performance degrades mainly on hard tasks like writing correct SQL.

  8. Universal and Context-Independent Triggers for Precise Control of LLM Outputs

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A single trained token pair inserted around any target text forces Qwen-2 7B and Llama-3.1 8B to output that text on 54 to 75 percent of unseen prompts.

  9. Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    Non-targeted merge-list-free BPE inference causes minimal downstream performance loss, unlike targeted merge-list corruption.

  10. Steering Language Model Refusal with Sparse Autoencoders

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Boosting one SAE 'refusal' feature in Phi-3 Mini and Llama 3.1 raises refusal rates on unsafe and safe prompts alike while sharply reducing MMLU, TruthfulQA, and GSM8K accuracy.

Pith tools