Pith. sign in

Prp: Propagating universal perturbations to attack large language model guard-rails

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

citation-role summary

method 1

citation-polarity summary

fields

cs.CR 2 cs.LG 2

years

2026 2 2024 2

roles

method 1

polarities

background 1

clear filters

representative citing papers

Distilling Safe LLM Systems via Soft Prompts for On Device Settings

cs.LG · 2026-06-08 · unverdicted · novelty 4.0

Soft prompt distillation with total variation and KL divergence transfers safety behaviors from guard models to on-device LLMs and outperforms LoRA adapters, steering vectors, and direct optimization in safety-usefulness trade-offs with minimal inference cost.

Agent Security is a Systems Problem

cs.CR · 2026-05-18 · unverdicted · novelty 4.0 · 2 refs

The paper argues that agent security is best addressed as a systems problem by applying principles from operating systems, networks, and formal methods rather than relying solely on model robustness improvements.

citing papers explorer

Showing 1 of 1 citing paper after filters.