LoRAScan detects trigger-bearing prompts for backdoored LoRA adapters by monitoring low-variance down-projection activation sites and rejecting outlier spikes at inference time.
Proceedings of the 2021 conference on empirical methods in natural language processing , pages=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes
LoRAScan detects trigger-bearing prompts for backdoored LoRA adapters by monitoring low-variance down-projection activation sites and rejecting outlier spikes at inference time.