LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing

Chunyi Zhou; Jiahao Chen; Junhao Li; Qingming Li; Shouling Ji; Tianyu Du; Yi Jiang; Yiming Wang; Yong Yang

arxiv: 2507.07056 · v2 · pith:YXBA6UUAnew · submitted 2025-07-05 · 💻 cs.CR · cs.LG

LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing

Jiahao Chen , Junhao Li , Yiming Wang , Yong Yang , Yi Jiang , Chunyi Zhou , Qingming Li , Tianyu Du

show 1 more author

Shouling Ji

This is my paper

classification 💻 cs.CR cs.LG

keywords loramodelslorashieldpersonalizedadversarialbenigncriticaldata-free

0 comments

read the original abstract

The proliferation of Low-Rank Adaptation (LoRA) models has democratized personalized text-to-image generation, enabling users to share lightweight models (e.g., personal portraits) on platforms like Civitai and Liblib. However, this "share-and-play" ecosystem introduces critical risks: benign LoRAs can be weaponized by adversaries to generate harmful content (e.g., political, defamatory imagery), undermining creator rights and platform safety. Existing defenses like concept-erasure methods focus on full diffusion models (DMs), neglecting LoRA's unique role as a modular adapter and its vulnerability to adversarial prompt engineering. To bridge this gap, we propose LoRAShield, the first data-free editing framework for securing LoRA models against misuse. Our platform-driven approach dynamically edits and realigns LoRA's weight subspace via adversarial optimization and semantic augmentation. Experimental results demonstrate that LoRAShield achieves remarkable effectiveness, efficiency, and robustness in blocking malicious generations without sacrificing the functionality of the benign task. By shifting the defense to platforms, LoRAShield enables secure, scalable sharing of personalized models, a critical step toward trustworthy generative ecosystems.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Customization under Fire: Plugin Poisoning in Text-to-Image Ecosystem
cs.CR 2026-06 unverdicted novelty 7.0

PoisonLoRA demonstrates ~100% attack success rates for stealthy LoRA poisoning via concept hijacking and task injection on real platforms, with robustness to base model transfer and multiple remixes.