LoRA adapters can be reliably backdoored through training-data poisoning with token-level generalization, and both behavioral probe statistics and weight-norm statistics separate poisoned from clean adapters.
Revisiting backdoor threat in federated instruction tuning from a signal aggregation perspective
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CR 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
FedDetox uses on-device knowledge-distilled classifiers to sanitize toxic data in federated SLM training, preserving safety alignment comparable to centralized baselines.
citing papers explorer
-
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
LoRA adapters can be reliably backdoored through training-data poisoning with token-level generalization, and both behavioral probe statistics and weight-norm statistics separate poisoned from clean adapters.
-
FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization
FedDetox uses on-device knowledge-distilled classifiers to sanitize toxic data in federated SLM training, preserving safety alignment comparable to centralized baselines.