ProjLens shows that backdoor parameters in MLLMs are encoded in low-rank subspaces of the projector and that embeddings shift toward the target direction with magnitude linear in input norm, activating only on poisoned samples.
Backdoorvlm: A benchmark for backdoor attacks on vision-language models
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CR 2years
2026 2representative citing papers
SkillTrojan backdoors skill-based agents by partitioning an encrypted payload across benign-looking skills that reassemble and execute only under a predefined trigger.
citing papers explorer
-
ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety
ProjLens shows that backdoor parameters in MLLMs are encoded in low-rank subspaces of the projector and that embeddings shift toward the target direction with magnitude linear in input norm, activating only on poisoned samples.
-
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
SkillTrojan backdoors skill-based agents by partitioning an encrypted payload across benign-looking skills that reassemble and execute only under a predefined trigger.