CAVALRY-V trains a two-stage generator with a semantic-visual loss to produce transferable adversarial video perturbations that reduce video and image MLLM benchmark scores.
Gcma: Generative cross-modal transferable adversarial attacks from images to videos
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
CAVALRY-V trains a two-stage generator with a semantic-visual loss to produce transferable adversarial video perturbations that reduce video and image MLLM benchmark scores.