Pith. sign in

REVIEW 3 cited by

FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.13507 v2 pith:TAMEUIHT submitted 2024-03-20 cs.CV

classification cs.CV
keywords adversarialllmsvideo-basedfmm-attackmulti-modalattackflow-basedinduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the remarkable performance of video-based large language models (LLMs), their adversarial threat remains unexplored. To fill this gap, we propose the first adversarial attack tailored for video-based LLMs by crafting flow-based multi-modal adversarial perturbations on a small fraction of frames within a video, dubbed FMM-Attack. Extensive experiments show that our attack can effectively induce video-based LLMs to generate incorrect answers when videos are added with imperceptible adversarial perturbations. Intriguingly, our FMM-Attack can also induce garbling in the model output, prompting video-based LLMs to hallucinate. Overall, our observations inspire a further understanding of multi-modal robustness and safety-related feature alignment across different modalities, which is of great importance for various large multi-modal models. Our code is available at https://github.com/THU-Kingmin/FMM-Attack.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Large multimodal models are shown to rely on prior knowledge and text cues rather than video order under temporal inconsistency, and a benchmark plus preference-optimization method partially correct this.

  2. Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Coordinated misleading text descriptions of video, audio, and meaning flip the appropriateness labels assigned by most multimodal LLMs in about 90% of test videos.

  3. Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A survey that organizes audio and video AI security research into adversarial, backdoor, and jailbreak attacks, with extra attention to multimodal large language models.

Pith tools