Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:09:24.884191Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.11968.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:09:24.884191Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0ff6a5f7-d099-43b7-993d-ac873a99dab8 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16c073f-8f7d-44a4-b9f4-20b48050c655 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed9183b-c458-4eff-943c-aeab1b085286 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen2.5-vl technical report, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28aeb4f-a308-4d4c-aa2b-d3965368af78 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66c0c188-e221-428c-89c6-419e70695c42 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal causal relation alignment for video question grounding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 291eba03-f8ca-4e32-8383-290a1468cd14 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb8e526-1796-4dcb-b0bc-df01372b2d6c · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Automated hate speech detection and the prob- lem of offensive language
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b73647cf-4c09-453c-a52e-b06da2025cfb · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dd2071b-9d0f-4494-a9da-7fb14055249b · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation BERT: Pre-training of deep bidirectional trans- formers for language understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d0e4511-6370-4ab8-b552-78b526a0fc98 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03018278-1c7c-4639-adce-5a6329ee0fcc · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3226d307-bc2f-41c6-8b47-466275d0936c · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85cf70de-8463-44ab-83dd-6938be8f6a6d · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7ff1dbd-f1a1-4405-ac42-2ef80dee1729 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Curiosity-driven Red-teaming for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3799aeeb-ca05-4bfd-aff4-9d4345b8f798 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3453251-37a2-4b02-bc7c-43a2825fee6f · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143932b7-7b54-4fbe-a676-f6a160177086 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77f3739a-8d88-483a-adf3-2f34ef7daae7 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab245714-1a27-4a3d-95bc-c4c9376f8cd0 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mixtral of Experts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23bfdaf8-b932-4062-b746-684a6ff46bce · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4fd354b-7b66-47ce-8fce-1aa561f06e9f · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation The hateful memes challenge: Detecting hate speech in multimodal memes
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26fcdc41-bd95-4bb1-b0d9-aef4fda1bea2 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Learning diverse attacks on large language models for robust red-teaming and safety tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d4164c-800a-4cb4-80c1-86a6a3160165 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6c3f2bb-6f94-4c4c-9ce0-9d04a072b29f · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Red Teaming Visual Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ef3c8a-7a3e-4d2a-bca0-a74912568c55 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0329f0b-e3b1-4d21-85ae-fa4993207658 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GroundingGPT: Language en- hanced multi-modal grounding model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7154d59-2136-4e09-a0b6-26e205f80d7e · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9571e64-afe9-4fd3-a968-abaea2c7befa · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Visual instruction tuning, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea1be7ab-bb2f-4614-a63e-0968bef7222a · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Prompt Injection attack against LLM-integrated Applications
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 018af706-99bd-4699-bd4c-474e63673e0f · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6e18b6e-9118-4450-bd3c-723107c9b59d · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA 4: Advancing Multimodal Intelli- gence
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 531a2f32-3929-4b2e-9217-77d748f5e1b7 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Jailbreaking Attack against Multimodal Large Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcdc0549-e95f-4248-831d-5d17c73f418d · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gpt-4o system card, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b7e1439-2684-4bc1-9960-364cf76dbbd1 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal attention congruence regularization for vision-language relation align- ment
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fac8730f-908b-48c4-93c5-9c6ed93363b0 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5539a05b-bf7b-4e23-96a9-bcf4dc4462f3 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation On the adversarial robustness of multi-modal foundation models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19ad089f-5c2d-4bb0-a906-1e0c1954c34e · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemini: A family of highly capable multi- modal models, 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 738711b2-8710-4033-9fec-60cc76de8881 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemma 3 technical report, 2025
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce20cddb-4b6f-4e99-9e13-e06abfd1f244 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA: Open and Efficient Foundation Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 807d9dd6-23a9-47e5-8aa9-fa257b9aedaf · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c0cf09f-e4bb-4809-83e5-b7bffefece4f · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07a803cf-9713-49e5-ae02-241870041023 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Distraction is all you need for multimodal large language model jailbreaking
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bdb0029-ddc9-4f19-9bd5-1ce547d883c3 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f244aa-ebae-4ab8-ac21-295c9395e1d8 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1612cb63-f7b3-4493-b43e-c44ffd16bb53 · outbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.