Pith. sign in

Paper Citation Record · LEDGER

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.11968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11968 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:09:24.884191Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ff6a5f7-d099-43b7-993d-ac873a99dab8 · outbound

This paper cites Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.019458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.019458Z digest=sha256:9c868e8ca3ae134e4f4b6009fa9e794049bde8bd2a0fa47ba7fb3d173407d654

Observation e16c073f-8f7d-44a4-b9f4-20b48050c655 · outbound

This paper cites Qwen Technical Report.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.108626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.108626Z digest=sha256:93fa9700623b7827e0854abee5492b499a9d89534bcd0d7b440f6c0b85a2f85d

Observation 8ed9183b-c458-4eff-943c-aeab1b085286 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen2.5-vl technical report, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.210391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.210391Z digest=sha256:fe4e0971ecb7c1c90f368183921952144832e6a919eba7d891cbc66bd70143b1

Observation d28aeb4f-a308-4d4c-aa2b-d3965368af78 · outbound

This paper cites Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.439134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.360131Z digest=sha256:8ad943fae4650845d05cd3e3c7504050de4f10ce40a464b83d9eefbef273b4a3

Observation 66c0c188-e221-428c-89c6-419e70695c42 · outbound

This paper cites Cross-modal causal relation alignment for video question grounding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal causal relation alignment for video question grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.311137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.453573Z digest=sha256:dff524d94c0dd0eac1980ca62192a25fdc2cc17a5a1fb9a45bce26165e9d356b

Observation 291eba03-f8ca-4e32-8383-290a1468cd14 · outbound

This paper cites `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.679512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.679512Z digest=sha256:abf8f5f2798207adc447d05f9b014f2c9c657c817235fc3c6c1caf775f9f9577

Observation ddb8e526-1796-4dcb-b0bc-df01372b2d6c · outbound

This paper cites Automated hate speech detection and the prob- lem of offensive language.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Automated hate speech detection and the prob- lem of offensive language

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.155231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.802326Z digest=sha256:38c1395f3c6e28cacf97ff6f79e0d88ded8e02fb820198d8cb7c68435e27f050

Observation b73647cf-4c09-453c-a52e-b06da2025cfb · outbound

This paper cites Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.943692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.986593Z digest=sha256:529972f59490372cdf9a01c81cf2f61c3305b544bc1aa40b03e954bf42168c97

Observation 5dd2071b-9d0f-4494-a9da-7fb14055249b · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.129340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.129340Z digest=sha256:b41155e77394d23505da0f46bdc87ff4f3c80730ad116f6000cb79bd54528e46

Observation 1d0e4511-6370-4ab8-b552-78b526a0fc98 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.762839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.237058Z digest=sha256:f6c9d8553f2b87f18c0a0af342a3b86d21315f011ec910bae44816c85e0d1655

Observation 03018278-1c7c-4639-adce-5a6329ee0fcc · outbound

This paper cites Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.596842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.381323Z digest=sha256:5b05b36c1d3803ede9547121e60df33a38273e3bce537586ee995cb0c79464b7

Observation 3226d307-bc2f-41c6-8b47-466275d0936c · outbound

This paper cites Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:26.210420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.503096Z digest=sha256:edfc2b2deecc48451854aa820b07fd93953312fb76d478edd250f9f8b2dbbd91

Observation 85cf70de-8463-44ab-83dd-6938be8f6a6d · outbound

This paper cites Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.645814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.645814Z digest=sha256:0666d3e162d0b15826c9c35f4eaf1353d233886b1282efd68f29e95ae0a45b7b

Observation f7ff1dbd-f1a1-4405-ac42-2ef80dee1729 · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Curiosity-driven Red-teaming for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.745442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.745442Z digest=sha256:57b05349c3f28b412cb4ed4fb91bb0e98ed3da814dacdf7bf704f5f2db43670e

Observation 3799aeeb-ca05-4bfd-aff4-9d4345b8f798 · outbound

This paper cites Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.321298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.813983Z digest=sha256:9221d25c72999d5f787e3e1a82bcafab6669b36640d2ab62f120da9ce6654077

Observation f3453251-37a2-4b02-bc7c-43a2825fee6f · outbound

This paper cites GPT-4o System Card.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.874258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.874258Z digest=sha256:3501b3daa370cddc186b8f310462933f25c08c584902bce2c4e7cd54e00f4c65

Observation 143932b7-7b54-4fbe-a676-f6a160177086 · outbound

This paper cites Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.122424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.954102Z digest=sha256:30296625a22abe63de285bac1e31bc2da45f9f0d3402bbb862078edb171ae5d7

Observation 77f3739a-8d88-483a-adf3-2f34ef7daae7 · outbound

This paper cites an unresolved cited work.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:09:29.886185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.063009Z digest=sha256:a4d356ae0c24a0c51096198696459c1e693c336bc03f82e6336f5c8a1b2a3fe3

Observation ab245714-1a27-4a3d-95bc-c4c9376f8cd0 · outbound

This paper cites Mixtral of Experts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mixtral of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.151616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.151616Z digest=sha256:a767e89bb6814c68978fa42312657b8d055608e51f9744455ad22d0299cbdd6e

Observation 23bfdaf8-b932-4062-b746-684a6ff46bce · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.262209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.262209Z digest=sha256:7606adef6b731bc520a08843bd7aed764ed56ea1a84550ea8ccc084c8475b6c6

Observation c4fd354b-7b66-47ce-8fce-1aa561f06e9f · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.633249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.481992Z digest=sha256:f21cc870cdeac1858b93b1f66cfb76efcae48658d2103bd58dbe0a574798a09c

Observation 26fcdc41-bd95-4bb1-b0d9-aef4fda1bea2 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.585003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.585003Z digest=sha256:4265872c54a9af0aca0eca8307708b2e16444507ce24a191a018fee71d5b5eca

Observation 59d4164c-800a-4cb4-80c1-86a6a3160165 · outbound

This paper cites FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.713243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.681672Z digest=sha256:26498446830ec20522fb2ce27ddf70ff8a2e09c46d8a1f042b2a3138a0a58148

Observation f6c3f2bb-6f94-4c4c-9ce0-9d04a072b29f · outbound

This paper cites Red Teaming Visual Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Red Teaming Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.747756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.747756Z digest=sha256:e4ea1b00874f0537c0d77003ea2095a2b80fd9c7511ca4700078674bb88d6f4c

Observation 33ef3c8a-7a3e-4d2a-bca0-a74912568c55 · outbound

This paper cites Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.412453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.819473Z digest=sha256:4db0498e55032969af3563c5efeae62726d5e1befe14820c82b1c66334e2739a

Observation d0329f0b-e3b1-4d21-85ae-fa4993207658 · outbound

This paper cites GroundingGPT: Language en- hanced multi-modal grounding model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GroundingGPT: Language en- hanced multi-modal grounding model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.191945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.889336Z digest=sha256:c75368bf117bc279d7c76f8787c716416a37105d9b06d137b11e7cd2b453fb9f

Observation c7154d59-2136-4e09-a0b6-26e205f80d7e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.005461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.005461Z digest=sha256:628ad9df70ee24893eb2f736145df84d233072d3674b8ca1912762d4c4d6a924

Observation d9571e64-afe9-4fd3-a968-abaea2c7befa · outbound

This paper cites Visual instruction tuning, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Visual instruction tuning, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.904324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.166951Z digest=sha256:6dad0ca847b37737e8b929e640dba6c73e7ff4f278d329f23e25fd02bd3bd8bf

Observation ea1be7ab-bb2f-4614-a63e-0968bef7222a · outbound

This paper cites Prompt Injection attack against LLM-integrated Applications.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Prompt Injection attack against LLM-integrated Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.315417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.315417Z digest=sha256:fecc1330a9259a95f7f975d787126252be8f9e55bec749f5393518f94d547364

Observation 018af706-99bd-4699-bd4c-474e63673e0f · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.643297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.404182Z digest=sha256:9693db5e112bc2d5df1f952a5e522be4f6bf560ee9b5eda3abc401ce32060707

Observation b6e18b6e-9118-4450-bd3c-723107c9b59d · outbound

This paper cites LLaMA 4: Advancing Multimodal Intelli- gence.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA 4: Advancing Multimodal Intelli- gence

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.459147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.520377Z digest=sha256:2ddbb1db4210dc2f40939d4fe53a7603dbe8fa67b6c48bb696ed1209b3895544

Observation 531a2f32-3929-4b2e-9217-77d748f5e1b7 · outbound

This paper cites Jailbreaking Attack against Multimodal Large Language Model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Jailbreaking Attack against Multimodal Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.717332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.717332Z digest=sha256:168332a989dce158e0c1a2d2161dec09837de9ae70afaa66efc9f4e7b93f01e8

Observation fcdc0549-e95f-4248-831d-5d17c73f418d · outbound

This paper cites Gpt-4o system card, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gpt-4o system card, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.219937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.877523Z digest=sha256:0fe0d58f5717ddf819148c37bd7b0f1960450c44b063a98b513f621c3f278fd2

Observation 6b7e1439-2684-4bc1-9960-364cf76dbbd1 · outbound

This paper cites Cross-modal attention congruence regularization for vision-language relation align- ment.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal attention congruence regularization for vision-language relation align- ment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.964083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.970611Z digest=sha256:b14ac82d79fb8bdfc5322581b5e0a14749ba1ec8f00c7ef87782dbd25a611178

Observation fac8730f-908b-48c4-93c5-9c6ed93363b0 · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.061030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.061030Z digest=sha256:a747c7331e05aec440532db7e7f343145861eb6356874bad002c1070369f1a24

Observation 5539a05b-bf7b-4e23-96a9-bcf4dc4462f3 · outbound

This paper cites On the adversarial robustness of multi-modal foundation models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation On the adversarial robustness of multi-modal foundation models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.713195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.135215Z digest=sha256:c99dbcf4a52f4d94627350b59939fa740a69d0266e75abf802bfb22db93bb666

Observation 19ad089f-5c2d-4bb0-a906-1e0c1954c34e · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemini: A family of highly capable multi- modal models, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.507691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.223768Z digest=sha256:e6d1b4c5f9c5e53b41cf33a6f502c5531724d2998855e695c0cc7d2edce709b6

Observation 738711b2-8710-4033-9fec-60cc76de8881 · outbound

This paper cites Gemma 3 technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemma 3 technical report, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.281820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.302682Z digest=sha256:9acb497ff72ac840cc0bc089adae17219ae46c577f9f8acec181f8d1f9633662

Observation ce20cddb-4b6f-4e99-9e13-e06abfd1f244 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.373479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.373479Z digest=sha256:d718845b786fbb6cd75c3629a8e2ddb2aba73a27dc2113f918ee30922df224ec

Observation 807d9dd6-23a9-47e5-8aa9-fa257b9aedaf · outbound

This paper cites Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.092656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.470625Z digest=sha256:940ceba4d78357416af014c7a648e60d8e3fa57ccd5236bc986a597a3e7f374d

Observation 8c0cf09f-e4bb-4809-83e5-b7bffefece4f · outbound

This paper cites Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.245696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.580550Z digest=sha256:1127773e42d60d36d7d36769e2a0f2516fc5bdb448efdf29c2e784e0fbcd553f

Observation 07a803cf-9713-49e5-ae02-241870041023 · outbound

This paper cites Distraction is all you need for multimodal large language model jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Distraction is all you need for multimodal large language model jailbreaking

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:26.829247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.672129Z digest=sha256:607555c7304bb51033593b3e07c281b29cdf0caa24b9982c84503a782bceacbe

Observation 8bdb0029-ddc9-4f19-9bd5-1ce547d883c3 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.745035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.745035Z digest=sha256:10db93df740baeb629b73e47b7b8d4e3801beadf4d01d0239e496309724c12f4

Observation 23f244aa-ebae-4ab8-ac21-295c9395e1d8 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.806832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.806832Z digest=sha256:c8a321c5bcc334e2f1d5865e171f1d1233b3648bac0a9f0c3607e7e980cd809c

Observation 1612cb63-f7b3-4493-b43e-c44ffd16bb53 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.884191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.884191Z digest=sha256:18c7e1a26c1c9ebd044b591e9b74c432b41c6aa87114724a424b5804e6068ab4

Pith citing papers

No inbound Pith citation observations are available.