Pith. sign in

Paper Citation Record · LEDGER

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.11968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11968 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:09:24.884191Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ff6a5f7-d099-43b7-993d-ac873a99dab8 · outbound

This paper cites Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.019458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.019458Z digest=sha256:5f8c054d4883cff01d78b51b92ee1b06b66b0cb815c77f41df0764d52bb4afe3

Observation e16c073f-8f7d-44a4-b9f4-20b48050c655 · outbound

This paper cites Qwen Technical Report.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.108626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.108626Z digest=sha256:884aeb144cd598aef4f50f14b22bf7b4605bf2a8cfa31ef3cce8654ff95c0ba9

Observation 8ed9183b-c458-4eff-943c-aeab1b085286 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen2.5-vl technical report, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.210391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.210391Z digest=sha256:8601b6cf76a2e0bad0e72acb9a446e932544a8eedde5f2c10c632b0c68d51195

Observation d28aeb4f-a308-4d4c-aa2b-d3965368af78 · outbound

This paper cites Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.439134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:20.360131Z digest=sha256:424b713627d454226e8f71a80d3175d977c578cdfa1dbb9e415274f685ec79d2

Observation 66c0c188-e221-428c-89c6-419e70695c42 · outbound

This paper cites Cross-modal causal relation alignment for video question grounding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal causal relation alignment for video question grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.311137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:20.453573Z digest=sha256:8e4431e411453a4f701f8f3b69abc86520ec4ea255c7da4cde1a511db047d347

Observation 291eba03-f8ca-4e32-8383-290a1468cd14 · outbound

This paper cites `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.679512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.679512Z digest=sha256:beaac82c2998bbff63797d54dd4207e68e3a8a8ed976787d5f559267b038c2e1

Observation ddb8e526-1796-4dcb-b0bc-df01372b2d6c · outbound

This paper cites Automated hate speech detection and the prob- lem of offensive language.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Automated hate speech detection and the prob- lem of offensive language

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.155231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:20.802326Z digest=sha256:7e5432d802fea4af8b392157b733d65ce17eb27b450b604ee8a50d816a03c150

Observation b73647cf-4c09-453c-a52e-b06da2025cfb · outbound

This paper cites Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.943692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:20.986593Z digest=sha256:7f0d03ef1313958cce56212582b6c729c025794f652bcd839192ab838b1b858f

Observation 5dd2071b-9d0f-4494-a9da-7fb14055249b · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.129340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.129340Z digest=sha256:4602c4ac1b379ccecaf2bd529fe8df70e31e0f10959e3b3f15682a805f72362e

Observation 1d0e4511-6370-4ab8-b552-78b526a0fc98 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.762839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:21.237058Z digest=sha256:55adaaf371ba0a746fe7a03bdeff4f27af24654ce5d64b74f8a227e7c2d20cd2

Observation 03018278-1c7c-4639-adce-5a6329ee0fcc · outbound

This paper cites Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.596842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:21.381323Z digest=sha256:ea2cdeb5ce59c810e9b72d8703781ab1c2dd9b787e7d811adba4a7cc908dc422

Observation 3226d307-bc2f-41c6-8b47-466275d0936c · outbound

This paper cites Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:26.210420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:21.503096Z digest=sha256:ea381fb5389837cf04f597cafe46a71494d6270cc0b993f91210703506a29b8b

Observation 85cf70de-8463-44ab-83dd-6938be8f6a6d · outbound

This paper cites Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.645814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.645814Z digest=sha256:4668978e961b6d79c14b3e2883e5a859f64d532f9cc67fe0841a7d7b0c1e1b94

Observation f7ff1dbd-f1a1-4405-ac42-2ef80dee1729 · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Curiosity-driven Red-teaming for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.745442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.745442Z digest=sha256:8c288727b3f138ea01535d77813edf8a3636194dc89fbf4c9569f5733876db90

Observation 3799aeeb-ca05-4bfd-aff4-9d4345b8f798 · outbound

This paper cites Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.321298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:21.813983Z digest=sha256:7a7e503ab72eddfe5b3e60b796ecf7313a134c21bd78412e311b954636e898c4

Observation f3453251-37a2-4b02-bc7c-43a2825fee6f · outbound

This paper cites GPT-4o System Card.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.874258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.874258Z digest=sha256:75694aa7d99d298ba6880034b2871a0d55583d11ca9cd879c47e8750180984fd

Observation 143932b7-7b54-4fbe-a676-f6a160177086 · outbound

This paper cites Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.122424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:21.954102Z digest=sha256:c6783f64c44479470f8c665fbbc71f746dc5576139a2d26a9c3a46fec56868f5

Observation 77f3739a-8d88-483a-adf3-2f34ef7daae7 · outbound

This paper cites an unresolved cited work.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:09:29.886185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:22.063009Z digest=sha256:d18a662f695875c80bd4f4ff8e35a47e874ddb13f2a7cbe5755ef7f9ecadbd37

Observation ab245714-1a27-4a3d-95bc-c4c9376f8cd0 · outbound

This paper cites Mixtral of Experts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mixtral of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.151616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.151616Z digest=sha256:87b7cb8fedf8e1806a1e7a88afad5de39267048822401128b7574b552b54217a

Observation 23bfdaf8-b932-4062-b746-684a6ff46bce · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.262209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.262209Z digest=sha256:94df52da376a82a8337ca28b738222e70cbd87d6cbab5450bb0f5d86de18cc5e

Observation c4fd354b-7b66-47ce-8fce-1aa561f06e9f · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.633249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:22.481992Z digest=sha256:1ee8db2c09fb151f5210e14d025b1911607f6196089f95e6f7c92d7de06af92b

Observation 26fcdc41-bd95-4bb1-b0d9-aef4fda1bea2 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.585003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.585003Z digest=sha256:684bce80322fdef1b8ae26712625a1f8e61d9c94b7c5441df8b3bf6a5b3e8863

Observation 59d4164c-800a-4cb4-80c1-86a6a3160165 · outbound

This paper cites FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.713243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:22.681672Z digest=sha256:aa2bd227e0cc1d5f2c8ac56cc6c8c26beeffe8495e74518c2707a72cf7cf1a68

Observation f6c3f2bb-6f94-4c4c-9ce0-9d04a072b29f · outbound

This paper cites Red Teaming Visual Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Red Teaming Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.747756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.747756Z digest=sha256:7ca9867d631339df500d7550e9c929ae52581adfd8189bfbc8ca83184ff851ba

Observation 33ef3c8a-7a3e-4d2a-bca0-a74912568c55 · outbound

This paper cites Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.412453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:22.819473Z digest=sha256:9bb1a169e606bae2d60f7985e70a4c656d239cc3c9b9cfb59f6b41e1574a6e42

Observation d0329f0b-e3b1-4d21-85ae-fa4993207658 · outbound

This paper cites GroundingGPT: Language en- hanced multi-modal grounding model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GroundingGPT: Language en- hanced multi-modal grounding model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.191945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:22.889336Z digest=sha256:44da2fa40620d917203ca6e6f25e856436f7fd860f5ed3701ec53c6d10f582e4

Observation c7154d59-2136-4e09-a0b6-26e205f80d7e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.005461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.005461Z digest=sha256:cb989e0481f1577d785c2265bef8a6c29a208b66cc17f800d603ab8718db618e

Observation d9571e64-afe9-4fd3-a968-abaea2c7befa · outbound

This paper cites Visual instruction tuning, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Visual instruction tuning, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.904324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:23.166951Z digest=sha256:15889164aa963463a661937ffc225880798e686ef4e1bd37a25578672cd69c9c

Observation ea1be7ab-bb2f-4614-a63e-0968bef7222a · outbound

This paper cites Prompt Injection attack against LLM-integrated Applications.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Prompt Injection attack against LLM-integrated Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.315417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.315417Z digest=sha256:f95d2f1066cb1e9dc56859df26bbb886c39c6d382a5b7a8812e4c2ac8219833b

Observation 018af706-99bd-4699-bd4c-474e63673e0f · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.643297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:23.404182Z digest=sha256:b2c633b468fa85e91861ec01661dc6ddfb5bb9640d66f72b5bd740eec3af2cf8

Observation b6e18b6e-9118-4450-bd3c-723107c9b59d · outbound

This paper cites LLaMA 4: Advancing Multimodal Intelli- gence.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA 4: Advancing Multimodal Intelli- gence

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.459147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:23.520377Z digest=sha256:048ccd24d53e41417fa7d16f0ddf9787598986f225ea5b13bf3de58cd9e625c6

Observation 531a2f32-3929-4b2e-9217-77d748f5e1b7 · outbound

This paper cites Jailbreaking Attack against Multimodal Large Language Model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Jailbreaking Attack against Multimodal Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.717332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.717332Z digest=sha256:dba53c779e8a82b94e83cb8e17e69e047e2313e83f2d67da968c7fae88560f13

Observation fcdc0549-e95f-4248-831d-5d17c73f418d · outbound

This paper cites Gpt-4o system card, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gpt-4o system card, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.219937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:23.877523Z digest=sha256:b5483e8dc5701074e980de64a8b215088157c5f4d8826ebf2012e8d82ee0bf63

Observation 6b7e1439-2684-4bc1-9960-364cf76dbbd1 · outbound

This paper cites Cross-modal attention congruence regularization for vision-language relation align- ment.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal attention congruence regularization for vision-language relation align- ment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.964083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:23.970611Z digest=sha256:962f1f20faea876ea4206b8104c51ab2ba3d559ecd20d9dff5a85fce7bbf5968

Observation fac8730f-908b-48c4-93c5-9c6ed93363b0 · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.061030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.061030Z digest=sha256:6f1a30ac525b3c751ce32dd4c2dcac6b38a74abac9d78bffc649a0756a46e26f

Observation 5539a05b-bf7b-4e23-96a9-bcf4dc4462f3 · outbound

This paper cites On the adversarial robustness of multi-modal foundation models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation On the adversarial robustness of multi-modal foundation models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.713195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:24.135215Z digest=sha256:5e8d67da0d11214205ae9a302ebd79c4edb1445e1fdc404219e964be29f7819f

Observation 19ad089f-5c2d-4bb0-a906-1e0c1954c34e · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemini: A family of highly capable multi- modal models, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.507691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:24.223768Z digest=sha256:41e6206c637f15574a428670e2453473ee5a567f5aef879d07cf43654597e25e

Observation 738711b2-8710-4033-9fec-60cc76de8881 · outbound

This paper cites Gemma 3 technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemma 3 technical report, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.281820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:24.302682Z digest=sha256:148b48acb9071e7d5e4c587c993010811246fcd777e072d05af8142ca63225ac

Observation ce20cddb-4b6f-4e99-9e13-e06abfd1f244 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.373479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.373479Z digest=sha256:e6c3b298b81a462ef74ceef316ccb2718b3388609d02fc9b2e58c3a5ca7d91ea

Observation 807d9dd6-23a9-47e5-8aa9-fa257b9aedaf · outbound

This paper cites Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.092656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:24.470625Z digest=sha256:262f9d4958d30583f2cd14f3a7fc8f69039c24a94c701e9c8c70d2275d31cb2a

Observation 8c0cf09f-e4bb-4809-83e5-b7bffefece4f · outbound

This paper cites Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.245696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:24.580550Z digest=sha256:6668a4890aa7cba13c9687e787998961d405abb564d73ba132920aa5879bf843

Observation 07a803cf-9713-49e5-ae02-241870041023 · outbound

This paper cites Distraction is all you need for multimodal large language model jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Distraction is all you need for multimodal large language model jailbreaking

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:26.829247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:09:24.672129Z digest=sha256:5f8a1eaa1fc6d530688e81413d409e88edc76db08756be7ba01e521068962924

Observation 8bdb0029-ddc9-4f19-9bd5-1ce547d883c3 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.745035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.745035Z digest=sha256:9840717db01dae19ed12d59e9216539c3154e3688c209e9900ea794d436d041f

Observation 23f244aa-ebae-4ab8-ac21-295c9395e1d8 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.806832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.806832Z digest=sha256:38a10dc3fb6234480b3b4de2b1e24dacf516479c650a844e81a860ad8fc8331e

Observation 1612cb63-f7b3-4493-b43e-c44ffd16bb53 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.884191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.884191Z digest=sha256:343c2443796dd555a87c8f88537dca3ad7a7304473c31ccc707f03b3153fe80b

Pith citing papers

No inbound Pith citation observations are available.