Pith. sign in

Paper Citation Record · LEDGER

Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2406.04031.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04031 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:55:40.145094Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4619acca-29f7-42f3-8aa5-bb477dad491e · inbound

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey cites this paper.

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T20:53:15.023557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:53:15.023557Z digest=sha256:f500a5405f68b165f3969e5592232d66b70e8e90b660a7e494587f4f781d5b79

Observation ff31550b-44b9-42cb-9b21-da7caa2da275 · inbound

Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models cites this paper.

Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T18:32:12.139219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:32:12.139219Z digest=sha256:bbc07cb11abf463aaff4aab85f1154bf54381e68eb254b5cf73046d382f190d8

Observation 30864464-a2bb-479a-9f5d-00825b8bef33 · inbound

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving cites this paper.

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:35:42.163051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T16:35:24.063578Z digest=sha256:35232e20ead2d4173c89b6d484629286b9f8e7b4d743c445b4d977a8c1b196a5

Observation d4ec7664-e5e3-4e5c-bb1f-23632bd73800 · inbound

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment cites this paper.

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T11:03:01.268738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:03:01.268738Z digest=sha256:20b5e3be5f663946553cecafcde195bb6b8cec0f5d4fdc1a3d3d97af397830c8

Observation 01229ed3-161a-4dc1-ac80-81bcd3404ff2 · inbound

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds cites this paper.

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T20:53:38.132216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:53:38.132216Z digest=sha256:a62aa3cb275cf5a0f75481ddc7a4be6b811bd381bf08743702aa83da2181ed0b

Observation 5aed803f-fa3d-4931-a465-9dc96b29a343 · inbound

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models cites this paper.

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:17:02.829992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:17:02.829992Z digest=sha256:e954a4cc1bbb5e1c25da272431bd166c99858bad443c8d7d227f6cd04590429a

Observation a407190c-d276-4913-9227-5357fa4a81ac · inbound

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models cites this paper.

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:28.408053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:28.408053Z digest=sha256:e6128890789e1d711c228d9e4aca8cd3082c60203cb98bd5892c50ea310f5261

Observation 80ed813d-ef3c-4083-beec-a5ac8fd8c19d · inbound

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models cites this paper.

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T18:05:52.403894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:05:52.403894Z digest=sha256:c36245691b74b1b2dd9eb3a11aa672daeab7b70fc0b93990a34f51fa7e3c45c7

Observation 17617b06-e3ce-4470-8cd3-2e1a01db0ce9 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.363130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.363130Z digest=sha256:89c4d23446653e7455435cfda25b5f965fdac5b9b3d4c13ae62538e82bd1cb22

Observation 682e0d26-172e-493c-aea5-14398add3938 · inbound

Manipulating Multimodal Agents via Cross-Modal Prompt Injection cites this paper.

Manipulating Multimodal Agents via Cross-Modal Prompt Injection Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:40.145094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:40.145094Z digest=sha256:c4575c6ca7e255ea094b189ab213da7343a06305160505d1ca0e7ad891dd0f0b

Observation 4ae8d774-f9cd-41b2-b352-61eb00eb7c9d · inbound

T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models cites this paper.

T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:11.092728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:11.092728Z digest=sha256:467b2ecba6381f77c10e27fe91a562dd1998795ea9c33fc64f575b7fe9424176

Observation c46a4fdc-382a-4cba-9e48-9af2e9b7f6be · inbound

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM cites this paper.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.148983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.148983Z digest=sha256:9de90bbd725743f97cbc93262abbb352cee80d2e88cbacb83d34830bbc7ddcc9

Observation c6d5bdc4-9725-47fe-bff1-392ceb50acf5 · inbound

POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models cites this paper.

POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:43:06.054228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:43:06.054228Z digest=sha256:18e10e5e63451e9d23a51bc60b828306a501cdaad6acab268166bd0f2112d5ea

Observation d41d7b91-e49d-4205-aaf6-873fbd4ebee2 · inbound

T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks cites this paper.

T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:40:28.650354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:40:28.650354Z digest=sha256:9899457a211398e1623d978861b1dd96f8c18cd286f7c5d69d8da37c943695e1

Observation 22d4788a-ca49-465a-bd61-75398dd055e2 · inbound

No Query, No Access cites this paper.

No Query, No Access Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:25:33.241825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:25:33.241825Z digest=sha256:fb8af6ef41799c696fbc7f4854888660a18d52b23f67926875f672b0dac1c83f

Observation 44fc0931-9434-46b7-8d0d-cdbcfe1a0931 · inbound

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models cites this paper.

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:05.926019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:11:05.926019Z digest=sha256:4de194c5483a691e06bce1efc79ba5b579f654334b4fade11be510fcd47b1d5d

Observation 0c0fbad8-071f-44db-a668-11ea5518a66c · inbound

Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers cites this paper.

Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:13.366909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:08:13.366909Z digest=sha256:58e47dca3a13348bec7c62a08457bf40c739c876540f5c92ceb6858e28f9b5d0

Observation 8cf43540-c992-4d50-a46c-a9d45d12c8bf · inbound

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space cites this paper.

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:36:00.334024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:36:00.334024Z digest=sha256:8f602b48e52677bc8a8a9fed0bf5f954b771cd9469a281bbdfb79aca66d9a0e2

Observation 2506fa70-0265-491c-9720-27f438d9d0da · inbound

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts cites this paper.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.165635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.165635Z digest=sha256:6c2a14aed618c79a58da892675fdd995f32067c557cbd6a304cf00454072498c

Observation 22298515-bd24-4400-84c9-790d5dad6e67 · inbound

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models cites this paper.

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:32.894561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:32.894561Z digest=sha256:553e80fc2357f6a84ca960c5130a773df5755034b28c6c61904e5f32a691398a

Observation 31c5f5a1-8f32-465e-965f-77fdacbc0fd4 · inbound

PRJ: Perception-Retrieval-Judgement for Generated Images cites this paper.

PRJ: Perception-Retrieval-Judgement for Generated Images Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:03.532221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:02:03.532221Z digest=sha256:063099c9bbbb6e50ec5ae1993039942b3e734697fc14a6821dc44fae419e5f4f

Observation 6195d26a-623c-434c-995a-39f5edcafdb9 · inbound

VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models cites this paper.

VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.686795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.686795Z digest=sha256:c428e43a197e63cb81fc416e8bd2ed83787aaf8daa045f1fb00653c7bd1e0347

Observation 667e004a-1aaf-4620-82b6-ff1b3d6d6ffe · inbound

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents cites this paper.

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:18.221976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:18.221976Z digest=sha256:2cf142b1a0a59c57788ed32b3221a905795f5eb5f08cd53f363649e37f4b29d8

Observation 630d645b-ce3a-474b-bc30-f6c34f4e2371 · inbound

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking cites this paper.

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:37:01.063367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T03:36:24.013477Z digest=sha256:b4e7a117b75be662e66b70bd542cdb251893f8e4adf58520c866702305e5c6f8

Observation 1d3abdd7-d341-43bb-b3d4-21e9f885d2f3 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.597526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.597526Z digest=sha256:4bccec1b6bc650fd9aa0573f79eaac23ad165563df9c6ac4e40ba1093e69263a

Observation 38b60e6e-d2e0-4de8-8ba2-edb6f8936389 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:06.610465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:06.610465Z digest=sha256:88f7d81b160d0cd59c911e03e2ecba3917b24e76b06f0af1dd0fdaf7e052dd4a

Observation d2d09ee6-0988-4a21-afbd-8110667ce07d · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:52.253832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:52.253832Z digest=sha256:fb0992dca3a2607185297b8a4cb6012581be278b827a848e718e6e862109cd02

Observation e4b34017-fd2e-46ab-84b7-de38e5014b22 · inbound

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models cites this paper.

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:40.894740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:58:40.894740Z digest=sha256:08c93e95f722134dc9c227ab341f0979355948d8ea4fa7c3e34bd5a5b3a7dd8e

Observation 172ab056-f2f2-4ede-aee0-38548d9ae382 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:58.141356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:58.141356Z digest=sha256:77af9f06904ca6ea2749b91c60dc680b856008acb72ebd2781261dcd0ba2ba63

Observation 7216e311-fdd7-4b39-96ea-783fb2e30473 · inbound

Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation cites this paper.

Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:27:03.863353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:27:03.863353Z digest=sha256:792120c3e45d973cbc68a603005e15d3c03acd1a88c1adadf3fa1266d7a11b8f

Observation caa29f76-9ae6-47f8-b5ef-2113bd0f7225 · inbound

A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models cites this paper.

A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:05:49.112628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:14:02.553313Z digest=sha256:a50df45fcbbe879127dff41b1ea35809e9efeb7a80291d6e8a1212559c021ef0

Observation 46391421-e264-4d6e-9dc2-d895c864b488 · inbound

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers cites this paper.

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:55:49.742793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:30:29.032353Z digest=sha256:5b537d11ba55a86ea47368460f9fe92df1bdba8b706ffe42cf345a83e03aaa55

Observation bf63562f-4f2e-4eda-8cea-7c1cd4710d3b · inbound

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models cites this paper.

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.129002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:11:01.039309Z digest=sha256:3bf38e7188fb8448472ac2f95d23adc87f29cf5504a6508c0c43bc60746c3234

Observation cebbc9fd-bf29-4a68-b47e-4bf63477d9cf · inbound

Laundering AI Authority with Adversarial Examples cites this paper.

Laundering AI Authority with Adversarial Examples Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:07.993730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:19:38.662062Z digest=sha256:4de65363bc01ea4d60462cd09469c74dced7580d656f7277b22238c8213b6268

Observation fc0bf83b-9254-4bf7-bdc2-d4148539f05c · inbound

Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack cites this paper.

Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:24:39.793465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T12:19:12.910048Z digest=sha256:c56119ea32a1563898c57671d0767cdf1f73494bf15375750fcf45cfd96d99d5

Observation a51ec3bc-a427-4317-886b-1db2d4d5f433 · inbound

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models cites this paper.

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-06-26T04:38:59.191152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T04:35:51.583460Z digest=sha256:08a200752ddee3c8bb3816eff60f1b0dcd89869721472590ba316fd0b821d9d7

Observation e2c429a6-3c47-4176-8030-153b619f6b56 · inbound

GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models cites this paper.

GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:27.555375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:04:27.555375Z digest=sha256:5edd0da5c3c31d077f352d06bf50a30dd78daf9e0ba8c1e69fa79ad69e1ebae5

Observation 62a1afb1-8ef7-4f5d-9bd2-81bbc9236c81 · inbound

SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense cites this paper.

SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:01:41.516193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:01:41.516193Z digest=sha256:42a59c6fdc0dfc601f34a55767608ae355e26d1b9f07b54faa800c1682ff1cc2