Pith. sign in

Paper Citation Record · LEDGER

FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 64 inbound Pith citation observations for arXiv:2311.05608.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05608 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 64 of 64 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:55:40.174005Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d6611326-d04b-4e0e-a9eb-3c0a32f7beac · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.644037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:2faf0614b0811caea9e1530e3f6c848362c0423a9b0243b5360f9e680086587d

Observation 20571498-ca2c-4d01-b1ad-cf4779b735d0 · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 261

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:49.827489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:bec58f7f280e67870d04b08f8f7747baf09ca78daf097350a534cba6aea7dd7e

Observation f590a976-a726-4e93-a67b-bcf63b63f027 · inbound

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense cites this paper.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.681816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.681816Z digest=sha256:61dd3852d2b5b43b0177d9ab0d6cdd66555f9f4aa083794d7b315d571adebdc3

Observation 3a0f850f-fcfb-4952-911f-9cfe9e31dfe5 · inbound

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey cites this paper.

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T20:53:14.927923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:53:14.927923Z digest=sha256:c156c84e36a2875a7417f32ce801cfa55f179ec3a12599fc5b66ad5bf80e2329

Observation e1073da5-1b48-4e04-98fd-b3114cc1f269 · inbound

Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents cites this paper.

Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T20:36:01.688800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:36:01.688800Z digest=sha256:af2b7bfd1db53fa4be63ab0a4d6f76815bdccc8a116e98cda915d0994700daa1

Observation 7f66642a-eb7d-443e-a0ea-ed7e5c5cbd49 · inbound

Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models cites this paper.

Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:32:12.008410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:32:12.008410Z digest=sha256:4f625aa68cb8bc3ea2ae7398211881f2109566753c41793ae15e244f2d58e886

Observation 455e54ab-a27b-4814-8c6b-0ccdb3868698 · inbound

Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks cites this paper.

Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:45.364618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:45.364618Z digest=sha256:7b6d8b052e5646f149659c450e0de67b371e2a6894334e4748db22494be5ff5e

Observation 0acfaf43-7260-48f6-a294-f1aa1ffd933a · inbound

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks cites this paper.

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:32.422535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:32.422535Z digest=sha256:51014383173dfde26cb7ef5e318d0841a5b01bd4678c4f4ae38029947efddae4

Observation 8d1bc228-aa78-4677-8435-e35e65ea3596 · inbound

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment cites this paper.

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:03:01.068838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:03:01.068838Z digest=sha256:d0e483d2ab40ccb2052c35c000cdb8732bab1716860d40407d7c3754653e00cd

Observation f1d59096-0e05-46ee-bf34-88367218808a · inbound

VLSBench: Unveiling Visual Leakage in Multimodal Safety cites this paper.

VLSBench: Unveiling Visual Leakage in Multimodal Safety FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:46.382144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:46.382144Z digest=sha256:8da6a5fd46d5a740787172b454b778c9645c3065a58bf42b5d154b8b8ae41599

Observation 6b1623d3-0fbe-4081-81a6-db20054458ab · inbound

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds cites this paper.

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:53:37.051669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:53:37.051669Z digest=sha256:37ff87cf8a6526652f096b144e3813c60434a04249dae3161acf3f80092135cd

Observation a10f859c-81cf-4f5d-bcbb-ef1ca7fee064 · inbound

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models cites this paper.

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:17:02.618814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:17:02.618814Z digest=sha256:98eb9ca5cb20d4f5eb82e3fab4703a8e287718b2a665f6aca33f7015a6c67808

Observation 7b1736b7-ac9a-4c0e-8d82-b797a7d691ff · inbound

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision cites this paper.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.451631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.451631Z digest=sha256:7539d1008aa32b706d21e1943c294f22f74ee02858e778bbc70168bd608c9935

Observation 5c2b8983-526d-4d18-8a3a-aeecf1f3d099 · inbound

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models cites this paper.

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:28.253863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:28.253863Z digest=sha256:ff536bf3f10ca297aaba4d932fbfd6f83b6684edc84db0178ec8c4dd5369dc28

Observation 86c6a18c-c7e6-4fe6-9688-fc3c4da0504d · inbound

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting cites this paper.

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:04.239283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:29:04.239283Z digest=sha256:860753f19f6a5548b04670f911e6ae883d7d9e85f68f6616487ce49483160917

Observation 07f0defe-be02-4549-bfa7-49a3947901f2 · inbound

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models cites this paper.

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:32.275326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:28:32.275326Z digest=sha256:5fa467b9dc3e8f78498866b2c74001dc100779742ea24ff308d14fdd0cf1d008

Observation 6af22920-5d93-4986-82a1-db96bc17f2c1 · inbound

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering cites this paper.

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:25.182762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:03:25.182762Z digest=sha256:174dbb9ba35576e40257786fc3e586ec8ca790924fedf18ae008152241999ab2

Observation 58de6437-1fcc-45b6-b975-e6141587a160 · inbound

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency cites this paper.

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:26:00.521572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:26:00.521572Z digest=sha256:b2a63c38d4c03c1086dfb0c18a9121acf0ecbd7e906b7510a9437a09716dbd3d

Observation 68faee73-b6c2-4a4d-be86-451e438c806a · inbound

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness cites this paper.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.539812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.539812Z digest=sha256:91fae3917a1cf810f7a7ac71248ce9a6d720d5192e4ae9197d5ac961ceab3415

Observation ff11f3ff-3dff-4ef0-93c9-6c0e284c5670 · inbound

MSTS: A Multimodal Safety Test Suite for Vision-Language Models cites this paper.

MSTS: A Multimodal Safety Test Suite for Vision-Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:42.222300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:27:42.222300Z digest=sha256:245c0757c90dfc6cf63501b182ab1cd8c0ba10c6ecff19bea59765c3be15de08

Observation e8daa2fc-b5c4-433c-9815-f0bc5301c395 · inbound

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink cites this paper.

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:29:27.955780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:29:27.955780Z digest=sha256:e9739117bac36eaca63e9b9ef140fd9aa016f8feb97d1162831ef807206a045b

Observation 8d5ddc72-7959-432a-9303-81335cb62a47 · inbound

Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update cites this paper.

Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:19:10.058952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:19:10.058952Z digest=sha256:bf691d88d3f9cdf4c5ce5378f92cd969ffa32539ed26a63b5abd83192b226f66

Observation 9286d120-4a50-4efb-b285-aeaf00e32d60 · inbound

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation cites this paper.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.087069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.087069Z digest=sha256:259577793cab8255653b8ac262144d0ea88fdad41c08218dba61b4fc534b0608

Observation 06db8f48-9059-46fa-bec0-a95bce650ee8 · inbound

`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs cites this paper.

`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T17:58:57.502267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:58:57.502267Z digest=sha256:361cf51a2ad31c2307a37ac268594012f3f95afacc033bac7ffa62494df258e1

Observation 0dcd471c-92c4-46e6-8864-473a96d2d4d7 · inbound

Peering Behind the Shield: Guardrail Identification in Large Language Models cites this paper.

Peering Behind the Shield: Guardrail Identification in Large Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:45:21.421463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T03:45:14.234545Z digest=sha256:7047a6dec0e021f6b537f2830c8dd06f71e4612c3102cfe3a95c94be54a1d7af

Observation 45db22fc-32bf-4669-aa6e-79235529bc6e · inbound

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models cites this paper.

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:58:53.648957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:58:53.648957Z digest=sha256:e5ff675fffe5db9eedba85deaf8ad552341458d2503c488b2134108ea7408cbf

Observation ee3a88df-54f0-4c89-9c74-a5b3a3237360 · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.245674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.245674Z digest=sha256:3585fdbdb3da712b3495c887c83ed62b81d924b25f7ae93db961fab0c4af4cde

Observation 8a28ffb6-f22f-4ade-aba3-4914e6ba9bfe · inbound

Universal Adversarial Attack on Aligned Multimodal LLMs cites this paper.

Universal Adversarial Attack on Aligned Multimodal LLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T11:17:01.577794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:17:01.577794Z digest=sha256:d55985ce628117fd86e99fa426363fc818c3e4805aa593ce41f938982b11f7d4

Observation 2b4a0320-2e8a-4784-bf05-44489b5580c7 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.513844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.513844Z digest=sha256:50db12b1d93bad7eeded169b67e5f7c785d143e8a09a96f7af4856f76623270e

Observation a447fef4-d575-4fbf-b5c6-439c13bc93bd · inbound

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion cites this paper.

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:15:14.901879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T00:13:08.603115Z digest=sha256:ca911996daf8aae22af0219b3c8b489feef54bed4ebb12016483f4ca925d1149

Observation 0f36ae66-b552-4414-be16-ba32471dcfee · inbound

Manipulating Multimodal Agents via Cross-Modal Prompt Injection cites this paper.

Manipulating Multimodal Agents via Cross-Modal Prompt Injection FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:40.174005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:40.174005Z digest=sha256:08246861a1740aaf4b8254a109f5b6c8d335009343fdf0da730a62d808e5da8c

Observation 5ecdc404-efe4-4b23-ba54-eefb57098d32 · inbound

DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models cites this paper.

DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:29:59.043297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:29:59.043297Z digest=sha256:094ec1fbe4e3885cbee083d8b30fc1e7b9e5274f8c5f2531851e5224d310de93

Observation 1d9547d3-307b-4888-9c84-c4832b70f0ad · inbound

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM cites this paper.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.995456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.995456Z digest=sha256:2aa52ce87d485387103d715238eeaeb10f71ef48ac259dede1162137a94f4ac2

Observation 48419eff-b2b3-43aa-8648-18699d8d518f · inbound

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning cites this paper.

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:17.372492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:05:17.372492Z digest=sha256:e4fe4f9f4fa6cc8c4f057bb58c22f8c7bd98a622c8d581dea72362544b0c3877

Observation 103793a0-decc-44e8-bd1a-2f4d876afd11 · inbound

Backdoor Cleaning without External Guidance in MLLM Fine-tuning cites this paper.

Backdoor Cleaning without External Guidance in MLLM Fine-tuning FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:50.049188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:50.049188Z digest=sha256:f4c042585f049222c88bb8d16a36cd854d6008791a94c73ec866c14b3097439e

Observation 0a8124b6-20c0-41e0-a7c7-ed3b7ee9bce7 · inbound

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework cites this paper.

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:05.699522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:05.699522Z digest=sha256:f89d803ea228a313c00fc1ca742f73615bee3addb6a4c5ae8a601cbda473adcb

Observation 4ff5cd9d-7983-411d-a328-87439cb68029 · inbound

VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration cites this paper.

VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:11.421412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:13:11.421412Z digest=sha256:5d53f736f44daf995811a4a34813680f67d30056c8cf8c141a32a43d577065fd

Observation 154a0da8-bb06-4901-b630-d98a7539b344 · inbound

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack cites this paper.

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:00.808067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:22:00.808067Z digest=sha256:95f33d7e95082ede3695e93d9eaa2bb117618fd8ab72406a9db19ec28a04e1fb

Observation c9a16fcf-f021-4a0c-a6cb-addc007f4cbc · inbound

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models cites this paper.

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:20.261903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:20.261903Z digest=sha256:9c76715b92b4452e57c591e2c23e09229d7e45050463b93e9a14c7ebdc95e97c

Observation 197b9f6a-0da3-4cf5-a240-b7adcd798e32 · inbound

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models cites this paper.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:37.104583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:37.104583Z digest=sha256:3f4cc04457a4e8858ba0c3599d0dcecca2bdb40b295a29811b4f1a3feb3c3574

Observation 15745a6e-fa80-4aa0-b09a-8879aa8366e6 · inbound

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders cites this paper.

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.878499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.878499Z digest=sha256:8574167873845f8654026fab9631e8fc56f3c3e2fe72beacd069c69c75f1dcf8

Observation dacb2f80-318f-4ac3-bbe4-bcd682011472 · inbound

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities cites this paper.

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:59.166840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:59.166840Z digest=sha256:27d6455b3904f806535d35f392a5b87c3525822cead5d30b640605734109e137

Observation 21cc80ab-5b0e-443a-801f-4c1316b01b37 · inbound

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem cites this paper.

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:09.795414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:45:09.795414Z digest=sha256:39d9ceba0ce0c98fb7361d0f9d947a02cfacb88473a45defdc270f6a2bc6d474

Observation 68b78950-31ea-4ae7-a11a-c446900e07a2 · inbound

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message cites this paper.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.203611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.203611Z digest=sha256:9c2c24ff85ae11b58235356e14a4519a6e147d2edf8f4f9d21ab2eee483ea7f7

Observation 646bbd30-5388-4171-9b02-5d93ffeb2371 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.544437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.544437Z digest=sha256:bdd7552f1dd38abf5486cb674c1f668db13d0a3a3c6c26e84843c440da5a9acb

Observation d549b25a-4dac-4f19-9750-9debf0272fd8 · inbound

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities cites this paper.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.208807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.208807Z digest=sha256:205217fd0e2b6eee5dbbbba530017b053620cdaca9ac68bfbf980837f02502f9

Observation 4f27bd73-024c-4e13-adfe-8123bb08ea74 · inbound

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models cites this paper.

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:40:05.291273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:40:05.291273Z digest=sha256:706e0698899570806c36bc8b4a7828da810bd34d59db34c342cbc936ad9d5c2b

Observation ebab30ad-0dee-4a65-835b-e9d800498ce6 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.668500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.668500Z digest=sha256:f3a3929a544bb2800f9244afec46981f51c8999d9adbc735ed1c15ad211fd830

Observation d510cf0d-ca3a-4b04-b357-685170176142 · inbound

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair cites this paper.

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:21:02.779195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:21:02.779195Z digest=sha256:76c878908b8e327d76e457c235677ffe43999c6e42dd59164a093226d4975106

Observation a9d54ccb-3003-4fc0-aa8a-ac732c816833 · inbound

Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models cites this paper.

Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:22:36.936309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:22:36.936309Z digest=sha256:27543aa93fc0ef57afad3b5412414b2d25dc77b467c24fa4763b859c96d4fa16

Observation 5740cbd4-b93f-4ade-b1a0-87b323e0e519 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.152848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.152848Z digest=sha256:5ff3e21fca056d4228019db75ff33469195ee236a484a7edb517b2ca0170fce5

Observation c4b7143d-a5e2-4c76-abb8-656d7930f615 · inbound

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models cites this paper.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.849334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.849334Z digest=sha256:f541ee464a44ab383aab4bbcf072743cf9ae1f9c2386a579e4234c807523754e

Observation 4c40d6e0-a570-4127-b746-4696d7e71c9f · inbound

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs cites this paper.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.482899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.482899Z digest=sha256:49ad0f8a976b2195256c6d3016598cc0dc71b3ba877bdbb19671ffccfe9c84b4

Observation 065152dd-2878-42c2-ae83-a853d098f144 · inbound

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models cites this paper.

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:22:20.345146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:22:20.345146Z digest=sha256:b7d1543f382ca0d7cfc222f3e2f2801fd9663559272c08aeaa1ca36cd947e3e3

Observation c782b4e9-3165-442d-a618-76714429a737 · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.717999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:351ae5ae13f5d302257d07c48d56970609262f467fd3b720f62b4f77cde3777f

Observation 33eee8d1-b51e-4d1d-ac56-a4ab45110693 · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.062154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:1bd9547099c892986a3dd34d485df16946c94ad639de1ce3a21c986c8bd0bfa5

Observation c4d9e311-7b61-4e74-90e8-cf1d6cfc511a · inbound

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models cites this paper.

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:08.044730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T14:24:48.999632Z digest=sha256:bfa1665f686b377130765ec8bf5ddda6cf1730e9451c4cf174639d30357c3d5f

Observation 1a793493-6053-4015-b0b5-66b60f0b1867 · inbound

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing cites this paper.

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:51:27.613031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T04:50:08.866969Z digest=sha256:d946a8a674a5a56037043ffc758f3bcf88f6c7a12ec271283d78fcb11174eb09

Observation 52479ce2-430d-4a5a-9cd6-745d42dcee2e · inbound

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges cites this paper.

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:31.023210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T16:33:28.848573Z digest=sha256:f974cf63919ae1c9d06d6493556d2bcb1d22980acf9b17938ebe3fd2c842fd1e

Observation cce14d1e-e0ed-4e91-be22-99cc7cca0d5a · inbound

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models cites this paper.

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-26T04:38:59.203237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T04:35:51.583460Z digest=sha256:9326734d79fe8b6d7836638ad50e49be0c84e46ab05ab5dedf9d8d24b41e07b3

Observation bfae213e-9029-4b58-954f-127a92ef4fd2 · inbound

Securing Multimodal AI through Internal Information Decomposition cites this paper.

Securing Multimodal AI through Internal Information Decomposition FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T15:02:59.139372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:02:59.139372Z digest=sha256:622b2c4c2b697c27ab846b6800ac90401176da6c8766c9fc02a888d625641600

Observation d3a5ca68-e5e4-475a-9d0b-7ca49e3997bb · inbound

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination cites this paper.

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.682360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:14:47.682360Z digest=sha256:c604620b4fa4d0f18cdece2f1bf695636dba1dda0e887e48fda146a819b149ff

Observation ef8609ee-5f43-433f-80a7-e685488daf32 · inbound

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use cites this paper.

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:32.663900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:40:32.663900Z digest=sha256:030134093f9b67f6350f1f4e22e34527197b30c961433784e05517ddd8f8991d

Observation 28abac41-ddee-48b4-b94f-af4761b95fc6 · inbound

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning cites this paper.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.599621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.599621Z digest=sha256:9342bbf4fd0acd8657ed84969eb951b55e9042de534940084ea7825952e24a85