Pith. sign in

Paper Citation Record · LEDGER

Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2403.09792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.09792 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:45:18.507649Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fa29efff-feab-44d5-850d-a51af125e94f · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.507649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.507649Z digest=sha256:952d519b3fb70ec2d25aa4365e26ffc43e81b9231c26fed31bcc3b62c54c3627

Observation 6bbb22b3-c3e5-4fb6-9e72-e6ca0e43ee60 · inbound

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion cites this paper.

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:11.420668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:11.420668Z digest=sha256:548d122283612c51ebedc21fc9e2d36683598c95f55fb8f1714670755cd8a0ca

Observation b98ecd05-61e7-4851-ac8c-c2ab05677a56 · inbound

Backdoor Cleaning without External Guidance in MLLM Fine-tuning cites this paper.

Backdoor Cleaning without External Guidance in MLLM Fine-tuning Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:51.412320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:51.412320Z digest=sha256:2f23df4702ce9def1616739d2d0326aab52d524de8dea453a3c336c2c1d5739a

Observation 2a43af3d-bcb8-4cfb-9482-2b6cd603c2ae · inbound

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack cites this paper.

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:01.312248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:22:01.312248Z digest=sha256:b37dc87eeb5cb9462400c1fe6a45890b45fd42757f6923cfe05bdfa59da1e210

Observation 43c65426-d610-4b47-a5df-062d796a065b · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.497004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.497004Z digest=sha256:d306e2100af393e866602e3d06180a24aaf1ac2db96c4c764a949b2060930495

Observation 3ce15541-625c-495a-ab08-158e850576cd · inbound

SALLIE: Safeguarding Against Latent Language & Image Exploits cites this paper.

SALLIE: Safeguarding Against Latent Language & Image Exploits Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:51.245712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:04:46.426969Z digest=sha256:0aee3ec6755ed9cb3a426a4123061f12e71872363b1f09f7d992440781d85b5f

Observation 2f9908e5-36b1-4c44-96d3-04107195ea7e · inbound

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation cites this paper.

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:10:26.165165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T13:09:35.407790Z digest=sha256:11f881e54c4bfed184b4238c92ced770083c9078965db2372fc078cbd279acb8

Observation 756304eb-f199-48c5-ad30-710f8f78ad83 · inbound

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models cites this paper.

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:56:08.176963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:24:48.999632Z digest=sha256:b2f4ff999c51e6db7533b03d7ef0108883c66f86c7bdd36701ba6c4b2f303270

Observation 8469b9c3-1e1f-498e-8828-843be58a7a9f · inbound

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing cites this paper.

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:27.478710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:50:08.866969Z digest=sha256:a12001cb6a630835e34add650d9bc738618879e9649d3c78f4b86d5d0f14cfa1

Observation 99fd1219-c007-4303-b66d-521a0bbd0d06 · inbound

EVA: Editing for Versatile Alignment against Jailbreaks cites this paper.

EVA: Editing for Versatile Alignment against Jailbreaks Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:35:47.008475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:40:51.290851Z digest=sha256:c3f24036703ded0a409e0f388a0200b0baec73a29d19ce685ddddecf10bc47c1

Observation ac30a608-caba-4506-af9e-8b6353e1c10a · inbound

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents cites this paper.

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:02:46.383448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T22:59:43.041678Z digest=sha256:da08c52d9fa2c61683490cec5b3cae00f5b659a3f685b714a17b36ed6d3c03ce

Observation f88c73f7-be16-451e-a3ac-d9fbafd23c98 · inbound

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models cites this paper.

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T04:38:59.195293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:35:51.583460Z digest=sha256:80b698dfde8a8e16dab770d5893c2ff255bf39db0d2b7f14bfeccf01b2247eae

Observation 56e80f9a-1140-4f7e-883e-db34e857a2e8 · inbound

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination cites this paper.

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 130

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.340224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:14:48.340224Z digest=sha256:d431fcb59b2167281b3d6f9175295481a9d8ec9b6cbcc79f94b5104df4ffdb52