Pith. sign in

Paper Citation Record · LEDGER

A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2407.07403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07403 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:11:04.716866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T12:24:39.804121Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8dd62ce0-0f80-4fd5-8f9e-40e6611c3926 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.029473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:4b3b51c11847de6b0110141b6cf7e3e3a9a587da8cc2fcc2e1ee38e0c0f31d5d

Observation c882a38d-0f9d-485d-b53a-65516bb50c15 · inbound

Hard-Label Black-Box Attacks on 3D Point Clouds cites this paper.

Hard-Label Black-Box Attacks on 3D Point Clouds A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:37:44.862049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T08:37:08.467026Z digest=sha256:17ff3f3eac49854c445748b807f6f96406b1014fe51838f764b9ef487ea1923f

Observation 32b9ac69-e3d6-4b00-839f-61bdc561c601 · inbound

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion cites this paper.

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:15:14.853288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T00:13:08.603115Z digest=sha256:a2b2cd92ba4a48557b2e0de432ee29fd08d4b706bb78c91bbfca2772d02cb4cd

Observation 76a19b3e-210a-4668-8a87-d24970af5263 · inbound

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models cites this paper.

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:04.716866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:11:04.716866Z digest=sha256:0d6a059ea67b6560db1bd54b3213b326a44c8ec55ea74a273bc15afae709b9b8

Observation dc243864-cd91-4efc-9c81-290fc99b70c8 · inbound

BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization cites this paper.

BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:59.647228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:59.647228Z digest=sha256:f7ca7a4750600bd80b45f5050c96006052a819503bf48b6ce86248e30ef3ff7b

Observation cedc2a29-f36b-4b2d-a00c-e406fd72200d · inbound

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection cites this paper.

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:03.798834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:03.798834Z digest=sha256:2ed20b572b5d407052c1066d5355b3b0190f7dd42669786f00c91a4bf408e1e7

Observation 7fb282ec-aef9-4cc8-9396-e252390f3be7 · inbound

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs cites this paper.

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:14.948476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:14.948476Z digest=sha256:5f8f6b848277f39b171b511f70a02d71f50059bd3ce2b15bf420d2f80e017593

Observation 311183e1-ba57-494f-b7ca-4fa9cfbb9f6b · inbound

Diffusion-based Cumulative Adversarial Purification for Vision Language Models cites this paper.

Diffusion-based Cumulative Adversarial Purification for Vision Language Models A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:59:18.054582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:59:18.054582Z digest=sha256:aa6ddc58545c4503c11f081f328af1378c7d9eb7b0533c1e1bb01e0e88ddee3f

Observation ed8edf02-b147-4fc6-ac77-8a5a5de05570 · inbound

One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models cites this paper.

One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:54.067388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:54.067388Z digest=sha256:0dc1695442d4866afb717a2edb9290ffa8b806f4c3f40f5d05c0917d45a7f875

Observation 63d6651f-e8da-4f2f-b415-e3cce06f15ba · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.862985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.862985Z digest=sha256:6166c3603480e9e4331545278a2d1dd123551c29d73e168247e555ef050f3093

Observation 7142993e-7a9f-488d-bfc9-66bc3f31608c · inbound

Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is cites this paper.

Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:03.941242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:23:03.941242Z digest=sha256:f2fc73c4d533fae9d8f454430ff5bcf073484cd0d66e5f6df1c0fbcee23a6b6b

Observation be887a7a-fa2c-4f19-b28d-02fa7b34b7ba · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.727842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.727842Z digest=sha256:aec803ba28905878a2c034e3b88af1280438150dd4e5c1d3d7fecfc7c30bcd27

Observation 88073051-5f25-467e-b525-4dcb62ae96af · inbound

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding cites this paper.

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:54:42.946018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:54:42.946018Z digest=sha256:09ba89029780831a2c774154bf5dd923cab69d98d8fbf0ef878b0bcdd099628e

Observation cd77eacb-4564-4c74-b403-f36417a61e67 · inbound

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair cites this paper.

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T22:21:02.801582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:21:02.801582Z digest=sha256:28708cfb5dd306216c6b74da2a749921f509c4c9e65063b00bdad37d53caa239

Observation 6ada323d-15db-4664-b52f-06f645d8c4c9 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 143

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:45.892114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:45.892114Z digest=sha256:d0312f2db8fcf9ce33d2eced9e9a96193953b8accb7e60725aa9a3ad0c52150a

Observation 3e27a6b0-49c7-4296-a9fd-67485edf0f45 · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:26:33.861908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T15:23:13.318310Z digest=sha256:39c51a3e75882099240ede6c62341c5d6195007344bf4b1191303461a636bebc

Observation e22cb968-630e-44fe-9025-d5ae556e722c · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:30:39.447565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T21:28:59.898029Z digest=sha256:60217a2bd22eb1fa06374f2a59ea2160139f83fe25990097eb8efee014197f9e

Observation 2846145d-382b-4083-8fa0-224cc3195575 · inbound

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models cites this paper.

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:15.314824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:15.314824Z digest=sha256:29866dc4486eb703b7d358c0b3c220ec2acd71a18d96909fbaaa413c2a4a7b19

Observation a29357e8-e064-431b-86bd-8b2f39bb2024 · inbound

FENCE: A Financial and Multimodal Jailbreak Detection Dataset cites this paper.

FENCE: A Financial and Multimodal Jailbreak Detection Dataset A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T22:03:49.351667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:03:49.351667Z digest=sha256:a128e8a9f7b559c12dbd96e1238e291f6c30db8b9cc32169a7155a455e3e4211

Observation a16e3c19-3995-4325-a328-36701531f6ed · inbound

When does learning pay off? A study on DRL-based dynamic algorithm configuration for carbon-aware scheduling cites this paper.

When does learning pay off? A study on DRL-based dynamic algorithm configuration for carbon-aware scheduling A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T14:09:20.967908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:09:20.967908Z digest=sha256:d14e6d5accecec3e8abf8bfc42af5f46d6e4c64d9316c2dec989d4f1396e5253

Observation f83ad23b-de4e-498a-be5c-1cef73021209 · inbound

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs cites this paper.

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.418356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T19:45:33.950587Z digest=sha256:37071692f2dd5191e09ac4aa048975487799602707162638b3424d96a590cd50

Observation c294da9f-1e67-4451-92f9-69a9a63ba734 · inbound

To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model cites this paper.

To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:40.903845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T02:38:37.358485Z digest=sha256:3b2af8215705578f1c010b4cbb75fbd517557c963df44c167d1191ff11fb3833

Observation e6c5f2d4-11c2-495d-9b25-64befe1c7f95 · inbound

Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack cites this paper.

Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.805804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T12:19:12.910048Z digest=sha256:84ea12f43a9018aee710a5f7700bcb2112dec2d8cac688aea9e2db6fa78fc5ee