Pith. sign in

Paper Citation Record · LEDGER

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs

As of 15 August 2026, this Paper Citation Record lists 100 of 130 outbound references and 0 inbound Pith citation observations for arXiv:2502.06390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06390 v2

Coverage vector

measured 100 of 130 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:38:17.453227Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 130 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 099f1e85-e121-4dd7-82fa-771525c9fef7 · outbound

This paper cites Efficient multimodal large language models: A survey,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Efficient multimodal large language models: A survey,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.028360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.028360Z digest=sha256:3bf58b2d8d0500508cdc3b79508dda483fbb7b1c2afa087b5dc0dc8c2dafe78d

Observation b736f5bf-c1a3-4c91-8983-9ffb7500cb1d · outbound

This paper cites Vision-language models for vision tasks: A survey,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision-language models for vision tasks: A survey,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.033499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.033499Z digest=sha256:98be0ba84bd501d457b43388ebce95e42f21427acfe2748007893696e59a4737

Observation 15d7f862-3768-4384-9218-bfcad2b6fdf1 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Learning transferable visual models from natural language supervision,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.038157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.038157Z digest=sha256:d9cb66fd9c9a9ec1e5b4b5b810915e8ea702f0248c71415442934bf4e7973c29

Observation 1f70bd67-0e81-4ffe-9075-4d11bb96dfdc · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision language models in autonomous driving: A survey and outlook,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.042884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.042884Z digest=sha256:00b0b6cfcab75a937b1066dd83649234c519eae7b757b9977d3592b372d2e2b5

Observation d46bf3ae-44a0-48e9-9bb1-4a01bf7215d1 · outbound

This paper cites Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.047450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.047450Z digest=sha256:4f69de00660c55329631af7fbb702c857822f3d64025f2cfc101d762a515f8f5

Observation 32dbaf8f-1567-4bbf-9a52-3ef8a9f551fc · outbound

This paper cites Visual instruction tuning,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.052425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.052425Z digest=sha256:0143a9b583f77a3a42add5b2736f1e84f572cc5a70f936f334ba9c4773b63e4f

Observation cd77a381-8ad5-4801-b6df-7d80726574a8 · outbound

This paper cites Listen, Think, and Understand.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.057432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.057432Z digest=sha256:a58847a0cd379be9cbdb18b1c0e8fa45b0774be9ab5073952c26b04f81d0a9da

Observation 746508a9-917e-431b-a74b-ea30b275419b · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Sonicvisionlm: Playing sound with vision language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.062234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.062234Z digest=sha256:c5a456ee1e3a1466305cb7843191c473a92fef2547c9407b06f3be35531f7155

Observation 46bd1cec-8f67-4472-80c8-b5a06e2115ba · outbound

This paper cites Grounded language-image pre- training,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Grounded language-image pre- training,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.067115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.067115Z digest=sha256:b75f1e2dcde7305a671d816e66870a9c4fa8fe429df05e479962698e7834a11d

Observation 00734e26-a685-47bb-991a-a66f5039580d · outbound

This paper cites Segment anything,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Segment anything,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.071570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.071570Z digest=sha256:d7e3deda1bcf91ba7e1cd1a833329c101bff8713c4d83b1d03b91a8587d23aac

Observation 845577d3-c834-4cf7-8a40-9e0b9b155257 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs On evaluating adversarial robustness of large vision-language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.076098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.076098Z digest=sha256:b1e0a45a56b6f615d2d318073fc122d98fe973f37443a2850de6fd7905816e26

Observation 067c0329-e16f-4c3e-88ac-a626ce7d55e7 · outbound

This paper cites Mma- diffusion: Multimodal attack on diffusion models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Mma- diffusion: Multimodal attack on diffusion models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.080542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.080542Z digest=sha256:80bd6f603762246940640a4f77a90fa73bb2e7d4edb9282446ebacfc709f1894

Observation d508c24d-5592-4712-a758-1cb1b657c72c · outbound

This paper cites Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.084735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.084735Z digest=sha256:c1052827ac375602ea860a510e0ef166c1a48d8e938465f41dec705370f21a6e

Observation 631adf9a-90c7-43e1-ae22-9610b9e7c1ea · outbound

This paper cites Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.089320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.089320Z digest=sha256:6bdb010e2fe641b92a052f2e1b95bec0816fa9e20af51a3a2179c922b6d8e2cd

Observation 3037f0c1-e04b-4d46-a13a-2a6106a1b05a · outbound

This paper cites Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.093802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.093802Z digest=sha256:c511990c9947de6dd9a51f32c261096dea6e4928e3e6c040ab7301a8503f1a33

Observation 555c05e6-2912-4c4c-82e6-9d5d054b1f52 · outbound

This paper cites Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.098397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.098397Z digest=sha256:09c858b544e5d86368bc0e893b324aaca27395ec3e0a794507edda4b37eeb662

Observation fd9d8eb5-393c-4e58-a585-61fa9825b70d · outbound

This paper cites Safety of Multimodal Large Language Models on Images and Texts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Safety of Multimodal Large Language Models on Images and Texts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.103048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.103048Z digest=sha256:9cef0a6005df79a06f2c35048250c8e5275a5e4a3015a07c676b40b461d4a77d

Observation 25248a1d-fd28-42f5-be5f-2c1319542a97 · outbound

This paper cites Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.107046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.107046Z digest=sha256:0bbd8c202b1a155086c1409d0cb9b52dc8205503285547e75798682a5193e1f5

Observation 0db4d2a4-38f5-4cce-ba9f-41abdcbb29b4 · outbound

This paper cites From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.110619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.110619Z digest=sha256:94094d16c04ae6331d6edcaef916ff534f8dd6e3ef3dee55f2d3cbbd4612cffc

Observation 0917c57e-d361-42ac-b959-9fe6f0a3b87c · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.114309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.114309Z digest=sha256:32294469b3191164def1c0322ff3266ce5c8757e45f452e9c86d7ed2823f4ae7

Observation dd5065d5-9f21-4841-aeea-08d517887c04 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.117952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.117952Z digest=sha256:ca1901ce95b146509cae77879c3d4ce95b773f5d076d62b92e0042d2c76a62f9

Observation 1d8946c5-1be5-4d0e-9386-e4773a6471da · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.121586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.121586Z digest=sha256:bea3bae00b0d4e2b4d7bb1b988d5441827bb3a136605cf0d9204f5a7ff0ab2ea

Observation 941f2746-7dbf-4864-bf01-cd80d9282a10 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.126010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.126010Z digest=sha256:28418f662fb6b473ea2b7cb9bf098882d702fe4cacfe0090836c3833d7639d19

Observation ef38d53c-ace8-43ce-8583-d49993e6e61f · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.130611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.130611Z digest=sha256:7b3d5d189ab5b1f991247f99394f000808c0c04d34652fb8115fd985eb9df979

Observation 003f4fc9-c90d-4d73-a018-ee564682fb36 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.134977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.134977Z digest=sha256:12876d4fffac2ded081d44168944173c54720b43682ff386173018fbb7ce9a44

Observation 11729ae4-7e6f-4177-b07f-e09cc9c8cd7b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.139189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.139189Z digest=sha256:543f7c06124135f7481ad79a896e524810488b8f21b701c900064427cfe4f534

Observation f32d0cc9-7686-4cc0-aa81-9e02a377b1de · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.143275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.143275Z digest=sha256:429a766652329752fd9297fc6a19862812a688ed1d92710ca1d6365d659132fa

Observation e4461dbd-432b-4e2f-b074-9f31f946ceee · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.147602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.147602Z digest=sha256:ede3e9471d28437ec89ca0f0e0d9150312e6373220f8ce268dbaf1281e70aec2

Observation 16c56381-9fb9-4393-9292-6c4bbff720c2 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs OPT: Open Pre-trained Transformer Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.152111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.152111Z digest=sha256:1d6b11836a774b10bfc2a1ce64bac9821fd5220a16721e049871da61f2908c19

Observation 983b002b-3f83-4f76-a676-5d9ef691c433 · outbound

This paper cites Scaling instruction-finetuned language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Scaling instruction-finetuned language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.156459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.156459Z digest=sha256:d038864519b33d4dfc9fe022469a862b48ac196dc0199fdb9a251d7d98eba033

Observation edddaa3f-a9ee-4945-bae4-2ae520539233 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.160454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.160454Z digest=sha256:6031ea0e0b4e5319e1c372cf2fc85f2d63ba1f62e3d09930129c120592a2e4be

Observation 5c14a08b-325f-4be1-aaf3-66564b0eb654 · outbound

This paper cites Introducing mpt-7b: A new standard for open- source, commercially usable llms,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Introducing mpt-7b: A new standard for open- source, commercially usable llms,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.164950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.164950Z digest=sha256:11853a96bcd841fdaaa21f624ecd960964d425d2ffa7764955f3ac09d2cdfd23

Observation 3416eb19-6004-46f8-980b-2cfd858a6705 · outbound

This paper cites Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned and chat models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned and chat models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.169487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.169487Z digest=sha256:ef0ebfe0e3b273aabdaa1ca218ebb6803c65feb90db8c3f4f7254a6a3c5e39e1

Observation bdc303a5-d975-4f94-b9d6-8dc809e39f17 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.173697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.173697Z digest=sha256:2ada6d8afda5441ca05b379260bc7d495608e898640e71fcc8ea159322de444f

Observation 4f146210-a6f6-425d-98f6-66988df75d7c · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.178121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.178121Z digest=sha256:27cbea39368968b72c60c4336228400eaed7647ca0aa3d9b74b4b88e23483379

Observation d114bf67-f0b6-4562-9331-e42f5946f563 · outbound

This paper cites Gpt-4v(ision) system card,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Gpt-4v(ision) system card,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.182522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.182522Z digest=sha256:73dad8e79af0e38202043b185ac6ed430ffc75aab8deb4bfc670f6ba5a8acb00

Observation 8c7cf713-d876-4f54-a548-41487f2a4a10 · outbound

This paper cites GPT-4 Technical Report.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs GPT-4 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.186790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.186790Z digest=sha256:e7f911e1f5e50e9918198ecaec4a80c6dd8a8c6a9773fec1dcd27053f4f9dbc3

Observation 0eb8dd0f-d1cc-4a0c-be49-0ee2c92b8340 · outbound

This paper cites Gemini: A family of highly capable multimodal models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Gemini: A family of highly capable multimodal models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.190925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.190925Z digest=sha256:e523c84dbc3f976625cc0a5746c7511ccf078bbdef973b655b36b3d64d9c6a3d

Observation a4d978bc-4cfc-4a3a-9fa0-4c829556a078 · outbound

This paper cites AI, “Bard,” 2023.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs AI, “Bard,” 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.194855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.194855Z digest=sha256:299e07e1afed38d560be15ea378b963823d6a939d6a4fc04dc133b1a96150dce

Observation b7fd3fdd-3840-4f90-9b56-b47225b8d540 · outbound

This paper cites Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.198969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.198969Z digest=sha256:ff6fa1d57b98f74b5e793401df3851484fd132861e281c440e5f2de39006b5f5

Observation 5625fcc6-37d1-415f-80bb-faa10be4830e · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs High- resolution image synthesis with latent diffusion models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.203461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.203461Z digest=sha256:d572061721836aa2e32343e2f346add41194193c58afacca883671fa319b84cc

Observation a177be31-e27d-4e68-8ddf-e0b80f2dcfb6 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.207019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.207019Z digest=sha256:2a6d0e863e9fb81ec3891b1378cb16d0b0895db5db9baec016382d1c6a45b87f

Observation ed837fdc-ba9a-4819-bc70-5edd775120a1 · outbound

This paper cites Zero-shot text-to-image generation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Zero-shot text-to-image generation,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.210902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.210902Z digest=sha256:f714128316c9f05c1ab8345b0f4dcb8756bd49b98fd3b5c8192094aef3a17f1c

Observation 1100fcc0-ad9e-4ec8-9b38-c0bf3599914a · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Imagenet: A large-scale hierarchical image database,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.214490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.214490Z digest=sha256:f1d382f27887ece86c191d78b73db7f60396f8aaa58ebce4d63e9ff46ca9c207

Observation e3f4e440-2293-4a9d-af2e-25122a822280 · outbound

This paper cites Microsoft coco: Common objects in context,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Microsoft coco: Common objects in context,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.218337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.218337Z digest=sha256:102d9376c1fe53aea1ab984d67d852ae01d59cd785e1d1e843485998d6715ef8

Observation ac85e91c-d502-4609-b080-0ce4f68c9b18 · outbound

This paper cites Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.221963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.221963Z digest=sha256:fd61d9ebe71e85c5f5850a78c9fc6519a7b5bafcf3747dc3a900e54d6649173e

Observation 32692a65-3f1f-48d9-a6b6-76a59f1c9991 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.225531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.225531Z digest=sha256:231e555c9259cbf328fd8baae7afc24daa6b222908b0589eab3e37ace95d15dc

Observation a1c67bea-2577-4d10-a065-eb1c50eb896a · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.229098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.229098Z digest=sha256:16283b690378e215315c3b0c381452d73f119c323d9a36b371e3c2e66de260b6

Observation 1e710829-82aa-40cb-95e4-a2246126d1d5 · outbound

This paper cites Laion-coco,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Laion-coco,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.233181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.233181Z digest=sha256:38b157ff270f16a9f185f220a9c09e2cf3f218bab6d398882f68f58a5daa8cf8

Observation bab26b47-e212-4fc4-9b81-845154b749c6 · outbound

This paper cites Stanford alpaca: An instruction- following llama model,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Stanford alpaca: An instruction- following llama model,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.237161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.237161Z digest=sha256:f3407dad38f8fc271dbb364fec4fc40c1c39f27d3e9723d92d0f9f5edfe70c4b

Observation 3aa30c57-54e4-4291-a717-b60c00693bfe · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.241278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.241278Z digest=sha256:05eabcc14f0226f196d2aa870f36ef9eb475142cc4ae9d2ce896e0cd865672a3

Observation ee3a88df-54f0-4c89-9c74-a5b3a3237360 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.245674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.245674Z digest=sha256:3bf347ff4f3825f12022513fd2ae1b15fa5c0e9786ef55a68269afca6c1c68a9

Observation 46bc22fa-4bd9-47a4-a302-53cf613a35be · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.250038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.250038Z digest=sha256:cbaf6f6a1b53ed1fc457dd3e6986451f5e3b442b50a55bc5486efde6f741b062

Observation f11ade04-d062-432e-aca9-7a1f0d77267a · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.254478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.254478Z digest=sha256:867ab83fcc0d327e3c83430c1f19a3d90b602a8139f6139e0ca6a110014371d2

Observation 521af43c-1569-4d88-9ae6-c5dfa1c93189 · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.258909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.258909Z digest=sha256:e017310455d0df81a80ffbe1c50f598f368380a6c0892aba0248ded8ccf07d20

Observation 41103c7b-89af-4865-9aa5-3ae0ef32b93e · outbound

This paper cites Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.263253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.263253Z digest=sha256:462c7889f9a7520f464e3f55c843d518d3a210dcb4d343f3329b1275407c7284

Observation 7ea2d64e-8301-4b41-9f70-5da442bf8d5c · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.268038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.268038Z digest=sha256:be46fbda77505b8be56a222e4c1d9b6fed9ba3ceb5af4050d99b224da69bb88c

Observation 8f76bc81-b1b7-4a23-8d96-39574b69da3a · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs White-box multimodal jailbreaks against large vision-language models,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.272280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.272280Z digest=sha256:b5d8ce8e4be90b3fc4c6afcd365938642b78f54c3c131f4ef57ad072f690ed0d

Observation 708e9679-eccd-4a47-a55c-56ee013447f4 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual adversarial examples jailbreak aligned large language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.276363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.276363Z digest=sha256:23a851ae588159b2664a603a3d4326bb471dccf8dccd2558c914b4d977c969a1

Observation 678e795e-b5be-45ba-8444-16542a03c308 · outbound

This paper cites To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.280459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.280459Z digest=sha256:8029f7256227ffd2a738c36bb28adea79b447dc06dbbef4e58eeae6da3d09040

Observation d5283b45-37d3-493e-b99f-6cd61d3b3c6c · outbound

This paper cites Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.285032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.285032Z digest=sha256:90793b586da6fd88ecb683cbbbcf79bd85ae516095062f833f82107652393d90

Observation 83fb702f-4757-4424-8500-8ad9bc8d54a9 · outbound

This paper cites Are aligned neural networks adversarially aligned?.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Are aligned neural networks adversarially aligned?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.289387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.289387Z digest=sha256:6a69e2b5305ac66718129230c0d28dd7105154af83ce553f206dd2595449f11a

Observation e90ebd9a-795b-42d7-8e1c-cc832f82d4dc · outbound

This paper cites Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.293414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.293414Z digest=sha256:711a76e1c45fc5d4a20ec45771d9f70da895636eade579d2c36b955aef0bd0ff

Observation fa808e8d-d9f1-40b2-9d2c-2f32762d6ecd · outbound

This paper cites Extracting training data from large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Extracting training data from large language models,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.297509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.297509Z digest=sha256:13a81582ba963dcf1811f861dc66f63c945ee21919487ab51439a08e003b51c6

Observation 1ee4c5d8-a577-4edc-ac4d-6a518faf0807 · outbound

This paper cites Va3: Virtually assured am- plification attack on probabilistic copyright protection for text-to- image generative models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Va3: Virtually assured am- plification attack on probabilistic copyright protection for text-to- image generative models,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.301812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.301812Z digest=sha256:d01f1235409ab93bc0c73b1ec8954f8d83f8ffa4d7229f68d91f493584eea057

Observation 623ce292-9288-4cca-968f-ae441b468e05 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.305857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.305857Z digest=sha256:fe93deead8c833a3e8ce27296240387ac17cf8ae0f8b32a14cd3b9512b646b7c

Observation cabab675-5987-49ed-ba20-d59e8a0236ca · outbound

This paper cites Large language models are zero-shot reasoners,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Large language models are zero-shot reasoners,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.309887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.309887Z digest=sha256:3c359dfd8d0c80e2e65ca2a19c2d9f802dc0a485f970ea9dd61d23c0da9fe014

Observation 24ac7ab2-87d9-43c5-881a-87d8a9933c22 · outbound

This paper cites Black-Box Prompt Optimization: Aligning Large Language Models without Model Training.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Black-Box Prompt Optimization: Aligning Large Language Models without Model Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.314113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.314113Z digest=sha256:ae8660462d8061f73c1f21b4504c4052b5eaf24f180250c33c18b57c59cfaae9

Observation fdf0ceed-5041-4a53-84f5-954b56bd6274 · outbound

This paper cites Training language models to follow instructions with human feedback,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Training language models to follow instructions with human feedback,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.318645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.318645Z digest=sha256:ac0a188988889eef80e5e57329d6d71b8f34c8c74edbbfe4d3a394a043179a87

Observation aca12052-b038-4c5e-ac9e-45d7fa2bc64b · outbound

This paper cites RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.323020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.323020Z digest=sha256:48e4efc114a426f9b0d66f2a71572b65f45bf532321c00b3575409f1e30d5b41

Observation 3020da8a-d357-4371-8332-36ae6a9b8ec9 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural language feedback,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Dress: Instructing large vision-language models to align and interact with humans via natural language feedback,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.327520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.327520Z digest=sha256:1a0ea9dddaa3598f7539224c543b922df52fac2bbae2a354c06a4bb059dad5b5

Observation cc4414d5-f3bf-409e-84a4-11a15b4ce7c6 · outbound

This paper cites MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.331678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.331678Z digest=sha256:6d9206cc623271a127284c99b7db3f483453d031b2c674d84e969225af43a3e3

Observation 5fa57646-fe9f-4fef-ae7e-95bd036355e5 · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.335773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.335773Z digest=sha256:a0fef8dccf95856e42d1083240fe5117e6166917c0ec362e62df922bb1a83a89

Observation 1f6b0708-eff5-43d3-b9f1-d5d481deebad · outbound

This paper cites Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.340113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.340113Z digest=sha256:bfe49ce65ca71bae4509b460b2142d79b3591361858f6b918fcb1261ba2bd704

Observation 99c2107b-9d2b-4ff9-b4ff-020b4345d0ef · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.343991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.343991Z digest=sha256:0c16a890cbcc427820880400402956217b71b09ef591ec46a3fa9aa10c7fc9d0

Observation fb549679-a2df-41f3-ba76-183bff6136c9 · outbound

This paper cites AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.348048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.348048Z digest=sha256:631a07b4b21aca13da79dbd505b16f38fa6b9777c8002ac1bb0874bedc137067

Observation 347259e9-cd97-49d3-8e8d-41f127fdb9cb · outbound

This paper cites Adversarial illusions in multi-modal embeddings,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial illusions in multi-modal embeddings,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.352864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.352864Z digest=sha256:bdda54c4bdf22f7ce30815d75e59f439ed0cb937ef8f578630ab939c7bf3e3ad

Observation e21fc647-f769-4fea-99a0-637a2626a322 · outbound

This paper cites Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.357496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.357496Z digest=sha256:1d8881e96f99040d00128b980fa1a1674195fc113201d9b8ee144cdb0a2fd386

Observation 6cd72d6f-3c28-4821-a83a-82c1c71a52e9 · outbound

This paper cites Badclip: Trigger- aware prompt learning for backdoor attacks on clip,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Badclip: Trigger- aware prompt learning for backdoor attacks on clip,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.361725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.361725Z digest=sha256:12dedd1c00605253cd4d88586d7fc040874e5ea15b1c9d2a1cc8bb29c964eb8e

Observation d5a2d880-df5d-4819-a307-813218e66dd1 · outbound

This paper cites Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.365929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.365929Z digest=sha256:f606aac53d5a6338342db66e3216492bf7d48738924b6a8a0d33d788526e23d5

Observation e43542ff-2dbe-4099-94ca-8c6b8de0e925 · outbound

This paper cites An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.370563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.370563Z digest=sha256:a9e273c56209980fb3d213ea0dd36db35ffb45813e6fc72c41af3de894e63058

Observation b6cfe0d7-418c-4597-842c-c21e74afd408 · outbound

This paper cites Break the visual perception: Adversarial attacks targeting encoded visual tokens of large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Break the visual perception: Adversarial attacks targeting encoded visual tokens of large vision-language models,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.374971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.374971Z digest=sha256:daea3db12e7b5e630119f3286c96670711b2449db394b6417fc1f6283a93669e

Observation 715f514c-6c7c-48bf-9423-830233b485fc · outbound

This paper cites Misusing Tools in Large Language Models With Visual Adversarial Examples.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Misusing Tools in Large Language Models With Visual Adversarial Examples

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.379148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.379148Z digest=sha256:59f6c6e9d001488657a3e410f153f13cb922d5b648339cbcd74294ee038de20d

Observation da0d21db-ba63-41b2-967b-01947986e400 · outbound

This paper cites Prompt-driven contrastive learning for transferable adversarial attacks,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Prompt-driven contrastive learning for transferable adversarial attacks,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.383721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.383721Z digest=sha256:fcb26c078f9e0bda9bf5bb2a32fd4586712d3dc2dbbd8609fbbb648f497a7677

Observation 395a51a3-4556-4e29-8cee-74623948076a · outbound

This paper cites Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.387834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.387834Z digest=sha256:c69871804fcb2c66ca340d95919c53e38b5eb7e1a41ccba70cc24650359c71a3

Observation 24d757c0-3786-4b74-83e2-fbd6c3eb2232 · outbound

This paper cites Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.392810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.392810Z digest=sha256:813332939fab3eabd301865a7aa95c1a848c06a809d520e3ebb497c9c71a4a25

Observation 5cc07c5e-937b-4b1a-8724-f3c748fb577e · outbound

This paper cites Slowlidar: Increasing the latency of lidar-based detection using adversarial exam- ples,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Slowlidar: Increasing the latency of lidar-based detection using adversarial exam- ples,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.397009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.397009Z digest=sha256:0ee9d8471e70452b3f379ea15dfb7161a51b9bbb94bd62091638f8fbddc4b789

Observation 3cc0603d-1d78-488a-b509-d0acdf803e9a · outbound

This paper cites The dark side of dynamic routing neural networks: Towards efficiency backdoor injection,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs The dark side of dynamic routing neural networks: Towards efficiency backdoor injection,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.401241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.401241Z digest=sha256:e4bd372e4bdde3f4ec7dc958bbe604dc707797238c426dfe29cf64b37f572723

Observation 27d46af8-06b9-41d9-8eca-990a446d7cb5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Chain-of-thought prompting elicits reasoning in large language models,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.405556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.405556Z digest=sha256:8f005625c086e3ce2b44f06dea3af2a328d078205b5d8e84a7173c562bfd7cb9

Observation 04bf9728-5647-4b68-8153-ba90901a86f0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.409426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.409426Z digest=sha256:9c4bc34ed77ca41c2b58d66cfbfdb9263d1625218ef4d95df3ebb48a0bc924be

Observation 842c88c7-17ed-4f21-8dbb-e531802b598b · outbound

This paper cites Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.413950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.413950Z digest=sha256:f925b8b0cc1e028a5c0482e7b0222f559923ef681e25fd05c5ec3159a928cbb9

Observation d3cb49b6-d9e9-4162-bbb9-7fc7ad644f0e · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Explaining and Harnessing Adversarial Examples

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.418146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.418146Z digest=sha256:9286dc3653310529b88925d3ff29ffdbacb970a092a9be8a9d510a9e91f3e5e6

Observation 5e13c29d-db2d-4fc5-b128-4ad787e5a1ba · outbound

This paper cites Adversarial examples in the physical world.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial examples in the physical world

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.422466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.422466Z digest=sha256:f52d73ae0fa6652210d9aeaaac7c68fe960018de03fa8046570925279c579325

Observation 63a5e170-3e13-4c20-9864-c3fefbd01bd3 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Towards deep learning models resistant to adversarial attacks,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.427191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.427191Z digest=sha256:0e7ff9cf15604cd3e21909b46789a79a0d4a890147ef868b9b7adb5a53908d21

Observation 96a42497-ffce-4860-ae27-4461d962c1b0 · outbound

This paper cites Boosting adversarial attacks with momentum,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Boosting adversarial attacks with momentum,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.431794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.431794Z digest=sha256:f3c2667d6e4f16b755502db87316acf326f442a78f6b2dd92a3a23056347e941

Observation 7e134122-f8e5-49f0-8480-91fcb4c8d387 · outbound

This paper cites On the adversarial robustness of multi- modal foundation models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs On the adversarial robustness of multi- modal foundation models,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.866336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T15:38:17.436070Z digest=sha256:e2279cc1cbbde9bc40d7ef34476ab163a40cc96df1a5594687caa9a9aa6e702f

Observation 796a9483-6eea-4aed-93bf-0e36a26ff584 · outbound

This paper cites Transferable multimodal attack on vision-language pre- training models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Transferable multimodal attack on vision-language pre- training models,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.852715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T15:38:17.440408Z digest=sha256:fdcaf180f0f2955c36514e0022f843d5a65a47f2f4493dcadfdbc92feb71ca8a

Observation 74fc4860-c25d-4864-acb5-a8d646622cc4 · outbound

This paper cites Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.838566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T15:38:17.444773Z digest=sha256:c02ecc2972e97731a7ee9cbf0c1857321774dd71b4da89c9d4d114c6a252de08

Observation f730f568-73ba-451a-927c-a27f96afe440 · outbound

This paper cites Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.449204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.449204Z digest=sha256:66e9c2dcb6bf5980a27529a8b211d643c1ef70ff8d60f7708d6d381ca59543d4

Observation 28158ef0-3b69-49a8-b5e6-5c1c8098beb6 · outbound

This paper cites Adversarial Patch.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial Patch

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.453227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.453227Z digest=sha256:b6a2bd000f83e27c06597858cfe5a0bff1493747bfb4bc33dee01d7c1d877118

Pith citing papers

No inbound Pith citation observations are available.