Pith. sign in

Paper Citation Record · LEDGER

Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2307.10490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.10490 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:38:15.346908Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8bcc565e-2fd1-4a78-95d7-86c2094b28be · inbound

Whispers in the Machine: Confidentiality in Agentic Systems cites this paper.

Whispers in the Machine: Confidentiality in Agentic Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:03:53.924075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T03:59:03.972043Z digest=sha256:7d42ce930b5974f2731d3466d294bcc761b5031e1e1637d4a40b82607bb3ae4e

Observation ba6f02e6-5b5e-4765-a2d5-7ce7ab6c3470 · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:50.756346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:69d8162ac5be2a9720ad6e7321cb2fd1bb09507c71c29d2f9d134bdf1d741f5c

Observation 2586924b-d8a2-40a5-8ac9-2d56e349ac96 · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 293

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.252950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:87347360b842b7343bf14536a84120f320661c86759d4685e4acdd133549d66f

Observation 5fbed452-07e8-4d5e-bcb8-94d749fcab45 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.262296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.262296Z digest=sha256:79e2557690b8207ec68942bb519aed2a0fb32e3e9b7cb4debc1236cff3a9250e

Observation 72c7d2c3-637d-49f7-9744-6d1eec08e422 · inbound

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion cites this paper.

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:15:14.829234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T00:13:08.603115Z digest=sha256:563bd734e5a192d469e652491bc270638d390787c82a2641b09db4a7749abb54

Observation e177162d-e68a-4e9f-a1e1-382d99dbc1e8 · inbound

Seven Security Challenges in Cross-domain Multi-agent LLM Systems cites this paper.

Seven Security Challenges in Cross-domain Multi-agent LLM Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:34.852642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:34.852642Z digest=sha256:1df3eef5726248ce283642401936d6aaeb1ecc4513ce2d96a2f36eee8d94fbe0

Observation e45d494d-1e18-483f-80c2-3892c392dc02 · inbound

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities cites this paper.

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:58.079163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:58.079163Z digest=sha256:251bd763d212d0101c816784256a534cf6682b883594f8800f1534f7eeec7304

Observation 75da2dfd-2a96-4e97-9eda-f29973741e30 · inbound

Normative Conflicts and Shallow AI Alignment cites this paper.

Normative Conflicts and Shallow AI Alignment Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:54.417585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:54.417585Z digest=sha256:34f7dd0059343106c596ac412f26f5680a717e694f6aba1b0a5d23605fffa910

Observation 80107157-42c9-4510-a17b-5ec0990ce669 · inbound

Prompt Injection 2.0: Hybrid AI Threats cites this paper.

Prompt Injection 2.0: Hybrid AI Threats Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:05.101373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:05.101373Z digest=sha256:0f4c0f0ce72738e95294083793265b2b6ac9fdf0fb48198c878a702acf331d4e

Observation 6cd493d6-d804-492f-b72c-8cde5d1c5bb4 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.603513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.603513Z digest=sha256:cfc0da3a7b5e883a1e80c9c19e18f15bc03375a24c2c8c40f4e98e49e2670d8b

Observation f12926ff-74ab-4146-9cad-112b3fb9769a · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.258454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.258454Z digest=sha256:3f3dbc57f72fff5cc6f638eee3e057269a28245b487d3f5ccb4ea87f1ec454ed

Observation b451794c-2a09-460f-9695-11d6ad1a2e58 · inbound

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous cites this paper.

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:09.526632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:09.526632Z digest=sha256:4dd5473292c05c7df473f281303388b0590ff9e7344ad0bff523a558074fd5f8

Observation d2734992-083f-4ae0-b78f-89aaeff0a021 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:42:22.045980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:0e3caca28c309cb4fdca2c779ee303a14f74ab466e7a8a8c9fb5330b1f62cba3

Observation f42a9e4d-5ec1-476b-9dd8-94feee168b2e · inbound

Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents cites this paper.

Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:01:14.163930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:01:14.163930Z digest=sha256:7716197aadcf625b70a4efd7eee960f19a6f7b20b33458a4853f06659f25e4f1

Observation 8463fab9-571f-4e8e-b7dd-dabc3bfac939 · inbound

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring cites this paper.

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:31:19.422099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:28:41.134253Z digest=sha256:e3059d5e968c61899530f2f5525de3eeba1413cbf08de1ea504271b504cb39f4

Observation 34480220-1d5d-4fee-98a5-820723d20d35 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:bdba15920a70330f454ba94ee046c2e7c8ced07b78573d8a2022908dce2caca6

Observation ba77251d-9df2-4d85-8ca2-bdeb9251ed60 · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.867051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:5aa7e0b7ef7f3292a7e3092352c7ab62fb6b16117020279b8ee79924c2dd52e9

Observation e7e78990-4ca0-4ece-913f-9385457e017b · inbound

Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents cites this paper.

Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:04.276860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:51:47.820649Z digest=sha256:3ec9e7e40f6532139051eed8091590abe29573927bb17f76617757d385b2baac

Observation f7c90804-d766-4002-a5b8-6af4d28cd976 · inbound

MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks cites this paper.

MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:31:07.944342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T21:37:28.671785Z digest=sha256:a5d209243c75157ff33e18eb338aac286231fc534b4fc7c56783ea867d818916

Observation 3d59fd30-10a5-4c33-a868-42015cdb1a8f · inbound

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents cites this paper.

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-08T22:39:20.664808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T08:08:24.524671Z digest=sha256:49518f6571c92e2eee0a4c01e40fe19fcb9d1f674633b252eabce178009a8ec4

Observation 86b07d54-192f-42a8-bf58-7f0780903955 · inbound

Semantic Denial of Service in LLM-controlled robots cites this paper.

Semantic Denial of Service in LLM-controlled robots Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:14.602650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T07:59:42.478294Z digest=sha256:c12b34ed4b6def08629598cfc86d2577b26fe1aa5cbdba09bb6c35512ddc6ed0

Observation 2bf34282-1d7d-4985-99ff-d4341cef654e · inbound

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems cites this paper.

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:26.904379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T09:44:14.993126Z digest=sha256:b53319f963ae53dda95981455a3b09027ac65ba371352e72d1acfee7b0002f9e

Observation 8b73f468-d843-4c22-a309-b934f46880ed · inbound

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models cites this paper.

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:08.007157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T14:24:48.999632Z digest=sha256:d061e6e2592d01214bbcd2dcb26de99225d58c59dfe96fe285dd4ae1900ca752

Observation dd5b62e4-8956-45c1-8185-8601b0578196 · inbound

Laundering AI Authority with Adversarial Examples cites this paper.

Laundering AI Authority with Adversarial Examples Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:08.017954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:19:38.662062Z digest=sha256:003f6d505bb331b2ceb563f945d5a596ef9e412cf170871fe74d65f4d5263d6e

Observation b3141f92-532f-4220-9c54-27138fe399e7 · inbound

Cross-Modal Backdoors in Multimodal Large Language Models cites this paper.

Cross-Modal Backdoors in Multimodal Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.541381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:51:11.432424Z digest=sha256:6af8f9ccf394a5afe4276e7fb22f1f234861c1a275a206fea2f55e7652735168

Observation a11d1780-b5a2-4e01-8618-3e765bf64265 · inbound

Hallucination as Exploit: Evidence-Carrying Multimodal Agents cites this paper.

Hallucination as Exploit: Evidence-Carrying Multimodal Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:48:11.621232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T09:46:42.413501Z digest=sha256:3a148b5e767dbfec3438a9de49673a56e0a5805897ac94b8b81b8a3254bda52a

Observation 37c91339-e6ee-4f0d-a953-64b7410774ae · inbound

Hallucination as Exploit: Evidence-Carrying Multimodal Agents cites this paper.

Hallucination as Exploit: Evidence-Carrying Multimodal Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.902347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T08:57:29.491043Z digest=sha256:159e870bd2445af6e1849900cc5a66c9056c8eb3d2fe92ff13f2cb216ba56814

Observation b2cc4d12-f2ca-4f45-ae83-c5cd23ffa2c7 · inbound

The Surface You Test Is Not the Surface That Breaks cites this paper.

The Surface You Test Is Not the Surface That Breaks Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.206657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T06:37:19.674012Z digest=sha256:5c2a537bc308d70920b9ee4e776befe9f4eb441b8acf47d15add98141867a5dc

Observation f1131dda-ded6-43fc-b755-4e2713447359 · inbound

HLL: Can Agents Cross Humanity's Last Line of Verification? cites this paper.

HLL: Can Agents Cross Humanity's Last Line of Verification? Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.607755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:57:57.218669Z digest=sha256:54650c9c55fee4e2e219a00122275438889d9c4c77c1a5c4953988d4f8df9300

Observation 10cc6aee-624b-410a-a368-a9a0c102c966 · inbound

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models cites this paper.

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:34.719751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:10:38.261635Z digest=sha256:dec6d2928e5a2a1b821005c733bde2aa78e0a5a9787aaf592713c64f0d68ffae

Observation acb5163e-3814-4a3b-85b2-f480b497ffa0 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:31:29.301820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:25:07.890796Z digest=sha256:943e5c9bd81c7328062ab1ced3f5540b339ba0c78993d4dd4566f5178850eb73

Observation bfa57969-40fc-4e2e-8e8c-8cc1183fea77 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T02:14:12.243970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:14:12.243970Z digest=sha256:b0f3557eaee67d2319858bd2f8fef86d42132571e61b8642651886c618332ae2

Observation e9c648c7-be23-4f06-805c-ed9d7d789a84 · inbound

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices cites this paper.

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T13:02:42.673767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:02:42.673767Z digest=sha256:26018915f6dc1b7e55c853362f76e3f7dc5f09ad7ffdb212b1d8fe65dda71fdc

Observation 3ca2b105-d4f9-4efc-a76a-fcd377754eae · inbound

Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents cites this paper.

Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:43:50.536410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:43:50.536410Z digest=sha256:31f25f63807e35ae95113bb6de0746c0d8d62b6b6174ed079511fc40ac89ec51

Observation 9b7750ee-0b58-4992-bd03-9af838bcdeb8 · inbound

Agent Security Needs Redefinition through a Holistic Framework cites this paper.

Agent Security Needs Redefinition through a Holistic Framework Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 295

Resolution
unresolved
no resolver link, observed 2026-08-01T06:04:46.731247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:04:46.731247Z digest=sha256:0b79b328f9fe5be767b15cda137cc257d9a0f82198a38ae5f0b91b901bcf1361

Observation 293a4c02-1cd1-473b-8c3d-18967550e9d9 · inbound

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails cites this paper.

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:20:22.213214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:20:22.213214Z digest=sha256:9ef93b2e3c44e04ebd5446f00518b4b0703208699607ecd8df4637ba07a86cfe

Observation 6e5903d9-8340-4ce3-87d2-e302054779fd · inbound

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots cites this paper.

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:15.346908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:15.346908Z digest=sha256:64395f786b7b573e9165860d8fff3ca37910e0d118587949666bb5cf868286f0