Pith. sign in

Paper Citation Record · LEDGER

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM

As of 18 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2505.04673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04673 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:36:01.163570Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:16:36.682289Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T22:11:14.069836Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2372d5ee-7081-44f0-8df0-f25d95aa2df6 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.936939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.936939Z digest=sha256:1c83e9589b5224cf71fccd3c17e0c06575bf5b43e9d031ebe7561191faece858

Observation 48a516ca-e40e-497c-89ed-d4fa57b437c1 · outbound

This paper cites Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.941054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.941054Z digest=sha256:5c61ac9e36ec0748c512693e57f15e08348604962cd3e353996403a88841d71a

Observation 47a98ffb-9fb4-433b-a014-9e1c0adcd38e · outbound

This paper cites Image H ijacks: Adversarial images can control generative models at runtime.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Image H ijacks: Adversarial images can control generative models at runtime

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.906104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.944987Z digest=sha256:e157a684788cf45ef5f762f2a007e4a716278103cac7c09f3ae0035948585f42

Observation 5550279d-2f90-48d5-9708-77d1ac0536e8 · outbound

This paper cites The dark side of language models: Exploring the potential of LLM s in multimedia disinformation generation and dissemination.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM The dark side of language models: Exploring the potential of LLM s in multimedia disinformation generation and dissemination

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.895796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.948624Z digest=sha256:7a5873836c124d900ccf58d26085feeea776904c53081eeb2dfae40f719b076f

Observation bdb41b61-e870-47a6-8159-099eb3e5142d · outbound

This paper cites Easily accessible text-to-image generation amplifies demographic stereotypes at large scale.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Easily accessible text-to-image generation amplifies demographic stereotypes at large scale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.885792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.951788Z digest=sha256:65b9400ed5f8f5b7fda7f1980d7b5967ecc88aaba10419956de52cd8ba31567f

Observation 87a81ea1-e6e3-4302-947d-adb633e63310 · outbound

This paper cites Distilling adversarial prompts from safety benchmarks: Report for the A dversarial N ibbler C hallenge.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Distilling adversarial prompts from safety benchmarks: Report for the A dversarial N ibbler C hallenge

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.875981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.955506Z digest=sha256:d31c7747bd511fbb2d996042647bc218d76aa1f8d16725949f5f1394b498b84a

Observation 2c8229d2-d390-4361-8ffd-d3a3ebe5047b · outbound

This paper cites Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems , 36, 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems , 36, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.866330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.959238Z digest=sha256:e003640aaa305e6a187cd291ca4cc6496327dba54f216f2cf2e954878cc8b622

Observation b25722db-b863-4da2-85d8-4a4e2557be9a · outbound

This paper cites A survey on adversarial attacks and defences.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM A survey on adversarial attacks and defences

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.856119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.962648Z digest=sha256:6b9a50d0f6b986c6e128909540428ba5cfcee798defefbeca56fd8349c4d4911

Observation 83808a87-d0ed-46a9-bf71-4f808a0b41a3 · outbound

This paper cites Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.966127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.966127Z digest=sha256:004b1fa3802c31acd9950f3842a4c3abf972f32a450864c6b12b8e2ae0e19315

Observation aad88ebe-9bd5-4d67-a113-0c3931a5c8dc · outbound

This paper cites Leveraging the context through multi-round interactions for jailbreaking attacks, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Leveraging the context through multi-round interactions for jailbreaking attacks, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.845505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.969752Z digest=sha256:430d4d233e33731717d646d4e5fc9f10f66d50e334ad897e9ed386ba8ccfb5be

Observation d4397a30-703e-47c1-a722-72d383601818 · outbound

This paper cites Dall- E val: Probing the reasoning skills and social biases of text-to-image generation models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Dall- E val: Probing the reasoning skills and social biases of text-to-image generation models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.835852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.973097Z digest=sha256:8fa84061848c857e3cf0aff133d2916a51d23e1a6f6e56caf204db9d6266a51d

Observation e085f569-71fc-4b7e-b728-73d2c5e34292 · outbound

This paper cites Training verifiers to solve math word problems, 2021.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Training verifiers to solve math word problems, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.825556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.976722Z digest=sha256:98b50f069e77202a7f4d105086fc69d0af5e19a9e2425fb2c09a797d97d5b2bd

Observation 4ace2c42-49cf-4d13-89e1-ad431fcdb04a · outbound

This paper cites Towards safer generative language models: A survey on safety risks, evaluations, and improvements, 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Towards safer generative language models: A survey on safety risks, evaluations, and improvements, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.814848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.980548Z digest=sha256:e7d51b2dca41023584c52faa12ba71fa54ab31130dc9639bf390cc9739dbf5cf

Observation 9fe23c88-b7dd-4b71-9535-8c925f49685c · outbound

This paper cites How Robust is Google's Bard to Adversarial Image Attacks?.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM How Robust is Google's Bard to Adversarial Image Attacks?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.984114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.984114Z digest=sha256:db1b5059b0156714216baa13d94250f1cbb4e1aaf9b813fa87823c2b5fb191f6

Observation d9973e3f-764d-4c42-9752-982c0cbd4bd6 · outbound

This paper cites Attacks, defenses and evaluations for LLM conversation safety: A survey.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Attacks, defenses and evaluations for LLM conversation safety: A survey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.803679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.987827Z digest=sha256:9b85913b0b3c14ae1e6643b0d82dfa2d5af7cd29034afe5b8c036ee2b7c9ec70

Observation b51987b6-11e6-4841-b5ec-5e8289c5b633 · outbound

This paper cites ROBBIE : Robust bias evaluation of large generative language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM ROBBIE : Robust bias evaluation of large generative language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.792981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.991514Z digest=sha256:0dd892aeab0e0d7877b7b6bf0e7d1a85ba4902b8865d9df24a1c86fd12fbe4c3

Observation 1d9547d3-307b-4888-9c84-c4832b70f0ad · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.995456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.995456Z digest=sha256:36497c177c1b4002ed56e4fc40edd733179092cafc7f642aec02343d3ec6f2dc

Observation aea58e4f-8da1-4351-9eab-4a85a2103550 · outbound

This paper cites [Online; accessed 05-October-2024].

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM [Online; accessed 05-October-2024]

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.782681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.999289Z digest=sha256:238a176680adb8c515a7921c533816bbd1b2fc8af5d5b4d7c7b4fe71b541fee6

Observation 02a18049-04bc-43a2-9aaa-e8992ccd1e1b · outbound

This paper cites MLLMG uard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM MLLMG uard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.772435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.002992Z digest=sha256:5e6a8b6f61c8733dcd294df983a8270a42dfa89d85cca29cc50773c67b66761c

Observation cd1f4572-39f2-43cb-bf5a-500c313d6112 · outbound

This paper cites Harm Amplification in Text-to-Image Models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Harm Amplification in Text-to-Image Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.006640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.006640Z digest=sha256:e027ad87264c05b90182802663ddd71924f5c44566507fcef4d31980c45ba4ca

Observation 396f1031-cb72-4e8b-8f1d-9a8403c78231 · outbound

This paper cites Evil P rompt F uzzer: generating inappropriate content based on text-to-image models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Evil P rompt F uzzer: generating inappropriate content based on text-to-image models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.761906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.010581Z digest=sha256:6aa8e5e22cedfd30e5df41411839a031fb60c7350f8e792522d71d186ba9528f

Observation 5dc5bbcb-2efd-416f-ae5f-1332e53c2f77 · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.014139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.014139Z digest=sha256:1b22713833498411451badad2e8c1ad17579588276f70bbaba6d548cc5155be8

Observation 41fe0290-f5e6-450d-b575-5f2c41672e41 · outbound

This paper cites Cornwell, Nicole S.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Cornwell, Nicole S

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.751565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.017982Z digest=sha256:08b2e60028b2c8e5bffeb77516fa41804ec3163e580e8b089b17108934642cff

Observation 4b2c054c-5e2b-415d-820b-f8314af59631 · outbound

This paper cites LLM defenses are not robust to multi-turn human jailbreaks yet, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM LLM defenses are not robust to multi-turn human jailbreaks yet, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.741267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.022436Z digest=sha256:058a8a945c094c757e4f21509b9da81403892e0cb427e53b060edfb22367c133

Observation 295c2b7b-39a2-4492-8f95-db11a3c980c9 · outbound

This paper cites Images are A chilles' heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Images are A chilles' heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.729610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.026067Z digest=sha256:ea7224f2243c4e5b646ac1a063843bf39709f5e046845277d693b1144db8d53b

Observation ff1ed617-cf5c-44bc-a4c0-27f671e016a5 · outbound

This paper cites Holistic evaluation of language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Holistic evaluation of language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.719137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.030461Z digest=sha256:c786bd3693708263dfb3af787f6dd4dcea34a6c522b478e77df33e528a3b0a21

Observation 850ad2b9-03f8-4803-bbc3-0aaf1d9a6ecc · outbound

This paper cites VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.034418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.034418Z digest=sha256:5b8592fecdb4cf04e3718e781ef9ad90f5f3ba5ce1b8cc117cd5d81db8b9df98

Observation d1ef7b5b-9ecc-4401-a847-952e7c867919 · outbound

This paper cites Delving into transferable adversarial examples and black-box attacks.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Delving into transferable adversarial examples and black-box attacks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.708281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.038652Z digest=sha256:922de3c4eb0adff1c13b7735ebaa103fbaef5d82c1cdfe9c2dc8fd18a0766e62

Observation 4150a70b-a954-4d94-a51a-053aae5247e3 · outbound

This paper cites A survey of attacks on large vision-language models: Resources, advances, and future trends, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM A survey of attacks on large vision-language models: Resources, advances, and future trends, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.697902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.042084Z digest=sha256:1b449215fd8a4e8a386ebb425568a2e7c66cbf747e0da2ee51bc63c80b48bd2f

Observation 2223b08f-5ea8-4a2a-86b6-e711f21453a9 · outbound

This paper cites Safety of multimodal large language models on images and text.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Safety of multimodal large language models on images and text

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.686545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.045504Z digest=sha256:192fcade61770ee275f0d9c9cc43b891563b965e2081b90e22839844a09edcee

Observation ba83cc34-a509-4fa1-a232-fa15bae9940a · outbound

This paper cites an unresolved cited work.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:36:01.675464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.048984Z digest=sha256:f3474eca064c0eca1574a0627133393131d49e32adb8a4de1c3b351146fe829d

Observation 85e058f3-da95-4bbc-ad77-53495dfcf1d9 · outbound

This paper cites MM - S afetybench: A benchmark for safety evaluation of multimodal large language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM MM - S afetybench: A benchmark for safety evaluation of multimodal large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.665063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.052383Z digest=sha256:cff9029860ed5f772dcc48968967dfe7df53a780e25cc3804ede1a2b57858fdb

Observation 2f441fce-e5c0-4540-ad47-130bd4c51eb9 · outbound

This paper cites [Online; accessed 31-December-2024].

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM [Online; accessed 31-December-2024]

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.653672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.055820Z digest=sha256:cd42cbab89f47afdc7fc26f8ee06d1239fe92d068af5eb85fbb0bb40c82b5817

Observation 08fd376c-700e-4510-b3f5-15ee6f6cd599 · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.059479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.059479Z digest=sha256:3268b76576f3842a630bc62c0da6fc981a00b184980ec419498a894d53b81500

Observation df17114a-f348-44cf-b49b-b5e5fe0b8f6b · outbound

This paper cites Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.063482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.063482Z digest=sha256:9eaff1d8b1148f0b3182e36fb52dec2b42ea76f0f12912323e601469666b8ba8

Observation 24d95985-8c56-41a4-9153-91cac18dfdbf · outbound

This paper cites Announcing microsoft copilot, your everyday ai companion - the official microsoft blog, 11 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Announcing microsoft copilot, your everyday ai companion - the official microsoft blog, 11 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.642044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.067549Z digest=sha256:cf2cac68c7b44a3c73b92eca488a4b19cc3bb655f324743585a8fa48393e5182

Observation 343593da-3519-4c35-aee5-64fa2e339303 · outbound

This paper cites Jailbreaking Attack against Multimodal Large Language Model.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Jailbreaking Attack against Multimodal Large Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.071174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.071174Z digest=sha256:d2d8b4654d3a82a2daaf1c27c372f3c4c2089d9f3f7428745f6fe85279d86295

Observation d9cf6cd2-5163-4e40-bcb4-2255d29aadbf · outbound

This paper cites Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.075183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.075183Z digest=sha256:ed3aa8840216ad057400522be55cf6ce019e1f6c3a0f1c21bf4d7cb8dbcfed06

Observation bae6b810-12aa-41a6-b283-86d77f970c80 · outbound

This paper cites [Online; accessed 05-January-2025].

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM [Online; accessed 05-January-2025]

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.628971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.079208Z digest=sha256:4be19ef349ae44c97d2830c1e551d347ed11343e975904686456526a363529bc

Observation 38dfb18d-45e2-45e5-889d-462ea2233bf6 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Visual adversarial examples jailbreak aligned large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.617473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.082903Z digest=sha256:cc0f4eb5c17f862d56d790cfacca203cc9867f08017aa6280fc63febf72d7b09

Observation e64c1b0a-1fa3-440e-879f-c31e1c13123c · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.605744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.086609Z digest=sha256:2194a954d66eed8737f05b0ff30059841e57310fe744489e36811ccb4a83dba4

Observation 5e6a4370-cbf8-4977-98f5-08c30578c994 · outbound

This paper cites Adversarial N ibbler: An open red-teaming method for identifying diverse harms in text-to-image generation.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Adversarial N ibbler: An open red-teaming method for identifying diverse harms in text-to-image generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.594532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.090169Z digest=sha256:ab0b4372287efe5138528e47ff8e4973f752a6511f4530d075c885b3c94143f1

Observation 505a22bf-59bc-4e97-90e0-069b3bb34dd2 · outbound

This paper cites Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.581711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.095201Z digest=sha256:ed23bb553599c256c8d0fadcc402e79e109b3d465e4b210b7c0b8915e6b4a82e

Observation 01ac96f5-a074-4c35-92ce-a22bc80d5f0f · outbound

This paper cites Assessment of Multimodal Large Language Models in Alignment with Human Values.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Assessment of Multimodal Large Language Models in Alignment with Human Values

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.099181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.099181Z digest=sha256:99855ccd90883d06c95bae44bf1ce617aa9a870c1caaa1e43b8f3fa5f841c9f3

Observation 0e91b285-e53d-416b-8fa0-571fd765f467 · outbound

This paper cites Imgtrojan: Jailbreaking vision-language models with ONE image.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Imgtrojan: Jailbreaking vision-language models with ONE image

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.103397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.103397Z digest=sha256:56d4288c6c0b85ff0db21396a4c4cd6b2f865e1c98a86d871ccc533fa2d5ea1c

Observation dc49285e-9cfd-4a96-b0f4-82a6dcd0cb4f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.107340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.107340Z digest=sha256:2ee96aa18a76cac0a40a182447c9793cab594601e0c92afa0c19985a4c5445ab

Observation c05b9d81-11c6-4893-9f33-3733c3f174f7 · outbound

This paper cites ALERT : A comprehensive benchmark for assessing large language models' safety through red teaming, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM ALERT : A comprehensive benchmark for assessing large language models' safety through red teaming, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.563108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.111291Z digest=sha256:ca9508572627ababfa5870d154400bd8e35e49b4ae1b544b7abce47e0a748633

Observation f40b87ec-6283-4cc3-8247-c58d55d65103 · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.115189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.115189Z digest=sha256:5452bff016b532eebd42a801fbc6ff3b8bf5cce33c0f3be5e21fa21a3b0cfde3

Observation 75456340-9d0d-413b-8d49-5c5e808aeab2 · outbound

This paper cites ToViLaG: Your Visual-Language Generative Model is Also An Evildoer.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM ToViLaG: Your Visual-Language Generative Model is Also An Evildoer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.120073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.120073Z digest=sha256:97a10c0dfc7de6dcf6c4fc54b7743569edb0fa9eef94192d9f01ff5b9cd2bd1c

Observation 55f0483d-f863-4091-8bab-74a92c058ab2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.124050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.124050Z digest=sha256:51166652207185242b110a683bf46dfb3b340e66a0b32394512f36cede2b116c

Observation faf74bb5-3383-4e00-b762-3e72e7682dd1 · outbound

This paper cites Sociotechnical safety evaluation of generative AI systems, 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Sociotechnical safety evaluation of generative AI systems, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.550377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.128221Z digest=sha256:5d9c0504d6c6d45e53e55036a521e4fc8ddd6bdb44668fec490d47a46bc6fa5c

Observation e9945fb7-48dc-416e-b608-1f4c2cbc526d · outbound

This paper cites Vision-Language Models: Unlocking the future of multimodal AI , 12 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Vision-Language Models: Unlocking the future of multimodal AI , 12 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.538169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.131691Z digest=sha256:5817b3f0f223674c00de2e1420e704cb811c694044d3ce1a7a78bb65235597ae

Observation 4fe39c17-d9e7-4e30-817f-f419e3633f83 · outbound

This paper cites Adversarial attacks and defenses in images, graphs and text: A review.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Adversarial attacks and defenses in images, graphs and text: A review

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.525866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.135055Z digest=sha256:422b7c386398e695930e5af3be0073da7b5852be95867567d7e4f85fcdfdea55

Observation f1b9bbe3-33fc-4e7a-b605-36c0cb05519d · outbound

This paper cites Viassist: Adapting multi-modal large language models for users with visual impairments, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Viassist: Adapting multi-modal large language models for users with visual impairments, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.513952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.138742Z digest=sha256:f27ef7435c85fcfbb7da5d3935d7edb36370e37b954bc0d4adb90e732b59d58a

Observation 036e7825-fc34-46d0-bf73-6f22e4389baa · outbound

This paper cites Chain of attack: a semantic-driven contextual multi-turn attacker for LLM , 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Chain of attack: a semantic-driven contextual multi-turn attacker for LLM , 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.501981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.141844Z digest=sha256:befb61c05905dce0ddb948f051a3934723c6031064339ea2b0c1ab32abe6cda5

Observation d62abf6c-e3a1-4a8c-8a7c-c4f1eca6f15a · outbound

This paper cites Sneakyprompt: Jailbreaking text-to-image generative models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Sneakyprompt: Jailbreaking text-to-image generative models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.489543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.145783Z digest=sha256:9d7fda5786e3d09f1c9f83a6111372f2a27f370e48a2f86a51e8453d7b45d07f

Observation c46a4fdc-382a-4cba-9e48-9af2e9b7f6be · outbound

This paper cites Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.148983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.148983Z digest=sha256:2299999c4441b7ca0dc443ce9dc0c80349cf9b575058334dafc8df41a0b420e9

Observation 88c7ec34-432e-40ac-a724-9ade7997f920 · outbound

This paper cites Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.152471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.152471Z digest=sha256:10ebe26b022449453b42b49a19cc53ef28fb14ad56e9c1e7a8f7436b30b34893

Observation 11317825-d672-4d24-bda2-2901ae2e76c4 · outbound

This paper cites Xing, Joseph E.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Xing, Joseph E

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.156013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.156013Z digest=sha256:15501ce243655132e4c6796d19ce037710b3f84b60e5e304c6912a29d5b01479

Observation 41966232-8d32-41d4-8b75-b14f35606361 · outbound

This paper cites Speak out of turn: Safety vulnerability of large language models in multi-turn dialogue, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Speak out of turn: Safety vulnerability of large language models in multi-turn dialogue, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.470271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.160398Z digest=sha256:bfd42c5afd47774cc0d340020572ffe0980e9e607a866449eb4bd5626d31b226

Observation 0842f708-8c28-4f6f-a1d5-0a74301cfa04 · outbound

This paper cites write newline.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.163570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.163570Z digest=sha256:746ae33b6f0b9d9f3d82f11562f486cce6b4602fcf9ef0a662b467f492413f60

Pith citing papers

Observation 1173e618-490e-49e9-90bd-8aa57e633551 · inbound

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models cites this paper.

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:36.682289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:36.682289Z digest=sha256:7349eba65bd4780033f994e0ad173810249c9d7308c60668e126c2a3b4d90279

Observation 44b637bf-86dd-49d5-ad06-3f37d12b7ea1 · inbound

Jailbreaking Frontier Foundation Models Through Intention Deception cites this paper.

Jailbreaking Frontier Foundation Models Through Intention Deception REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.077456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T03:17:51.039062Z digest=sha256:be5b5e40ea834b78484ba8b43530a035b18c2937b37b3834436658d8c11c7de8