Pith. sign in

Paper Citation Record · LEDGER

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM

As of 18 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2505.04673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04673 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:36:01.163570Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:16:36.682289Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T22:11:14.069836Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2372d5ee-7081-44f0-8df0-f25d95aa2df6 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.936939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.936939Z digest=sha256:6f11124b4705df9ad1229cbca7f1c071907959a38b342ce02108cf9cca1b5294

Observation 48a516ca-e40e-497c-89ed-d4fa57b437c1 · outbound

This paper cites Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.941054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.941054Z digest=sha256:a043370d8ac4e2f767912e29d435c323269096be7e7b3f2bc122f75f7b9422b8

Observation 47a98ffb-9fb4-433b-a014-9e1c0adcd38e · outbound

This paper cites Image H ijacks: Adversarial images can control generative models at runtime.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Image H ijacks: Adversarial images can control generative models at runtime

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.906104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.944987Z digest=sha256:879e77e6e9ce866237d83dff12f5887cda76a9ed4d7e5af1263701622081eef1

Observation 5550279d-2f90-48d5-9708-77d1ac0536e8 · outbound

This paper cites The dark side of language models: Exploring the potential of LLM s in multimedia disinformation generation and dissemination.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM The dark side of language models: Exploring the potential of LLM s in multimedia disinformation generation and dissemination

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.895796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.948624Z digest=sha256:4372f71638959e5a5b4f516159143322f41a7ea74f65daeaaa1179035d0f314d

Observation bdb41b61-e870-47a6-8159-099eb3e5142d · outbound

This paper cites Easily accessible text-to-image generation amplifies demographic stereotypes at large scale.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Easily accessible text-to-image generation amplifies demographic stereotypes at large scale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.885792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.951788Z digest=sha256:c80279b3b76c88396c4776582fde8f84d5e100c3d23d96ba0bda3d36dad7cfb7

Observation 87a81ea1-e6e3-4302-947d-adb633e63310 · outbound

This paper cites Distilling adversarial prompts from safety benchmarks: Report for the A dversarial N ibbler C hallenge.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Distilling adversarial prompts from safety benchmarks: Report for the A dversarial N ibbler C hallenge

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.875981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.955506Z digest=sha256:bd4c694e76b6a5588641e32a65d263a0cae27e5f0b82a1c9b584e93e7b31667f

Observation 2c8229d2-d390-4361-8ffd-d3a3ebe5047b · outbound

This paper cites Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems , 36, 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems , 36, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.866330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.959238Z digest=sha256:bddad332ebe3389d8d65e3de2c6255ed32263712be6d12e8917de64c0de21c81

Observation b25722db-b863-4da2-85d8-4a4e2557be9a · outbound

This paper cites A survey on adversarial attacks and defences.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM A survey on adversarial attacks and defences

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.856119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.962648Z digest=sha256:cded24bd9eb5508a9aa16571917df2c30d37345d1c5f0f9f09b1e260c0c5b62c

Observation 83808a87-d0ed-46a9-bf71-4f808a0b41a3 · outbound

This paper cites Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.966127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.966127Z digest=sha256:ef9dadbd61b9d5c6f143e115d716c4a427f1080bf8208272ead999d8a6e9eaba

Observation aad88ebe-9bd5-4d67-a113-0c3931a5c8dc · outbound

This paper cites Leveraging the context through multi-round interactions for jailbreaking attacks, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Leveraging the context through multi-round interactions for jailbreaking attacks, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.845505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.969752Z digest=sha256:5159d19afdd23fd47da7aaec6b33ec88d5cd2511faafd130ac0068bb889bd2d5

Observation d4397a30-703e-47c1-a722-72d383601818 · outbound

This paper cites Dall- E val: Probing the reasoning skills and social biases of text-to-image generation models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Dall- E val: Probing the reasoning skills and social biases of text-to-image generation models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.835852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.973097Z digest=sha256:fce3d25c533d313f4001f390bda19cd4a69fb2f418a28448e921abaf5c9852c5

Observation e085f569-71fc-4b7e-b728-73d2c5e34292 · outbound

This paper cites Training verifiers to solve math word problems, 2021.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Training verifiers to solve math word problems, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.825556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.976722Z digest=sha256:9d458d4994d82c66a01e948f960888895c7a85e106ffaa198260e7659bb6059c

Observation 4ace2c42-49cf-4d13-89e1-ad431fcdb04a · outbound

This paper cites Towards safer generative language models: A survey on safety risks, evaluations, and improvements, 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Towards safer generative language models: A survey on safety risks, evaluations, and improvements, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.814848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.980548Z digest=sha256:153d443bdcde95d6c5c0b8b7236823f36a767af2292a0e3e8f2a29823d065363

Observation 9fe23c88-b7dd-4b71-9535-8c925f49685c · outbound

This paper cites How Robust is Google's Bard to Adversarial Image Attacks?.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM How Robust is Google's Bard to Adversarial Image Attacks?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.984114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.984114Z digest=sha256:21026b9e41d19b89077b165131c34fd432305486a25d36b589bf48a1bbcf822b

Observation d9973e3f-764d-4c42-9752-982c0cbd4bd6 · outbound

This paper cites Attacks, defenses and evaluations for LLM conversation safety: A survey.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Attacks, defenses and evaluations for LLM conversation safety: A survey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.803679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.987827Z digest=sha256:ccfbc1c939b9cc982ee151a51b1ffef7872f2a95757a9867228b9f371dbdec39

Observation b51987b6-11e6-4841-b5ec-5e8289c5b633 · outbound

This paper cites ROBBIE : Robust bias evaluation of large generative language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM ROBBIE : Robust bias evaluation of large generative language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.792981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.991514Z digest=sha256:cf7b9a13411de8dc4e51ab958dbc3c2302969cb168c3f059b0e0b1c0da5e4021

Observation 1d9547d3-307b-4888-9c84-c4832b70f0ad · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:00.995456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:00.995456Z digest=sha256:d4b442a4c339fe7de0e443622d11da49be24e47eeccc0ee619f7705b457a77fc

Observation aea58e4f-8da1-4351-9eab-4a85a2103550 · outbound

This paper cites [Online; accessed 05-October-2024].

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM [Online; accessed 05-October-2024]

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.782681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:00.999289Z digest=sha256:71c7db398b0bf8fbc80bb4280670cf79929e5092a2552f8b9b8a59867b20186e

Observation 02a18049-04bc-43a2-9aaa-e8992ccd1e1b · outbound

This paper cites MLLMG uard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM MLLMG uard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.772435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.002992Z digest=sha256:deb061d027006c9d9b95045678c48290adcc4506a9c272bb851afcac54a6f533

Observation cd1f4572-39f2-43cb-bf5a-500c313d6112 · outbound

This paper cites Harm Amplification in Text-to-Image Models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Harm Amplification in Text-to-Image Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.006640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.006640Z digest=sha256:f55ee21d771eb7549470de6ee02142a6c57c379c63cb1ea52218bb30a0c3d143

Observation 396f1031-cb72-4e8b-8f1d-9a8403c78231 · outbound

This paper cites Evil P rompt F uzzer: generating inappropriate content based on text-to-image models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Evil P rompt F uzzer: generating inappropriate content based on text-to-image models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.761906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.010581Z digest=sha256:39b4db5b416eb80c324a5d54cd339235bbf5305f1a0c808685d69d500482f9dd

Observation 5dc5bbcb-2efd-416f-ae5f-1332e53c2f77 · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.014139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.014139Z digest=sha256:9614cbe1ef62e20d99d8fe828471311c65bd7c2094077c88b353ee74969d030c

Observation 41fe0290-f5e6-450d-b575-5f2c41672e41 · outbound

This paper cites Cornwell, Nicole S.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Cornwell, Nicole S

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.751565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.017982Z digest=sha256:f735031db57d7c5baf922a0f7197a5063695a1eb34bb1b2668c09f6f25b29ff9

Observation 4b2c054c-5e2b-415d-820b-f8314af59631 · outbound

This paper cites LLM defenses are not robust to multi-turn human jailbreaks yet, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM LLM defenses are not robust to multi-turn human jailbreaks yet, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.741267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.022436Z digest=sha256:0a5fab47f180dbe9e0be445f33eea510a4327378cd31c0fa51b9d2891723b979

Observation 295c2b7b-39a2-4492-8f95-db11a3c980c9 · outbound

This paper cites Images are A chilles' heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Images are A chilles' heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.729610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.026067Z digest=sha256:c1ab544fb277174d27817bcd134b4573207ebeb648913cf3b0d6bbab23569aa4

Observation ff1ed617-cf5c-44bc-a4c0-27f671e016a5 · outbound

This paper cites Holistic evaluation of language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Holistic evaluation of language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.719137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.030461Z digest=sha256:dc07ad8c4c8642440c92077c715ac168ab86d40a2a9e82c35ace34ab840ec67c

Observation 850ad2b9-03f8-4803-bbc3-0aaf1d9a6ecc · outbound

This paper cites VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.034418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.034418Z digest=sha256:64f583e78b3b691353a9aafadb08a870fa1bde10b00891b4c53c677ab6cba09c

Observation d1ef7b5b-9ecc-4401-a847-952e7c867919 · outbound

This paper cites Delving into transferable adversarial examples and black-box attacks.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Delving into transferable adversarial examples and black-box attacks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.708281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.038652Z digest=sha256:6d6119b569f9e29a521cea3f70b60b4e86ee00c75d56b57740afaa7aa1aabe7e

Observation 4150a70b-a954-4d94-a51a-053aae5247e3 · outbound

This paper cites A survey of attacks on large vision-language models: Resources, advances, and future trends, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM A survey of attacks on large vision-language models: Resources, advances, and future trends, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.697902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.042084Z digest=sha256:f87f569e9823672da8e712936e7b8633c23ad21cec8d4658a7efa057ca9d3ea0

Observation 2223b08f-5ea8-4a2a-86b6-e711f21453a9 · outbound

This paper cites Safety of multimodal large language models on images and text.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Safety of multimodal large language models on images and text

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.686545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.045504Z digest=sha256:9304ab940fcb0fbbf92c322df8d6e600165e4d9a4d45231bc14b061f914ec121

Observation ba83cc34-a509-4fa1-a232-fa15bae9940a · outbound

This paper cites an unresolved cited work.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:36:01.675464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.048984Z digest=sha256:283fd87657e4944b02e72599442428a66c022e624c28234b7000f92a43d33c37

Observation 85e058f3-da95-4bbc-ad77-53495dfcf1d9 · outbound

This paper cites MM - S afetybench: A benchmark for safety evaluation of multimodal large language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM MM - S afetybench: A benchmark for safety evaluation of multimodal large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.665063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.052383Z digest=sha256:ab28079aaf0cef62c21efc555537280aecd95f6e4de027b21ea0809a795684b1

Observation 2f441fce-e5c0-4540-ad47-130bd4c51eb9 · outbound

This paper cites [Online; accessed 31-December-2024].

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM [Online; accessed 31-December-2024]

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.653672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.055820Z digest=sha256:a7fa3f514e69beb67deeb8ffb6a151865688893da03f1d414e40c475170c87b1

Observation 08fd376c-700e-4510-b3f5-15ee6f6cd599 · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.059479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.059479Z digest=sha256:749b4477ab2866aa3d6e3ea028621a90cbf2a993c975219abf61d235fb8d71d4

Observation df17114a-f348-44cf-b49b-b5e5fe0b8f6b · outbound

This paper cites Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.063482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.063482Z digest=sha256:9ca70adfba27f90fd8f88a167b7d3b027593cc73775c6bb1f57c207fc7369bd9

Observation 24d95985-8c56-41a4-9153-91cac18dfdbf · outbound

This paper cites Announcing microsoft copilot, your everyday ai companion - the official microsoft blog, 11 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Announcing microsoft copilot, your everyday ai companion - the official microsoft blog, 11 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.642044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.067549Z digest=sha256:bdfb3199a913191c49d3b6e6280f2ec72d269120040f7522de46338488aa46b3

Observation 343593da-3519-4c35-aee5-64fa2e339303 · outbound

This paper cites Jailbreaking Attack against Multimodal Large Language Model.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Jailbreaking Attack against Multimodal Large Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.071174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.071174Z digest=sha256:02c80e0a96f33e298fe41d46fc7396ec2187c841dfa6ca03ce9e2f7b6dafd7df

Observation d9cf6cd2-5163-4e40-bcb4-2255d29aadbf · outbound

This paper cites Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.075183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.075183Z digest=sha256:7ee24c012a2c90e5d44e6d3de73f30cd74fb97872fa74fa11229069505928d7f

Observation bae6b810-12aa-41a6-b283-86d77f970c80 · outbound

This paper cites [Online; accessed 05-January-2025].

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM [Online; accessed 05-January-2025]

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.628971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.079208Z digest=sha256:3e4585c68557e4ecf56c95b4faf6680ab151c9b54db5909b693d4ece523559f0

Observation 38dfb18d-45e2-45e5-889d-462ea2233bf6 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Visual adversarial examples jailbreak aligned large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.617473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.082903Z digest=sha256:738e12accd24431a4722ea625373c76a7b9ef9ec3339b96094e837caf20a108a

Observation e64c1b0a-1fa3-440e-879f-c31e1c13123c · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.605744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.086609Z digest=sha256:29f6d8e13306e95972b16d4fe213d45fd05f096de8151335b280eae254d036f6

Observation 5e6a4370-cbf8-4977-98f5-08c30578c994 · outbound

This paper cites Adversarial N ibbler: An open red-teaming method for identifying diverse harms in text-to-image generation.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Adversarial N ibbler: An open red-teaming method for identifying diverse harms in text-to-image generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.594532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.090169Z digest=sha256:c2a147628d0ba85e3867b591f084890ecc883701e2611c88a5d12f6596cb6445

Observation 505a22bf-59bc-4e97-90e0-069b3bb34dd2 · outbound

This paper cites Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.581711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.095201Z digest=sha256:36f302105f03ffb025cf0341b18ceb0c1e666f6e6408863fb6819f109c289cd5

Observation 01ac96f5-a074-4c35-92ce-a22bc80d5f0f · outbound

This paper cites Assessment of Multimodal Large Language Models in Alignment with Human Values.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Assessment of Multimodal Large Language Models in Alignment with Human Values

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.099181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.099181Z digest=sha256:20a8d31246bf164005a67fbaa7a97f1cf43ad0cfe2f9bc325182f3187bcb2019

Observation 0e91b285-e53d-416b-8fa0-571fd765f467 · outbound

This paper cites Imgtrojan: Jailbreaking vision-language models with ONE image.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Imgtrojan: Jailbreaking vision-language models with ONE image

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.103397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.103397Z digest=sha256:6aeaf9fbe793ce1cf70445c1387124487ecb98da2afbf68e85942c2d1097676b

Observation dc49285e-9cfd-4a96-b0f4-82a6dcd0cb4f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.107340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.107340Z digest=sha256:b329c4eb7e7853eed952dcc684e62ccbc44082bd70d5aed9397bcff753ec2fdc

Observation c05b9d81-11c6-4893-9f33-3733c3f174f7 · outbound

This paper cites ALERT : A comprehensive benchmark for assessing large language models' safety through red teaming, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM ALERT : A comprehensive benchmark for assessing large language models' safety through red teaming, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.563108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.111291Z digest=sha256:2c0517a4cab3468b401c1d257f3493bb43d1541d172016a8a0b261a3fa30c776

Observation f40b87ec-6283-4cc3-8247-c58d55d65103 · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.115189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.115189Z digest=sha256:d10b8a8c9f06ae2d741f96a38f376ec8e00b41598f9256babf83522a3ac9330b

Observation 75456340-9d0d-413b-8d49-5c5e808aeab2 · outbound

This paper cites ToViLaG: Your Visual-Language Generative Model is Also An Evildoer.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM ToViLaG: Your Visual-Language Generative Model is Also An Evildoer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.120073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.120073Z digest=sha256:fe408b1564b32877977ef3bdd9c7e5b44977da377acc81bbc120a4db9721b400

Observation 55f0483d-f863-4091-8bab-74a92c058ab2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.124050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.124050Z digest=sha256:f1b337a742f0424bc736180993203def843dac48fc05a6f3877e148730b4a6b2

Observation faf74bb5-3383-4e00-b762-3e72e7682dd1 · outbound

This paper cites Sociotechnical safety evaluation of generative AI systems, 2023.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Sociotechnical safety evaluation of generative AI systems, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.550377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.128221Z digest=sha256:13f988b25c16c9d6bdd4fc94b418a5449d2f767a0fc47c23a1d675bd79fc4ca9

Observation e9945fb7-48dc-416e-b608-1f4c2cbc526d · outbound

This paper cites Vision-Language Models: Unlocking the future of multimodal AI , 12 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Vision-Language Models: Unlocking the future of multimodal AI , 12 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.538169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.131691Z digest=sha256:b3d3016154a89f8f4d06100c7b3ce1bd768a05952a204c3eaa73cd0616a28de0

Observation 4fe39c17-d9e7-4e30-817f-f419e3633f83 · outbound

This paper cites Adversarial attacks and defenses in images, graphs and text: A review.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Adversarial attacks and defenses in images, graphs and text: A review

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.525866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.135055Z digest=sha256:9bf32d4fa9176427baed0bf97b752b417db2148a7b1efff61232d7d02f8999f4

Observation f1b9bbe3-33fc-4e7a-b605-36c0cb05519d · outbound

This paper cites Viassist: Adapting multi-modal large language models for users with visual impairments, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Viassist: Adapting multi-modal large language models for users with visual impairments, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.513952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.138742Z digest=sha256:12d62002a938ec14daf86e63a4434359dfd07fc697724fdb1eb616f5635cda97

Observation 036e7825-fc34-46d0-bf73-6f22e4389baa · outbound

This paper cites Chain of attack: a semantic-driven contextual multi-turn attacker for LLM , 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Chain of attack: a semantic-driven contextual multi-turn attacker for LLM , 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.501981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.141844Z digest=sha256:a3efad3427ad547fa3136eb2340f273d5276e8dfb9acdf4ffe1e1b76f9ce9827

Observation d62abf6c-e3a1-4a8c-8a7c-c4f1eca6f15a · outbound

This paper cites Sneakyprompt: Jailbreaking text-to-image generative models.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Sneakyprompt: Jailbreaking text-to-image generative models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.489543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.145783Z digest=sha256:17c223b05afd978d1dc298d1f8e0f2e358c73f7e5990377b172a3b1e21d2cb1b

Observation c46a4fdc-382a-4cba-9e48-9af2e9b7f6be · outbound

This paper cites Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.148983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.148983Z digest=sha256:ebad3917292d3468acb4795b811c107ad75433de0b02716f93f47a970524e64a

Observation 88c7ec34-432e-40ac-a724-9ade7997f920 · outbound

This paper cites Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.152471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.152471Z digest=sha256:85f3eb0326d96faef35e211498a1561f3c3137be30f256bd17f25e99eb4c530e

Observation 11317825-d672-4d24-bda2-2901ae2e76c4 · outbound

This paper cites Xing, Joseph E.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Xing, Joseph E

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.156013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.156013Z digest=sha256:40735848055ff887cb95de464213700ac1a3e5ad8bda9b5b0f1933812086f007

Observation 41966232-8d32-41d4-8b75-b14f35606361 · outbound

This paper cites Speak out of turn: Safety vulnerability of large language models in multi-turn dialogue, 2024.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM Speak out of turn: Safety vulnerability of large language models in multi-turn dialogue, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:36:01.470271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:36:01.160398Z digest=sha256:0ae3e53c425622fa93090927acb1fb7d234f719389b1a72ec563a050ec51beef

Observation 0842f708-8c28-4f6f-a1d5-0a74301cfa04 · outbound

This paper cites write newline.

REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:01.163570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:36:01.163570Z digest=sha256:05748f0b27d6c7bc83ea9f5b1c3e312b03ae557ef697a859af333790507aeefd

Pith citing papers

Observation 1173e618-490e-49e9-90bd-8aa57e633551 · inbound

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models cites this paper.

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:36.682289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:36.682289Z digest=sha256:6035da5ceac067eede628ad61fb1ae48d6c176e87e7a29691c28423eff11b7cf

Observation 44b637bf-86dd-49d5-ad06-3f37d12b7ea1 · inbound

Jailbreaking Frontier Foundation Models Through Intention Deception cites this paper.

Jailbreaking Frontier Foundation Models Through Intention Deception REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.077456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T03:17:51.039062Z digest=sha256:2af80b221ad7a4d8372319f1c736a9f1d9331077c3c3fe0f3f68f90955d28440