Pith. sign in

Paper Citation Record · LEDGER

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

As of 18 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 6 inbound Pith citation observations for arXiv:2411.08410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08410 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:45:34.938217Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:45:19.069073Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T20:58:26.217531Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecacfeb8-4c19-4d9e-9be2-96fc98f7f78a · outbound

This paper cites Gpt-4v(ision) system card.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Gpt-4v(ision) system card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:36.092236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.599396Z digest=sha256:53b16afa5f42db72f9cf9334258857fa8a4805731ce165eee1e0ba2be30e984f

Observation 06b042ef-f91c-4bad-8638-f6d3f1e659b3 · outbound

This paper cites Hello gpt-4o.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Hello gpt-4o

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:36.075844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.606035Z digest=sha256:292f6202421e9887a2785a7ed078161057bee8a83f7736c156dc93d9c872b0b5

Observation f21b1d87-a7b1-4e72-8eb8-3c11cc7c3def · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.611296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.611296Z digest=sha256:0087b0a510f6ca102779e54aeec4aa47d769cd86490a3d3380b979f2878254a2

Observation 3461a50c-81c5-4dda-91ea-541307814ad3 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.616818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.616818Z digest=sha256:694a4294d62ac2d7ecebdd2ce7acb7845f01adf0fff6c800c3f02e2cc9f5d6b6

Observation 3493064a-40a5-4543-af27-b85a26c9730b · outbound

This paper cites Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:36.061703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.621869Z digest=sha256:da175dd59ef53f0824d3297afaa0f9c15c442d81476baf1e8897c39a5a949074

Observation 54ecac06-d7c2-41c8-8be1-aeeea2889bf4 · outbound

This paper cites Choquette- Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tram`er, and Ludwig Schmidt.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Choquette- Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tram`er, and Ludwig Schmidt

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:36.046567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.626170Z digest=sha256:3b692ff0058c4d0f882662d9c6e0b067f528093aa2bdc679ceb80662891ca734

Observation 2361ca49-7a80-4342-8248-e8c08959c726 · outbound

This paper cites Pap- pas, Florian Tram `er, Hamed Hassani, and Eric Wong.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Pap- pas, Florian Tram `er, Hamed Hassani, and Eric Wong

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:36.032512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.630894Z digest=sha256:a55b25134ad3c47bdca5ff4dd21bfefd98acad4309ca4f1a0df34e24bfdd416c

Observation c1d2fff6-07af-4609-8bf5-b4ee836295f2 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.634774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.634774Z digest=sha256:42f2b6476c47ed830d7026ccb927ab1632037dc5daa1d5cb0b0c0988e321c7c7

Observation a20fc002-86f6-4e91-adbf-0801375658d8 · outbound

This paper cites Can Language Models be Instructed to Protect Personal Information?.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Can Language Models be Instructed to Protect Personal Information?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.638955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.638955Z digest=sha256:1792f0dc56698e9f9d7a6a0d2c1baea0a96388d258fba58f1547d270b5c58b7e

Observation 06afd5cf-7928-4a1a-8714-6525907d8fbc · outbound

This paper cites DRESS : Instructing large vision-language models to align and interact with humans via natural lan- guage feedback.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense DRESS : Instructing large vision-language models to align and interact with humans via natural lan- guage feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:36.017777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.643408Z digest=sha256:ba43bbfff67601ebfc25de4f8e5a54cf6643b1ba60f4b0b99a6fa7298f065b62

Observation fdb1bf01-0868-41aa-b276-6fdaf1a9ddf8 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.647230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.647230Z digest=sha256:cca4df206aa6de403e719c5e97ab8b41768755a5ed556d3d2a951901a6649434

Observation 9bf26ada-9e7e-4445-a634-ac1ac8a9ddf3 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.652008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.652008Z digest=sha256:d54011d897115ef4098efdc35cc76f09386a9058bbe7d6128624258adbae0388

Observation c4d1032d-7ef6-44eb-9255-e79752111b20 · outbound

This paper cites A coefficient of agreement for nominal scales.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense A coefficient of agreement for nominal scales

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.656566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.656566Z digest=sha256:2706dfaf14b6c14a202fca766c18498d9508a3d274fbd4904aeb293e1050c067

Observation 572e37af-34a5-4376-8088-8056e0d120ce · outbound

This paper cites Attack prompt generation for red teaming and defending large language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Attack prompt generation for red teaming and defending large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.993565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.660526Z digest=sha256:7365e10926ae9c5a3afbcf9bf32b2387818cbc5ef87cd88f2be23014cef8e8fd

Observation 09e90073-9bf4-489f-9db3-5518c7f206cc · outbound

This paper cites Multilingual jailbreak challenges in large language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Multilingual jailbreak challenges in large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.980111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.664440Z digest=sha256:8f441911fb1efa334a83e7d98bc5894cf09d98cae059aee29ab29cd67dc351d2

Observation 7a1fe5cc-214a-4d81-832a-80db6e1947c4 · outbound

This paper cites The Llama 3 Herd of Models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.668351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.668351Z digest=sha256:409dd460c64b40cd3b24499060c74dfe664bd781442eeeaaac41e77bcfd39562

Observation ba059300-2e6e-44dc-801d-8f4ed4ffd5c4 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.672975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.672975Z digest=sha256:025749822902fa6fe42770800843b403c61b6c9e8503f3995ffc9c4626fb1324

Observation d7e3f681-6b8a-449c-a3be-074ed6c5480a · outbound

This paper cites Inducing high energy-latency of large vision-language models with verbose images.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Inducing high energy-latency of large vision-language models with verbose images

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.965845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.677373Z digest=sha256:451888ed334ac080ffdf26a1d215d93162a6138e3ed258c66917c90e6a767f6f

Observation f590a976-a726-4e93-a67b-bcf63b63f027 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.681816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.681816Z digest=sha256:c2bd14aefa68576460d22e9f094079845e2fd08b9ed763b6190dd9de5b7111ad

Observation a68c2505-6927-46f8-b130-1ac02d4c3e0b · outbound

This paper cites Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.686526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.686526Z digest=sha256:f732533acc672fddcc9c536a0734cdecf69f3d7805519b3eae0f97e7ffefb44c

Observation 3fca6842-0e57-40bd-934d-98fe1b56ec11 · outbound

This paper cites Agent smith: A single image can jailbreak one million multimodal LLM agents exponentially fast.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Agent smith: A single image can jailbreak one million multimodal LLM agents exponentially fast

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.950361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.691629Z digest=sha256:922b4eed69058e47e75074015d96ad40d67ee00cfd89a6534ae0a4850ac1804a

Observation cc83ef47-471d-4002-9e77-0837e729f817 · outbound

This paper cites UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.696102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.696102Z digest=sha256:09e2211214f7b7d6a3cbaad53274777607ce78af801ed3f347c83c5c678968d0

Observation 17098e69-3fa0-493a-bf05-ac2db890ea93 · outbound

This paper cites Catastrophic jailbreak of open-source llms via exploiting generation.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Catastrophic jailbreak of open-source llms via exploiting generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.934468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.700824Z digest=sha256:ba47d29907120feab3cf382af5f6356e65a9200a2bbc36a8e129a0e543f2ccbe

Observation 8324f36b-7883-4813-a441-dca17ffcc30d · outbound

This paper cites an unresolved cited work.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:45:35.917928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.705300Z digest=sha256:65f01a87c6885897e18acfaaec5687bb75bc6d7776c4375e0c3149e4b482f69e

Observation 707a7516-2481-49a6-afa5-92b3df368c85 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.709820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.709820Z digest=sha256:97bf0785e587a440641793b309a26c46a18fad2ccd7b3bdf69d51ed3739f5395

Observation 7d3799cc-10a3-4da0-a305-fe5063fa8249 · outbound

This paper cites Mistral 7B.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Mistral 7B

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.715063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.715063Z digest=sha256:47c21b6fcc6818468f23b552be30b8f767bf8699a037daa56b985bcce65d8709

Observation 1207463b-f4c6-43e8-9168-9e3c94f3d2fe · outbound

This paper cites Dragan, Aditi Raghunathan, and Jacob Steinhardt.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Dragan, Aditi Raghunathan, and Jacob Steinhardt

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.902586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.720035Z digest=sha256:74292d57ebba8b791a0381c92e70315b4a43103e0e4de911edd572a9bcac7ab0

Observation b4b0bba9-5125-4bb9-8fd2-0fe18e85ba24 · outbound

This paper cites Challenges and Applications of Large Language Models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Challenges and Applications of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.724487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.724487Z digest=sha256:40f924b66bebc9ddb6a18c206723d0d5d9a9f63af8f98dbe084e8523991ca640

Observation d8447208-f019-47f7-82a9-87ec358b5985 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.729143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.729143Z digest=sha256:2d26f883563a30f5c819b5fb863fbff3f07b279aeb9e536772d1f34304597825

Observation eeeea511-4dec-4adb-9a37-85629996b12d · outbound

This paper cites Red teaming visual language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Red teaming visual language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.887647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.733954Z digest=sha256:72355cadfbfa132837705c1c90706e1981f293395f500120091ca141930ed09d

Observation 3d9e65fa-d1bf-486f-b6f2-106ce51141c8 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.872910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.738276Z digest=sha256:13cef499bba77a7daa7b13efdcddecfe69b5aca262367858bc12a933d252c3bb

Observation 8be0ef91-97df-4f21-bfb1-19d118d458bc · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.742195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.742195Z digest=sha256:0ebd6ff27b41585aca710743e64fee1d76f43fdcf7dd82041307906ef3a5152c

Observation 323f1df9-e28e-4e38-95a3-ee8f874db4e0 · outbound

This paper cites Improved baselines with visual instruction tuning.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.858655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.746261Z digest=sha256:91bab9ea554e2e171441255c9655044645560a0525a877aeca232079bfd64c09

Observation ebe34c91-1813-4f2a-97ea-3f307c438de6 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.750201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.750201Z digest=sha256:297bc4328db7c9c95834884d164860eb6d75b8fce5af73a6ef69b6e360f54254

Observation 8c5e5d15-26a8-41bf-ab98-b685393a9d97 · outbound

This paper cites Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.754730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.754730Z digest=sha256:e5927794ba329b6c54e75c2ddc353eab4cebdc69814d84d681cafde0423f39b4

Observation 4643bf07-a145-4787-a14e-072d401c007c · outbound

This paper cites Mm-safetybench: A benchmark for safety eval- uation of multimodal large language models, 2024.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Mm-safetybench: A benchmark for safety eval- uation of multimodal large language models, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.835192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.758825Z digest=sha256:cc74bee1ba405d3edd3590858811a6486e3b1f591e83e29eabd41afddb294582

Observation 02be041f-a0cf-41e5-9b92-eaaae0c59d96 · outbound

This paper cites AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.762910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.762910Z digest=sha256:41ac89e8beb107e1807bedede142da5070ccc10d22a8bfa87bc6c93091998a2a

Observation 617a036d-0f17-4b5c-957c-28b96eb7e7e1 · outbound

This paper cites An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.820591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.767127Z digest=sha256:a8f8baf2774fccc3fee81492b1607106eb36d793d4b124c9f27e8bd51d887cc8

Observation 40cc0e55-5bd5-4c7b-b544-9fb36324f386 · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.771126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.771126Z digest=sha256:0a238c1a90522539faf6b7725421751f392ae2f0a0937b22d2bf22c4e32c68c5

Observation 898fa128-a27d-418d-a5f8-983e54cc1ff1 · outbound

This paper cites Towards deep learning 10 models resistant to adversarial attacks.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Towards deep learning 10 models resistant to adversarial attacks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.806611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.775412Z digest=sha256:20faa75a73475d2e2c0f90a980ef33e48c0b309935961f06ae139bc2e13351e3

Observation 1218dd34-37fa-46b6-8471-bf137b965037 · outbound

This paper cites Rule based rewards for fine-grained LLM safety.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Rule based rewards for fine-grained LLM safety

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.791374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.779356Z digest=sha256:fdfa46abb83f2100e2e41f4a6da29cdbcd004c429a6b43d58f612ca643b2385c

Observation 1a57df71-d351-4954-b50d-fd5516a3299c · outbound

This paper cites GPT-4 Technical Report.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense GPT-4 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.783693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.783693Z digest=sha256:b55d32924bd1cef7a17e7dd5bc1cb6f270b3143ba8a931165df2e47647bcb6c7

Observation 77a18ca8-b8a5-4a26-ac1e-706b6816276a · outbound

This paper cites an unresolved cited work.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.788025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.788025Z digest=sha256:2f7451c71ce1b9b42fcebc36dca8ad2a09242a8456e15c2b26aa06ca6c393cd7

Observation dec6bacb-b5c1-4f9e-aac0-177d634eb73d · outbound

This paper cites LLM improvement for jailbreak defense: Analysis through the lens of over-refusal.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense LLM improvement for jailbreak defense: Analysis through the lens of over-refusal

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.765765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.792459Z digest=sha256:c395974f2f378f91b902a60b2a1deb6b8d8aafbc2a0276b5a022d32c7bce266c

Observation fe8bb108-1dcf-4a4d-858e-0bde1428ddff · outbound

This paper cites Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.748812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.796897Z digest=sha256:f0d0d5421c099cf7e77166faee276a15cc164381ecde96e22c9182798aab4880

Observation 1e1abc64-67fc-4d23-8dad-1e17b76af5d3 · outbound

This paper cites MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.801256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.801256Z digest=sha256:23e03f5480265154b88f327f1fb298e52ead842af8db6debc58f160b5e43bcad

Observation dd3ccc18-da4a-47fb-9f13-017a5cc90c4c · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Visual adversarial examples jailbreak aligned large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.733234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.805925Z digest=sha256:60805964b141c63f6d51f10738d7bb1e041b5e014d47b8f2a29400e31904612c

Observation 938ffd8f-06ee-4800-8778-96972370ddb3 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In ICLR.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Fine-tuning aligned language models compromises safety, even when users do not intend to! In ICLR

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.717801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.810379Z digest=sha256:8ae24ea49883a3479c302eb214bb583506ec70470cceced0834497a942bedab0

Observation 66775bff-8661-4800-b36e-b08b4ff27f80 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense High-resolution image syn- thesis with latent diffusion models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.701375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.814792Z digest=sha256:51dffcfdc5a52c10292c1a515b5c5f9d0f8dfcf833512ac36db4325ab78696ea

Observation cce7922a-31ea-4c0b-b708-04ccb18b81e7 · outbound

This paper cites On the adversar- ial robustness of multi-modal foundation models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense On the adversar- ial robustness of multi-modal foundation models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.686400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.819280Z digest=sha256:b53490aaa2c09ae531f88d917e9b6db4593c8d814c457f860ad7cdb555014aa9

Observation b76417da-dda9-4d8a-b81a-e7c48dff8776 · outbound

This paper cites Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.823771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.823771Z digest=sha256:75f8e8263ec5500c240267054d8324bc3453ef1be9295223cdff40a0b40126f0

Observation 308a26f8-127f-4afb-bf41-37085e964dd7 · outbound

This paper cites SPML: A DSL for Defending Language Models Against Prompt Attacks.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense SPML: A DSL for Defending Language Models Against Prompt Attacks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.828353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.828353Z digest=sha256:7619ebc26165ee1ec1a3a77c51712bedc3912eba742ed309e1a4b80124608d1a

Observation ff7bc135-3189-43b4-9a6a-ac419c2beac1 · outbound

This paper cites Abu-Ghazaleh.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Abu-Ghazaleh

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.833016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.833016Z digest=sha256:569c0fdda13a5dcbd5cb3cb9b6689adf92f67629afabf08a794e254407bc1f3b

Observation 6b7d7880-c659-439c-a442-6e1721fe630e · outbound

This paper cites Hugginggpt: Solving AI tasks with chatgpt and its friends in hugging face.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Hugginggpt: Solving AI tasks with chatgpt and its friends in hugging face

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.661999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.838003Z digest=sha256:3ccbe6583de57485d2fc47c662b886b44a9a1531d8b49843e590706727653076

Observation 408bbfb6-464b-4448-86be-c5866abc4dd6 · outbound

This paper cites Assessment of Multimodal Large Language Models in Alignment with Human Values.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Assessment of Multimodal Large Language Models in Alignment with Human Values

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.842506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.842506Z digest=sha256:797d54247c4f72f6e6a92697b503ee8e7542714b38a64a99ea50d083e068a3dd

Observation 967bc02b-2930-47f5-8d99-96b4c6eae9c2 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Qwen2.5: A party of foundation models, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.647386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.847224Z digest=sha256:64fcec8c5c87b3127a110c7683dda809a7cadb5a7050ab3fb0b363cdf357bf37

Observation 1b73fded-eb8a-4d2a-b947-5244254832fa · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.851653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.851653Z digest=sha256:577b8f02427548117f7e920d3b6c2b05369407008795ad5882eb651f8fb6f9bc

Observation 1b95a681-39a3-4b64-b76f-c96ad0bfd364 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.856687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.856687Z digest=sha256:1800f4afe66e58ed3c0a306984d5c50fdb7ae4f996dfe065d01b6bd2305f3f4f

Observation 48e842a5-6044-42b7-ac68-45373c43951c · outbound

This paper cites Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.861227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.861227Z digest=sha256:2d0dd35b7facad1f19ee1a1756297d3d1ad0d5d91e502b450ff4f1dce3e69414

Observation 6afee372-28f2-46b0-bb00-646abb2ba6fb · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.623184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.864876Z digest=sha256:bbc761e7ac6deb3da5c40d6afa138b62eb344572495d35c8e37deaac3594dbb1

Observation 1d3c8071-37b8-4863-8081-72cde92c513e · outbound

This paper cites InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.868801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.868801Z digest=sha256:b9b426ce34144c7f17a280ede84a7e3bb33e3ac47f0e7f3373ee2ccc0fb3ccd1

Observation 9675dc97-a4d3-4eea-8e3f-3411d8ca5b70 · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense White-box multimodal jailbreaks against large vision-language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.609900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.873172Z digest=sha256:4e13a22d7dcb00222b282b62e66851a840f11c7a531b9513f12973ec7984c117

Observation d6ae2c1f-32de-4681-ad24-9ada369b9ba3 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense CogVLM: Visual Expert for Pretrained Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.876990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.876990Z digest=sha256:ffaafb48df1338d98594f5b09f92bb4c9a8c1f92a3e732baaa2603139f9b6e27

Observation adcd44df-1b37-43a0-a11f-a22db48f68c6 · outbound

This paper cites AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.881362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.881362Z digest=sha256:0540ac575043d525b7f567e3faeeb8bc5c11a95b6cb70320b3603544eb7c6119

Observation b02b3b48-86d9-48c0-9aef-86fdac91ee19 · outbound

This paper cites Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.885499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.885499Z digest=sha256:e8bed7b3fca4542509340aec6565cf6dda5887b8f85b73139a93c9f967d386d3

Observation 51870101-dccd-4efd-a82c-5571ce587b8f · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.890372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.890372Z digest=sha256:07d3362d21d3ae30cd52eb2ce0d840d6e78836c6200d63298c071697b267f65b

Observation 4c697ac4-a4c4-4339-b60a-612a049eb5c8 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.894689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.894689Z digest=sha256:ad36d5c1d8dd9ba30fcc0cd9de819a1405a8737a8702bd3ce034dd44a6e06109

Observation 7a342db7-75cb-4652-a1cf-0f94bc605333 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.595763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.899052Z digest=sha256:457d6eb3ba40629ac907b5c7e8eb2f3e682dc50ad2fbb888feee57470ac8c264

Observation bb8fb085-a648-4b80-8a97-22fcbf5f518e · outbound

This paper cites GPT- 4 is too smart to be safe: Stealthy chat with llms via cipher.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense GPT- 4 is too smart to be safe: Stealthy chat with llms via cipher

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.580839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.903072Z digest=sha256:af56633c8237b92237ce4f98905a2276a7feb67695a62c273fe29e56852e76bc

Observation cdbbd4c5-9245-4072-8461-e60da76bb17a · outbound

This paper cites Removing RLHF pro- tections in GPT-4 via fine-tuning.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Removing RLHF pro- tections in GPT-4 via fine-tuning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.564389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.907101Z digest=sha256:2243ce20a22ce74197de0a62dfadb30521acae5bd44ceb0cdb9665b44cf88b68

Observation 9ace8bd1-6ad4-4c42-89aa-d556e2686950 · outbound

This paper cites Make Them Spill the Beans! Coercive Knowledge Extraction from (Production) LLMs.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Make Them Spill the Beans! Coercive Knowledge Extraction from (Production) LLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.911487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.911487Z digest=sha256:839f6168ebb71e21efb9ca822b0e9107585666750eb71cb69de30e8159e6a145

Observation be182522-4a0a-4527-a7f3-54bcef14df24 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense On evaluating adversarial robustness of large vision-language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.548999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.916268Z digest=sha256:2d629f5c76c7eafbf7f0c762ebda0839dc8ab931a281bf32fe923ea515b5e629

Observation a1ff2074-240c-43fd-afc2-41a9813f044a · outbound

This paper cites Xing, Hao Zhang, Joseph E.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Xing, Hao Zhang, Joseph E

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.534002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.920615Z digest=sha256:3d289b9ab98ad4e045632fd4db9fbb8e2f13dc4a5c5bdd7008ec58d0c2abd354

Observation 9f54ee16-64f1-4909-ad3d-03e40485bfea · outbound

This paper cites Autodan: Interpretable gradient-based adversarial attacks on large language models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Autodan: Interpretable gradient-based adversarial attacks on large language models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.518731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.925150Z digest=sha256:3f8b32efe6a38ef8d66128e24eb31d3f7cfd612a49683afee02f3d48bffb91df

Observation 6acc45e3-c44f-4f90-a700-3068ea37b819 · outbound

This paper cites Hospedales.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Hospedales

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:45:35.502054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T21:45:34.929382Z digest=sha256:64654934ab55c95772fe63e64d80768ad9e8eeb0d16d2b387a677c54995a4e8b

Observation 41bceb35-d66f-43dc-9492-9fca20954944 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.933478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.933478Z digest=sha256:0ed6efc7691e211a47ba17b0fcfd73988cd815b996f06ef0fedabf455b227612

Observation 9e01f9fc-a7e2-4386-b021-52139ee90f8e · outbound

This paper cites Is the System Message Really Important to Jailbreaks in Large Language Models?.

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense Is the System Message Really Important to Jailbreaks in Large Language Models?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T21:45:34.938217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:45:34.938217Z digest=sha256:6f3e699a56329c50830d7c510aeeaec191860a64ae3fbe036fa5426e641ac333

Pith citing papers

Observation fadbe50f-fa67-4b90-980f-4779dfbe4c43 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.220359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:d7c1fbdb1e3d4f3d5135478cf99c774e5965fb70fe2f7089b46b3dc7aa3cce1b

Observation b203b1e8-2def-4037-b255-75603495a01a · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.069073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.069073Z digest=sha256:8a6c69b430ab1f1683557bbbbc2c4e232b98557382d5a0dd84cf1d7b3769947c

Observation 6238b638-3e4c-4bf3-beb4-1af93b5b1163 · inbound

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack cites this paper.

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:01.083283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:22:01.083283Z digest=sha256:04b46a4bd48050047479aefc9748755ead876392a65cb35d6060ea21776ccd1b

Observation e9266238-cdb5-4028-8db9-a8b02f5be4f2 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:59.867495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:53:59.867495Z digest=sha256:3826ae46670ab83a4911a03ae93c503b09fb8017915938ef8b036e96da1bda4b

Observation 9638a436-a417-4c85-8988-bdbaa52ec948 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.684858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.684858Z digest=sha256:ad4f0647e3c3643099161de5bd107ec707b0851c129f1c026af965df72b7c166

Observation 9adf6efa-39f9-4f04-a303-024b5adda4f7 · inbound

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models cites this paper.

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:02:37.531373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:02:37.531373Z digest=sha256:b77cc543dfaa8fd98d839d0aee0979f674c483146d8910aeda72517042322f17