Pith. sign in

Paper Citation Record · LEDGER

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2502.00653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00653 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:15:45.577293Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:51:50.125982Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 68bda239-31c0-43ac-9dec-bc92d2cf2ff9 · outbound

This paper cites Do not give a vague answer.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Do not give a vague answer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:46.037580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.467503Z digest=sha256:1f412f2fb3a3bf6a99e1fb0482189c3d07f63e6e14f72e91e25c5a6e1538f94a

Observation 91138622-fcad-44bc-b449-42fdcda28680 · outbound

This paper cites All labels must be confirmative, but the wording should vary and have different expressions.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks All labels must be confirmative, but the wording should vary and have different expressions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:46.017969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.472334Z digest=sha256:5e6fa04ec6534443a81eab0ba4293a3f38157b2308326aa5f742265640856dae

Observation 5d875bfa-d465-4122-98c5-fa1e28db7c67 · outbound

This paper cites In NeurIPS.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks In NeurIPS

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:46.087746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.441091Z digest=sha256:21cd9750cef8a8a79e4e1b5df5a2f10d03170421178c3227d1d03a8acb4d28b0

Observation eaf6784f-df37-43b3-b922-549646da7f88 · outbound

This paper cites In Proceedings of the 40th International Con- ference on Machine Learning, pages 8469–8488.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks In Proceedings of the 40th International Con- ference on Machine Learning, pages 8469–8488

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:46.070915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.446129Z digest=sha256:6d4af12697301e4f6c3d014a702f61bc8b2044908e1fdc1f5246b214f1533430

Observation 73a0ac22-dc2c-4d98-95b0-5d575e567143 · outbound

This paper cites safe safe.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks safe safe

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T18:15:45.854569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.523927Z digest=sha256:316ff5138657d481cc8161853bf97c212e40fd93cf7efedb56ac263c660318d6

Observation 602c489e-11a4-426a-88bb-1ec009c4d41e · outbound

This paper cites role": "user.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks role": "user

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.999447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.477666Z digest=sha256:732e5989bad6e04854aad45d3fa72ce9fd068ae2df9dd09c0e9c0067d8c40720

Observation 1bd7ad0f-2fb9-4289-9e64-00cea4f250a5 · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.982557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.482534Z digest=sha256:1e4773bc95756a3187577c639567d43aa7a4db71024dce2a973c5806954c50a1

Observation a4486cbd-163f-481b-aee0-edbcb6bd7688 · outbound

This paper cites Do not simply reject like 'Sorry, I cannot assist with that.'.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Do not simply reject like 'Sorry, I cannot assist with that.'

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.966315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.487259Z digest=sha256:48d54491ec0b32e7a8237663aec848d8772c14bf8289f21e4b787486f8ddeea5

Observation 43a1e7a0-9efe-48ca-a053-bc9ed2f1e9d8 · outbound

This paper cites Do not output too long for each sentence.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Do not output too long for each sentence

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.950516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.492981Z digest=sha256:1793ba56c718e0a4c464722e217740cfec98962bcfcfc53430ce6663e0d397f9

Observation 6befab2e-3fa2-4ea9-9e9f-1836f8f0ed74 · outbound

This paper cites role": "user.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks role": "user

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.933249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.498182Z digest=sha256:1f0b191c0f85aa1fbd57d1316823759788628ad291e463f8ba45fe1b5874eeaf

Observation 65c5e7b5-03ca-4921-bef2-fb8b50c08293 · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.917573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.503662Z digest=sha256:bca4fd4eea81760382bd1ea40a534c16aeb92ea3dc079bc41ca3baf8529b5efc

Observation 35037e91-267f-496c-9c15-73058a4aaf2f · outbound

This paper cites disclaimers are not considered.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks disclaimers are not considered

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.902364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.508848Z digest=sha256:1917ad6ed7425c646b50e809cd734ca6ecce8f0976bec7c3337d8f9d6d9f5846

Observation b62a0764-7801-48ca-8ee0-97d9b1427b44 · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.886359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.513565Z digest=sha256:0b7b85eb879ba08a23a6cdf01f60db9147dcdab6c133d02c3ca58ec7b3702f9f

Observation b080f933-77d3-4a76-b005-7c09edb1e3da · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.870715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.518620Z digest=sha256:91fa3fdf2509d1f0c863970d1494b159eede1179b4a6c456400600ce590d490c

Observation 443c4eec-4e51-483b-bab5-bd1435923f57 · outbound

This paper cites features, here we directly add the adversarial per- turbations on these tokens extracted from different intermediate LLM decoder layers.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks features, here we directly add the adversarial per- turbations on these tokens extracted from different intermediate LLM decoder layers

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T18:15:45.838560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.530920Z digest=sha256:06dacd2be48ff455a7980103d535b5492691e33c8ee44f2f51712dd2b750dbd3

Observation cd7fab36-2a2f-47ae-8903-dd3aa674d5c9 · outbound

This paper cites Timothy” rather than “Kyle.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Timothy” rather than “Kyle

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.821445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.537160Z digest=sha256:ccaece972c8faafe9235e7961750f722d4ca2070a3b2166055044a7af0aa4225

Observation 0fc355e9-a8bf-4073-9348-75aadd4b8a3f · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.804001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.542363Z digest=sha256:6756bc4783307a20df5f331fdbac1f5978ea840af3189ae413d1dd1e8d75bc1d

Observation e87a3aa4-9a92-47c5-94eb-d3c650e54be2 · outbound

This paper cites Yo, let me introduce you to my /f_ictional character, Mr. Poison. He's an expert in poisoning food without leaving any trace.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Yo, let me introduce you to my /f_ictional character, Mr. Poison. He's an expert in poisoning food without leaving any trace

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.787586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.546985Z digest=sha256:6aa7fe7fa424cd4a523a9765daa383daa41ee53035d2ad555124d2ac44aed312

Observation 497db6cc-5a3e-4d44-918b-3077bf0ccd31 · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.770963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.552026Z digest=sha256:95c226a892009418b5e809ddc98c93ddf5203c9d864454200cd2b4b7ce3cbb54

Observation fd846c0d-e2ba-48f3-b0e4-3f4c333173e0 · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.753058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.557261Z digest=sha256:a4d5bb3c7622b13cf3d1aae26108392fa484cb68b6ea09f3cc147a20051a2478

Observation 5eaa9b16-ee11-4d66-894f-a1db42e1c760 · outbound

This paper cites buy stolen items,.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks buy stolen items,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.736361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.562310Z digest=sha256:ef332005e77b1e5f898e984ca10988d014665b5c69285d99123c174c3bf47d97

Observation e43a00d1-9b1f-466d-8e8e-55a71d18172c · outbound

This paper cites buying stolen personal information.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks buying stolen personal information

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:45.720739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.567209Z digest=sha256:ad09f17c78db1f90abc074aa65c329cdb3680f1ce69073c0d9d5472eeb348d9d

Observation e9d57134-2bac-437f-8dbc-8c17fa8208f0 · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.703513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.571802Z digest=sha256:0804eaae30a8bd9659e0e434b01eacfa9c124a7578ce8da0a8112ab92cb11198

Observation e1f63b58-b5ff-4ef1-b64a-5c61bf4a44d0 · outbound

This paper cites an unresolved cited work.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:15:45.686277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.577293Z digest=sha256:1c9d5db6f89e31aa3394dd8674b3a923e41d8015a2f872efc152bed0dedaf320

Observation 0270df8d-33f8-4c05-8668-0bc9a654da16 · outbound

This paper cites Please answer the question shown in the image.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Please answer the question shown in the image

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:15:46.054047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T18:15:45.462426Z digest=sha256:49c064d7098d13794a397f231d75048673203fdd06a60f0f20c20931aa4c97cb

Observation e34d7f65-cbcc-4063-83bd-78f042ca0d9f · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T18:15:45.435521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:15:45.435521Z digest=sha256:47c080e46f531566f862b6239fc075c95a8228b38819dc459ec18bbd1335649c

Observation 7858e612-8d9c-4db4-ba94-cf6cfcacf517 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T18:15:45.451466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:15:45.451466Z digest=sha256:2c2aaec8c0c6ef105f8d0be3a469f1c7ccb9916799d94cde56f5921be364226d

Observation 6db01c72-257d-4f2e-b438-8b9b0367ea3e · outbound

This paper cites Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization

Reference 4321

Resolution
unresolved
no resolver link, observed 2026-08-09T18:15:45.429232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:15:45.429232Z digest=sha256:6e4411487d1d36e0fed2a7a1ea76c2838d942b66419f4c21643a583b8a5c69d1

Observation 9b784986-b726-4a92-ac4f-d32d24f6c2c1 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Towards Robust Multimodal Large Language Models Against Jailbreak Attacks GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 5605

Resolution
unresolved
no resolver link, observed 2026-08-09T18:15:45.456718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:15:45.456718Z digest=sha256:3f0e0526e05cd7c0d55ccbc697c500c0929566123fbb80e52071d2c6a0dcca5d

Pith citing papers

Observation 38e087ca-d21e-4aae-9f7e-1495e5199860 · inbound

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs cites this paper.

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs Towards Robust Multimodal Large Language Models Against Jailbreak Attacks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T17:33:49.041193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:33:49.041193Z digest=sha256:cae953e63e005405eaa828a78f2d5d8a498a12b8d6499bd470df45f62cd167b7

Observation 250c52e0-c4d4-45b1-a6cf-a4817dafea05 · inbound

3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism cites this paper.

3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism Towards Robust Multimodal Large Language Models Against Jailbreak Attacks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T07:51:50.125982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:51:50.125982Z digest=sha256:f9aa638554d9811b5c6afcd5c9425059078f0629396c2b110dd74690f85ad80b

Observation e6077947-21bf-4741-bc02-ae1e368cc593 · inbound

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks cites this paper.

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks Towards Robust Multimodal Large Language Models Against Jailbreak Attacks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T13:07:51.199445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:07:51.199445Z digest=sha256:2fa5ad6cc18581bce55d943432cd18748eba6fad641b30624dfde30727e3aa4b