Pith. sign in

Paper Citation Record · LEDGER

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

As of 5 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2605.18915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18915 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:19:11.536555Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact6
  • verified fuzzy33
  • unresolved3
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec833870-3cea-4bcd-b075-dd15f1e6cad4 · outbound

This paper cites Aho and Jeffrey D.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Aho and Jeffrey D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.492539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:a8946c06598ca740ccb8069c743ca7fdf194a9cda91aba9e5e3108cf227cf3a6

Observation 79193c3b-ac5d-4b9d-8bad-39b0d1d0ab3f · outbound

This paper cites an unresolved cited work.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:23:26.494054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:f20df494d0354a14413f3765b1706bd70e81c64e1c78b00c6a7dad749897de61

Observation 7113f324-37e9-4df1-bf07-caade71fa203 · outbound

This paper cites Chandra and Dexter C.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Chandra and Dexter C

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:23:11.772475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:65aa29f6555a8b3763250887fed5b23d28ce54eec794794c5cdfc103ae7bed81

Observation 7a000a03-929b-480b-9106-aa00d1765cc9 · outbound

This paper cites Scalable training of.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Scalable training of

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.484520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:cf68081199264255ce4544fe48017fa396e24d72ec78e8eb047850050e85d352

Observation 52f4efe5-2e0e-44e4-9039-5d78165c0bc8 · outbound

This paper cites an unresolved cited work.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:23:26.486261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:aa6e2986308bb7f2c2ac9d6a83ed507fed82dd3f52043d7138cf63c32b9129c4

Observation 42cdbe90-f3a1-4148-a81a-1ef4c8c075c6 · outbound

This paper cites Tetreault , title =.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Tetreault , title =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.487906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:f70d367967ee90b11f55e674852b2b93bab198a6eedb8488f9e8c2a63fd80f8a

Observation 1ca387f9-ec6b-4b47-901f-563039bbf318 · outbound

This paper cites A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.479027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:dca2f046b8a679f4d621c20674af3a7084b97c281c63f502269c391cb977492d

Observation 212d290f-e352-44ea-9aee-2d9925cbe0e8 · outbound

This paper cites European Conference on Computer Vision , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs European Conference on Computer Vision , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.482455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:5f60b5a40cd99c757bd705f58d874fef88c7b2e164fc9899830957abab8b737f

Observation 6c68bde5-bab7-4b3d-ad4b-84e531927faa · outbound

This paper cites European Conference on Computer Vision , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs European Conference on Computer Vision , pages=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.474055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:af3e3b899128d9def5e82d59d1f090f71f331d64e2f082d67f9ce64290f52723

Observation daf6d30b-753e-41e6-b334-f49604270399 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.470072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:15c0806c3e850949fcba88f952c768461a38e636bab1156bc90043df18f5353f

Observation 3a5cbc8b-f5ab-4f1c-b380-655d531801a5 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.468557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:0fa7a5f57ded63bc43e30adbcb833ce20f3c1623879d776300d15cfa7afa5680

Observation 04c9ff00-c8ed-4208-a287-2a1745b32ba8 · outbound

This paper cites 2023 , month =.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2023 , month =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.472007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:5f97795bdfc7f65bd140f7b47b45d32b60756837baeb76b48ead210b49308434

Observation 92027816-29b2-425e-a9be-d6d4accd744d · outbound

This paper cites 2024 , eprint=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2024 , eprint=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.475725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:f8bd4cde62ebb8edbd82b60b83feffb1dfd56d15be6a6688847ac7f540c586d3

Observation f0a59ca5-6f6e-424e-b750-3ba847213ff2 · outbound

This paper cites 2025 , month =.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , month =

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.477414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:78b9a267d368718878e72d1ad8fcce9836c6ad92e6c261f82639b06ea3dd8ab3

Observation af47920a-c85e-42cf-b6f6-f306a79dfb90 · outbound

This paper cites 2025 , month =.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , month =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.480651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:0d67398c42b859b458adf87c7fd8947bef9b869c6988e477ed3633f7ec2f25c5

Observation 28ebeb4e-5a43-4ca4-bd6e-ae73a9e45f1b · outbound

This paper cites 2025 , eprint=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , eprint=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.491024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:854cec85b813910684f4c69c1206ea2ca585ee12b3ddf024e21711291eadf66e

Observation 0b76c8ce-da19-4d0d-9cb7-70696faf3f7b · outbound

This paper cites an unresolved cited work.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:23:26.465231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:356ff81683e3a96ad0a962044163d57738e7bdcb8d0976c1e788b38460cd5fb7

Observation 1dd29265-ae95-4de3-9c6f-ccb23be7c4d5 · outbound

This paper cites 2025 , eprint=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , eprint=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.466899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:e84ce891250247f20cd1ec967195abe8b6914794edd86f90f763e26c9909b261

Observation e156e8d7-d72e-4eae-9600-cf251c9f10cb · outbound

This paper cites 2025 , eprint=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.463626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:36b4b6de874f4ef032d0d9a84f250f500d86a7e8b2807eb92039f0459b828447

Observation 4af692b8-4571-4308-89e7-5ff4b6bb7a29 · outbound

This paper cites an unresolved cited work.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Unresolved cited work

Reference 20

Resolution
parse uncertain
raw_fallback, observed 2026-05-20T10:23:26.458661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:48eaadc09be111428ee16cf46ab2b2db2c977a37bf7d5413f9971d4be18dd724

Observation 0c1cbece-d06e-49a4-8c8f-254417c90bcd · outbound

This paper cites 2025 , month = aug, day =.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , month = aug, day =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.457131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:ee64e17512037ac9047bc0a3ff3f6f6f57795a1ce688126123542f049e7ce591

Observation 7736e1a6-da99-473e-a1ce-e4c42e9f474a · outbound

This paper cites GPT-4 Technical Report.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs GPT-4 Technical Report

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:23:12.287037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:9afaedd9d3966871fd703fe3363d93f30bff207b4dca0884bcf77f57f540eab4

Observation 172da944-bb0b-4d54-9c68-878d13a3ba2c · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.460405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:912cd7693ebf7c20a8b8c4bc5777aed260740f72d058cef9043193eb272da7d6

Observation 14433505-a26b-4f64-99d2-177b355a984e · outbound

This paper cites arXiv preprint arXiv:2510.21189 , year=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs arXiv preprint arXiv:2510.21189 , year=

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:23:12.292585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:dc526b86229e06ae8ce765a705c89c8b26e922269c33e4e1d493d829ac257bcb

Observation 46ada190-3360-477b-86f0-b3d9dd9fe1b3 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.461924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:029f05786a4f7005c93e085d441897101a656b22d85cbe71888d75760872ba2d

Observation a00bb4cf-63c8-4ece-8365-61c517e9644d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:23:12.295195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:e7c4f0f45a32e928851269df014a4ee7dd5f624fb3d9fd3fbb42234bab718f6a

Observation a416fb78-2955-4b81-9a74-ef8c8306ecb8 · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs The Thirteenth International Conference on Learning Representations , year=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.489504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:41e02d5fe7bb8047d56ede44e9d27a7eb31d96c775c7d4726a9018a613120b24

Observation 55178eb6-937d-4822-ace7-c49303eddfbd · outbound

This paper cites MM-IFEngine: Towards Multimodal Instruction Following.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs MM-IFEngine: Towards Multimodal Instruction Following

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:23:12.304574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:b4d8a31b6e111e73ebdbac17ed8d14601d564d7c127808c662a4d07787658716

Observation 069e2bdd-8ea7-4940-b7da-cbd77e4a7fda · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.453784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:5bd02b784ef409e9f6a22c4c36d425a4d2574dd84b727e89c86728a801adfa28

Observation c7a24a65-7923-4420-99d3-eb68ec015f9c · outbound

This paper cites European conference on computer vision , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs European conference on computer vision , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.433081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:0fc6d425b33e6d2e174ddade03033ae8e94204cc369a156ce9797dac7aeae9cd

Observation 4119c6b9-bb63-4279-9663-ac275fbb2cad · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.448322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:3e9e23e38ff8df421e3702aad375596494bc7e01e92d13ee38183a349c42f4f4

Observation a8ac6898-9b60-4ea8-b1be-d6300fa42d62 · outbound

This paper cites Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:23:12.298703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:d230ae8e8b7abea2d56142baf87bdbeac0eae1f80120d0f0a69d2c765527bdf7

Observation 6236a050-80fa-44b6-9f30-d0e05e6c610c · outbound

This paper cites How Robust is Google's Bard to Adversarial Image Attacks?.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs How Robust is Google's Bard to Adversarial Image Attacks?

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:23:12.313384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:6eddeb3234357163c02134cedb8a68412cf355c243e844dedb86ad456cac1f9a

Observation 83612874-f71e-43c7-9cd0-1cb2b5c6d125 · outbound

This paper cites Failures to Find Transferable Image Jailbreaks Between Vision-Language Models.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Failures to Find Transferable Image Jailbreaks Between Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:23:12.310429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:a9f76e5b9a15588c762ba273bc7fdb42ed77790724fbfb5b126f298a9a0c97c9

Observation 58f560d1-2eab-4fe6-bb87-4b2c4ffd9ba4 · outbound

This paper cites 2024 , eprint=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2024 , eprint=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.451734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:70abbd4dee21eea836e15858b25541a44bc87d7acbd29a6ff19c862ca271577b

Observation c3241354-9dcd-4cb9-81f6-bfd51bc4c788 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.437261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:3ef8c48a99857c16005ed05bae17c3c1f432c12554451cb390443bea2ed3592d

Observation 62af2491-9e3f-4975-8a7a-ab3780282637 · outbound

This paper cites 2025 , eprint=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , eprint=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.444321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:e865d8cc31750ce1616e643093e651174af46bf6617b4176272273508bd4fe64

Observation 31bea0db-a6ee-4716-8b81-a6b91ef6f6bd · outbound

This paper cites Nature Machine Intelligence , volume=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Nature Machine Intelligence , volume=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.435256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:3912a31fc9a3fa34440017f93307415d8f4590a674e25ae7c2c96eb7f5150c7d

Observation aeeb9844-7185-43fb-b4b5-24c50cd4e172 · outbound

This paper cites European Conference on Computer Vision , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs European Conference on Computer Vision , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.442533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:3b9da9935bae15afd7c19e6583409435c24a1cbc8b6974990642de43cb70879f

Observation 74047749-9f99-49f8-8f4e-14b8bd379560 · outbound

This paper cites Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:23:12.301594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:32086216035e3ac68d4063e11eab331b87c7fa4b9a8f2b2e9db0b1d27db171a8

Observation ba7d8269-1b89-4eed-a1cd-3046e7b1f719 · outbound

This paper cites European Conference on Computer Vision , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs European Conference on Computer Vision , pages=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.446538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:590e721c46ee5c3382c97ce0db795973bb9641bbbd80e70210d488ef02979e91

Observation 6b7f2b4c-9b57-4904-94dd-a40f890aade3 · outbound

This paper cites Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:23:12.307541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:c4c8bc9f6ced6d99a2468f9d54efdd234d0f9c1fdea90b5f34ecd1b8e7d717e3

Observation ef3c8ebb-ef83-4c58-b83e-1c68aab78042 · outbound

This paper cites 2025 , title =.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2025 , title =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.450101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:3e7f9900d47721f11011ee54ee6a2089a1a06484b18b199f8a2fe1856c8eafab

Observation f87a275c-6094-44fc-8311-e358333ca8c8 · outbound

This paper cites ICLR 2025 Workshop on Building Trust in Language Models and Applications , year=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs ICLR 2025 Workshop on Building Trust in Language Models and Applications , year=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.455529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:b98ffef46d0cc909f93cdec4483a1ff758557f48bc22c0d57ba2e104d141c0e2

Observation 5f17f2ca-6e54-4d39-a36a-65c712a641c4 · outbound

This paper cites NeurIPS 2024 Competition Track , year=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs NeurIPS 2024 Competition Track , year=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.431023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:368b1ebac3ca79a9f4bc61bef818bea511c9197a6a376455d5431e669ef7eddd

Observation 25a0cf93-1aaa-4371-8f00-38b4c56b6602 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:23:12.289767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:41443381ed416181af1108405bef5aae8f23658a80deca3c00b55161d3a4765d

Observation 81e2419d-e454-4b14-ab8e-973c3f3c8c01 · outbound

This paper cites 2024 , eprint=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs 2024 , eprint=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.440673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:83d4b23b66c1628828f57b084580f2697c9f9053b81d62050874f95ae11668c6

Observation 15d80182-900f-4972-b9d4-6e8afbc0906b · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:23:26.438966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:6a0155dfe601dcfde6d7fadd2c0f0bbb385961dc2a3c3786ab45c35c11ee6263

Pith citing papers

No inbound Pith citation observations are available.