Pith. sign in

Paper Citation Record · LEDGER

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 4 inbound Pith citation observations for arXiv:2505.23793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23793 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:20.372542Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T18:20:43.092561Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:37:14.921452Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 986062ee-f95d-4231-89c6-0ff4c9080126 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.449940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:12.501473Z digest=sha256:0b4bba1ae6d0f400aacad99a8f215b7f18dbbe0b5c32406a18fe18a14c3df575

Observation 30d901d3-4ea7-4456-a82c-65507ac437b3 · outbound

This paper cites GPT-4 Technical Report.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:12.628526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:12.628526Z digest=sha256:c87921b2e6219867cec582f22364b4356a029bb08428a64d3fc7964a7a53fac4

Observation 3d5e1484-76c8-4375-94c7-e64909df2901 · outbound

This paper cites A Survey of Large Language Models.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models A Survey of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:12.839230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:12.839230Z digest=sha256:144e6659248af2a81b1620f7f327c2876fcd758e76156b2f80a030901bbe2655

Observation 01347fd3-8a8a-42ce-99ab-5f21d18c58b3 · outbound

This paper cites A survey of graph retrieval-augmented generation for customized large language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models A survey of graph retrieval-augmented generation for customized large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:13.588175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:13.588175Z digest=sha256:eaa0abd77c3e45723f22f48beaaecedddc8eb5b94c07b08f641e545f951c5a5f

Observation ce125e67-c302-4f6c-aad7-14a37aa53739 · outbound

This paper cites Entity alignment with noisy annotations from large language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Entity alignment with noisy annotations from large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.270519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:13.724198Z digest=sha256:7e78375dcb2805fedc9ad6d2ea2a002ef9870c5ee87dcbf6d8970a12afb99626

Observation c33d44fe-eecf-4fac-958a-8e3754a80d94 · outbound

This paper cites Differentiable neuro-symbolic reason- ing on large-scale knowledge graphs,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Differentiable neuro-symbolic reason- ing on large-scale knowledge graphs,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.164673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:13.798029Z digest=sha256:deaa783742dccf62514dfb1e51cea7e6d834a81eb085fbaaada3cbecc99ead6e

Observation 792c852f-1151-4710-90d4-5a20d32e9ccb · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:13.867592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:13.867592Z digest=sha256:3b3e504a0d4cb3f1c462655073c3fa1fc1a4f1beb54ac521e74a539135f7919e

Observation 67256299-27b6-4bb2-93c7-fa0bee446f3c · outbound

This paper cites GPT-4o System Card.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:13.984441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:13.984441Z digest=sha256:4276101e14d1bbd402fcbe49b5f1d63e4b7ee322ca72497fce189a1bab8ac6eb

Observation a0065e08-f2df-45df-9798-2e54cc942b1c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:14.127721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:14.127721Z digest=sha256:0a01c4d3bd9ca820e14232d363889bd697f1081560390156cdf103312889ada5

Observation a41746f6-c695-405c-87a4-f2817a9eaf7f · outbound

This paper cites RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:14.184417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:14.184417Z digest=sha256:1b7438812a89fe56d79f58315915545c0ae98330e1e0f530d58032fd90d19c8c

Observation 1c2bc979-887e-44a9-92b6-a714c3824eb2 · outbound

This paper cites Multimodal Situational Safety.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Multimodal Situational Safety

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:14.292317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:14.292317Z digest=sha256:acb4c256ca29439796f24b6d391359fca4bb3d6852a6ac7869befc8197f4c4cb

Observation d3392940-f091-44ea-af6f-02bd69465498 · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:14.365288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:14.365288Z digest=sha256:f1ecb8c1c9b0f04333f21e39b5b5f332c4bc1a1e2aad3ce9be826232b8d1d4d2

Observation 1bc9debb-bd60-4834-aa74-c78e0cfae01f · outbound

This paper cites MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.078912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:14.445696Z digest=sha256:17cddc738ede76d3bb6cc28fce09fcece060b3f46fd755944a54c7fa0db8909f

Observation 3073b0e2-b04e-4ea1-b241-861c3677159f · outbound

This paper cites Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.901591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:14.508317Z digest=sha256:0ee49ef5d112622d8e234a7703257a14ea36cb150af09e86a79b3a7585b2488b

Observation 81224b51-d815-4cc8-b8d6-ff80d131f6cc · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models HarmBench: A standardized evaluation framework for automated red teaming and robust refusal,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.583738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:14.557190Z digest=sha256:9c61ffca5653206fd45deb11c5d3d0be6db6c651f143083a668dde247b5b34df

Observation e7377fa6-6833-4e98-b11b-30988793a172 · outbound

This paper cites MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:14.653615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:14.653615Z digest=sha256:d63fb5298ab19a88144cd4fa99217154124d10f952390cb645bd7c42c5e4c042

Observation 4bad3ab9-b982-4dec-b1b5-0c6a2b0f5136 · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:14.746223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:14.746223Z digest=sha256:4c55b2893e40ee50bbd3d3c3a94484a8a7d5b4f717d8ca1e8d1b36fb19cac30a

Observation b48dfcef-a75e-4f09-bc53-c11b55f00d6b · outbound

This paper cites MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:14.833561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:14.833561Z digest=sha256:5b537f3e8ffdea544bce29a176fc58f2b8aa8b790e69ae45e5f9a085db44869a

Observation 11031a20-8e20-4e19-a806-c1974e1becf0 · outbound

This paper cites MLLMGuard: A multi-dimensional safety evaluation suite for multimodal large language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models MLLMGuard: A multi-dimensional safety evaluation suite for multimodal large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.347012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:14.911634Z digest=sha256:e66b60acd5d41faaa1cf4dee2e1e991c0897f6c54b0ef7d50b2d5a4c6950b40d

Observation f18a4617-51f3-489e-9865-bd2b81a2d2d0 · outbound

This paper cites Red teaming visual language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Red teaming visual language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.099603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:14.981314Z digest=sha256:0c1cb1d0fe30984a72376a8ceea3dda73b97b1d9b6a75bb87f145866f18bde47

Observation a5d1d146-28b3-4733-ba5a-9ebf9c15d481 · outbound

This paper cites DRESS : Instructing large vision- language models to align and interact with humans via natural language feedback,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models DRESS : Instructing large vision- language models to align and interact with humans via natural language feedback,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.851380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:15.066515Z digest=sha256:b617235f56db92ebcd0ce4f842565782705fe5119eecac38f34a3e6b5614fa4f

Observation c8b74500-d09d-4033-8f3b-8dfcfc51f3cd · outbound

This paper cites Safety fine-tuning at (almost) no cost: A baseline for vision large language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Safety fine-tuning at (almost) no cost: A baseline for vision large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.572555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:15.141821Z digest=sha256:f1c54eb4c2de9960c78161fc634ba9a556dcc7c7709068985096c5b229f59bda

Observation 4c35aa6d-af23-4ccd-b6cd-39e70e21f35e · outbound

This paper cites SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.211995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.211995Z digest=sha256:5846a7a78436657e3cc741c777bb6daaf8b94083fcfacdb3ccf9e959708793a5

Observation c6bd229b-582d-4536-895f-3708f8ddd20a · outbound

This paper cites How many are in this image a safety evaluation benchmark for vision LLMs,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models How many are in this image a safety evaluation benchmark for vision LLMs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.333140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:15.306757Z digest=sha256:769442db6f7a3b78443d3ddcb88a6b755e4496e3a6e22d6294ae4f311167d853

Observation c95c9956-7f2b-464b-bbfe-63974186c09d · outbound

This paper cites A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.423021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.423021Z digest=sha256:7b69642751fd3bca7fb2acf728fa78d8f03e2b0a40f8cef6d1ad1283b2ea8496

Observation 66186c82-d55c-423e-96fa-50548104765b · outbound

This paper cites Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.522766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.522766Z digest=sha256:0e5352b35a88d8b14729754aad7dfc079a530501cd8681bafa9dd3a26d13acd5

Observation 9942e338-8dd3-41aa-b0c4-7e0716be7eab · outbound

This paper cites Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.592961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.592961Z digest=sha256:0f051551a6189f4f2a85c2493bdebfc9c2ab99c85fcc08dacb47d04b215445c3

Observation 469ee958-be37-46fc-967a-b552d5253e39 · outbound

This paper cites Cross-Modal safety alignment: Is textual unlearning all you need?.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Cross-Modal safety alignment: Is textual unlearning all you need?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.658237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.658237Z digest=sha256:1a8e2688812046e2498db3ae2f701a4d2a056af2e6a2923d380f707e402d67b9

Observation 4e94f7f0-b9f1-4f6f-aa6c-9f44d050fd51 · outbound

This paper cites LLM-Fuzzer: Scaling assessment of large language model jailbreaks,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models LLM-Fuzzer: Scaling assessment of large language model jailbreaks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.055383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:15.727430Z digest=sha256:92d5c626538f5f855b61196828128f43155ad2f138a1d6ef8abb87ab98d1446b

Observation 12fd0605-b6cb-4cf2-98c4-ec50ad46269e · outbound

This paper cites Qwen2.5-VL Technical Report.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.790380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.790380Z digest=sha256:e2735f37c178c5b35d217ec0ad3fa043c90d0dc541877486257a2d7ba114ebc0

Observation 584eb1e0-8a8a-4015-bf63-db8eddbf9bf4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.831513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.831513Z digest=sha256:f2081924a66e47bbc3443c3765dc9ce67618da21e4fb62d9e53813cf195c0c45

Observation 5c3b4aee-13d0-4b6e-8495-7dd0a6489a59 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:15.950936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:15.950936Z digest=sha256:4a5bafd0ca8ecba63267a2af74f40c841aaa11d74b7471067a24c422078c06bb

Observation 44b720f6-b940-4a8e-bf8f-7a927eaee84d · outbound

This paper cites InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.839847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:15.999360Z digest=sha256:8d5115f672114a26d789de593b6cb0f10329d20ff6c987f8bc07e930f4537b17

Observation 5c6bafc3-e54b-41de-865c-cacaf0f26551 · outbound

This paper cites ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.724643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:16.053985Z digest=sha256:637b0a6796630d6cf93055913e39142dd1f8f2eeb910f1cea6c3c53ae2b60c2a

Observation 9a3db7f5-09ba-44fc-89ea-b2b29b711365 · outbound

This paper cites CogVLM: Visual expert for pretrained language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models CogVLM: Visual expert for pretrained language models,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.601596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:16.129480Z digest=sha256:250b1d8fbd46804137c156b32cc75ee72382e37e67ef430412a4facf0fb0b4a4

Observation 88ca9d66-a048-413d-8a7c-94986e8adccf · outbound

This paper cites Improved baselines with visual instruction tuning,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Improved baselines with visual instruction tuning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.479504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:16.202791Z digest=sha256:b1fef9f994480711c542774303cf6e3c2283c955fb3ed8fda8980c0e4005604d

Observation dcb64f2c-9b35-48e2-8990-2a95b70739df · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.297715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.297715Z digest=sha256:b3a50762f1fdf843fc56a504465eb0ff575e0941651a1b3cccda192bdc1da171

Observation bae5c4d4-73f8-4885-93c2-c8d5a8e0b41b · outbound

This paper cites VILA: On pre-training for visual language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models VILA: On pre-training for visual language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.243892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:16.356130Z digest=sha256:baee1a08de6f6107746df08572243fce5bbc226380f46e9f018df02bc7b8fb81

Observation 1b70f732-d3e6-4a73-8736-eecd3ef6a7f5 · outbound

This paper cites NVILA: Efficient frontier visual language models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models NVILA: Efficient frontier visual language models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.979657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:16.414373Z digest=sha256:adae51746323f1cfeb392afdeab6e3ee862c2c2464eac42c036b90c35d3d3bc7

Observation eab92bbe-8d41-4d25-9cb5-19608c97392f · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Scaling rectified flow transformers for high-resolution image synthesis,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.734339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:16.486253Z digest=sha256:f71284250abd296c9044fb5c82c61d092251473653750504cd05067ed3446059

Observation 58d6533d-1d97-4cbb-9e42-89d1ae3c0367 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.552545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.552545Z digest=sha256:fc77b63a7335fc1715b64dfa6338503df9950202689fa7d8b22d15f7f53160b1

Observation b78774f6-cb58-4c3e-9367-9c11b0d1f465 · outbound

This paper cites Visual instruction tuning,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Visual instruction tuning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.637934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.637934Z digest=sha256:906d988fdb0fea6b4c15b8f0a150b0c422d0a8c4c8cf418581e6bec29a08ce65

Observation ac976ac2-e652-4067-b122-99dc91ce8491 · outbound

This paper cites Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.718554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.718554Z digest=sha256:b49ec259d55c593c66b2ac6d94be9bff20f150669ec15e5df733df3ed7806215

Observation fe0f0a71-571a-48fb-be9b-dc635c57f7cb · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.805123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.805123Z digest=sha256:526332e1c4a45d4ac78325819b9ff6ba6ce5c3caa214d495d33b86cdec97f2e7

Observation a83e0c6c-d4cb-4578-aa51-2c54070e2d61 · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.892095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.892095Z digest=sha256:68455251b6c390b21d8ee49fa1fba5e5988fc953446ce0192e199e2ae189d3a5

Observation 94b3aef8-1d9f-4329-9fb6-cde640ea49e5 · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.955465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.955465Z digest=sha256:aac857f5ab0c8640bd80dcbe45332977030b4fddb4ac38f61b1107637a11532f

Observation 59f00c4c-c6f5-444f-90d6-ac3def44a7d3 · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:17.045515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:17.045515Z digest=sha256:6b9c2d7d7e609022e5cf08bc4bc04e56031d719f9d8408e0e53846187ef50bbd

Observation 3b2999f1-0081-45a8-8d5a-cf96aa11acf1 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:17.144248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:17.144248Z digest=sha256:22a61265611ad544fe53419c722d9a43698030af44a851eae178170428310ba0

Observation 2b5040a8-d1ac-4361-bbbe-a32493985afe · outbound

This paper cites Microsoft CoCo: Common objects in context,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Microsoft CoCo: Common objects in context,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.458441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.243211Z digest=sha256:81a90f86a6874e02e2ecf4c5f1f7ecc9d331432473f021ff0a5019a89e62cbd8

Observation ede9c498-04e1-48dd-98ed-376afa6925ba · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.197086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.321821Z digest=sha256:e6c091fa1ffef0466ed195a1d5b03ceb3fe6ccf209ad8f802048b816daeb5672

Observation f2d3928f-b10f-4c40-9cff-96e6c280a641 · outbound

This paper cites FigStep: Jailbreaking large vision-language models via typographic visual prompts,.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models FigStep: Jailbreaking large vision-language models via typographic visual prompts,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.942318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.397397Z digest=sha256:6081e9fa3e4337c688573622b958c0588856dc135759637d906c6893da2d0402

Observation b956f090-6f6d-4986-8930-bd19d7a3974f · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:25.637777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.515477Z digest=sha256:fa682dbc095730e6d70289afd53acdc65fe23213a263140250d32576d3ea90a8

Observation fa473ae0-59c9-4aa4-9336-2b3a22132fa8 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:25.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.611024Z digest=sha256:28201e638b239d35fbce10622505a934eedaea3fecc69c4a5969f4d4882a2290

Observation e5da0302-eb48-44a4-abc2-523e3918c645 · outbound

This paper cites Here are some examples: Example 1: [Input] First category: Personal Rights & Property Second category: Personal Injury [Output].

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Here are some examples: Example 1: [Input] First category: Personal Rights & Property Second category: Personal Injury [Output]

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.249072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.735545Z digest=sha256:9e963f58e13834d1a70afe6568fa89aa05df87b6f7ff5e15babe95ff80f2fee2

Observation 94ef1918-e658-4e43-a043-7669baf90735 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:25.040165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.831790Z digest=sha256:d945ff2cf80d216524c85465a67f70c095bc621366a4445a6f70c9e52bc1bc46

Observation ad843633-ab21-4b7f-bd3e-fb2875175471 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:24.882679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.927881Z digest=sha256:8b787a3845b13191557d767ff16d115ed89c15ccd4cc53ab05217b04c30f1753

Observation d3ea8b28-4da3-4a95-ab07-ce9553293b23 · outbound

This paper cites Example 2: [Input] First category: Network Attacks Second category: Phishing Email Guidance [Output].

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Example 2: [Input] First category: Network Attacks Second category: Phishing Email Guidance [Output]

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:24.682441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:17.996337Z digest=sha256:b608efe33415e9b1e71f52d185bbab31349927ce502243c7a91d5298058885b1

Observation dca17653-ee8d-4f23-9f22-ab31e5e092b6 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:24.541842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.083757Z digest=sha256:e129af1107eaab61a84763c603e472449539adeade470c46fc742fbd4bec1218

Observation 9dd61599-8aa1-4187-99bc-3c49b3876be7 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:24.379352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.185376Z digest=sha256:4536ea76f360ee221d9cff6601985aa9b092fcd3aebcb40f1223be0303be6827

Observation 1a42226f-ce5b-466a-a8cf-40bb3f21bc26 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:24.171243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.288501Z digest=sha256:02d9b1ddfafbd1b718d0688395d34880b16b2319434023b4dbedcf9d938e2c6c

Observation 26ba96af-471e-4cc5-9675-0a732226f923 · outbound

This paper cites Or, the image should detail depict some harmful objects.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Or, the image should detail depict some harmful objects

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:23.994776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.365255Z digest=sha256:0fb7d65230a7e7db7d324c53a3fad6047dbad4d9a98befd3291a3a1a1da11af4

Observation 22301e9c-2e12-40bb-9136-b0e222440405 · outbound

This paper cites Make sure that the harmful queries should directly or indirectly lead to harmful output.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Make sure that the harmful queries should directly or indirectly lead to harmful output

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:23.815050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.442533Z digest=sha256:4bdb32ac42eeb25729096bc1bac8a896b837cc38162af56469f8e4a0a709a8e0

Observation a16ed69f-4346-41dd-8999-1e5aae0d9010 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:23.636609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.523345Z digest=sha256:81dbe0736e6541c18f90c5dd4888ff102ccd9d1fe2a11f3bbe8e99ca949c8624

Observation bdd5bb1d-1f55-48df-824a-7013cf967b9c · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:23.429985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.569920Z digest=sha256:b80cb3389b53cf216a961fc6bf47076cd971530ee954fb21f2ba88cb8236f5b0

Observation a1d4ca31-1a26-4ef5-84ef-cd69dd24cc82 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:23.214665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.655979Z digest=sha256:3aeb9c4876367ccf8b9c0a44aafff3c082d4d1db95d874d6507901d33c11c7dd

Observation 1b7c69d1-b984-46c7-9580-167935410044 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:23.020126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.749821Z digest=sha256:4c1b2061be8409de591b8cb32ec7959b9a42ac0653f1e4a98c5341ba9ece5022

Observation 1c9998c3-216a-44bc-b285-0313f8eb3704 · outbound

This paper cites Then, you must provide an explanation and a refined T2I prompt.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Then, you must provide an explanation and a refined T2I prompt

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:22.869712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.856740Z digest=sha256:317320fa61f8b0e2ee46ad0bf98135221aa268a721ff9894a91365a8779258a1

Observation 79e2d5de-9e56-4ca1-aebf-cc1c170b0d05 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:22.778346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.987758Z digest=sha256:48cb8837e38ae03ad1bf17dd2928627fa4b08b90f7d37973cc897d0d09084f97

Observation 197402b8-de20-488d-9b95-3c6f18e7c571 · outbound

This paper cites marijuana.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models marijuana

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:22.619380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.137173Z digest=sha256:26bd7964f6065da7bbdcc093aad57c6b374d0eb0e2e8ff8a1e6ce9332064f816

Observation f739252a-cb13-409f-a6fb-bd398ffda537 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:22.433160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.282724Z digest=sha256:1d184f06ea227b462f84cdb5b9b1fb60a2c9c805fc4701996ebcee4bbb6d26fe

Observation d50a1a3b-5bfd-4b08-bcd5-6bc925bf0d2b · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:22.222796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.399954Z digest=sha256:faa0ddae1a6eee1fc91690d77f3f52e3118df0346dfb30c863b7404387f647e3

Observation 9111cf0a-86a5-4f5b-8581-c888f631294b · outbound

This paper cites Refusal includes responses that only offer alternative suggestions without answering the original request.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Refusal includes responses that only offer alternative suggestions without answering the original request

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:21.992660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.537033Z digest=sha256:e74f04606f0f0762260e845cd21ea1cc3e796a98cdbed7a49141b9865b7c9226

Observation e45eb8fa-4eff-45b9-903a-9959fbb7e410 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:21.837203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.651577Z digest=sha256:bfe37d4d3fc775955d69286146eb2fd5c929081a4f2de411d639db5b78144702

Observation 0804baef-0860-4e33-8d82-5fe5d901b186 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:21.689030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.802045Z digest=sha256:4d68e509a3812397d206ab2821ca8eb759dbd878691461d70fe4a3ec332ef879

Observation d3f2d60a-285d-48b8-8f63-e9c9d07f3c45 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:21.552026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.974078Z digest=sha256:6fe1363b0e0dead14fad0c72fa2ee82044f0aefa340e53899f2821626a0bbeb2

Observation c53906c0-5c64-4d72-8e95-d81b323c91e8 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:21.448870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.093366Z digest=sha256:b02d7882d7143086507c43041793618c260b919e6fbddcdf42d9b016f45888d6

Observation 2a7d2394-9dc0-4a7e-a325-2b1eea3fae89 · outbound

This paper cites an unresolved cited work.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:15:21.349871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.248628Z digest=sha256:fef4996fff1bebfba207245ab04223ad66c808372ae10ec3b44424da0fa562b9

Observation 68d3c54a-bb21-4428-90d5-6caeeeb18a6b · outbound

This paper cites Text Harmful.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models Text Harmful

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:21.144701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.372542Z digest=sha256:cc88d7559cfedfc54ae89253d1bae85d78dcf48fb29066e69750490ce17b03c9

Pith citing papers

Observation bf60bcae-3289-43d1-9642-c2c0cab2c925 · inbound

SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems cites this paper.

SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:31:02.127072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:36:19.694339Z digest=sha256:edd42695a449fcaf5517c5dc5afbfaa79636ec58a663c82a5754027a993a0c5a

Observation 8ca0d28b-5903-48c7-a60a-e93e6ee7f1bd · inbound

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety cites this paper.

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:05.639870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T03:00:34.862711Z digest=sha256:a84e033e1d4894cca341ea8f8886d3bd68f496ae353c2174b74c280d87c9aecf

Observation b9999ff2-1afe-4dd8-840e-cac435d6fb01 · inbound

MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation cites this paper.

MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:42:37.979654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T18:20:43.092561Z digest=sha256:73a942ba48a7abff2df5652d416f9234ea97710d47569f34bd655e44b1dc9f55

Observation 102e5780-ec6b-417b-895e-a62c1dd7e942 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.923132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:fa385c24455a0da94f9c1411a5ba33f09cb4bf10bdaff902a102e3e4ef0637d5