Pith. sign in

Paper Citation Record · LEDGER

Safety of Multimodal Large Language Models on Images and Texts

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2402.00357.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.00357 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:31:25.694237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:27:30.991837Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6eee69c-3885-4fef-b765-55b7dd97713e · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Safety of Multimodal Large Language Models on Images and Texts

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.275081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.275081Z digest=sha256:8c680340abf006c189d647a2e43510cb24fd7fe9ec0c956fdac020d44ad4b266

Observation 001a728f-c1d1-479f-9d59-743bc432ad1a · inbound

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting cites this paper.

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting Safety of Multimodal Large Language Models on Images and Texts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:04.285198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:29:04.285198Z digest=sha256:f8d4848e90d6ba6aa03075262dbf755a4f2852511ca146cc811a559a877f5eea

Observation 1b17aab4-b0ab-4e96-8c3f-41bdf1be6cf6 · inbound

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models cites this paper.

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models Safety of Multimodal Large Language Models on Images and Texts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:32.306786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:28:32.306786Z digest=sha256:2ce3a3e3d861f74713aa83b77d755a731616146645c205f89ddc7f9cc66d5be2

Observation c919b9d8-7b51-4ed8-9944-1158ad58de36 · inbound

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models cites this paper.

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models Safety of Multimodal Large Language Models on Images and Texts

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-09T14:58:53.796847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:58:53.796847Z digest=sha256:aac4f7380d5daab892127a621dff64454add169f16267b82735bbf70ec01873a

Observation fd9d8eb5-393c-4e58-a585-61fa9825b70d · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Safety of Multimodal Large Language Models on Images and Texts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.103048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.103048Z digest=sha256:22301108711101be1643997c8129b3077c085ed19e81fff1660e44ef8745b1dd

Observation f71db071-708a-45be-900d-a66b1762284c · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Safety of Multimodal Large Language Models on Images and Texts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.435520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.435520Z digest=sha256:b21c299fe669cc8094bf596013c46f14d3a805047870282379894edda9a8556b

Observation ba15c73c-8609-4fb2-a3be-41e2bb43b959 · inbound

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects cites this paper.

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects Safety of Multimodal Large Language Models on Images and Texts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:10:22.971662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:10:22.971662Z digest=sha256:423150390a91cedee7518fdc24fc18ae90fd710501b60b05789933da4ddff2a8

Observation cccdc95a-89aa-4648-9e47-aaf4bb2cc0e4 · inbound

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment cites this paper.

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment Safety of Multimodal Large Language Models on Images and Texts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.646908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.646908Z digest=sha256:06fbfc9aed3043d48306a45567a096ebb0f5020389b27d310f791c5376167786

Observation fae58dc4-feec-41ee-ac7a-2ab4edfe5819 · inbound

Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks cites this paper.

Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks Safety of Multimodal Large Language Models on Images and Texts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:31:52.585142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:31:52.585142Z digest=sha256:1bb6f81a18850bf58c1dc927b35a453f7dc507d15eb5f458f1223c56433b85cb

Observation d2bbf271-6201-44be-a588-6223836e8b4e · inbound

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair cites this paper.

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair Safety of Multimodal Large Language Models on Images and Texts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:21:02.805074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:21:02.805074Z digest=sha256:07c18a6c2c11d6ddfd410c0345e2d99cfbba666c527e670849297801408f7850

Observation cbad9768-587a-4bb5-8ba6-3e50a50fa5e4 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Safety of Multimodal Large Language Models on Images and Texts

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:46.157970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:46.157970Z digest=sha256:10c1ab6683d8462e595da5f5695ce47787faadaa0aafac4e62f674f5ea007911

Observation 3894a040-3054-4e18-a724-87ac4b6121f9 · inbound

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection cites this paper.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Safety of Multimodal Large Language Models on Images and Texts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.384240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.384240Z digest=sha256:8812cafecd5a3f57ee85e3a1d36fba0ba223eaa112d061b5f304998f051423bc

Observation 97a40a27-a51b-49e8-b125-c291234dd07f · inbound

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing cites this paper.

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing Safety of Multimodal Large Language Models on Images and Texts

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:27.469126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T04:50:08.866969Z digest=sha256:c632cdc184d4783aef98e00fc9536b8c61d583e12a06a3686788ebd1bbf0a242

Observation a3431651-ae9c-4ad6-9d92-b0bb7f6ad56a · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models Safety of Multimodal Large Language Models on Images and Texts

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.639652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:ce36239ad7300a7bfa3e7e0c034f64cce3434cf3644cd45fcda669d8f893c70d

Observation 339e86b5-919c-4603-ab5f-e6107eca8397 · inbound

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges cites this paper.

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges Safety of Multimodal Large Language Models on Images and Texts

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.993178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T16:33:28.848573Z digest=sha256:a862b42b8f7dba3b9f9f33a7676dda6d5d950a62389307820a5872d7fab269f5

Observation d54a7677-0c02-45c0-a740-7977675c926f · inbound

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure cites this paper.

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure Safety of Multimodal Large Language Models on Images and Texts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:23:25.032757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:23:25.032757Z digest=sha256:b371da1e669cce6535d3130e6a67a9a5dcd63482f92c912b7dbad37a15749087

Observation 718fda89-e059-4c42-8d1f-89c3c1e8b2bc · inbound

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment cites this paper.

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment Safety of Multimodal Large Language Models on Images and Texts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:25.694237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:25.694237Z digest=sha256:9a0524293c44a7a68d3872e98d8ec3b5e4633365425138f52e316c5d508b94af