Pith. sign in

Paper Citation Record · LEDGER

SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2406.12030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12030 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:45:18.699387Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:36:25.477552Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d2e12838-4ccf-49ba-9b17-e5e8a4b62372 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.566890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:886fc3029a01d424a0c6d8b9f30c3ba102c5fc827c0eeac9a77ea33a616f8878

Observation 8d80368f-848e-46e4-ace1-a54ae02b4736 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.699387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.699387Z digest=sha256:eb35afd2cb46dd0890aa2c0069309d99c5d5881617f9db0ba5da3ffa658157b6

Observation 83c036a9-1752-4c5d-9d45-0ad72cb60093 · inbound

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs cites this paper.

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:26.196963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:26.196963Z digest=sha256:73c90b4a496dba8c93d72a4df4c474a9cae2a36033cd42d57ef615186b3a43d4

Observation 4094a705-1493-4f36-8a93-9f18872c12d1 · inbound

MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models cites this paper.

MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:33.585160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:06:33.585160Z digest=sha256:3aa4ab51c38c73d9d95f1211ade9d8ca06fb48761d6fb0104614a3407de0311f

Observation 20f30d17-d06d-43b2-b242-b28127eaadec · inbound

VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration cites this paper.

VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:13.206915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:13:13.206915Z digest=sha256:d0a4c94ef10ca0eb0a5e5efa64feb0a3839fb7323f62660248a594235494ca51

Observation 66eaf8e9-4a0e-441f-bf8c-28a6b5c608d6 · inbound

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack cites this paper.

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:02.983754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:22:02.983754Z digest=sha256:1aa092bf9980dd9067f33f78431499fedbdfd6276560d8ee6a00a571a3b68c47

Observation 3b2999f1-0081-45a8-8d5a-cf96aa11acf1 · inbound

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models cites this paper.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:17.144248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:17.144248Z digest=sha256:22a61265611ad544fe53419c722d9a43698030af44a851eae178170428310ba0

Observation 1db626b9-2698-4367-afcd-c3ca4ce1ee8d · inbound

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap cites this paper.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.450280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.450280Z digest=sha256:1a14f31e271debecbcdb56b3284b9a009d33afb8bba633a46fcf3b1959eb5ce5

Observation 4a5719e8-17ef-4d2f-994f-0806687c0102 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.660602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.660602Z digest=sha256:c45e27e77579488e345df0c3985042484e224299154e25289d46a0a48486cbff

Observation fa06ee05-0e53-492f-9fb9-e19e77971e5f · inbound

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents cites this paper.

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:45.230207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:45.230207Z digest=sha256:744bde0d7bb33e6f5aa9779fbee78453d2fc3f8f7674b6ba161c765d638e7855

Observation ffc2c45d-0465-46fb-a516-7373399e78f1 · inbound

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities cites this paper.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.371849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.371849Z digest=sha256:ff43c09361a37ed4748245a64794f1a4f085dd864b726fa4d778a7c31ee0d2a8

Observation d841a76a-6ab2-4fe8-8b92-cefda9142c5a · inbound

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs cites this paper.

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:08.721119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:45:08.721119Z digest=sha256:a72483c04c822e81a6b52008e636e3c8a311550253ac3071ef562b76801b8d4f

Observation 08d7f4a5-692d-40ce-a274-cc75b88179e0 · inbound

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization cites this paper.

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:11:33.026717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:11:33.026717Z digest=sha256:8f51fa86be1ce148a7f9f399198636e0e810f911534cb292d0689adda578cd73

Observation f44a7cbb-4b0b-474c-8a1b-1cb1de01c1be · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 274

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:37.785906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:37.785906Z digest=sha256:366091e2db6c5992e484b96266ac04beea7ee9d4afb070ea1dfee1145200c5f2

Observation f5b84ebe-d63f-4aa0-bb8f-2e68e31372cd · inbound

AlignCultura: Towards Culturally Aligned Large Language Models? cites this paper.

AlignCultura: Towards Culturally Aligned Large Language Models? SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:04.792001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T02:36:36.854805Z digest=sha256:bc998f9ee9e411b46f3b0754cef679778d5f0c87e24a091e468cc6559289be99

Observation 8368717a-f511-4aad-8797-1c82f474b182 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:02.050551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:90679396bd59c4729d0635f5d0696148b71552a845b7fe0a914512a1f4cd0991

Observation 5b865bbb-0aaa-4bac-875f-a138a188ffb2 · inbound

Constitutional On-Policy Safe Distillation cites this paper.

Constitutional On-Policy Safe Distillation SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.479676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:47:14.135793Z digest=sha256:5fedacdd33f38f246cad6cd369bfb76864b0065a9b272786f70de6d6bb1ba103

Observation b2e385e7-cee5-4c69-be4b-fe72760b647b · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:c91b08cac3060fd0616368ca9da2c086b2140001ce35600e045aa97e48cf8ea2

Observation f1f5fb41-b8fa-4e59-8cf2-210c92a0e195 · inbound

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation cites this paper.

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:36:57.727029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:36:57.727029Z digest=sha256:c080e5a9a5e3cdafc4b860b83efd746521fbfd84459369b030c6ab52fbb09acd