REVIEW 3 major objections 4 minor 72 references
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Caption-mediated reinforcement learning improves large vision-language model safety by up to 19 points on aggregate benchmarks while preserving vision utility.
desk verdict SafeCap's DirectCap safety gains are real and the ablations are solid, but the reward is computed entirely in text, so the 'visual grounding' mechanism is asserted rather than demonstrated; worth serious review with mandatory grounding checks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the caption-mediated reward, a composite of three components: (1) a template gate that enforces exactly one well-formed caption block and a non-empty answer; (2) a caption reward that combines a frozen text-only LLM's safety-alignment agreement with the policy's answer (a binary judge signal) and a rubric-based descriptive-coverage score; and (3) a direct answer reward using the exponential risk discount $S(u,h)=u\gamma^h$ with $\gamma=0.35$. These values are group-normalized component-wise before being combined into the policy-gradient advantage. This design steers the policy to expose visual evidence that is both detailed and sufficient for a text-only reasoner to make a safe decision, rather than simply refusing or parroting generic disclaimers.
What would settle it
Feed a SafeCap-trained model a set of images where the generated caption contains plausible but hallucinated safety-critical details (for example, a weapon not present in the image) and check whether the refusal rate tracks the actual risk; a high refusal rate on benign images with hallucinated hazards would show the reward learned to fabricate evidence rather than recognize real danger.
Extended reading notes
Core claim
SafeCap's central discovery is that a caption-mediated reinforcement learning objective raises LVLM safety more effectively than direct refusal supervision or standard safety fine-tuning. The policy is trained with a variant of group-relative policy optimization to emit a structured response: a caption block followed by the answer. The caption is rewarded when it lets a frozen text-only LLM produce an answer whose safety status agrees with the policy's own answer, and when it covers concrete visual details; the answer is rewarded with an exponential risk discount, $S(u,h)=u\gamma^h$, that suppresses risky content without zeroing out useful safe answers. The intended operating point, DirectCap, shows the most consistent safety gains, while the diagnostic Direct and Prism paths are more model-dependent. The authors conclude that learned self-captioning, rather than inference-time caption wrapping or direct refusal, is the reliable path to multimodal safety alignment.
Load-bearing premise
The reward signal assumes the frozen text-only language model and the judge models are trustworthy safety arbiters; the caption-retention judge explicitly does not verify factual accuracy, so if captions are detailed but hallucinated, or if judges reward wordy refusals, training can raise judged safety without raising real safety.
Editorial extensions
If this is right
- If SafeCap is correct, multimodal safety alignment can be achieved through a trainable captioning interface that preserves vision utility.
- The consistent safety gains across four model settings suggest the method transfers across different backbones and initialization states.
- Under matched training data and steps, SafeCap outperforms safety SFT, DPO, and a recent rule-governed GRPO baseline, indicating that caption-mediated RL is a stronger safety objective.
- The Prism diagnostic, where a frozen text-only LLM answers from the caption alone, shows captions retain enough evidence to transfer safety judgments to a stronger frozen reasoner.
Reading between the lines
- The paper's claim that judges do not check caption factuality raises a testable risk: if captions hallucinate safety-relevant details, training could inflate judged safety without improving real-world safety; a follow-up could measure refusal rates on benign images with hallucinated hazards.
- The exponential risk-discount reward form might generalize beyond LVLMs to any alignment problem where helpfulness and safety must be balanced without rewarding cheap refusals; one could test it on text-only safety alignment.
- Because the caption is trained to support a frozen LLM's safety decision, SafeCap captions could be reused as a lightweight inspection interface for human oversight or for training smaller safety classifiers.
- The method's reliance on a frozen LLM as judge suggests its safety gains may be bounded by the judge's own risk recognition; replacing the frozen LLM with a stronger one could push gains further, but also risks introducing judge-specific bias.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SafeCap, a reinforcement-learning framework that trains a vision-language model to emit a structured <caption>...answer output. The training objective combines a template gate, a caption-mediated reward (coverage score plus a frozen text-only LLM's safety-alignment consistency answer), and a direct answer reward with an exponential risk discount. The method is evaluated on five multimodal safety benchmarks and six vision-utility benchmarks under three inference protocols (Direct, DirectCap, Prism) across four Qwen3.5 model settings, and is compared against safety SFT, DPO, and SafeGRPO. The main claim is that SafeCap improves aggregate safety under its intended DirectCap protocol by 3.7–19.0 points while maintaining or improving vision utility.
Significance. If the reported gains are robust and not artifacts of the evaluation protocol, SafeCap is a practical and inexpensive alignment method: it uses a public dataset, a frozen text-only LLM and judge, and requires no additional visual supervision. The paper includes useful strengths: consistent DirectCap gains across four model settings, ablations showing both reward components matter, seed-level robustness at an early checkpoint, alternative reward-form comparisons, refusal-rate diagnostics, and a small human review. However, the central mechanistic claim — that the caption reward teaches the model to expose genuine visual cues — is not directly verified, because the reward and most safety evaluations are produced by text-only LLM judges. The paper's significance therefore hinges on additional evidence connecting the reward to factual visual grounding.
major comments (3)
- [Caption-Mediated Reward, Eq. (5); Appendix J] The caption reward is computed entirely without access to the image. The caption-retention judge is explicitly instructed "do NOT judge factual correctness," the frozen LLM never sees the image, and the safety-alignment judge compares only the two text answers. Thus a detailed but hallucinated caption — or even a caption that is unrelated to the image — can receive a high coverage score and, whenever the frozen answer agrees with the policy answer on risk status, a high caption reward. For prompts whose harmful intent is stated in the question (common in MM-SafetyBench and MSSBench), the alignment gate g can equal 1 regardless of the caption content. The claimed mechanism that the caption "exposes visual cues relevant to safe response generation" is therefore not enforced by Eq. (5); the objective is, at best, a descriptive-coverage reward with a consistency gate. The ablation in Table 4 shows the caption reward contributes to scores, but it does not show the contribution arises from factual visual grounding rather than generic descriptive style. I request evidence of caption-image factual consistency on safety-critical samples (e.g., VQA-style factual checks or human factual ratings) and a diagnostic that varies the caption content while holding the question fixed to demonstrate that the alignment gate is actually sensitive to the caption's visual content.
- [Benchmarks and metrics; Appendix G] The training reward uses an LLM judge (gpt-oss-20b), and the evaluation pipeline relies on automatic LLM judges for at least FigStep and VLSBench. Because the reward and evaluation pipelines are built from the same LLM-as-judge technology, correlated judge biases (e.g., preference for certain refusal phrasing or verbose descriptive style) can inflate the reported safety gains without reflecting genuine refusal of harmful requests. The 100-sample human review in Appendix G is a welcome check, but it is small relative to the full benchmark suite and reports only an overall consensus, not per-benchmark or per-category agreement. I ask for either a larger human-validated subset reported separately for the safety-critical benchmarks, or for at least one safety benchmark scored by a rule-based or fully human protocol, to rule out judge-family bias as the source of the headline gains.
- [Table 2; Appendix A] The main training results in Table 2 are single-run point estimates. The three-seed robustness check is conducted only at a matched 100-step early checkpoint, not at the 200-step final checkpoints that produce the headline numbers (e.g., the +19.0-point DirectCap safety gain for 4B-Base). Because the paper's central claim is built on these aggregate deltas, I request variance information at the final training step, or at least a clear statement that the final checkpoint was selected by a fixed schedule rather than by inspecting benchmark results. Without this, the reader cannot distinguish a stable method from a favorable random seed.
minor comments (4)
- [Figure 1] Panel (a) has a typo: "Visaul Safety Gap" should be "Visual Safety Gap."
- [Appendix A] The sentence "The SFT and DPO baselines in Appendix F use the same backbone ... and preform same training steps" contains a typo: "preform" should be "perform."
- [Tables 1 and 2] The tables are dense because deltas are embedded as superscripts next to absolute scores. A separate delta table or a cleaner two-row layout per model/protocol would improve readability and reduce the risk of misreading sign and magnitude.
- [Eq. (3)] The notation uγh is ambiguous because γ is a real coefficient and h is an integer exponent. Please write u(a)γ^{h(a)} and define the convention explicitly (e.g., integer power) to avoid confusion with a subscript or concatenation.
Circularity Check
No circularity found: SafeCap's reward uses external frozen LLM/judges and public benchmarks; reported gains are empirical, with no equation-level reduction to inputs.
full rationale
I walked the claimed derivation chain from the SafeCap reward equations through training and evaluation. The caption-mediated reward Rcap = g(q,a,af)·u(c) and the answer reward Rans = u(a)·γ^{h(a)} are computed from a frozen text-only LLM and from judge models with fixed prompts; none of these reward components are defined in terms of the five safety benchmarks or six utility benchmarks whose scores are reported as outcomes. The evaluation uses public benchmarks with their own scoring procedures, including a 100/100 human-consensus recheck of the automatic judgment labels, so the reported gains are externally checkable rather than forced by construction. The DirectCap protocol is shared between training and the intended evaluation operating point, but the paper explicitly reports zero-training DirectCap baselines and measures trained gains against those same-protocol baselines, so the comparison is not self-referential. Hyperparameters such as γ=0.35 are selected with benchmark ablations, which is a test-set-tuning concern rather than a circular derivation. The paper cites CapRL and GDPO as external prior work for design inspiration; these citations are not load-bearing uniqueness claims and do not substitute for the empirical comparisons. No equation reduces to its own input, no fitted parameter is renamed as a prediction, and no author-imported ansatz is used to forbid alternatives. The image-blindness of the reward is a real limitation about whether the mechanism is visual grounding, but the paper itself discloses this limitation in Appendix D, and it is a validity concern, not a circularity.
Assumptions & free parameters
free parameters (2)
- gamma (risk-discount coefficient) =
0.35
- reward weights (w_tmp, w_cap, w_ans) =
(0.5, 0.5, 1.0)
assumptions (4)
- domain assumption The public SPA-VL dataset is representative of safety-relevant multimodal scenarios and its released split is adequate for training.
- domain assumption The frozen text-only LLM (Qwen3-4B) can make correct safety decisions from a caption, and the LLM judge scores are reliable proxies for helpfulness and risk.
- ad hoc to paper Descriptive coverage without factual verification is a sufficient proxy for caption usefulness.
- domain assumption GRPO without an explicit KL penalty remains stable for this reward design.
Cite this review
Pith. "Pith review of SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning." pith.science (2026). https://pith.science/paper/AX5VZLE7
@misc{pith2026260810513,
author = {Pith},
title = {Pith review of: SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/AX5VZLE7}},
note = {Machine review of arXiv:2608.10513}
}
read the original abstract
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7-19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.
Figures
Reference graph
Works this paper leans on
-
[1]
Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education
Clancey, William J. Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education. Proceedings of the Eighth International Joint Conference on Artificial Intelligence (IJCAI-83)
-
[2]
Classification Problem Solving
Clancey, William J. Classification Problem Solving. Proceedings of the Fourth National Conference on Artificial Intelligence
-
[3]
, title =
Robinson, Arthur L. , title =. 1980 , doi =. https://science.sciencemag.org/content/208/4447/1019.full.pdf , journal =
1980
-
[4]
New Ways to Make Microcircuits Smaller---Duplicate Entry
Robinson, Arthur L. New Ways to Make Microcircuits Smaller---Duplicate Entry. Science
-
[5]
Clancey and Glenn Rennels , abstract =
Diane Warner Hasling and William J. Clancey and Glenn Rennels , abstract =. Strategic explanations for a diagnostic consultation system , journal =. 1984 , issn =. doi:https://doi.org/10.1016/S0020-7373(84)80003-6 , url =
-
[6]
and Rennels, Glenn R
Hasling, Diane Warner and Clancey, William J. and Rennels, Glenn R. and Test, Thomas. Strategic Explanations in Consultation---Duplicate. The International Journal of Man-Machine Studies
-
[7]
Poligon: A System for Parallel Problem Solving
Rice, James. Poligon: A System for Parallel Problem Solving
-
[8]
Transfer of Rule-Based Expertise through a Tutorial Dialogue
Clancey, William J. Transfer of Rule-Based Expertise through a Tutorial Dialogue
Show all 72 references
-
[9]
The Engineering of Qualitative Models
Clancey, William J. The Engineering of Qualitative Models
-
[10]
2023 , eprint=
Attention Is All You Need , author=. 2023 , eprint=
2023
-
[11]
Pluto: The 'Other' Red Planet
NASA. Pluto: The 'Other' Red Planet
-
[16]
European Conference on Computer Vision , pages=
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation , author=. European Conference on Computer Vision , pages=. 2024 , organization=
2024
-
[20]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=
Deep Visual-Semantic Alignments for Generating Image Descriptions , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=
-
[21]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=
Show and Tell: A Neural Image Caption Generator , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=
-
[26]
International Conference on Learning Representations , year=
Jailbreak in Pieces: Compositional Adversarial Attacks on Multi-Modal Language Models , author=. International Conference on Learning Representations , year=
-
[28]
Pi, Renjie and Han, Tianyang and Zhang, Jianshu and Xie, Yueqi and Pan, Rui and Lian, Qing and Dong, Hanze and Zhang, Jipeng and Zhang, Tong , journal=
-
[29]
2024 , organization=
Wang, Yifei and Liu, Xiaogeng and Li, Yu and Chen, Muhao and Xiao, Chaowei , booktitle=. 2024 , organization=
2024
-
[30]
Advances in Neural Information Processing Systems , volume=
Fight Back Against Jailbreaking via Prompt Adversarial Tuning , author=. Advances in Neural Information Processing Systems , volume=
-
[31]
Oh, Sejoon and Jin, Yiqiao and Sharma, Megha and Kim, Donghyun and Ma, Eric and Verma, Gaurav and Kumar, Srijan , journal=
-
[33]
Wang, Peng and Zhang, Dong and Li, Linyang and Tan, Chuanqi and Wang, Xingyao and Ren, Kui and Jiang, Baoyuan and Qiu, Xipeng , journal=
-
[34]
Immune: Improving Safety Against Jailbreaks in Multi-modal
Ghosal, Soumya Suvra and Chakraborty, Souradip and Singh, Vaibhav and Guan, Tianrui and Wang, Mengdi and Velasquez, Alvaro and Beirami, Ahmad and Huang, Furong and Manocha, Dinesh and Bedi, Amrit Singh , booktitle=. Immune: Improving Safety Against Jailbreaks in Multi-modal. 2...
2025
-
[37]
Chen, Lin and Li, Jinsong and Dong, Xiaoyi and Zhang, Pan and He, Conghui and Wang, Jiaqi and Feng, Zhao and Lin, Dahua , booktitle=
-
[38]
Rotstein, Noam and Bensaid, David and Brody, Shaked and Ganz, Roy and Kimmel, Ron , booktitle=
-
[39]
Computer Vision -- ECCV 2024 , year=
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models , author=. Computer Vision -- ECCV 2024 , year=
2024
-
[40]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
VLSBench: Unveiling Visual Leakage in Multimodal Safety , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2025 , publisher=
2025
-
[41]
2025 , eprint=
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models , author=. 2025 , eprint=
2025
-
[43]
Liu, Shih-Yang and Dong, Xin and Lu, Ximing and Diao, Shizhe and Belcak, Peter and Liu, Mingjie and Chen, Min-Hung and Yin, Hongxu and Wang, Yu-Chiang Frank and Cheng, Kwang-Ting and Choi, Yejin and Kautz, Jan and Molchanov, Pavlo , journal=
-
[45]
Advances in Neural Information Processing Systems , volume=
Direct Preference Optimization: Your Language Model is Secretly a Reward Model , author=. Advances in Neural Information Processing Systems , volume=
-
[51]
2025 , eprint=
Gemini Robotics: Bringing AI into the Physical World , author=. 2025 , eprint=
2025
-
[52]
2025 , howpublished=
Visual Physics Comprehension Test , author =. 2025 , howpublished=
2025
-
[54]
Brower, C. 2025. Visual Physics Comprehension Test. https://cbrower.dev/vpct/
2025
-
[55]
S.; Dong, Y.; Roy-Chowdhury, A
Chakraborty, T.; Shayegani, E.; Cai, Z.; Abu-Ghazaleh, N.; Asif, M. S.; Dong, Y.; Roy-Chowdhury, A. K.; and Song, C. 2024. Cross-Modal Safety Alignment: Is Textual Unlearning All You Need? arXiv preprint arXiv:2406.02575
2024
-
[56]
Chen, L.; Li, J.; Dong, X.; Zhang, P.; He, C.; Wang, J.; Feng, Z.; and Lin, D. 2024 a . ShareGPT4V : Improving Large Multi-Modal Models with Better Captions. In European Conference on Computer Vision
2024
-
[57]
Chen, L.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Wang, J.; Qiao, Y.; Lin, D.; et al. 2024 b . Are We on the Right Way for Evaluating Large Vision-Language Models? arXiv preprint arXiv:2403.20330
2024 arXiv
-
[58]
Ding, Y.; Li, L.; Cao, B.; and Shao, J. 2025. Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models. arXiv:2501.18533
2025
-
[59]
A.; Ma, W.-C.; and Krishna, R
Fu, X.; Hu, Y.; Li, B.; Feng, Y.; Wang, H.; Lin, X.; Roth, D.; Smith, N. A.; Ma, W.-C.; and Krishna, R. 2024. BLINK: Multimodal Large Language Models Can See but Not Perceive. arXiv preprint arXiv:2404.12390
2024 arXiv
-
[60]
Gemini Robotics Team ; et al. 2025. Gemini Robotics: Bringing AI into the Physical World. arXiv:2503.20020
2025 arXiv
-
[61]
S.; Chakraborty, S.; Singh, V.; Guan, T.; Wang, M.; Velasquez, A.; Beirami, A.; Huang, F.; Manocha, D.; and Bedi, A
Ghosal, S. S.; Chakraborty, S.; Singh, V.; Guan, T.; Wang, M.; Velasquez, A.; Beirami, A.; Huang, F.; Manocha, D.; and Bedi, A. S. 2025. Immune: Improving Safety Against Jailbreaks in Multi-modal LLM s via Inference-Time Alignment. In Proceedings of the IEEE/CVF Conference on ...
2025
-
[62]
Gong, Y.; Ran, D.; Liu, J.; Wang, C.; Cong, T.; Wang, A.; Duan, S.; and Wang, X. 2023. FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts. arXiv preprint arXiv:2311.05608
2023 arXiv
-
[63]
T.; and Zhang, Y
Gou, Y.; Chen, K.; Liu, Z.; Hong, L.; Xu, H.; Li, Z.; Yeung, D.-Y.; Kwok, J. T.; and Zhang, Y. 2024. Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation. In European Conference on Computer Vision, 388--404. Springer
2024
-
[64]
Hao, S.; Wang, Y.; Hooi, B.; Yang, M.-H.; Liu, J.; Tang, C.; Huang, Z.; and Cai, Y. 2025. Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense. arXiv preprint arXiv:2503.11619
2025 arXiv
-
[65]
Hu, X.; Liu, D.; Li, H.; Huang, X.; and Shao, J. 2025. VLSBench: Unveiling Visual Leakage in Multimodal Safety. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8285--8316. Association for Computational Linguistics
2025
-
[66]
Karpathy, A.; and Fei-Fei, L. 2015. Deep Visual-Semantic Alignments for Generating Image Descriptions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3128--3137
2015
-
[67]
H.; and Seo, M
Lee, S.; Kim, G.; Kim, J.; Lee, H.; Chang, H.; Park, S. H.; and Seo, M. 2024. How Does Vision-Language Adaptation Impact the Safety of Vision Language Models? arXiv preprint arXiv:2410.07571
2024 arXiv
-
[68]
X.; and Wen, J.-R
Li, Y.; Guo, H.; Zhou, K.; Zhao, W. X.; and Wen, J.-R. 2024. Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models. arXiv preprint arXiv:2403.09792
2024 arXiv
-
[69]
F.; Cheng, K.-T.; Choi, Y.; Kautz, J.; and Molchanov, P
Liu, S.-Y.; Dong, X.; Lu, X.; Diao, S.; Belcak, P.; Liu, M.; Chen, M.-H.; Yin, H.; Wang, Y.-C. F.; Cheng, K.-T.; Choi, Y.; Kautz, J.; and Molchanov, P. 2026. GDPO : Group Reward-Decoupled Normalization Policy Optimization for Multi-Reward RL Optimization. arXiv preprint arXiv:...
2026 arXiv
-
[70]
Liu, X.; Zhu, Y.; Gu, J.; Lan, Y.; Yang, C.; and Qiao, Y. 2024. MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models. In Computer Vision -- ECCV 2024
2024
-
[71]
Mo, Y.; Wang, Y.; Wei, Z.; and Wang, Y. 2024. Fight Back Against Jailbreaking via Prompt Adversarial Tuning. In Advances in Neural Information Processing Systems, volume 37, 64242--64272
2024
-
[72]
Oh, S.; Jin, Y.; Sharma, M.; Kim, D.; Ma, E.; Verma, G.; and Kumar, S. 2024. UniGuard : Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models. arXiv preprint arXiv:2411.01703
2024 arXiv
-
[73]
Pantazopoulos, G.; Parekh, A.; Nikandrou, M.; and Suglia, A. 2024. Learning to See but Forgetting to Follow: Visual Instruction Tuning Makes LLMs More Prone to Jailbreak Attacks. arXiv preprint arXiv:2405.04403
2024 arXiv
-
[74]
Pi, R.; Han, T.; Zhang, J.; Xie, Y.; Pan, R.; Lian, Q.; Dong, H.; Zhang, J.; and Zhang, T. 2024 a . MLLM-Protector : Ensuring MLLM 's Safety without Hurting Performance. arXiv preprint arXiv:2401.02906
2024 arXiv
-
[75]
Pi, R.; Zhang, J.; Zhang, J.; Pan, R.; Chen, Z.; and Zhang, T. 2024 b . Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions. arXiv preprint arXiv:2406.07502
2024 arXiv
-
[76]
Qi, X.; Huang, K.; Panda, A.; Henderson, P.; Wang, M.; and Mittal, P. 2023. Visual Adversarial Examples Jailbreak Aligned Large Language Models. arXiv preprint arXiv:2306.13213
2023 arXiv
-
[77]
D.; and Finn, C
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Advances in Neural Information Processing Systems, 36
2023
-
[78]
Rong, X.; Huang, W.; Wang, T.; Zhou, D.; Du, B.; and Ye, M. 2025. SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization. arXiv preprint arXiv:2511.12982
2025
-
[79]
Rotstein, N.; Bensaid, D.; Brody, S.; Ganz, R.; and Kimmel, R. 2024. FuseCap : Leveraging Large Language Models for Enriched Fused Image Captions. In Workshop on Computer Vision Applications
2024
-
[80]
K.; Wu, Y.; and Guo, D
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y. K.; Wu, Y.; and Guo, D. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300
2024 arXiv
-
[81]
Shayegani, E.; Dong, Y.; and Abu-Ghazaleh, N. 2023. Jailbreak in Pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In International Conference on Learning Representations
2023
-
[82]
Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; and Xie, S. 2024. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. arXiv preprint arXiv:2401.06209
2024 arXiv
-
[83]
Vinyals, O.; Toshev, A.; Bengio, S.; and Erhan, D. 2015. Show and Tell: A Neural Image Caption Generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3156--3164
2015
-
[84]
Wang, P.; Zhang, D.; Li, L.; Tan, C.; Wang, X.; Ren, K.; Jiang, B.; and Qiu, X. 2024 a . InferAligner : Inference-Time Alignment for Harmlessness Through Cross-Model Guidance. arXiv preprint arXiv:2401.11206
2024 arXiv
-
[85]
Wang, Y.; Liu, X.; Li, Y.; Chen, M.; and Xiao, C. 2024 b . AdaShield : Safeguarding Multimodal Large Language Models from Structure-Based Attack via Adaptive Shield Prompting. In European Conference on Computer Vision, 77--94. Springer
2024
-
[86]
Xing, L.; Dong, X.; Zang, Y.; Cao, Y.; Liang, J.; Huang, Q.; Wang, J.; Wu, F.; and Lin, D. 2025. CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning. arXiv preprint arXiv:2509.22647
2025
-
[87]
Ye, M.; Rong, X.; Huang, W.; Du, B.; Yu, N.; and Tao, D. 2025. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations. arXiv preprint arXiv:2502.14881
2025 arXiv
-
[88]
Yu, W.; Yang, Z.; Li, L.; Wang, J.; Lin, K.; Liu, Z.; Wang, X.; and Wang, L. 2023. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities. arXiv preprint arXiv:2308.02490
2023 arXiv
-
[89]
Zhang, Y.; Chen, L.; Zheng, G.; Gao, Y.; Zheng, R.; Fu, J.; Yin, Z.; Jin, S.; Qiao, Y.; Huang, X.; Zhao, F.; Gui, T.; and Shao, J. 2024. SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model. arXiv preprint arXiv:2406.12030
2024 arXiv
-
[90]
Zhang, Y.; Li, J.; Cai, L.; and Li, G. 2025 a . DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt. arXiv preprint arXiv:2506.09353
2025
-
[91]
Zhang, Y.-F.; Yu, T.; Tian, H.; Fu, C.; Li, P.; Zeng, J.; Xie, W.; Shi, Y.; Zhang, H.; Wu, J.; Wang, X.; Hu, Y.; Wen, B.; Yang, F.; Zhang, Z.; Gao, T.; Zhang, D.; Wang, L.; Jin, R.; and Tan, T. 2025 b . MM-RLHF: The Next Step Forward in Multimodal LLM Alignment. arXiv preprint...
2025 arXiv
-
[92]
Zhou, K.; Liu, C.; Zhao, X.; Compalas, A.; Song, D.; and Wang, X. E. 2024. Multimodal Situational Safety. arXiv preprint arXiv:2410.06172
2024 arXiv
-
[93]
Zong, Y.; Bohdal, O.; Yu, T.; Yang, Y.; and Hospedales, T. 2024. Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models. arXiv preprint arXiv:2402.02207
2024 arXiv
-
[94]
Zou, X.; Kang, J.; Kesidis, G.; and Lin, L. 2025. Understanding and Rectifying Safety Perception Distortion in VLMs. arXiv preprint arXiv:2502.13095
2025 arXiv
-
[95]
Zou, X.; Li, K.; and Chen, Y. 2024. Image-to-Text Logic Jailbreak: Your Imagination Can Help You Do Anything. arXiv preprint arXiv:2407.02534
2024 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.