Pith. sign in

Paper Citation Record · LEDGER

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 0 inbound Pith citation observations for arXiv:2608.10513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10513 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:24:49.761179Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6025fc5e-cb13-4b9e-a1f7-bcc2fbfedd4b · outbound

This paper cites Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.286087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.286087Z digest=sha256:50414a563718e9dfd84434841f3b6bc0ca82bef33e303f1982ee92b799d8def6

Observation aafe544d-f17a-4557-a3b7-65bcaa860b07 · outbound

This paper cites Classification Problem Solving.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Classification Problem Solving

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.291857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.291857Z digest=sha256:0252c0a8c136c5b233017d9c03ad3a144aec853e13d9e94aea359e9e1a42c1b8

Observation 03f39d12-37c9-475e-8996-782a28b02c46 · outbound

This paper cites , title =.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning , title =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.297126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.297126Z digest=sha256:9fcba8e1786fe23cd9aa89841badb7ec43756d48e348b4736524669ba29fbc13

Observation 9b684ccd-2da1-4a2e-bd08-f0b86098ad39 · outbound

This paper cites New Ways to Make Microcircuits Smaller---Duplicate Entry.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning New Ways to Make Microcircuits Smaller---Duplicate Entry

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.302190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.302190Z digest=sha256:1c4199cf5e3d81485fd69df9ad1c03952142af767e1d4866f5e31ec117fdbd06

Observation 6e7c5bc3-0353-4ce4-a8bb-05eedb4e6db7 · outbound

This paper cites Clancey and Glenn Rennels , abstract =.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Clancey and Glenn Rennels , abstract =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.307325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.307325Z digest=sha256:d8b8b1b349535a1ce7cb8ffab08b106ed9784087330da7df873baf96cea03374

Observation b370780d-21ed-4d8e-922a-391605fab0d1 · outbound

This paper cites and Rennels, Glenn R.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning and Rennels, Glenn R

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.312726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.312726Z digest=sha256:32deb6cd6d5e17ffebe091601a2d529b16164681e54c976f53bc55e6e27935e2

Observation 7c2213ad-8c12-4d7b-a616-e70bc73431ad · outbound

This paper cites Poligon: A System for Parallel Problem Solving.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Poligon: A System for Parallel Problem Solving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.318979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.318979Z digest=sha256:9cc305f0d75ca3d691eb307eff4d96590d6ca2cef787a896d832ef4b7df38891

Observation 1af58472-120b-4390-a2f6-3e45ba91d558 · outbound

This paper cites Transfer of Rule-Based Expertise through a Tutorial Dialogue.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Transfer of Rule-Based Expertise through a Tutorial Dialogue

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.324337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.324337Z digest=sha256:efe64e0f522907cc28e0333829f574182717739062277c45d22840dac710bfd8

Observation 30e37e8f-9630-4387-a5dc-7562ba936e85 · outbound

This paper cites The Engineering of Qualitative Models.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning The Engineering of Qualitative Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.330017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.330017Z digest=sha256:2d11cc475db3acab22667b8ae38b0a3047500e222423fad66c411c3c419914a1

Observation 8dd50b5f-f6b2-47c0-b120-2f58e7e202f2 · outbound

This paper cites 2023 , eprint=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning 2023 , eprint=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.334945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.334945Z digest=sha256:866a7d5a751d5327a4c5aa21bfa68807d96d29692d452fd90399a5295598a311

Observation 8412b579-2ad8-4563-9fc2-71154bb189a0 · outbound

This paper cites Pluto: The 'Other' Red Planet.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Pluto: The 'Other' Red Planet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.340125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.340125Z digest=sha256:257b972229423f180e47ed38a30de46463d2123da4bd645b4e353c102b191382

Observation 228938b6-472c-4002-bb00-eaa25e47f0d7 · outbound

This paper cites European Conference on Computer Vision , pages=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning European Conference on Computer Vision , pages=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:51.161453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.366719Z digest=sha256:3a2d991e1a9c127d668a8c5889155ac17744acdc573253ec5fa9810b981c6af5

Observation 08e7d26c-4bf3-49e6-8c3b-0f8a7c03056b · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:51.144189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.387989Z digest=sha256:e088f59d25a1f1ac766107a962a5eaeaeb4c5e95127800cdf199068f08d6cc88

Observation 83099ffa-5463-4801-8505-c206be410151 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:51.127311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.392957Z digest=sha256:b682baf6ee51d8ba2caed3f9309349b038fbab2e1d38adca16fd56855311fce0

Observation 3dd945eb-064e-4277-833e-943d8cc322aa · outbound

This paper cites International Conference on Learning Representations , year=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning International Conference on Learning Representations , year=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:51.110729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.419372Z digest=sha256:6756b5ad8a9b4606b575f9d91506b0a9c1123df194093bc7102481743be8ae5e

Observation c0c9e85f-a084-4bbb-96e2-1570013db0a5 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:51.093297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.429522Z digest=sha256:83a78c1408e03e763aab7a8e397c74ad2140ce7256b417070f11fe1eb73a5c57

Observation 4e7208e8-6538-4df6-b419-8ca1cf767783 · outbound

This paper cites 2024 , organization=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning 2024 , organization=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:51.076144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.434452Z digest=sha256:c18980ae9f6bad50973e82669ca83090b35880d6380e9d0a0c22c862e655167f

Observation 3b98b798-fc97-4541-bc10-1eec8dbbd7f8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:51.058351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.440055Z digest=sha256:a36fea0ab7764274ea1c086be1912d64ea40a9b3db4599eaaac9a3754cecf6a8

Observation af1397e5-5deb-4e3a-8ee1-3a02f280d7c2 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:51.038568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.445335Z digest=sha256:03be94044d61332237968ae52a326bb3ef975bcb3750de885363f02e03995ee1

Observation bd1c348e-53c4-4dfd-ac82-5f65f525fdf2 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:51.019929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.455750Z digest=sha256:5e6aa529ebaab04fb2fc68a63e346cd9c5415f0e4bec5e5994afdead94674120

Observation 7c801524-ce3b-49e5-af20-dae48fee7937 · outbound

This paper cites Immune: Improving Safety Against Jailbreaks in Multi-modal.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Immune: Improving Safety Against Jailbreaks in Multi-modal

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:51.003543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.460698Z digest=sha256:f0bdf8a9c2e595721cb6cc61cba4b7daf729de8c09ba71294d8580cbb812cd4c

Observation 217032c7-0fc6-4961-983f-9dfe298829f1 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.985880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.477037Z digest=sha256:374ce1a38d56651269a99597c020eaf08bf2a9981a0ad557c016f416a0bff90c

Observation ec4efce1-7f79-4e07-9aa5-0179b1da6d55 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.968545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.481819Z digest=sha256:e10092fadbf26311a5df83edc813b3a07f13e3f2c60c6f5e9489129bc7bc8efa

Observation a63deb8f-bea4-4fa6-b41b-7cc2d9cdeff8 · outbound

This paper cites Computer Vision -- ECCV 2024 , year=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Computer Vision -- ECCV 2024 , year=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:50.952194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.486867Z digest=sha256:8897ca088a2738e9d3abb51727869174a50df76cfa6223f1679adaaa309e5dca

Observation 953d98a7-bbe0-4ebd-b7b5-21f38ce009f0 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:50.935088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.491449Z digest=sha256:dd672bddeb0b457136a298b80697260db69969cba95d8872f97e9802af45fa4f

Observation 2c142d4a-a0a0-47df-913c-31da69c8f934 · outbound

This paper cites 2025 , eprint=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning 2025 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:50.918592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.496220Z digest=sha256:7f5977961b7f59bca21c4dfd1bbb7188b737bdcaf31556fb2e4998f12345803a

Observation 034ce379-cce3-4d28-9fb0-a351d1e9e384 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.901331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.506246Z digest=sha256:5754056931936c837f95ad668069767c76140cfb6ba53fb03dfd9dc9eb66dc02

Observation 4d562cb7-31a3-43b9-a9dc-e10ec02ef7d5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.516080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.516080Z digest=sha256:e7b05b4e50d0a090fd5a5e6827777f9727e95692eecc983c209da5222eeaf951

Observation 3d8d9528-f3c7-4256-862a-3c5a8ccba9fa · outbound

This paper cites 2025 , eprint=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning 2025 , eprint=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.546118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.546118Z digest=sha256:45e3d5ab69df4209fb67a6d1679218b7ac9310069dd9772344b73bc539d56dc4

Observation 42b41b8a-ddef-45ae-a798-53896366a1fd · outbound

This paper cites 2025 , howpublished=.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning 2025 , howpublished=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:50.863889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.551124Z digest=sha256:56c2c8eb6f5e489ba40a18c1e6dd152f098e9c422c60ac72977ded3747fa137a

Observation b99c6442-d704-4675-9462-6145ec9c58e7 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.848497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.561077Z digest=sha256:69b95a49088a758586e3dcf80413238dc5a3eff132358f130b5663b388c4892c

Observation 7af71618-73f6-4217-a0d6-926fb2b4d874 · outbound

This paper cites S.; Dong, Y.; Roy-Chowdhury, A.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning S.; Dong, Y.; Roy-Chowdhury, A

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.566112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.566112Z digest=sha256:86da29e1a304e784da17a37a1b0866254b0d38557b42bb42876376699de5d71c

Observation fcac46c4-dee3-4c64-9f6c-298a3dc99766 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.833135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.570974Z digest=sha256:ff736a13a6cc8a2241ec0351dba6be910d2b3834291864a56f73575fff886574

Observation ead0cd05-33b7-4a54-9afc-9b2b4d1ec861 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.575768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.575768Z digest=sha256:fb39b1a0496fdba0be26a4894364c2472105f29336359f71cc3013a41ec6f794

Observation 59813561-4a5a-4fbc-88bb-df1b70ebd447 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.580381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.580381Z digest=sha256:860d57f3fb63b834cb957807c5a44d7b78af58bec6825e303b7c7b468a2df084

Observation 22e5aa0d-b540-4990-9f7d-b19fcaec8339 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.585173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.585173Z digest=sha256:dbff8713cc420ca2e78fabe929441e210661a4a701dcdd7f51e1a2a6c6404156

Observation 5948ede7-3092-4529-985d-e8cfa1c54525 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Gemini Robotics: Bringing AI into the Physical World

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.589605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.589605Z digest=sha256:b07f720adfe3ee964f16261f24b833c72025e167e9fc4d858516311985a67d87

Observation 31627a61-ebd4-4a9b-ab3a-ae7869c2a350 · outbound

This paper cites S.; Chakraborty, S.; Singh, V.; Guan, T.; Wang, M.; Velasquez, A.; Beirami, A.; Huang, F.; Manocha, D.; and Bedi, A.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning S.; Chakraborty, S.; Singh, V.; Guan, T.; Wang, M.; Velasquez, A.; Beirami, A.; Huang, F.; Manocha, D.; and Bedi, A

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:50.817152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.594546Z digest=sha256:fe261c9661d4d548a8faf88566e9feb480e6c5927102b31508b516985fa5903b

Observation 28abac41-ddee-48b4-b94f-af4761b95fc6 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.599621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.599621Z digest=sha256:7f576911a38012df1fe642f808a8b04a5bea34f033d4dbafb481948192d7158f

Observation 25918594-f841-43f1-8862-2212b3e99ad6 · outbound

This paper cites T.; and Zhang, Y.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning T.; and Zhang, Y

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:50.800740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.604537Z digest=sha256:8dda54adc980557b1b969f8d46097737a0fbf5f0b212554da187884c0b29fb15

Observation 4dd94b19-98cd-4cea-987c-4013b0862739 · outbound

This paper cites Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:24:50.323213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.609578Z digest=sha256:f5bb192be0edb8e61c84005967f4c513c8c0461a08505863feda5a683f182f21

Observation 5ac197ff-34c5-4570-82fe-84b6764a67ae · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.784835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.614405Z digest=sha256:4142ee04c08409c4fe87427e38e354ef286826102caa77393fa7b4f07e25fe15

Observation 2dc6e744-3892-41f5-815d-f8b9a5e4ff39 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.769285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.619496Z digest=sha256:3ef1191e8cabf77fe2efe712fc20e036622654a2c13ae65680b9cda1261b8686

Observation b9bf0e8c-3aaa-4677-8afe-b2c4d7fd2db9 · outbound

This paper cites How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:24:50.599930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.624109Z digest=sha256:b7e4f2a4043637ed5034bda023912b0cd6433f4df2bbdbb99a9a134603470c5a

Observation b330890a-6d8b-4c62-a945-a091af957280 · outbound

This paper cites Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.628739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.628739Z digest=sha256:f3aad049f6a3e9a389304f729084b4e0d5d6dd3576a2a5e6e327f47ea4d71bb4

Observation 8c6746d9-388b-4d77-992a-e3c437860635 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.633399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.633399Z digest=sha256:396bfa3d59a30506ddb3af47ccdb1ec36a359735b6b616870a475e03bef55513

Observation a1c14e02-0921-4471-83ab-86e644c7fb38 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.753489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.638422Z digest=sha256:3fef2632209f15569931f80914d0b324662d6a059d655b70000836aaec4255eb

Observation 269c60e0-8b68-425e-9851-c51d0534576f · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.736673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.643490Z digest=sha256:6f811ff86b6172b86555f0d84dad3aedfaf0e687b66fb92e36b1e5ee5d79fc3f

Observation 0566dc98-f71a-4d25-95ae-b9bf57f6730c · outbound

This paper cites UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.648328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.648328Z digest=sha256:102537f11d27f87ec59604f20b4a6a3a1c4c4cfe97d2eb99502aa49008b56894

Observation 490d3cb1-abf5-456a-8af5-94c4a05e0634 · outbound

This paper cites Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.653520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.653520Z digest=sha256:ca7ae94642f6d0028baeef058f8c9d0772d3539c312a12bf0d716bb0aabcd877

Observation f1bdd4e5-71c7-40b5-a5a0-d9479f72aa14 · outbound

This paper cites MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.658164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.658164Z digest=sha256:cfd9c35804bbdb0946ff4737a195cb39045b99ae994567318d4622d2f4173801

Observation b468061a-8828-426a-a25a-39e7fb2a1094 · outbound

This paper cites Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.663318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.663318Z digest=sha256:6cb4c804f21ebf6a5b3d629dd592ab3c1fa6076d38d7489067bd1032fb968396

Observation ff2c9902-8178-4ac7-b812-54158f74d295 · outbound

This paper cites Visual Adversarial Examples Jailbreak Aligned Large Language Models.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Visual Adversarial Examples Jailbreak Aligned Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.667991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.667991Z digest=sha256:8478ed86bdbf92b334ef6e0109da78bc9cd699b92dec9703a8ab501796b7cd80

Observation 80bcc7fb-bb9c-4d73-a9a3-fce47d6851a4 · outbound

This paper cites D.; and Finn, C.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning D.; and Finn, C

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:24:50.719088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.672709Z digest=sha256:c8a9f50bd59d5be336d8d0d48dc703fe6f252e8e73286d540a31a56ccf05da5d

Observation a084824b-141e-4c09-b436-0b33272b7c61 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.678351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.678351Z digest=sha256:90bab435841c74c4c8a5e0341b6e04be69e1c734ae6e0609c8cf45b5d77d802f

Observation ee26d16d-fc61-46ae-9fbc-4cce92410aa5 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.702794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.683390Z digest=sha256:d7f2b324a1cfcd0a08dd9dc2b366ded2686f14f43ccbb48a4727ab48636fc534

Observation f619761b-f2a6-471e-88ec-11824f613c4f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.688116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.688116Z digest=sha256:b3cec3f458f6a1c1deb55175c9d4719551ecf37d4d7d7691c198fb9c067f5cef

Observation fc5301ea-d58a-41b1-b04c-c916f780f957 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.685891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.692606Z digest=sha256:b14b909f9e04f26b428a245c091b7d0229deada8616b52ba797b88e1372d0a46

Observation 35a00d1f-c63b-41cb-a266-798754d1fa60 · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.697447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.697447Z digest=sha256:1389c39960dec8795cf391f22c2082fbbb9a6d315532617824937514c293f72e

Observation 319c5e91-143d-4896-964a-00b3ea1e1ccb · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.669292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.702249Z digest=sha256:14243f311789ad577df10150fc2b1b0576f8027aa7e9de1435033a39cc022e99

Observation 60a371c4-1909-450e-ae9c-02d9c15c21ed · outbound

This paper cites InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.707221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.707221Z digest=sha256:ed7e234e2fdcf505deee26b89d5fcca8219b35939eaf91abfbbec573c9a0174b

Observation b252a0b5-1046-4f9f-8096-c4836462f329 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:24:50.652790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:24:49.712315Z digest=sha256:ce71b94fc4c6713dd50e2aace4d446be7387348e0a6b89b2219425332d4f9578

Observation 47cbe7ae-0820-4309-b7a4-a69d2c725db3 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.717212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.717212Z digest=sha256:ca4681f33b74284de47037dee2621a762518dc59eb5bfeccda28ea239d3fa6ad

Observation 2c287b8c-e440-414c-8279-c9fa5299dea6 · outbound

This paper cites A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.722202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.722202Z digest=sha256:b2179e84e8c0504d9a99de309be74887016f1917059a03aa995f209ace912914

Observation d9b1a76d-8748-4a6d-ad38-cde6c6a62f53 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.727045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.727045Z digest=sha256:17f9a57bbe199356d188aeee28561d776643ae610254afa4477b6b1853996261

Observation 0124c92d-8536-4a13-87e3-6d0eaf31fbb5 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.731853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.731853Z digest=sha256:8e7f90412f703cc0a52fff1ad3a505a95b9aeecb3d5fc62946c3dbdfd1915da2

Observation bff65539-c313-49db-a9ae-bb5b9155bc60 · outbound

This paper cites an unresolved cited work.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.736633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.736633Z digest=sha256:94963b7bada93ae0aba80c80aa7e0e9ebfee4fe3ae31609e0b3e5d0d897a5f36

Observation c10d224c-4731-4766-8304-f2d4adcf8da4 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.741420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.741420Z digest=sha256:7c9ce3eff767d646ae9381f19f755d1a976472a533df290213cf9a48a79c3637

Observation 96f4e123-ea8d-4458-b644-8b37f42e79f0 · outbound

This paper cites Multimodal Situational Safety.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Multimodal Situational Safety

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.746172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.746172Z digest=sha256:e41d1a3f6e072fa5af83e21e7db542505169f7fcf1eedbf1669b9fd822678566

Observation c33304f9-7846-4db1-818e-036cdd66476b · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.751283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.751283Z digest=sha256:79f1a2b6a04636c11e3a7326a9b7c8a8716299bb549160beb914199904da50a3

Observation 0d7bae29-0166-43e4-84d7-a3053f47ca88 · outbound

This paper cites Understanding and Rectifying Safety Perception Distortion in VLMs.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Understanding and Rectifying Safety Perception Distortion in VLMs

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.755989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.755989Z digest=sha256:3b3176f60fe572ade7c725a9bb3bfbdd67c131206a756da718a1695c5e7ec978

Observation fcbbe27f-914f-4b57-9fd5-cf34162e1e2e · outbound

This paper cites Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.761179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.761179Z digest=sha256:3fb218d346c9a012480a65aba68ca0b5041eebe0775abb57bf6f2a19958432b2

Pith citing papers

No inbound Pith citation observations are available.