Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:44:56.708066Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2509.00373.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:44:56.708066Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T19:22:43.082083Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T19:23:40.845848Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b0708b9c-3417-4db4-a5a7-af4755228e2e · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3227f213-a1aa-4f19-81e2-84ec332bb600 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abffdf20-d6b4-49db-b979-225b93a0e9c2 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 348f2d69-e9be-4c67-a1e5-92c1f97f128c · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models org/abs/2502.01042
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ea9379-7816-4e1a-ae76-8f8b17bb589c · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a534b28-1ab2-4c06-b50e-4a969a466e5e · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ee2fb0-ff21-425e-be10-c99c8bf9b066 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b940c23-c3cf-4361-8f61-dc2aa2003cf5 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Training language models to follow instructions with human feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a1b601-a548-4dff-be26-deed374c2d07 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models The Linear Representation Hypothesis and the Geometry of Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5973cb-a02f-4be9-a475-36a9752362cd · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Visual Adversarial Examples Jailbreak Aligned Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7401d2b-0a18-4308-9b08-2dd5c8c7f73a · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models On the Adversarial Robustness of Multi-Modal Foundation Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd9d865-c49b-4625-bf38-8cd50f50e119 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Learning to summarize from human feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 646136b6-51c0-4fa4-95bb-153dfd9e1ad3 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models LaMDA: Language Models for Dialog Applications
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c001538-34a8-42f6-9972-7f97dea0a2fb · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Fine-Tuning Language Models from Human Preferences
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5373c59-f4d4-4435-8a6e-e0ef15286754 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fbaf813-d142-41e0-8e51-156b1b7b5221 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Understanding and Rectifying Safety Perception Distortion in VLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fedac94-aaaf-46b5-9e1d-666e7eed10a6 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models semanticscholar.org/CorpusID:125209808
Reference 1952
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7705a784-fd13-4e0a-bbfd-754c2a849208 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11528131-98cc-44e4-9a0f-d85f6e415ec6 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe0083e-acf1-4012-91ef-9c0ad18e8301 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models A General Language Assistant as a Laboratory for Alignment
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997c3dc8-d86b-472d-abb7-23ae05a9a131 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd7a290-1e3c-4ead-8c4d-74ec6625fbee · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc7a34f-f324-447c-82fc-6cdea44c5e2d · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Refusal in Language Models Is Mediated by a Single Direction
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b7143d-a5e2-4c76-abb8-656d7930f615 · outbound
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4736f01-8815-4b13-9e3e-564cff6477db · inbound
ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.