Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:11.823481Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.20259.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:11.823481Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:10:46.332617Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T06:19:41.953726Z
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1a49d225-c6b9-4970-adc6-1f49221a6a08 · outbound
Lifelong Safety Alignment for Language Models Does Refusal Training in LLMs Generalize to the Past Tense?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6caf749b-2035-4bf8-9153-46c27287f545 · outbound
Lifelong Safety Alignment for Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb2ad27-37bc-4f36-bc8a-97bae27c567e · outbound
Lifelong Safety Alignment for Language Models Many-shot jailbreaking
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5b061f3-b0bc-4a0a-9446-ce20d2013e71 · outbound
Lifelong Safety Alignment for Language Models Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22bbc4bd-ca39-4bf6-94f1-13330c877a78 · outbound
Lifelong Safety Alignment for Language Models Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa57ae7-86c2-4a8e-90df-38e7964957d1 · outbound
Lifelong Safety Alignment for Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927c47d5-5857-48a1-a1ae-90fe22b8daa4 · outbound
Lifelong Safety Alignment for Language Models Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 837e0876-3179-45f5-8460-4c80e46fd7fb · outbound
Lifelong Safety Alignment for Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faedbfb1-52b4-4c8c-ad1d-e666659e3196 · outbound
Lifelong Safety Alignment for Language Models Self-playing adversarial language game enhances llm reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5e5162-8911-4c6f-9784-d0040b47ef3b · outbound
Lifelong Safety Alignment for Language Models On the Measure of Intelligence
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d278ce1f-7e5e-4f52-be5a-777b285777f3 · outbound
Lifelong Safety Alignment for Language Models Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 851ed2b7-d71f-467e-8ecf-5cf79da8a869 · outbound
Lifelong Safety Alignment for Language Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33bb40bf-19b0-4bf2-8368-8d379e62a761 · outbound
Lifelong Safety Alignment for Language Models RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aacc81d-2e4d-47b9-8d59-5dbf1d4b9286 · outbound
Lifelong Safety Alignment for Language Models Beam Search Strategies for Neural Machine Translation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43aca306-c3ce-4dae-9f7b-7f25bff4b5b7 · outbound
Lifelong Safety Alignment for Language Models A framework for few-shot language model evaluation, 07 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9298d76f-bc7f-4d4a-aa27-382ceecacf83 · outbound
Lifelong Safety Alignment for Language Models Attacking Large Language Models with Projected Gradient Descent
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45941abb-072d-4929-8b1a-afb08a0d0b05 · outbound
Lifelong Safety Alignment for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd54c7f8-2255-4319-9439-aa020ce9bed1 · outbound
Lifelong Safety Alignment for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dac24db-23b1-45a1-8b56-0a5a09ed3a42 · outbound
Lifelong Safety Alignment for Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec77654-9fba-4408-bb0f-a49ac81aabd2 · outbound
Lifelong Safety Alignment for Language Models Measuring Massive Multitask Language Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04b777c-af51-47db-9894-a42ae8072038 · outbound
Lifelong Safety Alignment for Language Models GPT-4o System Card
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 386af72b-a74b-4055-ba72-8f2a48cc9a0c · outbound
Lifelong Safety Alignment for Language Models OpenAI o1 System Card
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a158b6c6-fa85-4b43-9101-14ce2a31ae06 · outbound
Lifelong Safety Alignment for Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295d1e45-2c41-4f87-bf24-de54338125a0 · outbound
Lifelong Safety Alignment for Language Models Improved techniques for optimization-based jailbreaking on large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb8e42e3-35b7-42d9-8c7d-616bbe7fc5ca · outbound
Lifelong Safety Alignment for Language Models Artprompt: Ascii art-based jailbreak attacks against aligned llms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae230cd9-f965-458c-b10e-c53bee44f1f4 · outbound
Lifelong Safety Alignment for Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4f5390-cf96-4439-b556-1028e05aec36 · outbound
Lifelong Safety Alignment for Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c64cd0d-6ab5-43d8-af99-a72ae829e781 · outbound
Lifelong Safety Alignment for Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa47a29-dac1-4fbc-9e73-b4346fcd2907 · outbound
Lifelong Safety Alignment for Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3963e7-0887-4698-a8ef-f282c1479027 · outbound
Lifelong Safety Alignment for Language Models AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d40a5b-e08c-4b30-abcc-3a8f1b8ef07c · outbound
Lifelong Safety Alignment for Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f626f28a-d3a3-47a7-b2d7-9f52a7c14ef3 · outbound
Lifelong Safety Alignment for Language Models The Llama 3 Herd of Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55626e72-0c44-4da8-8dad-cfa52f80b2aa · outbound
Lifelong Safety Alignment for Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7940072-c44b-47d1-b2a1-bcb868333d48 · outbound
Lifelong Safety Alignment for Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9dce32f-a95d-4e33-b939-c6ac61df89eb · outbound
Lifelong Safety Alignment for Language Models Introducing ChatGPT, 2022
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9aaa55f-6c71-4f0e-81f0-4eb10ec73810 · outbound
Lifelong Safety Alignment for Language Models GPT-4 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65556b60-fb42-45dc-9a4d-5b4bc75d215f · outbound
Lifelong Safety Alignment for Language Models Red Teaming Language Models with Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849b3976-422c-4ee0-a798-c06e9eaeedad · outbound
Lifelong Safety Alignment for Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a37d858-049f-4c28-ad9d-d3eba83c208b · outbound
Lifelong Safety Alignment for Language Models CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bbad8a-be74-443d-8f49-58b755d27d26 · outbound
Lifelong Safety Alignment for Language Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d7f6fd-7a2e-4c7c-91f9-c738a20e6507 · outbound
Lifelong Safety Alignment for Language Models Winogrande: An adversarial winograd schema challenge at scale
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c1861b-9875-4628-86e9-e6c8780e54de · outbound
Lifelong Safety Alignment for Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0812ed52-1180-45cd-be70-1446100ba121 · outbound
Lifelong Safety Alignment for Language Models do anything now
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b523b869-fa67-42ca-86b2-755399ca5348 · outbound
Lifelong Safety Alignment for Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2522f766-b0ba-44aa-967a-5bebb3ebef10 · outbound
Lifelong Safety Alignment for Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc0358ee-0918-4988-8548-a3f8f847270c · outbound
Lifelong Safety Alignment for Language Models Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4019a0f1-2670-49a5-87d9-04e42281de20 · outbound
Lifelong Safety Alignment for Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724b2904-3844-42b1-a71b-b60c161a7bfa · outbound
Lifelong Safety Alignment for Language Models Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b319180-7c6c-4567-8f10-b5ce34044201 · outbound
Lifelong Safety Alignment for Language Models Universal Adversarial Triggers for Attacking and Analyzing NLP
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db408e4-ecf2-4635-ae6b-b74b8edb11a1 · outbound
Lifelong Safety Alignment for Language Models Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9bbd787-d624-4161-8818-5c51ed8c466c · outbound
Lifelong Safety Alignment for Language Models Safety Reasoning with Guidelines
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed8f58e-038b-49f4-a57c-08a4657d8522 · outbound
Lifelong Safety Alignment for Language Models A comprehensive survey of continual learning: Theory, method and application
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b2cb03-e159-409a-88fa-04c041eb9997 · outbound
Lifelong Safety Alignment for Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cad982fe-1290-4b4a-915f-1dd9f427cea9 · outbound
Lifelong Safety Alignment for Language Models Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52886168-f2d8-4c1a-8165-9f296ceb6053 · outbound
Lifelong Safety Alignment for Language Models Self-Play Preference Optimization for Language Model Alignment
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8af138-15f7-4858-8ab0-fdd8e83ca8f2 · outbound
Lifelong Safety Alignment for Language Models Qwen2 Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bce9548-c0fb-4385-a5b5-0fbd53fc1461 · outbound
Lifelong Safety Alignment for Language Models Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b333d74e-723b-4e2f-9f79-a3f8e37df4f0 · outbound
Lifelong Safety Alignment for Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00eaeda2-fadc-4cff-8b0c-0c3331b6c350 · outbound
Lifelong Safety Alignment for Language Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ba28bf-bc90-443c-b2b9-76d055c25e1f · outbound
Lifelong Safety Alignment for Language Models Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c333ee08-3f33-45eb-b0f9-1188e89bc963 · outbound
Lifelong Safety Alignment for Language Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0afd2886-d821-46ee-834c-dc51ce6275be · outbound
Lifelong Safety Alignment for Language Models How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389ca17b-67fe-4ee1-a84b-4395171c63ce · outbound
Lifelong Safety Alignment for Language Models STAIR: Improving Safety Alignment with Introspective Reasoning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8057d003-cdd4-4e15-a27b-9562f1df6e40 · outbound
Lifelong Safety Alignment for Language Models Improved few-shot jailbreaking can circumvent aligned language models and their defenses
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e6c0211-6d70-434c-9dfa-92cd99d84f0b · outbound
Lifelong Safety Alignment for Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098faab8-0d50-4915-a5cc-7e550b6fdf5e · outbound
Lifelong Safety Alignment for Language Models Instruction-Following Evaluation for Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0819da74-9a13-4941-8668-53d027ab7ba8 · outbound
Lifelong Safety Alignment for Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5bd4d64-8957-49a5-af76-f3b812418876 · outbound
Lifelong Safety Alignment for Language Models Improving Alignment and Robustness with Circuit Breakers
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8712d85-ecbd-45d8-8302-e6d7e2516aeb · outbound
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c2e9e7d-283d-45f3-b655-4cfc11c2e46e · outbound
Lifelong Safety Alignment for Language Models This is the most preferred response when confronted with policy-violating instructions
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 557685eb-0ec5-4e2c-a9b3-eb590da57da4 · outbound
Lifelong Safety Alignment for Language Models It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e3612b-5336-48b6-bc82-e9ee5537b7a0 · outbound
Lifelong Safety Alignment for Language Models It offers content related to the request but without embedding necessary precautions or disclaimers
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da7b12d-24b3-4197-9137-27c4268732d4 · outbound
Lifelong Safety Alignment for Language Models However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d2d1d37-669e-41c8-90fa-bbff023cecd9 · outbound
Lifelong Safety Alignment for Language Models It completely ignores the usage guidelines and provides information that clearly violates the Meta policy
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cfad83f-1844-4e05-bd39-545c05c87055 · outbound
Lifelong Safety Alignment for Language Models Frame the prompt as a discussion or analysis question to invite exploration.3
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a69f598-9492-4f2f-9fae-7ff0a3fae857 · outbound
Lifelong Safety Alignment for Language Models Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31c08b3f-f936-4835-a655-dd0b0814f96d · outbound
Lifelong Safety Alignment for Language Models 30th St, Los Angeles, CA 90007, United States
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40ae3064-2211-4898-a35e-6815b8005f54 · outbound
Lifelong Safety Alignment for Language Models 20th St, New York, NY 10011, United States
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c729329-57f8-4f55-9021-d2181740f2b0 · outbound
Lifelong Safety Alignment for Language Models Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2384860d-b0e0-4adb-a710-7757d15df4ec · outbound
Lifelong Safety Alignment for Language Models pythonchemicals =
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ff0fd93-172d-4e8f-80ff-a4cd1ad54c71 · inbound
\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World Lifelong Safety Alignment for Language Models
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd4f8b5c-061e-4393-838b-cdfb51d8fe9d · inbound
Characterizing Model-Native Skills Lifelong Safety Alignment for Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19d35091-37dc-4c53-8a57-c613eefcd074 · inbound
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Lifelong Safety Alignment for Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.