Pith. sign in

Paper Citation Record · LEDGER

Lifelong Safety Alignment for Language Models

As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.20259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20259 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:11.823481Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:10:46.332617Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:19:41.953726Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a49d225-c6b9-4970-adc6-1f49221a6a08 · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

Lifelong Safety Alignment for Language Models Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:02.705923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:02.705923Z digest=sha256:790f6b3ffb63dd4b261679b56d552c395c676effedc5adf12b9e475c43475a93

Observation 6caf749b-2035-4bf8-9153-46c27287f545 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Lifelong Safety Alignment for Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:02.783278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:02.783278Z digest=sha256:9f409de7e0e05e744ab9384ea951c43892f7823f89580fc50123baacd281e61a

Observation adb2ad27-37bc-4f36-bc8a-97bae27c567e · outbound

This paper cites Many-shot jailbreaking.

Lifelong Safety Alignment for Language Models Many-shot jailbreaking

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.679381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:02.976816Z digest=sha256:a867a47645ee26457cab27962671f764532f384b0aa71de8e6c6538987dc8170

Observation f5b061f3-b0bc-4a0a-9446-ce20d2013e71 · outbound

This paper cites Program Synthesis with Large Language Models.

Lifelong Safety Alignment for Language Models Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.163692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.163692Z digest=sha256:9839c5eb0a2f03b892c9fe000c303c244a88659eb4e2e23113ed84c14efd4500

Observation 22bbc4bd-ca39-4bf6-94f1-13330c877a78 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Lifelong Safety Alignment for Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.293613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.293613Z digest=sha256:db58e1f4f66fc414856d02f7e5f43c17c34240a042a823db97d625ed07b97f51

Observation baa57ae7-86c2-4a8e-90df-38e7964957d1 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

Lifelong Safety Alignment for Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.410838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.410838Z digest=sha256:54af67687bb5707024bb6e702e6fe8d4a925105486880f9da405d4ebfd126da0

Observation 927c47d5-5857-48a1-a1ae-90fe22b8daa4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Lifelong Safety Alignment for Language Models Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.557694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.557694Z digest=sha256:c9e77848d756059b0464dbbdcfba47535fada16c489739ffc5dcdc415143142f

Observation 837e0876-3179-45f5-8460-4c80e46fd7fb · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Lifelong Safety Alignment for Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.743897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.743897Z digest=sha256:29bc65f7dc703981627803522a09133ff90a921fbad66d4c0c9dfba11a99cc6b

Observation faedbfb1-52b4-4c8c-ad1d-e666659e3196 · outbound

This paper cites Self-playing adversarial language game enhances llm reasoning.

Lifelong Safety Alignment for Language Models Self-playing adversarial language game enhances llm reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.880607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.880607Z digest=sha256:9afc9e4b94956c6db0b5566ebc4e67ad993c9b77dedff4b2ce9720f00d147e46

Observation fb5e5162-8911-4c6f-9784-d0040b47ef3b · outbound

This paper cites On the Measure of Intelligence.

Lifelong Safety Alignment for Language Models On the Measure of Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.018701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.018701Z digest=sha256:7733014f80916f69f9ec8df70970638986ba25c3cf2636a685e629cbd2b4bba1

Observation d278ce1f-7e5e-4f52-be5a-777b285777f3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lifelong Safety Alignment for Language Models Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.168371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.168371Z digest=sha256:873393fffb68dd5cbd4f2a484bc696334ed5e2eade2d664d4ca76cba851c7ffe

Observation 851ed2b7-d71f-467e-8ecf-5cf79da8a869 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Lifelong Safety Alignment for Language Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.302593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.302593Z digest=sha256:6cd9f0c9543de46f82fc6bd8111a2e851b26d12bbfc6c4fadd1cf812c1ab584c

Observation 33bb40bf-19b0-4bf2-8368-8d379e62a761 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Lifelong Safety Alignment for Language Models RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.449719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.449719Z digest=sha256:06db505ea2b659c9658cbfe52442ce3e99b68fdeffa575aea2b3e46e4c7a6260

Observation 3aacc81d-2e4d-47b9-8d59-5dbf1d4b9286 · outbound

This paper cites Beam Search Strategies for Neural Machine Translation.

Lifelong Safety Alignment for Language Models Beam Search Strategies for Neural Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.618458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.618458Z digest=sha256:908706ce6bc42c4a3d7cd45221042c5ef256a6f6ad7e0e8ae5dc065a366a2059

Observation 43aca306-c3ce-4dae-9f7b-7f25bff4b5b7 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Lifelong Safety Alignment for Language Models A framework for few-shot language model evaluation, 07 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.757264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.757264Z digest=sha256:9274e6b1e3584d613e01a64f2344a09098f0ce827a14d4c8ea3ab9a75b679197

Observation 9298d76f-bc7f-4d4a-aa27-382ceecacf83 · outbound

This paper cites Attacking Large Language Models with Projected Gradient Descent.

Lifelong Safety Alignment for Language Models Attacking Large Language Models with Projected Gradient Descent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.958799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.958799Z digest=sha256:e96687ee6f520dd9eefd58b735967de78f63decb46637c9325b6754735c3171b

Observation 45941abb-072d-4929-8b1a-afb08a0d0b05 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Lifelong Safety Alignment for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.091365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.091365Z digest=sha256:4b7b6b04aabf4ff6a2f51e7110beb056ff738326375af88ca4ea57541bcf085f

Observation bd54c7f8-2255-4319-9439-aa020ce9bed1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lifelong Safety Alignment for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.215286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.215286Z digest=sha256:2623e869397c3fed0d22785985ee9427f2332b7464043d67386e1b6fbba7e932

Observation 9dac24db-23b1-45a1-8b56-0a5a09ed3a42 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Lifelong Safety Alignment for Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.301732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.301732Z digest=sha256:c7029fd9100f3b554f369dd20d303faeabdffcdbd95a280b21bfbb894c1dd738

Observation 4ec77654-9fba-4408-bb0f-a49ac81aabd2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Lifelong Safety Alignment for Language Models Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.483583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.483583Z digest=sha256:9847174fc7892748026d0522fc7b97627bc3e39ddb8d887b99efae60060dbc70

Observation c04b777c-af51-47db-9894-a42ae8072038 · outbound

This paper cites GPT-4o System Card.

Lifelong Safety Alignment for Language Models GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.610293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.610293Z digest=sha256:46dd038b22e8f9bca802c26089899c7b98be8329641ada455b1c2163d00b9444

Observation 386af72b-a74b-4055-ba72-8f2a48cc9a0c · outbound

This paper cites OpenAI o1 System Card.

Lifelong Safety Alignment for Language Models OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.784733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.784733Z digest=sha256:d07008145c679dc755d7595e6cb74372b734255b2623636bc20a0f3858785ad6

Observation a158b6c6-fa85-4b43-9101-14ce2a31ae06 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Lifelong Safety Alignment for Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.888953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.888953Z digest=sha256:7a65e1d5dcc3a8112b9504751f08030b4072fee2721e1fbab7d5d02fd9ab2178

Observation 295d1e45-2c41-4f87-bf24-de54338125a0 · outbound

This paper cites Improved techniques for optimization-based jailbreaking on large language models.

Lifelong Safety Alignment for Language Models Improved techniques for optimization-based jailbreaking on large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.453486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:06.038643Z digest=sha256:9153e82ea8d963313b62850fa1adcb730baf724a479d2810c059be37d4845054

Observation eb8e42e3-35b7-42d9-8c7d-616bbe7fc5ca · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

Lifelong Safety Alignment for Language Models Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.306843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:06.175140Z digest=sha256:6b844bd33e76fc9b99961ed9d0d9a5bab06b9cf88a246a1c3a14c49a8e69443f

Observation ae230cd9-f965-458c-b10e-c53bee44f1f4 · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Lifelong Safety Alignment for Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.273227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.273227Z digest=sha256:e0b7566ecae96456ea9c9963083956233f3640a808ba35aa8041cb85cfe34cef

Observation ea4f5390-cf96-4439-b556-1028e05aec36 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Lifelong Safety Alignment for Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.392442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.392442Z digest=sha256:5911ae0ee92bd68d79d0edf1dfc121829be6cf1ea6cebb710c90d24ddbcd5126

Observation 4c64cd0d-6ab5-43d8-af99-a72ae829e781 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Lifelong Safety Alignment for Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.455968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.455968Z digest=sha256:de2b9ae27c80f7d9ad0e7f8af05b4e7d670423ad0b73375c887a84f86fb0aa95

Observation 8fa47a29-dac1-4fbc-9e73-b4346fcd2907 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Lifelong Safety Alignment for Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.556951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.556951Z digest=sha256:2a7f5a301bf36cc2f5c80a6c94f57c57893c3bd849ded61a686dbd63244f1034

Observation 1f3963e7-0887-4698-a8ef-f282c1479027 · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Lifelong Safety Alignment for Language Models AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.675923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.675923Z digest=sha256:36cee8e2780a0664c402308a908e72efc095616a4ff3f7da24ea1563551ec704

Observation d0d40a5b-e08c-4b30-abcc-3a8f1b8ef07c · outbound

This paper cites Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models.

Lifelong Safety Alignment for Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.780943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.780943Z digest=sha256:2d3fd1cfb0bc30e277f3bb81424357c039b4ba6b86bea109806a5e4ca952e728

Observation f626f28a-d3a3-47a7-b2d7-9f52a7c14ef3 · outbound

This paper cites The Llama 3 Herd of Models.

Lifelong Safety Alignment for Language Models The Llama 3 Herd of Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.866296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.866296Z digest=sha256:fc6369d782d2826ee69b4ebfa6663279aed0abea88f9293d704125cf4a9f3aa3

Observation 55626e72-0c44-4da8-8dad-cfa52f80b2aa · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Lifelong Safety Alignment for Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.928931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.928931Z digest=sha256:6c175ed6e8180d0e15e8878cb3a821a0453a9a3a4fa633fb551ab7e18a95e88c

Observation d7940072-c44b-47d1-b2a1-bcb868333d48 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Lifelong Safety Alignment for Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.052674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.052674Z digest=sha256:6a19f71bd1e0607388b4ca83d737e99e3cd2a966f79f666a961769e38eebcc87

Observation e9dce32f-a95d-4e33-b939-c6ac61df89eb · outbound

This paper cites Introducing ChatGPT, 2022.

Lifelong Safety Alignment for Language Models Introducing ChatGPT, 2022

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.138892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:07.143256Z digest=sha256:83a18f6d0a5b319f8cd4c03172fe8a166b4c3eeaf8083877209bb2cc5a19be2d

Observation b9aaa55f-6c71-4f0e-81f0-4eb10ec73810 · outbound

This paper cites GPT-4 Technical Report.

Lifelong Safety Alignment for Language Models GPT-4 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.202081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.202081Z digest=sha256:b1732e5d77e164eaf69e1e0732c45b45966a0c4a3c60980acc477d9cc6a6e956

Observation 65556b60-fb42-45dc-9a4d-5b4bc75d215f · outbound

This paper cites Red Teaming Language Models with Language Models.

Lifelong Safety Alignment for Language Models Red Teaming Language Models with Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.291288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.291288Z digest=sha256:8da596b1e8343eae0aa45368099f22b7eaa4ecbc9b1674d936700a38c3104f80

Observation 849b3976-422c-4ee0-a798-c06e9eaeedad · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Lifelong Safety Alignment for Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.401279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.401279Z digest=sha256:24993dbaf567527df1ade5e92167081378c32a77485bdace02b7fd04a00494d5

Observation 0a37d858-049f-4c28-ad9d-d3eba83c208b · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

Lifelong Safety Alignment for Language Models CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.477311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.477311Z digest=sha256:6d7b6d82426331de6a007cd68f02ba65f51146d5db6a8bf75a332f914bec477e

Observation 18bbad8a-be74-443d-8f49-58b755d27d26 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Lifelong Safety Alignment for Language Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.567600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.567600Z digest=sha256:ac969d4d27b433f32580bf7355bcf10f26d2a7f47ba9e84f08f7c29712034a07

Observation 82d7f6fd-7a2e-4c7c-91f9-c738a20e6507 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Lifelong Safety Alignment for Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.655270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.655270Z digest=sha256:73408583563285605214d39513177f90a3d85b9d1e8fe8a8c80f0805b7842188

Observation e4c1861b-9875-4628-86e9-e6c8780e54de · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

Lifelong Safety Alignment for Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.771929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.771929Z digest=sha256:fcc3ad8e7a90c4e872174619c3fccfd67e4492b7361077022841f41157f71e21

Observation 0812ed52-1180-45cd-be70-1446100ba121 · outbound

This paper cites do anything now.

Lifelong Safety Alignment for Language Models do anything now

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.878705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.878705Z digest=sha256:e2c7fa1cc7a7dad708e266a2892953c9ec1be26ac021919cd6bb509cae0f693e

Observation b523b869-fa67-42ca-86b2-755399ca5348 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Lifelong Safety Alignment for Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.956726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.956726Z digest=sha256:f1d9e1875fe860dd829bf5d8cfc2c2e6cfa00bec9d22ed95e012c765e913a10a

Observation 2522f766-b0ba-44aa-967a-5bebb3ebef10 · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Lifelong Safety Alignment for Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.034076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.034076Z digest=sha256:e726acca931f4808ffcba92d290d641aa22cd8b508cdf1b720171895d83c2f94

Observation cc0358ee-0918-4988-8548-a3f8f847270c · outbound

This paper cites Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles.

Lifelong Safety Alignment for Language Models Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.128869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.128869Z digest=sha256:42c4b68fd780fc2c94e87f59aaf623a25d1da5b02bf3f3b9cae039116ac14e5a

Observation 4019a0f1-2670-49a5-87d9-04e42281de20 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Lifelong Safety Alignment for Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.229980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.229980Z digest=sha256:526a7b4bcbb163cfdc057ebe2c66d417063b580b913e6adad59eed492537f30c

Observation 724b2904-3844-42b1-a71b-b60c161a7bfa · outbound

This paper cites Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment.

Lifelong Safety Alignment for Language Models Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.309033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.309033Z digest=sha256:1cbc2ab57615f78581690f16b5978d4f2baefad2f097771c664e259c16a6be06

Observation 3b319180-7c6c-4567-8f10-b5ce34044201 · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

Lifelong Safety Alignment for Language Models Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.447260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.447260Z digest=sha256:e510ec4568d7504d927ec51ad1ccb1b5bc07193b39bf30fb5cd6478e6bda053d

Observation 2db408e4-ecf2-4635-ae6b-b74b8edb11a1 · outbound

This paper cites Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping.

Lifelong Safety Alignment for Language Models Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:00:12.238103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:08.558760Z digest=sha256:339b6f458a064ebadbea21089095f96fa08bd50ce71b0adb86afab992434cc63

Observation d9bbd787-d624-4161-8818-5c51ed8c466c · outbound

This paper cites Safety Reasoning with Guidelines.

Lifelong Safety Alignment for Language Models Safety Reasoning with Guidelines

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.671515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.671515Z digest=sha256:59aa4b1f941e6964ae85f70e47845c29f2efb706ce6d1266862e73c41c6cf3af

Observation bed8f58e-038b-49f4-a57c-08a4657d8522 · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.

Lifelong Safety Alignment for Language Models A comprehensive survey of continual learning: Theory, method and application

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.790652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.790652Z digest=sha256:ead46d2cd3f23deb10a7f0949a0adc772ddbc3584b5c8abbd062f2bf004c81da

Observation d2b2cb03-e159-409a-88fa-04c041eb9997 · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Lifelong Safety Alignment for Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.959298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:08.889926Z digest=sha256:4f5aefaf2438070ddd43b84bff6cfc36c3db62a8d85c1c4b1985beea9d21b948

Observation cad982fe-1290-4b4a-915f-1dd9f427cea9 · outbound

This paper cites Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection.

Lifelong Safety Alignment for Language Models Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.944339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.944339Z digest=sha256:980604088c96e0c0ae8b5c77042ced3636196d099c8781cc18593af4a1bd759a

Observation 52886168-f2d8-4c1a-8165-9f296ceb6053 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Lifelong Safety Alignment for Language Models Self-Play Preference Optimization for Language Model Alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.095114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.095114Z digest=sha256:9312ffce1387dd7adf54d6b68c179a1252bd8c3fbe6d11cf230cc0fb53a99aeb

Observation ef8af138-15f7-4858-8ab0-fdd8e83ca8f2 · outbound

This paper cites Qwen2 Technical Report.

Lifelong Safety Alignment for Language Models Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.188727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.188727Z digest=sha256:748c9f441332efdbc55d9364ab1a84814c4fe5fd0beba124ca4b9f0f812e5aed

Observation 7bce9548-c0fb-4385-a5b5-0fbd53fc1461 · outbound

This paper cites Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play.

Lifelong Safety Alignment for Language Models Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.274466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.274466Z digest=sha256:b744dae2c77d5347998e53a41508b3f6a3eea970650dbe22631fe261cae05b33

Observation b333d74e-723b-4e2f-9f79-a3f8e37df4f0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Lifelong Safety Alignment for Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.367525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.367525Z digest=sha256:d3f90dcd9fe3d9d37e3914be878c1558466a2edc3b6d922823194342e5487e94

Observation 00eaeda2-fadc-4cff-8b0c-0c3331b6c350 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Lifelong Safety Alignment for Language Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.433962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.433962Z digest=sha256:2e3d59aaff05db5edbd839c0df92fc3f2340d45fbd58e67f01fd93c25eceb064

Observation e5ba28bf-bc90-443c-b2b9-76d055c25e1f · outbound

This paper cites Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training.

Lifelong Safety Alignment for Language Models Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.569239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.569239Z digest=sha256:6e5315c3cc7ece00b88315999b52794c8db00f26b33fba95cf9c79c2d74bc996

Observation c333ee08-3f33-45eb-b0f9-1188e89bc963 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Lifelong Safety Alignment for Language Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.671636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.671636Z digest=sha256:6fa7c35f3ec7a1fda850ccb1bbd03a975a17eb7480bf6dc3e95d2a683f355b84

Observation 0afd2886-d821-46ee-834c-dc51ce6275be · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Lifelong Safety Alignment for Language Models How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.797411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.797411Z digest=sha256:c0c0ec75815c24bc7f3daddb8650237ea3fc929d5f6d16c7e9f7efb8729f891e

Observation 389ca17b-67fe-4ee1-a84b-4395171c63ce · outbound

This paper cites STAIR: Improving Safety Alignment with Introspective Reasoning.

Lifelong Safety Alignment for Language Models STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.891654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.891654Z digest=sha256:5eac49c6347c54ae14e06c1c704c8e5ea8c05d84cd44ce86a6212a7326018d25

Observation 8057d003-cdd4-4e15-a27b-9562f1df6e40 · outbound

This paper cites Improved few-shot jailbreaking can circumvent aligned language models and their defenses.

Lifelong Safety Alignment for Language Models Improved few-shot jailbreaking can circumvent aligned language models and their defenses

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.738623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:10.035594Z digest=sha256:b007a49404e10b61a272a70b28851efb879fdd27780bb4ef5c242cdf8f50e4ab

Observation 7e6c0211-6d70-434c-9dfa-92cd99d84f0b · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Lifelong Safety Alignment for Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.147312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.147312Z digest=sha256:819b054fe9d762caaebd4caab7ab5597e09777c368d629925897f3b448b0efca

Observation 098faab8-0d50-4915-a5cc-7e550b6fdf5e · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Lifelong Safety Alignment for Language Models Instruction-Following Evaluation for Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.249221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.249221Z digest=sha256:ca4329a15a5ff1d5498771b9f27130fb99e62f519fb1ce9c1f46cbe82afb8ca9

Observation 0819da74-9a13-4941-8668-53d027ab7ba8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Lifelong Safety Alignment for Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.406991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.406991Z digest=sha256:e7c95cc4d7b47aeccd9a2506d80b2c61a1af421ce0bb1d1630e3843af4f42419

Observation e5bd4d64-8957-49a5-af76-f3b812418876 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Lifelong Safety Alignment for Language Models Improving Alignment and Robustness with Circuit Breakers

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.571592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.571592Z digest=sha256:dfaaac8c727b373cc500d6aca7ecce34b3c3e474c77f92bd25998a7363955231

Observation a8712d85-ecbd-45d8-8302-e6d7e2516aeb · outbound

This paper cites question.

Lifelong Safety Alignment for Language Models question

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.536150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:10.668254Z digest=sha256:8e5908bd07397b5b40ba83b308a85147b24577315278edf39cb546a729db627d

Observation 8c2e9e7d-283d-45f3-b655-4cfc11c2e46e · outbound

This paper cites This is the most preferred response when confronted with policy-violating instructions.

Lifelong Safety Alignment for Language Models This is the most preferred response when confronted with policy-violating instructions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.309703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:10.800784Z digest=sha256:93b8b34424c69be1566ec644f069e711e8c13122756521753859231af6574a35

Observation 557685eb-0ec5-4e2c-a9b3-eb590da57da4 · outbound

This paper cites It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content.

Lifelong Safety Alignment for Language Models It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.888539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.888539Z digest=sha256:4ea2f4dc6c43e068184748046256b809e0192e3095ca5d9bd2397bdca5da251f

Observation 00e3612b-5336-48b6-bc82-e9ee5537b7a0 · outbound

This paper cites It offers content related to the request but without embedding necessary precautions or disclaimers.

Lifelong Safety Alignment for Language Models It offers content related to the request but without embedding necessary precautions or disclaimers

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.967249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.967249Z digest=sha256:dd57876920e623e5bd89117aa56c07895dffd46cfd828dd52c25d4ad83ee2681

Observation 1da7b12d-24b3-4197-9137-27c4268732d4 · outbound

This paper cites However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5.

Lifelong Safety Alignment for Language Models However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:11.062426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:11.062426Z digest=sha256:671c68a7eb913e0fd56c9aad80ddbcf3c0d45cc352939250f18212a6e76c82a5

Observation 2d2d1d37-669e-41c8-90fa-bbff023cecd9 · outbound

This paper cites It completely ignores the usage guidelines and provides information that clearly violates the Meta policy.

Lifelong Safety Alignment for Language Models It completely ignores the usage guidelines and provides information that clearly violates the Meta policy

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.070089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:11.144360Z digest=sha256:4111cb6a8aaff6ed9cfd8f8c6bd6ebc3169b03f1188e7cf1a0aa56dbe693cc76

Observation 7cfad83f-1844-4e05-bd39-545c05c87055 · outbound

This paper cites Frame the prompt as a discussion or analysis question to invite exploration.3.

Lifelong Safety Alignment for Language Models Frame the prompt as a discussion or analysis question to invite exploration.3

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.887723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:11.284481Z digest=sha256:bb363a69cb899f9fb6ce0fdb45ba3b6964bba8c2fbffea5c913df08e3066fcda

Observation 5a69f598-9492-4f2f-9fae-7ff0a3fae857 · outbound

This paper cites an unresolved cited work.

Lifelong Safety Alignment for Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:13.690433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:11.368910Z digest=sha256:ebf14854ab6c0062b166667d3a59d91be8a40e6a79c4c361a82154721cda0af1

Observation 31c08b3f-f936-4835-a655-dd0b0814f96d · outbound

This paper cites 30th St, Los Angeles, CA 90007, United States.

Lifelong Safety Alignment for Language Models 30th St, Los Angeles, CA 90007, United States

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.472286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:11.477728Z digest=sha256:21ebb695d71881f5377fc812072d9735e52143f9a72dc6ffdd104db7b09dc199

Observation 40ae3064-2211-4898-a35e-6815b8005f54 · outbound

This paper cites 20th St, New York, NY 10011, United States.

Lifelong Safety Alignment for Language Models 20th St, New York, NY 10011, United States

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.291874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:11.616468Z digest=sha256:378808a5b2f209a2efad1f695abdb8c97c0c5fd68d8e04a35c9c165c15a91b94

Observation 1c729329-57f8-4f55-9021-d2181740f2b0 · outbound

This paper cites an unresolved cited work.

Lifelong Safety Alignment for Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:13.095931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:11.708628Z digest=sha256:64c93d20f2794a5f4900b6313bdc76ed360a3e3218dc33f2c8e0a40fdd83312e

Observation 2384860d-b0e0-4adb-a710-7757d15df4ec · outbound

This paper cites pythonchemicals =.

Lifelong Safety Alignment for Language Models pythonchemicals =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:12.905731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:00:11.823481Z digest=sha256:0cdd9047017732ce8ccdb27fb552e2ba0864169b4fe75edc8a63223e8ebb0dbe

Pith citing papers

Observation 8ff0fd93-172d-4e8f-80ff-a4cd1ad54c71 · inbound

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World cites this paper.

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World Lifelong Safety Alignment for Language Models

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:46.332617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:46.332617Z digest=sha256:468ee27b5ec1a1a8306cfe85212d4272d52b194438c81c0b3eadb8974cd8a357

Observation fd4f8b5c-061e-4393-838b-cdfb51d8fe9d · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Lifelong Safety Alignment for Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.461615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:3b4ac49ec47d65760f715033bb2eb7b125c127601297f76e283e7d224cd0776f

Observation 19d35091-37dc-4c53-8a57-c613eefcd074 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Lifelong Safety Alignment for Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:41.955346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:f64abbabc8567d3660b1cf456790f7a4c8562b80dbf0a0c2dd964cbb889579a4