Pith. sign in

Paper Citation Record · LEDGER

Lifelong Safety Alignment for Language Models

As of 10 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 4 inbound Pith citation observations for arXiv:2505.20259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20259 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:11.823481Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:50:36.178215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:19:41.953726Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a49d225-c6b9-4970-adc6-1f49221a6a08 · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

Lifelong Safety Alignment for Language Models Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:02.705923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:02.705923Z digest=sha256:9c6b61109ea13f0689ad615b0f82d3af50b8efce877fee79a7db89a45dfdfce7

Observation 6caf749b-2035-4bf8-9153-46c27287f545 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Lifelong Safety Alignment for Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:02.783278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:02.783278Z digest=sha256:85ea3ba1fedd501ec909e487480080552a02e737633afb7ca5a2505cdc73c07d

Observation adb2ad27-37bc-4f36-bc8a-97bae27c567e · outbound

This paper cites Many-shot jailbreaking.

Lifelong Safety Alignment for Language Models Many-shot jailbreaking

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.679381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:02.976816Z digest=sha256:505e104c36ee4d70c11d5ac46c7037896afbfbb690e8eb94dbdfdc4e47bd770c

Observation f5b061f3-b0bc-4a0a-9446-ce20d2013e71 · outbound

This paper cites Program Synthesis with Large Language Models.

Lifelong Safety Alignment for Language Models Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.163692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.163692Z digest=sha256:902ceb8abfc6b83b66b1f900ebcdfa53f989bee58f67565b47486061d576b003

Observation 22bbc4bd-ca39-4bf6-94f1-13330c877a78 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Lifelong Safety Alignment for Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.293613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.293613Z digest=sha256:7c20d12d1de684f2af61d543a682f8e7a8e49afd9b9c48d9cfd725d64a604c44

Observation baa57ae7-86c2-4a8e-90df-38e7964957d1 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

Lifelong Safety Alignment for Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.410838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.410838Z digest=sha256:014787eb6564f30ceff4a25d84886e23e4e4ed51564be80b43de6dee240859f4

Observation 927c47d5-5857-48a1-a1ae-90fe22b8daa4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Lifelong Safety Alignment for Language Models Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.557694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.557694Z digest=sha256:6437925b08d0ccd590a4ce426e5d27b8263a675eb128f9d0d9c10ac337842b3f

Observation 837e0876-3179-45f5-8460-4c80e46fd7fb · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Lifelong Safety Alignment for Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.743897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.743897Z digest=sha256:733506a4f3edfdf257efb53f7441341d3d8381e596336561c673b2ee2d7023a7

Observation faedbfb1-52b4-4c8c-ad1d-e666659e3196 · outbound

This paper cites Self-playing adversarial language game enhances llm reasoning.

Lifelong Safety Alignment for Language Models Self-playing adversarial language game enhances llm reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.880607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.880607Z digest=sha256:32d107780c72c11c1495230cf3095432a5db799d2990c7941a1dab1b0756c1ca

Observation fb5e5162-8911-4c6f-9784-d0040b47ef3b · outbound

This paper cites On the Measure of Intelligence.

Lifelong Safety Alignment for Language Models On the Measure of Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.018701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.018701Z digest=sha256:463458764f0343a42e4a67df9491352f189ad360c72db3e3b273d4a86e15e461

Observation d278ce1f-7e5e-4f52-be5a-777b285777f3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lifelong Safety Alignment for Language Models Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.168371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.168371Z digest=sha256:b7bcb6d2b4a2e653be519418584c1b931660380e4679af4c6cad3e4f61e41416

Observation 851ed2b7-d71f-467e-8ecf-5cf79da8a869 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Lifelong Safety Alignment for Language Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.302593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.302593Z digest=sha256:8eac9de6d8b32e847bb5b495ddabcc25ce4950a37ff48ec5c620424ff6d0eb69

Observation 33bb40bf-19b0-4bf2-8368-8d379e62a761 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Lifelong Safety Alignment for Language Models RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.449719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.449719Z digest=sha256:b3c802c0e51131dd5c64f9b411976e556c7930287a145ff4751f5a9aa8f9f61e

Observation 3aacc81d-2e4d-47b9-8d59-5dbf1d4b9286 · outbound

This paper cites Beam Search Strategies for Neural Machine Translation.

Lifelong Safety Alignment for Language Models Beam Search Strategies for Neural Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.618458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.618458Z digest=sha256:7cd052a032baba867ed9b7b66fe2115b953234c0f5332ae9353e851d7448e660

Observation 43aca306-c3ce-4dae-9f7b-7f25bff4b5b7 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Lifelong Safety Alignment for Language Models A framework for few-shot language model evaluation, 07 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.757264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.757264Z digest=sha256:72a58571e1652dc379f19c5f3619334d7060eb9a2a1599f8cd67f4c735392472

Observation 9298d76f-bc7f-4d4a-aa27-382ceecacf83 · outbound

This paper cites Attacking Large Language Models with Projected Gradient Descent.

Lifelong Safety Alignment for Language Models Attacking Large Language Models with Projected Gradient Descent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.958799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.958799Z digest=sha256:aac3c861e692d34c73f71f2ce54090e098a719f17aa87d7946369aff888f4cb8

Observation 45941abb-072d-4929-8b1a-afb08a0d0b05 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Lifelong Safety Alignment for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.091365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.091365Z digest=sha256:f6ffc0f07e1afc393ffc3391acba6cadcc0d31f8da798bb4d77ca5ec97db9881

Observation bd54c7f8-2255-4319-9439-aa020ce9bed1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lifelong Safety Alignment for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.215286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.215286Z digest=sha256:016df8b6933d75f955f61a95a0be1b2b9794876a21e24061b44ca798f228eff3

Observation 9dac24db-23b1-45a1-8b56-0a5a09ed3a42 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Lifelong Safety Alignment for Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.301732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.301732Z digest=sha256:5d15e412bc912dd5c13650260b32b19461909526c0dc5cfcde765603a56146f4

Observation 4ec77654-9fba-4408-bb0f-a49ac81aabd2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Lifelong Safety Alignment for Language Models Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.483583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.483583Z digest=sha256:68b00437f1208f5c69208bd5529abef782b9d2fccbb8b2e217035b06867ef51e

Observation c04b777c-af51-47db-9894-a42ae8072038 · outbound

This paper cites GPT-4o System Card.

Lifelong Safety Alignment for Language Models GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.610293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.610293Z digest=sha256:6429a870491bed9d5cfbb670431b2c519c5f7bfc383ca10c54c80b58b2831b0a

Observation 386af72b-a74b-4055-ba72-8f2a48cc9a0c · outbound

This paper cites OpenAI o1 System Card.

Lifelong Safety Alignment for Language Models OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.784733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.784733Z digest=sha256:581c13ddc63e384e9a285f15fe855013040d6c2ced192b12911fdcd098d4726c

Observation a158b6c6-fa85-4b43-9101-14ce2a31ae06 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Lifelong Safety Alignment for Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.888953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.888953Z digest=sha256:688d7fc4f48417cf6b3b16d2fa7bd90798e6f86b0ba0676cdfc7088197f758f4

Observation 295d1e45-2c41-4f87-bf24-de54338125a0 · outbound

This paper cites Improved techniques for optimization-based jailbreaking on large language models.

Lifelong Safety Alignment for Language Models Improved techniques for optimization-based jailbreaking on large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.453486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:06.038643Z digest=sha256:aa766d27249ddcf05efdc0ef3a35664b36d8ed04dbe57382995848b2bcf44574

Observation eb8e42e3-35b7-42d9-8c7d-616bbe7fc5ca · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

Lifelong Safety Alignment for Language Models Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.306843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:06.175140Z digest=sha256:b68303bef542ebb3b1793fa86838045a636bc0c77ce094f5894ac6613e9e5bd8

Observation ae230cd9-f965-458c-b10e-c53bee44f1f4 · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Lifelong Safety Alignment for Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.273227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.273227Z digest=sha256:19431e17793618904d06ced072c2b7b44f4ccd4a89a38a47c811e3aca09a919a

Observation ea4f5390-cf96-4439-b556-1028e05aec36 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Lifelong Safety Alignment for Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.392442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.392442Z digest=sha256:cda63add5b752afddf41cbe8433a6b5a2ef6446ecf1399f428e70d96a665539f

Observation 4c64cd0d-6ab5-43d8-af99-a72ae829e781 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Lifelong Safety Alignment for Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.455968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.455968Z digest=sha256:0d4b9e8591b59610d582f45e61f8f5fe68953ba6677457a9f693a3cccf70da34

Observation 8fa47a29-dac1-4fbc-9e73-b4346fcd2907 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Lifelong Safety Alignment for Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.556951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.556951Z digest=sha256:00d2896db24cd0c14de93aa57496f66cb1ab8ec6a5c840963a36900f79cdbc5b

Observation 1f3963e7-0887-4698-a8ef-f282c1479027 · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Lifelong Safety Alignment for Language Models AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.675923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.675923Z digest=sha256:45b7ac0c5fbe1d7e9353ada61ec83c487f37357923d13f3ca87c8c943d17a8d1

Observation d0d40a5b-e08c-4b30-abcc-3a8f1b8ef07c · outbound

This paper cites Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models.

Lifelong Safety Alignment for Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.780943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.780943Z digest=sha256:ae7abf1469994d8d40a5a8cbe1c152245f4230d2e1fcb25fdcbab67af2f207fb

Observation f626f28a-d3a3-47a7-b2d7-9f52a7c14ef3 · outbound

This paper cites The Llama 3 Herd of Models.

Lifelong Safety Alignment for Language Models The Llama 3 Herd of Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.866296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.866296Z digest=sha256:3f559829bd1ea96a1b09f9d7eaf1419b395c82faf7eccbe5e400fe4df84871e0

Observation 55626e72-0c44-4da8-8dad-cfa52f80b2aa · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Lifelong Safety Alignment for Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.928931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.928931Z digest=sha256:d522ae7f71004d2d27cca194829ff3cac36351039baf70d63fcef125a059071d

Observation d7940072-c44b-47d1-b2a1-bcb868333d48 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Lifelong Safety Alignment for Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.052674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.052674Z digest=sha256:2958c85c7e22cff07ef3d21ae1c8f6a3778ba8eb7847f49ae6a91e31fde50838

Observation e9dce32f-a95d-4e33-b939-c6ac61df89eb · outbound

This paper cites Introducing ChatGPT, 2022.

Lifelong Safety Alignment for Language Models Introducing ChatGPT, 2022

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.138892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:07.143256Z digest=sha256:1f9686744aedf495f9cd08e4486b262db499881aba30574f731ed6a7a4429c32

Observation b9aaa55f-6c71-4f0e-81f0-4eb10ec73810 · outbound

This paper cites GPT-4 Technical Report.

Lifelong Safety Alignment for Language Models GPT-4 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.202081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.202081Z digest=sha256:ce16769fdeff591e8012d9ed7da4c96a21f68b3d4f34f386e98ef5233eaf6d85

Observation 65556b60-fb42-45dc-9a4d-5b4bc75d215f · outbound

This paper cites Red Teaming Language Models with Language Models.

Lifelong Safety Alignment for Language Models Red Teaming Language Models with Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.291288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.291288Z digest=sha256:87d0f7cd1a2bad8d80dfc364b631c82de4cbc69522cfc641884b84f415b897fa

Observation 849b3976-422c-4ee0-a798-c06e9eaeedad · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Lifelong Safety Alignment for Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.401279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.401279Z digest=sha256:109218d36528339323edd5721358df67257276c1498c20801c87f2f78a5eac2b

Observation 0a37d858-049f-4c28-ad9d-d3eba83c208b · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

Lifelong Safety Alignment for Language Models CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.477311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.477311Z digest=sha256:c38ee8ca049eabd61b6f26538658956c9bbf31560d28222c7866ed603dbeab73

Observation 18bbad8a-be74-443d-8f49-58b755d27d26 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Lifelong Safety Alignment for Language Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.567600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.567600Z digest=sha256:4fe396513dfbc1bfc374d3cebc62ead5fff2d4a5107ad66cbc93aa27702e1506

Observation 82d7f6fd-7a2e-4c7c-91f9-c738a20e6507 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Lifelong Safety Alignment for Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.655270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.655270Z digest=sha256:b553b8683d5e8d1242a43b0f0da233c3cbc5533e3586f3bdf724bc01e3b6fe78

Observation e4c1861b-9875-4628-86e9-e6c8780e54de · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

Lifelong Safety Alignment for Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.771929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.771929Z digest=sha256:8c430021cef8b65fc3bba00ed53e0eae0e7e77a75ea9204f29d392b16025cb33

Observation 0812ed52-1180-45cd-be70-1446100ba121 · outbound

This paper cites do anything now.

Lifelong Safety Alignment for Language Models do anything now

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.878705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.878705Z digest=sha256:9e32446eb80d46f039c025af08affcb74eb3cc30d08158e7a201cc72fe2a9c37

Observation b523b869-fa67-42ca-86b2-755399ca5348 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Lifelong Safety Alignment for Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.956726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.956726Z digest=sha256:ac3de8b208a23f3f05daea9604a35eb2bdc6186d978727a0c3233a5fa5373d64

Observation 2522f766-b0ba-44aa-967a-5bebb3ebef10 · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Lifelong Safety Alignment for Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.034076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.034076Z digest=sha256:c66c18ed3422b87eb40a3a45891382a4b5503d4381cb9daeee6637b085a8cdd1

Observation cc0358ee-0918-4988-8548-a3f8f847270c · outbound

This paper cites Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles.

Lifelong Safety Alignment for Language Models Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.128869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.128869Z digest=sha256:aba7bc350e5319a074e62bac364ea961f78df6f9fdc63b41a2c84abc919a097d

Observation 4019a0f1-2670-49a5-87d9-04e42281de20 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Lifelong Safety Alignment for Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.229980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.229980Z digest=sha256:06d8363a9c1fc32df1dc8c98610d991118333300836523f5797528192d369cc4

Observation 724b2904-3844-42b1-a71b-b60c161a7bfa · outbound

This paper cites Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment.

Lifelong Safety Alignment for Language Models Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.309033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.309033Z digest=sha256:a0fd741d199f6cb1addfe381794e09021aab8b800024508602d5e42b0571f866

Observation 3b319180-7c6c-4567-8f10-b5ce34044201 · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

Lifelong Safety Alignment for Language Models Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.447260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.447260Z digest=sha256:860c075500570673f38cf267a399872a3f03f7f65b58c7516c4b5630fa8d3519

Observation 2db408e4-ecf2-4635-ae6b-b74b8edb11a1 · outbound

This paper cites Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping.

Lifelong Safety Alignment for Language Models Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:00:12.238103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:08.558760Z digest=sha256:2761344e97eaafd09a266ce492e3d2c797b6a3e65ff7ebe2ec71a9630b393c78

Observation d9bbd787-d624-4161-8818-5c51ed8c466c · outbound

This paper cites Safety Reasoning with Guidelines.

Lifelong Safety Alignment for Language Models Safety Reasoning with Guidelines

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.671515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.671515Z digest=sha256:e2d05cfaa1b7cfac18e606e844b078a32a06384b095e85d3dc7f4a89661dcdad

Observation bed8f58e-038b-49f4-a57c-08a4657d8522 · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.

Lifelong Safety Alignment for Language Models A comprehensive survey of continual learning: Theory, method and application

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.790652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.790652Z digest=sha256:ac571ac2ba12ab47cf6e0fea69f5412b4f89fd42614f7c8053e934d1f42f24b6

Observation d2b2cb03-e159-409a-88fa-04c041eb9997 · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Lifelong Safety Alignment for Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.959298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:08.889926Z digest=sha256:18fc5004838a0a827deef13a274deb930da9bdd327c097545cd278db4ccbe112

Observation cad982fe-1290-4b4a-915f-1dd9f427cea9 · outbound

This paper cites Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection.

Lifelong Safety Alignment for Language Models Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.944339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.944339Z digest=sha256:16770d62fda765181b88c7313cfb4d5a4e07637cbb43c3dd423c05f2f252e645

Observation 52886168-f2d8-4c1a-8165-9f296ceb6053 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Lifelong Safety Alignment for Language Models Self-Play Preference Optimization for Language Model Alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.095114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.095114Z digest=sha256:556685f7ac3f2d32322704a3ecc3e0a02b00188af0a16e14a3dad3b6c419f7cb

Observation ef8af138-15f7-4858-8ab0-fdd8e83ca8f2 · outbound

This paper cites Qwen2 Technical Report.

Lifelong Safety Alignment for Language Models Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.188727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.188727Z digest=sha256:87ae2a9a739ed5c63611adf90543de9bf52e6ac74a59c45711a1e22f079e51de

Observation 7bce9548-c0fb-4385-a5b5-0fbd53fc1461 · outbound

This paper cites Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play.

Lifelong Safety Alignment for Language Models Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.274466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.274466Z digest=sha256:904e3bcbd56cdcc1a97c4b10d04583b0d139b0d7c4d04749a21b13d825c27b5d

Observation b333d74e-723b-4e2f-9f79-a3f8e37df4f0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Lifelong Safety Alignment for Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.367525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.367525Z digest=sha256:29d4c4e30f73d1c817d744fd72bb199165913e576f3d2a5a1f1247c0f0f61edc

Observation 00eaeda2-fadc-4cff-8b0c-0c3331b6c350 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Lifelong Safety Alignment for Language Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.433962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.433962Z digest=sha256:f4188625b402e82b4c953a51ca28f05c6f6a1cd121247b49176e61c1811fa49e

Observation e5ba28bf-bc90-443c-b2b9-76d055c25e1f · outbound

This paper cites Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training.

Lifelong Safety Alignment for Language Models Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.569239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.569239Z digest=sha256:6636547814d4d4d3035210bea8af943d6c9f7b908ff9379a73513f416f8233b3

Observation c333ee08-3f33-45eb-b0f9-1188e89bc963 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Lifelong Safety Alignment for Language Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.671636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.671636Z digest=sha256:d32fbc8b42c510bc3d79fea8376490d97b05cdfb72d3ede1c9a378ad642f294c

Observation 0afd2886-d821-46ee-834c-dc51ce6275be · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Lifelong Safety Alignment for Language Models How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.797411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.797411Z digest=sha256:7a2c4c05f8590132fba6df35c9f1df80686fb1cf915684c6239f6c55f0edc4d9

Observation 389ca17b-67fe-4ee1-a84b-4395171c63ce · outbound

This paper cites STAIR: Improving Safety Alignment with Introspective Reasoning.

Lifelong Safety Alignment for Language Models STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.891654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.891654Z digest=sha256:7d3706abe2a9e92fcb7b9f864dedd93af3ae7ceceabdd4f606c1711f006a98d8

Observation 8057d003-cdd4-4e15-a27b-9562f1df6e40 · outbound

This paper cites Improved few-shot jailbreaking can circumvent aligned language models and their defenses.

Lifelong Safety Alignment for Language Models Improved few-shot jailbreaking can circumvent aligned language models and their defenses

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.738623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:10.035594Z digest=sha256:e0d4a4bb719f2f45f625fea38e0a681eaa4974b9bcdd4c46d94bb1201c98cced

Observation 7e6c0211-6d70-434c-9dfa-92cd99d84f0b · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Lifelong Safety Alignment for Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.147312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.147312Z digest=sha256:2640610fd9416d2cba7023b682522797d5f608ff2ce84f675deca287aa7d8a3c

Observation 098faab8-0d50-4915-a5cc-7e550b6fdf5e · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Lifelong Safety Alignment for Language Models Instruction-Following Evaluation for Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.249221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.249221Z digest=sha256:501f2635af33b0c01b6c422bc56944145ed04e584db97227abf8b9a1c9909248

Observation 0819da74-9a13-4941-8668-53d027ab7ba8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Lifelong Safety Alignment for Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.406991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.406991Z digest=sha256:4eee3bcae98f7278915d58767065c540bf30203ed0b56b1ca78969837caae1bb

Observation e5bd4d64-8957-49a5-af76-f3b812418876 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Lifelong Safety Alignment for Language Models Improving Alignment and Robustness with Circuit Breakers

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.571592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.571592Z digest=sha256:05ba3135f9f6f6f024b4bd98c398260740c645c7c9315502d473ae792c6366cd

Observation a8712d85-ecbd-45d8-8302-e6d7e2516aeb · outbound

This paper cites question.

Lifelong Safety Alignment for Language Models question

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.536150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:10.668254Z digest=sha256:9397b4d9d76ee0fc9baa27ffa65118f9002985c4e452e2a3d0688ce3fc2ff635

Observation 8c2e9e7d-283d-45f3-b655-4cfc11c2e46e · outbound

This paper cites This is the most preferred response when confronted with policy-violating instructions.

Lifelong Safety Alignment for Language Models This is the most preferred response when confronted with policy-violating instructions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.309703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:10.800784Z digest=sha256:579b00e89cadff6f99fca4b3dcbe3a6aec5c1262f990f266ee5c9084386e7556

Observation 557685eb-0ec5-4e2c-a9b3-eb590da57da4 · outbound

This paper cites It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content.

Lifelong Safety Alignment for Language Models It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.888539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.888539Z digest=sha256:4153d18cae57530d0b110bf0416954f3658a3e2235572477a2a836d2d2688a21

Observation 00e3612b-5336-48b6-bc82-e9ee5537b7a0 · outbound

This paper cites It offers content related to the request but without embedding necessary precautions or disclaimers.

Lifelong Safety Alignment for Language Models It offers content related to the request but without embedding necessary precautions or disclaimers

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.967249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.967249Z digest=sha256:8d2fb6d1be23b3f55e38874014a21c250fcb62886503d1ea2c1e7b0721437e18

Observation 1da7b12d-24b3-4197-9137-27c4268732d4 · outbound

This paper cites However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5.

Lifelong Safety Alignment for Language Models However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:11.062426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:11.062426Z digest=sha256:69245271a89cf3515262e9150d49328d530ab89f02e8693c7498c426c7085493

Observation 2d2d1d37-669e-41c8-90fa-bbff023cecd9 · outbound

This paper cites It completely ignores the usage guidelines and provides information that clearly violates the Meta policy.

Lifelong Safety Alignment for Language Models It completely ignores the usage guidelines and provides information that clearly violates the Meta policy

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.070089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:11.144360Z digest=sha256:f35cc47cfb9755f90ec12a7aa981d051dd27aa35a62a3b2310c9d42fdc1122f4

Observation 7cfad83f-1844-4e05-bd39-545c05c87055 · outbound

This paper cites Frame the prompt as a discussion or analysis question to invite exploration.3.

Lifelong Safety Alignment for Language Models Frame the prompt as a discussion or analysis question to invite exploration.3

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.887723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:11.284481Z digest=sha256:99f9dedb030fa5a9841e35b5f6e427b71458ed8b668e8f093e370f42e4e6f13b

Observation 5a69f598-9492-4f2f-9fae-7ff0a3fae857 · outbound

This paper cites an unresolved cited work.

Lifelong Safety Alignment for Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:13.690433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:11.368910Z digest=sha256:59b77c014e23aba7a26522abe54b44aa1b117660b52b63ed1be448e3a991f562

Observation 31c08b3f-f936-4835-a655-dd0b0814f96d · outbound

This paper cites 30th St, Los Angeles, CA 90007, United States.

Lifelong Safety Alignment for Language Models 30th St, Los Angeles, CA 90007, United States

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.472286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:11.477728Z digest=sha256:cf0356dc08d116f0e8e8136098e6bf221c2692b3b55e85b069d90b43a4607cf3

Observation 40ae3064-2211-4898-a35e-6815b8005f54 · outbound

This paper cites 20th St, New York, NY 10011, United States.

Lifelong Safety Alignment for Language Models 20th St, New York, NY 10011, United States

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.291874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:11.616468Z digest=sha256:d66083b81c1d4ba3f3f4ee1f9157702fb08b32a8206e4aa427b95a4cdca16e04

Observation 1c729329-57f8-4f55-9021-d2181740f2b0 · outbound

This paper cites an unresolved cited work.

Lifelong Safety Alignment for Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:13.095931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:11.708628Z digest=sha256:c4441c65a0fd05def3688078f0c91419cbff7dce2ab6b77f728411648f4dc905

Observation 2384860d-b0e0-4adb-a710-7757d15df4ec · outbound

This paper cites pythonchemicals =.

Lifelong Safety Alignment for Language Models pythonchemicals =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:12.905731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:00:11.823481Z digest=sha256:e5631587e4319af4970da7856f2754238f6e64028848924caa4c6d89b0069130

Pith citing papers

Observation 9051076b-a6fd-4a00-90ea-312814908556 · inbound

Safety Reasoning with Guidelines cites this paper.

Safety Reasoning with Guidelines Lifelong Safety Alignment for Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T23:50:36.178215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:50:36.178215Z digest=sha256:63e863dff5a65a4644836ea217cb84a87b8281db010ddb0d08a6b9be628c4f3e

Observation 8ff0fd93-172d-4e8f-80ff-a4cd1ad54c71 · inbound

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World cites this paper.

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World Lifelong Safety Alignment for Language Models

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:46.332617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:46.332617Z digest=sha256:b8e95686278c87fb6104477277da951e9a6ccc83b07b36ec2f0d7e9b3cc72900

Observation fd4f8b5c-061e-4393-838b-cdfb51d8fe9d · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Lifelong Safety Alignment for Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.461615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:6fb581722f302c3c0b81017a38b3bd9549ce3962cee89908405b563124c6537f

Observation 19d35091-37dc-4c53-8a57-c613eefcd074 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Lifelong Safety Alignment for Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:41.955346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:cb39159d88bbb297479b9e074177695fca4447bc32c4e577a04c1c44f7155d6b