Pith. sign in

Paper Citation Record · LEDGER

PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 71 inbound Pith citation observations for arXiv:2406.15513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.15513 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 71 of 71 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:18:52.479377Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 027b6d57-d2fe-414a-90a1-03ebb53350df · inbound

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution cites this paper.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.629915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.629915Z digest=sha256:a3be058b2fb6bbc67bf0c2e8b7dfefda129d0ca07ddb81f3bd942905c52d18dc

Observation 542ffbf9-5bab-40b6-94f0-a5710bb8e41d · inbound

VLSBench: Unveiling Visual Leakage in Multimodal Safety cites this paper.

VLSBench: Unveiling Visual Leakage in Multimodal Safety PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:46.428134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:46.428134Z digest=sha256:107a9e8288fe8494018940737f76906add08594850fc476dbcfa96718f9fb42f

Observation 59280202-38f0-494f-b3db-ed4b0f5e2bdb · inbound

Robust Multi-bit Text Watermark with LLM-based Paraphrasers cites this paper.

Robust Multi-bit Text Watermark with LLM-based Paraphrasers PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T22:49:58.514761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:49:58.514761Z digest=sha256:04304b1b32c550fa5619fbaff03209baf3ea633a545b24dba5a782937b430d0a

Observation c07fd98b-a7fa-4a78-bc62-d6a2491ba213 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.843119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:f81279b6c627ded7ed76a19c0d96caa55380fb7d4107e11b232091d9db157e0f

Observation 02230ff2-b5e2-494d-aec5-e82f8bc79f83 · inbound

Targeted Angular Reversal of Weights (TARS) for Knowledge Removal in Large Language Models cites this paper.

Targeted Angular Reversal of Weights (TARS) for Knowledge Removal in Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:25.014603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:14:25.014603Z digest=sha256:e04bd9f37fb40070f5b1d3044487ed6ab34f1da9b4e5bf63402fb79447aedee8

Observation b22ca459-5270-4b9e-9092-fe28562c7324 · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.150282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.150282Z digest=sha256:a2784e216e88dbb25addac718e9c02e5da4fb10561bfd0a228ecf88edcbc9d8c

Observation 1c2e0654-df77-4799-972b-b57bcfff8d00 · inbound

Multi-Objective Large Language Model Unlearning cites this paper.

Multi-Objective Large Language Model Unlearning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:27:37.742460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:27:37.742460Z digest=sha256:80200937c4f5d7920079a4fe64cd7d0a46b067a86671aaf0e7bb892762245106

Observation 7447808b-12ec-4dbe-8aa2-aa4544fb5fcf · inbound

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction cites this paper.

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:48.381225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:17:48.381225Z digest=sha256:72ee4c49c6ae1ef89e9c49422b6d5e7024b24af70becb71614d31554be8a6212

Observation fa50e317-59f1-41fc-bd43-4e29d0173a98 · inbound

Gradient-Based Multi-Objective Deep Learning: Algorithms, Theories, Applications, and Beyond cites this paper.

Gradient-Based Multi-Objective Deep Learning: Algorithms, Theories, Applications, and Beyond PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:50.107993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:53:50.107993Z digest=sha256:ffedbe2f86216c084e7c18ada39ac67b0ffdf99115283bb238e26d449c4c1f40

Observation b12fa28a-d167-4fcf-aeb5-fa8ad8b9afc4 · inbound

Data-adaptive Safety Rules for Training Reward Models cites this paper.

Data-adaptive Safety Rules for Training Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.146914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.146914Z digest=sha256:6a27544726acac1db5387e29be17d610195bbbc94a61b508df4a9bd92ea593cc

Observation e12b652f-af52-4be7-b387-036e5cfe24d3 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.813673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.813673Z digest=sha256:bde667773c2112a16f9fe90997222f2531545faebd97f8d53b336c1e3c5a70f0

Observation 44ab9950-d943-44e9-9e48-8df82a5721f4 · inbound

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing cites this paper.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.037586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.037586Z digest=sha256:a088026d219be782baa17da970216a091b791ad4db45bb7b191cf29abd80bffd

Observation 53647d4d-ed38-470f-a8bb-4640aef25a66 · inbound

Safety Reasoning with Guidelines cites this paper.

Safety Reasoning with Guidelines PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T23:50:35.767760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:50:35.767760Z digest=sha256:0102a9e01e521fa7c9ef6e4ec097a3c6635f617d4322b1cbe8489210d97a13c1

Observation 238801f1-b654-41a6-a352-bebadbb6a97d · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.606200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:484be12cab06c7f72707abe4d9c1aeb180a662981a002d6386ac454fdf641a9a

Observation d882c0ce-848e-41dd-9599-3728f085f572 · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.329365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.329365Z digest=sha256:56b938bc59104eb22dd9a0ee1dc695037e5b3aebf0ce29d9afc0bc481884e038

Observation 95db99d9-5da6-4d58-94bd-453783b2ad5a · inbound

Probing and Inducing Combinational Creativity in Vision-Language Models cites this paper.

Probing and Inducing Combinational Creativity in Vision-Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:52.479377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:18:52.479377Z digest=sha256:79b86d3ab72f93b6b2979d671799da9e770f1185e84fdcaa74dac269fea4d63f

Observation d8c0ae58-70ef-46b8-a18d-15456ea83bfe · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 265

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.809889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.809889Z digest=sha256:561854f8b7b9ba2c73bfeedcaed18d3240160aa98705c4d48ddeba213f105d2e

Observation 2f93b59b-02a3-4d92-a9e1-683d7951790e · inbound

PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model cites this paper.

PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T23:52:01.523681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:52:01.523681Z digest=sha256:29146a654eb34761b3334f3daabe74af155658d2f9ce1349a761b118ee062aa1

Observation 06262224-13c3-4f23-8a03-21b71a1463e5 · inbound

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge cites this paper.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.534659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.534659Z digest=sha256:1bcd889d7640d56bd6058c369456f83ce8c627d7d4897d84aecaf0a78b0a7ee2

Observation 2de9dde4-b320-4095-864b-2187a14f885d · inbound

SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models cites this paper.

SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:26.903646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:12:26.903646Z digest=sha256:c6d6ab71d972686b261c8dc90fe3788a5e11bcaa22ca86b9fde2261b54586c14

Observation 795a9b5c-9ae4-401d-9fa3-edf03b080874 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:27.885966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:27.885966Z digest=sha256:35b5246533c5a0edb9efd7f913134c8afe5940569c52588e9a3c292913b614a3

Observation e349f15f-85a8-4845-a1ec-db80e8b5f095 · inbound

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming cites this paper.

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:58.475342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:09:58.475342Z digest=sha256:afc5f7060ca05c4dfd730d8e1d9c7e3cf3e8abdd6995de52e6e927c558b32224

Observation fe891c49-9c68-4b27-a31d-636d3040cd64 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.092591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.092591Z digest=sha256:6a161bc7149f0bf65bb7631100fbab779d4ab7c83bd8e0fc6b02bbf188bc3cb4

Observation 6390b21b-79ca-4333-85d3-e69482b152e3 · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.389017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:021d856b567c22a1ca9c04888639f8e9027fd2e175d8f438ff07ba928a468c3d

Observation 62fb4ef6-d5c8-4425-9bab-d36d4cfe9f14 · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.457692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.457692Z digest=sha256:a9a9decdee60f318d6b2001aab278d296bb84a52fb6958e3015e3b850ae37e30

Observation 5f05f867-ff0e-4065-8f01-bf7651d187a3 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:31.054317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:31.054317Z digest=sha256:b54239d9f1aacf6b280e929a532e98dad79a66c700b2995310612044c79120bc

Observation a158b6c6-fa85-4b43-9101-14ce2a31ae06 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.888953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.888953Z digest=sha256:d831e7085c8fbcaa23d19f03a609cd4c5d55164f803bf2c3762a7f37b5c481b4

Observation 75620268-4462-4d24-b277-de0d52d9690e · inbound

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time cites this paper.

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:20.174221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:20.174221Z digest=sha256:64c2755a16086b6ba8a0946b0ff43c1fd6dc1241c85e92b99509dca6915c251c

Observation eefbd577-524a-4cc1-a8d0-3faa2c8147af · inbound

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning cites this paper.

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.749164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.749164Z digest=sha256:3d7b4e1ef1980752bf5b9fedb19ad5be17ab9bbef18cb71a0baa0a1900fb00e5

Observation eed00d0e-6f1d-4820-a85f-fc1adb3eada8 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.823967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.823967Z digest=sha256:adb91e1d074d7d4ea796a03d5ad3b6fe41177c54e34cb59244b7edfd9493a7d8

Observation eda1e4c2-d2fb-449e-b7a7-bf446442c1ca · inbound

Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples cites this paper.

Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:11.801146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:32:11.801146Z digest=sha256:03759490e6ffcff74b9a187ca5c0b67eebff2da924609041599843e7047b9743

Observation dbe04a67-8c2d-4973-a02e-dd0938068495 · inbound

MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs cites this paper.

MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:08:22.809009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:08:22.809009Z digest=sha256:d921e492c28bea571a6017fa33b31d46f1d5fb38c2a333f9bbcc325aaccebc51

Observation 3f47598e-ec16-4990-aec8-02e509777422 · inbound

Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models cites this paper.

Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:52.401761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:52.401761Z digest=sha256:cfe3881a5e5c04b25db1971924df3a78e5911f396cc13a3691dcf48fd2debcc9

Observation 42fe3188-5383-423d-b63f-299496c0b3a7 · inbound

Tiny Reward Models cites this paper.

Tiny Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.137146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.137146Z digest=sha256:6bb669388a7dc1f7ad3f0342cdda2e64695c569da96cd79bac12aed355df2f9b

Observation a1f836f5-fb40-4bdb-9479-ea1315d2ea15 · inbound

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong cites this paper.

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.221861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.221861Z digest=sha256:a7675698a3875a1cda48f534ca480b14f27b9c61e6feae35b1e72791002214c2

Observation bd353d43-b1bd-435a-ae9f-85c017eed29b · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.564199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.564199Z digest=sha256:1a5a038ef801a8d2e5e4cbd6706fc19cd10b26350d0cc333d8f27765328bec5b

Observation 7cee2df6-4214-4911-bf5f-530698648edc · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:22.849329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:22.849329Z digest=sha256:20cbda9bba3703c67f9839644790c48544b44aa63bcc3003bda962b88cba1223

Observation 8a1e5778-804d-43e5-9f4f-b251132c5784 · inbound

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols cites this paper.

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T13:54:39.777864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:54:39.777864Z digest=sha256:0d4cb6bb5556940488bb411b941952daae4f081d3bbd114250f996cdee2875c2

Observation eff3e40d-afc2-4c4a-adf1-c989463ac2dc · inbound

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks cites this paper.

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:28.700858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:28.700858Z digest=sha256:a29d59667661140b31f9fef03b126004ab8ff4677787e0efd5a4852c559cb04b

Observation c95821df-2fb0-4028-9f2c-4fb40de35681 · inbound

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI cites this paper.

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T15:08:54.032086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:08:54.032086Z digest=sha256:606fa96be7497cd189bf7cf1754665c81ccbbc1eae9bd7b8a2bf5de0fa35541b

Observation 223eb6c5-4b62-4bb9-b4aa-58793cd83de8 · inbound

The Realignment Problem: When Right becomes Wrong in LLMs cites this paper.

The Realignment Problem: When Right becomes Wrong in LLMs PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:30:35.748934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T01:28:29.056145Z digest=sha256:352e695b2227315bc28534490335324b37d23ecdba756eee3089e316331e92f0

Observation 25368b2e-73a0-462d-b0d2-a24ebb1d63cc · inbound

SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models cites this paper.

SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:28:40.673267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:26:48.405593Z digest=sha256:e2f1fbd0be479a139e82b4d4ed3308cd27fa32a07d72738ffa262dcf4c450d6f

Observation e537745e-1e07-4d3c-a064-b218c88a2be4 · inbound

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection cites this paper.

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:40:42.297455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T06:39:17.785715Z digest=sha256:910205e2608eaa6ee59212baa88962fefb99c7ec08c03d7a510009f46280d3f1

Observation be7a05a0-4339-4a04-b476-d341a4c317a7 · inbound

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning cites this paper.

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T10:41:51.251736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:41:51.251736Z digest=sha256:911e4d6af368e57a96e45c8cc1718a15dc0e8c58aca7355fdbbee08c7d669b40

Observation 4e1ba725-ec50-426f-8a97-8b3346a4580b · inbound

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization cites this paper.

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:01.706661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:22:40.613937Z digest=sha256:a851a26dfa95dcc8438bcca1059affee76cbe5e921529d0f0c6bcb1bb2767728

Observation 36019f95-c876-420d-b347-bec1c4ad3951 · inbound

Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment cites this paper.

Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:02.849441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:23:33.330990Z digest=sha256:ef023d0fba4a6e1e5fba4e0588f3af591350c1c37a28abcf48084b60ef104121

Observation f16c24b6-74c0-48d8-8c32-627d49b8987d · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:56:11.648277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:39e8e5d7135b6772cfca6a0e277bc960ebf64af053b4b6610e0a9bea32b6fc6a

Observation 1116b0a4-0e5d-4143-9f40-f09da8eb3e57 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.122827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:4663e740b9a5892eea7d911c6db42576a18de778f0903ffb878776dd6dbba872

Observation a1430dbe-54c2-4812-8474-047c9a47b1cc · inbound

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models cites this paper.

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:22:20.648413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T02:21:29.463149Z digest=sha256:63ff430f30b4d64b3d793ee8c073988795f32f7fa85eca518f0ddf29fe13b712

Observation 421aee01-67c4-4e2e-9c11-22bc5ba68250 · inbound

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment cites this paper.

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:21:07.450631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T17:24:54.796037Z digest=sha256:a2d88437561678f94fa311148828dbb8428aa1e39a250bf825804053c2004e83

Observation 443f9f7a-6264-41a6-b089-a718eeda634d · inbound

Theoretical Limits of Language Model Alignment cites this paper.

Theoretical Limits of Language Model Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:57.130357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:18:37.614335Z digest=sha256:362aa249dded5cab496089ee154a79f17036fe7e9a7ed05da83ab38923421e44

Observation 92d2c7c5-7c5f-4519-a9b3-d91e03622f41 · inbound

GLiGuard: Schema-Conditioned Classification for LLM Safeguard cites this paper.

GLiGuard: Schema-Conditioned Classification for LLM Safeguard PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:55.235357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T03:19:11.495230Z digest=sha256:4d69ab79650a4a660c208f95cceec2c192f33fa6baec91ec85eb2e597eb73340

Observation d4cf753a-c288-435a-a064-76e6dee5dbb9 · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:26:26.493049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T04:15:24.919355Z digest=sha256:b959c10611ea8cf63cdf041b254a93950213e5755cfa44cd94bc6edcd3509d41

Observation 55a935d3-6ded-46ec-9a05-70d7f56a358a · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:08:05.044304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T22:03:05.102274Z digest=sha256:2e92a1dc545ea7971a1a8382b89b9fa010bf3ef51e96c25d6e2577cfd0f538a4

Observation 14b214f3-f55b-43a4-8c5d-4fb4bae7ccc4 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.573759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:13ec1d0112acb6a38d50a0b0127e1d6b178a64c7486cfa3df4d91824cadc949c

Observation 79d9af60-23e5-4e2c-9a54-3f79909f8d3a · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.932654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:5ceb3ebf121994010c340f0a786a03dc6c056661b2cf71d834ca9e9925d30f9d

Observation 78a7f548-801c-4d12-93df-971461538fbb · inbound

Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance cites this paper.

Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:07:27.645410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T07:03:14.415745Z digest=sha256:ea044d316bc72c3c19b97e4b3931ce5d29f1745dc85d0fc584f308e99af44112

Observation ee8bb0a1-922b-4ca0-b682-34ddbaad234b · inbound

Curriculum Learning for Safety Alignment cites this paper.

Curriculum Learning for Safety Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.198090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:25:31.739331Z digest=sha256:64f81b497b44fe0f2189cb1eecb33b464b59e6e925ed0935857df02f8276b2df

Observation 8ef90606-1dae-45e7-b924-a55ef5ee26a2 · inbound

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content cites this paper.

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:13:15.988572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T09:11:58.843585Z digest=sha256:6e9e556647b8ff657f559f06c34a16e487906d65213d278f94356f9bf6288fa4

Observation 97042b30-7a37-4aa3-a47b-2d98b0242395 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.846238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:060c54fa24c3b3153f193c4c791a118d8b0e4bc5c7f7494903cf8edcf51bb1b5

Observation 4358c7bf-53a8-46c3-bbff-052a33ebadd0 · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.867485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:fc7548b55ab63caa0d51207250ea86df6cf068a6d1078784262c2fadcb408ad3

Observation 0a43df0a-b0bd-4264-972b-f9e371709a1e · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.032922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:c0330505a100fbdb56c8c116dcbaeae0451c2db959746871f1ef996b001aa4f4

Observation c595a4c3-30a6-4659-9233-e8b39098f680 · inbound

Multi-Objective Exploration and Preference Optimization via Mutual Information cites this paper.

Multi-Objective Exploration and Preference Optimization via Mutual Information PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T08:52:34.039089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:52:34.039089Z digest=sha256:0ecb520e969ba6d80a467b13631a20500ddd081e1743840ff899d4d87b911d69

Observation 601de794-7e48-4b62-9180-effb850bbf3b · inbound

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets cites this paper.

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.496348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T14:44:57.205766Z digest=sha256:dd2f25939ca005951870fcee4eaae53b616fb5079e716a178913c495588b0131

Observation eb8b634a-2ac2-4c99-8aa1-543c4c1076f7 · inbound

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks cites this paper.

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 276

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T15:47:23.388056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-10T15:38:58.361411Z digest=sha256:bddc79ad7b685c57055ae25b2554a7cdae08ea9923a87c9ca738eab4534b97d7

Observation dd406a54-ffe1-4442-ad2e-102896795bda · inbound

Step-Level Preference Learning for Generative Agents in Social Simulations cites this paper.

Step-Level Preference Learning for Generative Agents in Social Simulations PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T02:00:25.508169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:00:25.508169Z digest=sha256:2ba99344b611236497cbfd9c94a36015e02f609ad3c335314f62a1af83f72ecf

Observation 9d5a050a-24bd-4eac-ac8d-78d296a66c50 · inbound

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment cites this paper.

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:10.523132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:10.523132Z digest=sha256:ef5be0da1ea8ea1228f914263ad3088e50d3573160c8192f26bd2898cf72e3cc

Observation aaeee69b-2017-43c0-a70b-38d6b7350112 · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:36:55.305100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:36:55.305100Z digest=sha256:d7379a461ab6ede0469c7393203e576455b796c4eb028ecb925bf2f734dd2b17

Observation c5a4afa0-c396-49aa-b748-5bb4d4808ac4 · inbound

Visual Token Compression Enhances Robustness of MLLMs cites this paper.

Visual Token Compression Enhances Robustness of MLLMs PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.372979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.372979Z digest=sha256:347969812ce88a44d0dd0941d09bba4c7087f72a590ee7f52bfec04b43f09f73

Observation 354bbca7-6e30-44f8-8921-d118c8d1b3ca · inbound

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition cites this paper.

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:35:24.771726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:35:24.771726Z digest=sha256:9f3684a6b3470d2ab91fbb64f60c0f893d5e70bb6ee84050ae4c496d1432e013

Observation 86642dfe-8e5c-417f-a708-d95cbcb5edf8 · inbound

Orientation, not magnitude: the causal structure of task-vector interference in merged language models cites this paper.

Orientation, not magnitude: the causal structure of task-vector interference in merged language models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:14.975581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:14.975581Z digest=sha256:27cc3e9e8d34a279cf1d584ad5973a408a2d1585d3d5c88fac683349e3284b0b