Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 71 inbound Pith citation observations for arXiv:2406.15513.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:18:52.479377Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 027b6d57-d2fe-414a-90a1-03ebb53350df · inbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 542ffbf9-5bab-40b6-94f0-a5710bb8e41d · inbound
VLSBench: Unveiling Visual Leakage in Multimodal Safety PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59280202-38f0-494f-b3db-ed4b0f5e2bdb · inbound
Robust Multi-bit Text Watermark with LLM-based Paraphrasers PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07fd98b-a7fa-4a78-bc62-d6a2491ba213 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02230ff2-b5e2-494d-aec5-e82f8bc79f83 · inbound
Targeted Angular Reversal of Weights (TARS) for Knowledge Removal in Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22ca459-5270-4b9e-9092-fe28562c7324 · inbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2e0654-df77-4799-972b-b57bcfff8d00 · inbound
Multi-Objective Large Language Model Unlearning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7447808b-12ec-4dbe-8aa2-aa4544fb5fcf · inbound
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa50e317-59f1-41fc-bd43-4e29d0173a98 · inbound
Gradient-Based Multi-Objective Deep Learning: Algorithms, Theories, Applications, and Beyond PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12fa28a-d167-4fcf-aeb5-fa8ad8b9afc4 · inbound
Data-adaptive Safety Rules for Training Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e12b652f-af52-4be7-b387-036e5cfe24d3 · inbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ab9950-d943-44e9-9e48-8df82a5721f4 · inbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53647d4d-ed38-470f-a8bb-4640aef25a66 · inbound
Safety Reasoning with Guidelines PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 238801f1-b654-41a6-a352-bebadbb6a97d · inbound
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d882c0ce-848e-41dd-9599-3728f085f572 · inbound
AI Alignment at Your Discretion PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95db99d9-5da6-4d58-94bd-453783b2ad5a · inbound
Probing and Inducing Combinational Creativity in Vision-Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c0ae58-70ef-46b8-a18d-15456ea83bfe · inbound
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 265
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f93b59b-02a3-4d92-a9e1-683d7951790e · inbound
PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06262224-13c3-4f23-8a03-21b71a1463e5 · inbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de9dde4-b320-4095-864b-2187a14f885d · inbound
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795a9b5c-9ae4-401d-9fa3-edf03b080874 · inbound
MPO: Multilingual Safety Alignment via Reward Gap Optimization PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e349f15f-85a8-4845-a1ec-db80e8b5f095 · inbound
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe891c49-9c68-4b27-a31d-636d3040cd64 · inbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6390b21b-79ca-4333-85d3-e69482b152e3 · inbound
Incentivizing High-Quality Human Annotations with Golden Questions PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62fb4ef6-d5c8-4425-9bab-d36d4cfe9f14 · inbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f05f867-ff0e-4065-8f01-bf7651d187a3 · inbound
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a158b6c6-fa85-4b43-9101-14ce2a31ae06 · inbound
Lifelong Safety Alignment for Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75620268-4462-4d24-b277-de0d52d9690e · inbound
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eefbd577-524a-4cc1-a8d0-3faa2c8147af · inbound
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed00d0e-6f1d-4820-a85f-fc1adb3eada8 · inbound
RewardAnything: Generalizable Principle-Following Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eda1e4c2-d2fb-449e-b7a7-bf446442c1ca · inbound
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe04a67-8c2d-4973-a02e-dd0938068495 · inbound
MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f47598e-ec16-4990-aec8-02e509777422 · inbound
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fe3188-5383-423d-b63f-299496c0b3a7 · inbound
Tiny Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f836f5-fb40-4bdb-9479-ea1315d2ea15 · inbound
HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd353d43-b1bd-435a-ae9f-85c017eed29b · inbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cee2df6-4214-4911-bf5f-530698648edc · inbound
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1e5778-804d-43e5-9f4f-b251132c5784 · inbound
Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff3e40d-afc2-4c4a-adf1-c989463ac2dc · inbound
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c95821df-2fb0-4028-9f2c-4fb40de35681 · inbound
Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223eb6c5-4b62-4bb9-b4aa-58793cd83de8 · inbound
The Realignment Problem: When Right becomes Wrong in LLMs PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 25368b2e-73a0-462d-b0d2-a24ebb1d63cc · inbound
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e537745e-1e07-4d3c-a064-b218c88a2be4 · inbound
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation be7a05a0-4339-4a04-b476-d341a4c317a7 · inbound
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1ba725-ec50-426f-8a97-8b3346a4580b · inbound
FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 36019f95-c876-420d-b347-bec1c4ad3951 · inbound
Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f16c24b6-74c0-48d8-8c32-627d49b8987d · inbound
Characterizing Model-Native Skills PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1116b0a4-0e5d-4143-9f40-f09da8eb3e57 · inbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a1430dbe-54c2-4812-8474-047c9a47b1cc · inbound
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 421aee01-67c4-4e2e-9c11-22bc5ba68250 · inbound
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 443f9f7a-6264-41a6-b089-a718eeda634d · inbound
Theoretical Limits of Language Model Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 92d2c7c5-7c5f-4519-a9b3-d91e03622f41 · inbound
GLiGuard: Schema-Conditioned Classification for LLM Safeguard PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d4cf753a-c288-435a-a064-76e6dee5dbb9 · inbound
Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55a935d3-6ded-46ec-9a05-70d7f56a358a · inbound
Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14b214f3-f55b-43a4-8c5d-4fb4bae7ccc4 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 79d9af60-23e5-4e2c-9a54-3f79909f8d3a · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 78a7f548-801c-4d12-93df-971461538fbb · inbound
Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee8bb0a1-922b-4ca0-b682-34ddbaad234b · inbound
Curriculum Learning for Safety Alignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ef90606-1dae-45e7-b924-a55ef5ee26a2 · inbound
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 97042b30-7a37-4aa3-a47b-2d98b0242395 · inbound
Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4358c7bf-53a8-46c3-bbff-052a33ebadd0 · inbound
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a43df0a-b0bd-4264-972b-f9e371709a1e · inbound
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c595a4c3-30a6-4659-9233-e8b39098f680 · inbound
Multi-Objective Exploration and Preference Optimization via Mutual Information PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601de794-7e48-4b62-9180-effb850bbf3b · inbound
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eb8b634a-2ac2-4c99-8aa1-543c4c1076f7 · inbound
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd406a54-ffe1-4442-ad2e-102896795bda · inbound
Step-Level Preference Learning for Generative Agents in Social Simulations PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5a050a-24bd-4eac-ac8d-78d296a66c50 · inbound
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaeee69b-2017-43c0-a70b-38d6b7350112 · inbound
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a4afa0-c396-49aa-b748-5bb4d4808ac4 · inbound
Visual Token Compression Enhances Robustness of MLLMs PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 354bbca7-6e30-44f8-8921-d118c8d1b3ca · inbound
When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86642dfe-8e5c-417f-a708-d95cbcb5edf8 · inbound
Orientation, not magnitude: the causal structure of task-vector interference in merged language models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.