Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T13:14:34.157748Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2502.02153.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T13:14:34.157748Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5e99a51a-78c5-41c9-b065-1cd0e964302c · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e8b5c1-bf95-4261-af75-94fec11633a3 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Concrete Problems in AI Safety
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ec5777-ad42-4e52-a73e-8723cdbbbe3c · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing A general theoretical paradigm to understand learning from human preferences
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c7e0272-d3be-44bc-bfff-8c652f2e3c3f · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184e74a8-9b08-463c-850f-5e5d9544af9b · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing The ethics of artificial intelligence
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7db00aeb-e6a8-4640-8b25-34ecbf99858e · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e750a8bd-f1da-499c-9799-0934458e5bd8 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c5e9c0-6400-4e2e-94bb-ab5198eadf68 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Deep reinforcement learning from human preferences
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc7cf301-b09c-47d7-b707-001fba0ffc63 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfab5b30-930c-4324-b9f7-71afa4d2da0a · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Safe RLHF : Safe reinforcement learning from human feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb5db034-2f19-4d33-9cc9-24dbfe1d4142 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b13d94c2-4825-4279-b487-e61718bb53ce · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e1a8042-a0c2-4117-852c-1aca7cb31e64 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing KTO: Model Alignment as Prospect Theoretic Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f96e0cd-a1b2-423d-9820-ca7046ffb9a6 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Pal: Program-aided language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9101effa-532c-40d7-b501-1987bd7cde2e · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af0096dc-7c00-40b2-abc7-e90bd0b509c7 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Measuring massive multitask language understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 902cc25e-67ee-4d32-b70d-0278c7fc7eeb · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing ORPO: Monolithic Preference Optimization without Reference Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb16407-657e-4b5d-b21f-9dcc3433fc89 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing One-Shot Safety Alignment for Large Language Models via Optimal Dualization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c68720-f744-4ff2-869d-311c1b393377 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1d3a8a-76af-43e3-993a-6aaa39a579f6 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Reference point specification in hypervolume calculation for fair comparison and efficient search
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 33c186d5-3f4b-4479-ad01-e69a2adf1ffd · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing AI Alignment: A Comprehensive Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ab9950-d943-44e9-9e48-8df82a5721f4 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4afe23c0-593d-431b-9087-6b094fb3db7d · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3de26fcb-5bcd-45b2-afd1-e7c91484d7ce · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2953044-edaf-4691-82c5-bebcce94a697 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc250669-7669-49bd-8277-958dea84a5d3 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Controllable Text Generation for Large Language Models: A Survey
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c54c4e0-4462-4e33-8941-af2e662b11c2 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9328f6-2725-4373-b282-8c910e68d096 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 319a4ab5-96ea-43be-a1f6-df6a5040b93f · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1a9168-3123-4793-9419-953e9e725e33 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Enhancing LLM Safety via Constrained Direct Preference Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07339d1c-e2e0-412f-af1f-cc06439acf89 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Meta llama guard 2
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ff434a2-881c-4014-822f-0c9e403eee70 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cabeddd7-716a-4913-be82-5143ffd63a3c · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Controlled decoding from language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45c92630-c3b6-4cf6-9000-1558fc446376 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Training language models to follow instructions with human feedback
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2952507-47d7-4919-8b83-911c110e49ec · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bfc7c9d-2f7f-41eb-b1cd-3363dd6a452d · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ddc9fd-24d8-469a-b31b-710ddb4b2787 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Direct preference optimization: Your language model is secretly a reward model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceac96ca-22cb-4c6f-af82-1ec1d217ac9c · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d36f4813-4d84-4cbe-9bed-3d4848a17769 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Proximal Policy Optimization Algorithms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2fcfa10-a98d-45f9-9406-c10bcee36b09 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing LM-Nav : Robotic navigation with large pre-trained models of language, vision, and action
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1166d32d-20bd-4c3b-b1e8-4d4ce1f46fc9 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Hashimoto
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01242aff-dd32-464f-872b-43b3309c24e5 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Large language models in medicine
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df906c89-96cd-4fb8-a106-c46f4b520c8f · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing TRL : Transformer reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 306839f3-de22-4a95-8c25-eb9eaf73ecb4 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Stepwise Alignment for Constrained Language Model Policy Optimization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b219fbce-6ee7-4844-887d-2171c5551971 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aebf8f1a-3c1c-4fe8-a5c7-c90c725ec6d2 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bddb868-48fb-4ea8-b6c4-115df35357a5 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8161e2b2-6649-4b2c-8fa1-387350c8f91d · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Defending chatgpt against jailbreak attack via self-reminders
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ad5719c-32d5-46eb-b9c6-9bcddfe21f03 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36981ff1-9520-433a-ad4b-23e1a8cb120f · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Uncovering Safety Risks of Large Language Models through Concept Activation Vector
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b9a09f-61db-463b-8b6d-d0322382fd37 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 964664c9-7fef-45f7-92a2-028c2ff6a4a3 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing On the vulnerability of safety alignment in open-access llms
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7c1fbb45-dc78-4bd2-b6b9-a3ed47dc4c18 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Wordcraft: story writing with large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d7069d77-86be-4656-864c-0050591e34ce · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Prompting large language model for machine translation: A case study
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5f3243ae-972c-4be3-991e-2052a8ea7f29 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Panacea: Pareto Alignment via Preference Adaptation for LLMs
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96d3a07-5cec-41bb-a301-0f48fbcc06ff · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d0fdc2-7580-4405-948e-bdcdcc2e787e · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Fine-Tuning Language Models from Human Preferences
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ee1b5c-135b-4add-a6e1-7b061bef9de6 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdba3956-662d-498f-b783-e492e8e943b5 · outbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Improving Alignment and Robustness with Circuit Breakers
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.