Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:35:34.616461Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 20 inbound Pith citation observations for arXiv:2508.09224.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:35:34.616461Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T01:12:44.184834Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
6 of 6 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 6fae85d6-77d2-416f-8c51-c429779c4d98 · outbound
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2be2c8c-8ac3-4957-8e84-5c49decdb846 · outbound
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training [14]OpenAI
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81b07747-e930-4ec7-94d4-ab893594fb53 · outbound
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c3d15c9-ca69-4187-af43-76ddc9f34c93 · outbound
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05cc146-80ba-4697-b710-292789aa3a39 · outbound
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a63e92e-2b3c-4b20-afb6-6ea4a7b1b4db · outbound
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training GPT-4o System Card
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5e3815c-4641-4358-ba47-d92df3d6a102 · inbound
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5896c4be-d5ba-4f1d-b9d4-3591bf8b6ef7 · inbound
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36ecda09-0937-4c16-ab76-00ee4d62cc54 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07612579-4de2-4a67-a42e-7fd7c0e3ed24 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66929448-40e1-4ddf-be66-fc7de901521d · inbound
Cat-DPO: Category-Adaptive Safety Alignment From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41e0bb4f-39d5-48a6-83fc-c31191e12e91 · inbound
Using large language models for embodied planning introduces systematic safety risks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0bec346-0217-440a-9c94-502e42612725 · inbound
Jailbreaking Frontier Foundation Models Through Intention Deception From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aec6227f-8076-41dc-94a9-9d0babd9efd4 · inbound
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27dddf4d-e254-4f60-953b-8ff3632f9d83 · inbound
Internalizing Safety Understanding in Large Reasoning Models via Verification From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d489190-aa42-498a-9084-49fc23f6703c · inbound
Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27d5e496-392f-4511-a39f-bf7b49040453 · inbound
Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a845245-3a9c-46d1-9b9b-cdb20a841c83 · inbound
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d966b265-c9f3-4000-b32b-f8565878dee3 · inbound
Investigating and Alleviating Harm Amplification in LLM Interactions From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fed5df3e-33d3-415e-967a-4a6a64015832 · inbound
Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28e34803-b869-4ca7-93d5-9fb1e5c85b01 · inbound
Understanding Censorship in Large Language Models: From Mechanisms to Governance From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8032e4dd-60b3-47f1-a891-572cd3fd65ad · inbound
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13e43977-b44e-4383-bfa6-66195151621c · inbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b09427-a1e8-4a06-a09e-bb0a28d19412 · inbound
GPT-Red: Automated Red Teaming via Self-Play at Scale From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a138c4e-30e0-4b6a-8e2f-7bead74052f4 · inbound
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9486a6b5-410d-471b-afd1-06341419d6b5 · inbound
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.