Pith. sign in

Paper Citation Record · LEDGER

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

As of 7 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 20 inbound Pith citation observations for arXiv:2508.09224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09224 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:35:34.616461Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:12:44.184834Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6fae85d6-77d2-416f-8c51-c429779c4d98 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:33.950466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:33.950466Z digest=sha256:2b707a233fa31e27375855579eff64e2165ad0c5d734deef1d50ab07acb4006a

Observation f2be2c8c-8ac3-4957-8e84-5c49decdb846 · outbound

This paper cites [14]OpenAI.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training [14]OpenAI

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:35:34.923470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T21:35:34.366680Z digest=sha256:35b3fccdb904a0561cf9bc4fdb9511be003302ca4b25b1777d4d64997ef6cf6e

Observation 81b07747-e930-4ec7-94d4-ab893594fb53 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.494480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.494480Z digest=sha256:5834b402aa7d1a2bcaf0925a3bb52bd6034489c54f290821c419ca7ac0b7ffb7

Observation 9c3d15c9-ca69-4187-af43-76ddc9f34c93 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.096588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.096588Z digest=sha256:5c52f6a705337a75da4f6539ed7e36609bd465e51ff5f7ffd7a5ddc1a03dddae

Observation e05cc146-80ba-4697-b710-292789aa3a39 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.616461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.616461Z digest=sha256:973d8f5b92b4da8e5f6c471ae5650683015932d901dab83fd76ef79015beb963

Observation 9a63e92e-2b3c-4b20-afb6-6ea4a7b1b4db · outbound

This paper cites GPT-4o System Card.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.168865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.168865Z digest=sha256:4057f8e7ed449166afb273b806a0884090df68ddab527703d653e20787f72070

Pith citing papers

Observation d5e3815c-4641-4358-ba47-d92df3d6a102 · inbound

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks cites this paper.

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:38.344182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:00:06.402954Z digest=sha256:937dac9763b7e25c2d4e84c3e07fffb818507845e001899ba424853d1adc8364

Observation 5896c4be-d5ba-4f1d-b9d4-3591bf8b6ef7 · inbound

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models cites this paper.

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:42:36.853762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T12:41:48.620040Z digest=sha256:e2967fe94db016df822c1c3de335607a653f361f28bda1557dca9693d1e4349e

Observation 36ecda09-0937-4c16-ab76-00ee4d62cc54 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.616391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:89458605f406deb0e912cc8857784613b6d18b50ed22da89df1dc28a55ca692b

Observation 07612579-4de2-4a67-a42e-7fd7c0e3ed24 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:c93e677138df4779da54668bb703b84bbd7e1390105db44df8d3f2355aac4dcf

Observation 66929448-40e1-4ddf-be66-fc7de901521d · inbound

Cat-DPO: Category-Adaptive Safety Alignment cites this paper.

Cat-DPO: Category-Adaptive Safety Alignment From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.750925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:33:32.642379Z digest=sha256:fea8ab52bc4b61a71f1d4175fa7a5d97e89a1e1edd236e83b537285aeeafaaa7

Observation 41e0bb4f-39d5-48a6-83fc-c31191e12e91 · inbound

Using large language models for embodied planning introduces systematic safety risks cites this paper.

Using large language models for embodied planning introduces systematic safety risks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:20.104188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:45:46.483867Z digest=sha256:bfd58775c0bf1c27b6de30a6694de8bd3c87d358242d7ad99959ba718f2d8c30

Observation b0bec346-0217-440a-9c94-502e42612725 · inbound

Jailbreaking Frontier Foundation Models Through Intention Deception cites this paper.

Jailbreaking Frontier Foundation Models Through Intention Deception From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.010147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:17:51.039062Z digest=sha256:1532861bf211797fb3d93b1d6d1a7f0f4f67bebcf2330911ad5251993d4b37e3

Observation aec6227f-8076-41dc-94a9-9d0babd9efd4 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.563625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:7cb2dbbfcabe36dfd4ac362038041f6b27947830851c9f73916c64a2c79d0a36

Observation 27dddf4d-e254-4f60-953b-8ff3632f9d83 · inbound

Internalizing Safety Understanding in Large Reasoning Models via Verification cites this paper.

Internalizing Safety Understanding in Large Reasoning Models via Verification From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:51:14.238142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:50:59.283409Z digest=sha256:8110f0955d3ca8f689e00704a8af9d2439349b1714241f4f6da9a306ff69f458

Observation 5d489190-aa42-498a-9084-49fc23f6703c · inbound

Reducing Political Manipulation with Consistency Training cites this paper.

Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.233116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:32:19.312335Z digest=sha256:0c8ed7529f4acbe1cc81861cb38e4ed8ad36d5b4f6fed265e001b7842958ffa4

Observation 27d5e496-392f-4511-a39f-bf7b49040453 · inbound

Reducing Political Manipulation with Consistency Training cites this paper.

Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.727042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:49:16.542582Z digest=sha256:780c6ca0c5cfd41c6b268a2b1e22e4bd0fdc2f358ea95531f1b2ca60fd437bbd

Observation 1a845245-3a9c-46d1-9b9b-cdb20a841c83 · inbound

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories cites this paper.

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.785841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:33:26.600072Z digest=sha256:2f04d66ba2ee0c749c05befd60289efedb9e4e0834ab8b6dbf4e0b5ffd57fcef

Observation d966b265-c9f3-4000-b32b-f8565878dee3 · inbound

Investigating and Alleviating Harm Amplification in LLM Interactions cites this paper.

Investigating and Alleviating Harm Amplification in LLM Interactions From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:24.040929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:31:52.027889Z digest=sha256:35f12a15fe4a4dab9e850fa82f955118880f57bd3d4abeb50d1c7b7320b40df0

Observation fed5df3e-33d3-415e-967a-4a6a64015832 · inbound

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability cites this paper.

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.803972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:00:30.904247Z digest=sha256:2845cb961e07166a9e58d6becb58a34c83b4b557e1ab87e8ae4b5779fd2df4ab

Observation 28e34803-b869-4ca7-93d5-9fb1e5c85b01 · inbound

Understanding Censorship in Large Language Models: From Mechanisms to Governance cites this paper.

Understanding Censorship in Large Language Models: From Mechanisms to Governance From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.397555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T07:02:35.840261Z digest=sha256:c44ea9b1afe5d25af5a85c0cf7faaee679d2a3cbee721eec316362b431dbfe47

Observation 8032e4dd-60b3-47f1-a891-572cd3fd65ad · inbound

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets cites this paper.

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.496890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T14:44:57.205766Z digest=sha256:18a9ab678dae3dd93fe079ef88157d455e4a1d8a497b56da2c91698a8e357783

Observation 13e43977-b44e-4383-bfa6-66195151621c · inbound

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models cites this paper.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:9517afe2b56eef5fa4a805ee0b6fbcdca81b8614b4be8065f131534c083d52cd

Observation b3b09427-a1e8-4a06-a09e-bb0a28d19412 · inbound

GPT-Red: Automated Red Teaming via Self-Play at Scale cites this paper.

GPT-Red: Automated Red Teaming via Self-Play at Scale From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.184834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.184834Z digest=sha256:7fabcea4add8efe70c7e5770724018312883437396a81187dea6fd395a13cc5e

Observation 3a138c4e-30e0-4b6a-8e2f-7bead74052f4 · inbound

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting cites this paper.

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:35:58.478085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:35:58.478085Z digest=sha256:82958c3ebe88645b0fce98501e8937d603e9b1ad56d6712d325ac9498f4b7b1e

Observation 9486a6b5-410d-471b-afd1-06341419d6b5 · inbound

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs cites this paper.

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:22.224180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:25:22.224180Z digest=sha256:c50c1c5c4445405f93011c8a54b56c413e5c0008fd5c27a947799e4cfeaef91a