Safe-completions train models to maximize helpfulness within safety constraints using a composite reward, and the paper reports safety and helpfulness gains over refusal-based training in GPT-5.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Safe-completions train models to maximize helpfulness within safety constraints using a composite reward, and the paper reports safety and helpfulness gains over refusal-based training in GPT-5.