Reasoning models from SFT, RL post-training and distillation exhibit alignment regressions versus matched instruction-tuned baselines on safety, toxicity, bias, ethics, privacy and robustness.
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
2 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.
2
Pith papers citing it
5
external citations · OpenAlex
citation-role summary
background 1
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
Activation steering is cast as constrained optimization that minimizes collateral damage by weighting perturbations according to the empirical second-moment matrix of activations instead of assuming isotropy.
citing papers explorer
-
Does Reasoning Preserve Alignment? On the Trustworthiness of Large Reasoning Models
Reasoning models from SFT, RL post-training and distillation exhibit alignment regressions versus matched instruction-tuned baselines on safety, toxicity, bias, ethics, privacy and robustness.
-
Minimizing Collateral Damage in Activation Steering
Activation steering is cast as constrained optimization that minimizes collateral damage by weighting perturbations according to the empirical second-moment matrix of activations instead of assuming isotropy.