Pith. sign in

REVIEW 9 cited by

Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11555 v1 pith:6AYQ6KWS submitted 2025-02-17 cs.AI

classification cs.AI
keywords safetydataalignmentapproachhelpfulnessllmsmodelsrlhf
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Fine-tuning large language models (LLMs) based on human preferences, commonly achieved through reinforcement learning from human feedback (RLHF), has been effective in improving their performance. However, maintaining LLM safety throughout the fine-tuning process remains a significant challenge, as resolving conflicts between safety and helpfulness can be non-trivial. Typically, the safety alignment of LLM is trained on data with safety-related categories. However, our experiments find that naively increasing the scale of safety training data usually leads the LLMs to an ``overly safe'' state rather than a ``truly safe'' state, boosting the refusal rate through extensive safety-aligned data without genuinely understanding the requirements for safe responses. Such an approach can inadvertently diminish the models' helpfulness. To understand the phenomenon, we first investigate the role of safety data by categorizing them into three different groups, and observe that each group behaves differently as training data scales up. To boost the balance between safety and helpfulness, we propose an Equilibrate RLHF framework including a Fine-grained Data-centric (FDC) approach that achieves better safety alignment even with fewer training data, and an Adaptive Message-wise Alignment (AMA) approach, which selectively highlight the key segments through a gradient masking strategy. Extensive experimental results demonstrate that our approach significantly enhances the safety alignment of LLMs while balancing safety and helpfulness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Step-Level Preference Learning for Generative Agents in Social Simulations

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Step-level human preference data collected via SimPref, then SFT+DPO, improves long-horizon social-simulation behavior of open-weight LLM agents on held-out events.

  2. SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems

    cs.CL 2026-03 conditional novelty 6.0 of 10

    A two-stage Safe-SFT + Safe-GDPO training framework reduces personalized safety violations in conversational movie and game recommendation to near-zero on the authors' new SafeRec benchmark.

  3. Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

    cs.LG 2025-06 conditional novelty 6.0 of 10

    HC-RLHF returns an aligned language model only after a held-out safety test certifies, with probability at least 1-delta, that expected harm (as judged by a learned cost model) is below a chosen threshold.

  4. Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A new benchmark shows that top reasoning models identify all relevant risks in under 40% of cases even when their final answers look safe.

  5. USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

    cs.CR 2025-05 conditional novelty 6.0 of 10

    USB-SafeBench is a unified MLLM safety benchmark with 61 risk categories, 4 modality combinations, and dual-language vulnerability and oversensitivity tests.

  6. Safe Inference-Time Alignment via Lagrangian Reward Augmentation

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Dualizing Safe RLHF yields a one-dimensional convex calibration of λ that defines a drop-in safety-aware reward for Best-of-N and token-level inference-time decoders.

  7. Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation

    cs.CR 2026-02 conditional novelty 5.0 of 10

    A unified benchmark comparing RAG knowledge-extraction attacks and defenses, showing query diversity boosts extraction, embedding attacks fail to transfer, and graph indexing raises per-token leakage.

  8. Get Experience from Practice: LLM Agents with Record & Replay

    cs.LG 2025-05 reject novelty 4.0 of 10

    AgentRR is a proposed paradigm that records agent traces, generalizes them into multi-level experiences, and replays them under safety checks to make LLM agents cheaper, faster, and more reliable.

  9. Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI

    cs.LG 2025-08 reject novelty 3.0 of 10

    A survey of reinforcement learning in healthcare that frames RL as a paradigm shift from prediction to agentive clinical intelligence.

Pith tools