Pith. sign in

S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Preference-based reinforcement learning (PbRL) stands out by utilizing human preferences as a direct reward signal, eliminating the need for intricate reward engineering. However, despite its potential, traditional PbRL methods are often constrained by the indistinguishability of segments, which impedes the learning process. In this paper, we introduce Skill-Enhanced Preference Optimization Algorithm (S-EPOA), which addresses the segment indistinguishability issue by integrating skill mechanisms into the preference learning framework. Specifically, we first conduct the unsupervised pretraining to learn useful skills. Then, we propose a novel query selection mechanism to balance the information gain and distinguishability over the learned skill space. Experimental results on a range of tasks, including robotic manipulation and locomotion, demonstrate that S-EPOA significantly outperforms conventional PbRL methods in terms of both robustness and learning efficiency. The results highlight the effectiveness of skill-driven learning in overcoming the challenges posed by segment indistinguishability.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

unclear 1

representative citing papers

Preference-based Multi-Objective Reinforcement Learning

cs.LG · 2025-07-18 · reject · novelty 4.0

Pb-MORL learns a multi-objective reward model from preference comparisons and claims to achieve Pareto-optimal policies, outperforming an oracle in energy and highway tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Preference-based Multi-Objective Reinforcement Learning cs.LG · 2025-07-18 · reject · none · ref 26 · internal anchor

    Pb-MORL learns a multi-objective reward model from preference comparisons and claims to achieve Pareto-optimal policies, outperforming an oracle in energy and highway tasks.