Pith. sign in

REVIEW 1 cited by

Listwise Reward Estimation for Offline Preference-based Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04190 v1 pith:TMIJ5XJM submitted 2024-08-08 cs.LG cs.AI

classification cs.LGcs.AI
keywords feedbacklirepbrlrewardlearningofflinepreferencedataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Reinforcement Learning (RL), designing precise reward functions remains to be a challenge, particularly when aligning with human intent. Preference-based RL (PbRL) was introduced to address this problem by learning reward models from human feedback. However, existing PbRL methods have limitations as they often overlook the second-order preference that indicates the relative strength of preference. In this paper, we propose Listwise Reward Estimation (LiRE), a novel approach for offline PbRL that leverages second-order preference information by constructing a Ranked List of Trajectories (RLT), which can be efficiently built by using the same ternary feedback type as traditional methods. To validate the effectiveness of LiRE, we propose a new offline PbRL dataset that objectively reflects the effect of the estimated rewards. Our extensive experiments on the dataset demonstrate the superiority of LiRE, i.e., outperforming state-of-the-art baselines even with modest feedback budgets and enjoying robustness with respect to the number of feedbacks and feedback noise. Our code is available at https://github.com/chwoong/LiRE

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries

    cs.LG 2025-05 conditional novelty 6.0 of 10

    CLARIFY uses contrastive learning on preference data to embed trajectories, then rejection-samples queries that humans can distinguish clearly, improving offline preference-based RL.

Pith tools