Pith. sign in

REVIEW 4 cited by

B-Pref: Benchmarking Preference-Based Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.03026 v1 pith:PC6D6PGM submitted 2021-11-04 cs.LG cs.AIcs.HC

classification cs.LGcs.AIcs.HC
keywords b-prefpreference-basedbenchmarklearningrewardalgorithmsfunctionhuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) requires access to a reward function that incentivizes the right behavior, but these are notoriously hard to specify for complex tasks. Preference-based RL provides an alternative: learning policies using a teacher's preferences without pre-defined rewards, thus overcoming concerns associated with reward engineering. However, it is difficult to quantify the progress in preference-based RL due to the lack of a commonly adopted benchmark. In this paper, we introduce B-Pref: a benchmark specially designed for preference-based RL. A key challenge with such a benchmark is providing the ability to evaluate candidate algorithms quickly, which makes relying on real human input for evaluation prohibitive. At the same time, simulating human input as giving perfect preferences for the ground truth reward function is unrealistic. B-Pref alleviates this by simulating teachers with a wide array of irrationalities, and proposes metrics not solely for performance but also for robustness to these potential irrationalities. We showcase the utility of B-Pref by using it to analyze algorithmic design choices, such as selecting informative queries, for state-of-the-art preference-based RL algorithms. We hope that B-Pref can serve as a common starting point to study preference-based RL more systematically. Source code is available at https://github.com/rll-research/B-Pref.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    DROPJ trains a world-model-based MPC agent from one-shot human preferences plus safety justifications, cutting training cost and improving deployment safety in car-racing simulations.

  2. Noisy Pairwise-Comparison Random Search for Smooth Nonconvex Optimization

    math.OC 2026-01 conditional novelty 6.0 of 10

    Noisy-comparison random search reaches ε-stationarity in O(k/(p²ε²)) comparisons for smooth nonconvex objectives with k-dimensional active subspace.

  3. CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries

    cs.LG 2025-05 conditional novelty 6.0 of 10

    CLARIFY uses contrastive learning on preference data to embed trajectories, then rejection-samples queries that humans can distinguish clearly, improving offline preference-based RL.

  4. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

Pith tools