Pith. sign in

REVIEW 4 cited by

Preference Transformer: Modeling Human Preferences using Transformers for RL

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.00957 v1 pith:6553DB6Q submitted 2023-03-02 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords humanpreferencetransformerpreferencesapproachesarchitecturemodelpreference-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Preference-based reinforcement learning (RL) provides a framework to train agents using human preferences between two behaviors. However, preference-based RL has been challenging to scale since it requires a large amount of human feedback to learn a reward function aligned with human intent. In this paper, we present Preference Transformer, a neural architecture that models human preferences using transformers. Unlike prior approaches assuming human judgment is based on the Markovian rewards which contribute to the decision equally, we introduce a new preference model based on the weighted sum of non-Markovian rewards. We then design the proposed preference model using a transformer architecture that stacks causal and bidirectional self-attention layers. We demonstrate that Preference Transformer can solve a variety of control tasks using real human preferences, while prior approaches fail to work. We also show that Preference Transformer can induce a well-specified reward and attend to critical events in the trajectory by automatically capturing the temporal dependencies in human decision-making. Code is available on the project website: https://sites.google.com/view/preference-transformer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SARA extracts rewards from cosine similarity to a contrastively learned latent of preferred trajectories and outperforms or matches baselines under label noise in continuous control benchmarks.

  2. SimulPL: Aligning Human Preferences in Simultaneous Machine Translation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    SimulPL adds latency-aware preference optimization to simultaneous machine translation and reports better human-aligned quality at low latency on three language pairs.

  3. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    cs.AI 2026-07 conditional novelty 5.0 of 10

    LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.

  4. Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective

    cs.LG 2025-02 reject novelty 5.0 of 10

    The paper derives convergence rates for sigmoid gating mixture-of-experts with quadratic scores and uses them to argue sigmoid self-attention is more sample-efficient than softmax, but the link to attention is an unpr...

Pith tools