SARA extracts rewards from cosine similarity to a contrastively learned latent of preferred trajectories and outperforms or matches baselines under label noise in continuous control benchmarks.
Sutton and Andrew G
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
SARA extracts rewards from cosine similarity to a contrastively learned latent of preferred trajectories and outperforms or matches baselines under label noise in continuous control benchmarks.