Pith. sign in

Surf: Semi-supervised reward learning with data augmentation for feedback- efficient preference-based reinforcement learning

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

years

2026 3 2025 2

verdicts

UNVERDICTED 5

representative citing papers

MAPL: Multi-Objective Preference Learning for Robot Locomotion

cs.RO · 2026-06-24 · unverdicted · novelty 6.0

MAPL trains quadruped locomotion policies from LLM-generated multi-objective trajectory preferences and matches or exceeds expert-designed reward performance in four environments without manual reward engineering.

citing papers explorer

Showing 5 of 5 citing papers.