Non-asymptotic regret bounds for deep-network reward estimators from pairwise comparisons, with a Tsybakov-style margin condition that accelerates the rate.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Learning Guarantee of Reward Modeling Using Deep Neural Networks
Non-asymptotic regret bounds for deep-network reward estimators from pairwise comparisons, with a Tsybakov-style margin condition that accelerates the rate.