Non-asymptotic regret bounds for deep-network reward estimators from pairwise comparisons, with a Tsybakov-style margin condition that accelerates the rate.
Without loss of generality, we also define classes of sub-networks ofFDNN, that is,{F1,F 2,···,F |A|}, with non-sharing hidden layers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Learning Guarantee of Reward Modeling Using Deep Neural Networks
Non-asymptotic regret bounds for deep-network reward estimators from pairwise comparisons, with a Tsybakov-style margin condition that accelerates the rate.