URDP couples LLM-based reward component design with uncertainty-weighted Bayesian optimization, reporting better reward quality and efficiency than Eureka and Text2Reward on three benchmarks.
(9) By performing coordinate scaling transformation on the sample points, we obtain new sample points p(i) = (˜p(i) 1 /l1,··· , ˜p(i) d /ld), i= 1, 2,··· ,n
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Uncertainty-aware Reward Design Process
URDP couples LLM-based reward component design with uncertainty-weighted Bayesian optimization, reporting better reward quality and efficiency than Eureka and Text2Reward on three benchmarks.