Pith. sign in

Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles.arXiv preprint arXiv:2401.00243

5 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

5 Pith papers citing it
1 external citations · external index

years

2026 4 2024 1

representative citing papers

A Unifying Lens on Reward Uncertainty in RLHF

cs.LG · 2026-06-08 · unverdicted · novelty 6.0

A distributional reward model p(r|x,y) yields the closed-form effective reward ilde r(x,y) = eta ext{log} ext{E}_p[e^{r/eta}] (pessimistic branch) that unifies prior RLHF aggregation heuristics under Bayesian or KL-DRO views.

citing papers explorer

Showing 5 of 5 citing papers.