Establishes stability of belief filters to model error in log-linear and neural-softmax POMDPs under mixing conditions and derives finite-sample guarantees for preference-based reward learning that decouple statistical error from model-mismatch bias.
and Yu, Bin , Date-Added =
4 Pith papers cite this work, alongside 398 external citations. Polarity classification is still indexing.
years
2026 4verdicts
UNVERDICTED 4representative citing papers
EM algorithm for two-component exponential mixtures converges at sub-exponential rate in O(log n) iterations under generalized separation assumptions.
A gradient method alternating short GD steps and long Polyak steps achieves local linear convergence for overparameterized GMMs under mixture-weight assumptions.
Entropic optimal transport yields a clustering loss with the same global optimum as log-likelihood but a better-behaved optimization surface, outperforming standard EM in experiments.
citing papers explorer
-
Preference-Based Reward Learning under Partial Observability with Inexact Dynamics
Establishes stability of belief filters to model error in log-linear and neural-softmax POMDPs under mixing conditions and derives finite-sample guarantees for preference-based reward learning that decouple statistical error from model-mismatch bias.
-
Global convergence analysis of mixtures of Exponential densities
EM algorithm for two-component exponential mixtures converges at sub-exponential rate in O(log n) iterations under generalized separation assumptions.
-
Local linear convergence of gradient methods for overparameterized Gaussian mixtures
A gradient method alternating short GD steps and long Polyak steps achieves local linear convergence for overparameterized GMMs under mixture-weight assumptions.
-
On Model-Based Clustering With Entropic Optimal Transport
Entropic optimal transport yields a clustering loss with the same global optimum as log-likelihood but a better-behaved optimization surface, outperforming standard EM in experiments.