The paper claims entropy regularization forces the optimal LQ feedback policy to be Gaussian and uses that to solve a mean-variance asset-liability problem, but the proof of the main theorem contains a correlation error that invalidates the general claim.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Reinforcement Learning for a Discrete-Time Linear-Quadratic Control Problem with an Application
The paper claims entropy regularization forces the optimal LQ feedback policy to be Gaussian and uses that to solve a mean-variance asset-liability problem, but the proof of the main theorem contains a correlation error that invalidates the general claim.