For continuous-time policy evaluation, the LSTD estimator's H1 error scales as the square root of (approximation error plus m/T), with a trajectory length that can be nearly linear in the number of basis functions when the diffusion semigroup is regular.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Statistical guarantees for continuous-time policy evaluation: blessing of ellipticity and new tradeoffs
For continuous-time policy evaluation, the LSTD estimator's H1 error scales as the square root of (approximation error plus m/T), with a trajectory length that can be nearly linear in the number of basis functions when the diffusion semigroup is regular.