Pith. sign in

REVIEW

Reinforcement-Learning-Guided Data-Driven Estimation of Spectral Properties of Stochastic Koopman Semigroups

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2509.04265 v3 pith:2EO43TNK submitted 2025-09-04 math.DS

Reinforcement-Learning-Guided Data-Driven Estimation of Spectral Properties of Stochastic Koopman Semigroups

classification math.DS
keywords koopmansdmdstochasticregionsspectralanalysisapproximatedata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Koopman spectral analysis turns nonlinear stochastic dynamics into a linear evolution of observables and gives access to decay rates, oscillatory modes, and metastable behavior. In practice, however, EDMD, SDMD, and related estimators depend strongly on where the trajectory data are collected. If most trajectories start in regions that carry little spectral information, the leading eigenvalues and eigenfunctions can be poorly estimated even with a rich dictionary. We propose \emph{Reinforced SDMD}, a data-acquisition method that couples Stochastic Dynamic Mode Decomposition with reinforcement learning. The RL agent chooses trajectory-initialization regions, SDMD updates the Koopman approximation, and a spectral-consistency reward evaluates the estimated eigenpairs on the newly generated data. An exploration bonus is added to avoid repeatedly sampling only a small part of the state space. We test multi-armed bandits, DQN, and PPO on stochastic double-well, Duffing, and FitzHugh--Nagumo systems. The learned policies place more samples in regions that are useful for estimating the leading Koopman eigenpairs. We also give an error-propagation analysis showing how SDMD operator error enters the corresponding bandit, approximate value-iteration, and approximate policy-iteration bounds.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.