REVIEW 3 cited by
Active Learning for Nonlinear System Identification with Guarantees
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While the identification of nonlinear dynamical systems is a fundamental building block of model-based reinforcement learning and feedback control, its sample complexity is only understood for systems that either have discrete states and actions or for systems that can be identified from data generated by i.i.d. random inputs. Nonetheless, many interesting dynamical systems have continuous states and actions and can only be identified through a judicious choice of inputs. Motivated by practical settings, we study a class of nonlinear dynamical systems whose state transitions depend linearly on a known feature embedding of state-action pairs. To estimate such systems in finite time identification methods must explore all directions in feature space. We propose an active learning approach that achieves this by repeating three steps: trajectory planning, trajectory tracking, and re-estimation of the system from all available data. We show that our method estimates nonlinear dynamical systems at a parametric rate, similar to the statistical rate of standard linear regression.
Forward citations
Cited by 3 Pith papers
-
Adaptive Estimation of the Transition Density of Controlled Markov Chains
An adaptive histogram estimator with a data-driven penalty achieves oracle risk bounds for transition densities of controlled Markov chains with continuous states and actions, without smoothness or control-distributio...
-
Non-Asymptotic Bounds for Closed-Loop Identification of Unstable Nonlinear Stochastic Systems
Introduces regional excitation and proves high-probability, non-asymptotic RLS error bounds for sub-exponentially unstable nonlinear closed-loop systems, with O(sqrt(ln t/t)) convergence under global excitation.
-
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
MaxInfoRL augments Boltzmann exploration with an auto-tuned information-gain bonus and reports consistent gains over SAC, DrQ, and DrQv2 baselines on continuous control tasks.
Discussion (0). Continue with ORCID to comment.