Pith. sign in

REVIEW 1 cited by

How more data can hurt: Instability and regularization in next-generation reservoir computing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08641 v3 pith:LI6O7QGS submitted 2024-07-11 cs.LG cs.NEmath.DSnlin.AO

classification cs.LGcs.NEmath.DSnlin.AO
keywords datainstabilityngrcregularizationcomputingdata-drivendynamicalhurt
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

It has been found recently that more data can, counter-intuitively, hurt the performance of deep neural networks. Here, we show that a more extreme version of the phenomenon occurs in data-driven models of dynamical systems. To elucidate the underlying mechanism, we focus on next-generation reservoir computing (NGRC) -- a popular framework for learning dynamics from data. We find that, despite learning a better representation of the flow map with more training data, NGRC can adopt an ill-conditioned ``integrator'' and lose stability. We link this data-induced instability to the auxiliary dimensions created by the delayed states in NGRC. Based on these findings, we propose simple strategies to mitigate the instability, either by increasing regularization strength in tandem with data size, or by carefully introducing noise during training. Our results highlight the importance of proper regularization in data-driven modeling of dynamical systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A CFL-type Condition and Theoretical Insights for Discrete-Time Sparse Full-Order Model Inference

    math.DS 2025-05 conditional novelty 6.0 of 10

    For 1D linear advection, a discrete-time sparse full-order model inferred by least squares is guaranteed stable only if the training data satisfy Δt/Δx ≤ (m+1)/(3c), a 'sampling CFL' bound.

Pith tools