{"id":"17cbe89c-a198-40b2-adc5-92bc695225d3","arxiv_id":"1909.00076","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A Koopman-based method that feeds known nonlinear terms as external forcings forecasts chaotic spatiotemporal systems for roughly 4 to 8 Lyapunov timescales.","lead":"Researchers built a data-driven model that forecasts chaotic flows by treating the nonlinear part of the physics as an external input, and tested it on benchmark chaotic systems and a lid-driven cavity flow. The model predicts the exact evolution for several Lyapunov times before errors grow, and it beats a standard Koopman-based method.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported forecast horizons depend on the user supplying the correct nonlinear forcing dictionary; with a misspecified one (e.g., Lorenz-96 with J=0) the skill drops sharply, so the 'data-driven' claim is conditional on a physical prior.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the method assumes the user can specify forcing functions that capture the nonlinearity, and for Lorenz-96 the choice X_j^2 alone gives short horizons. The paper itself supplies the strongest evidence for this concern in Sec. II.B after Eq. (12) and in Fig. 4(a), where the highest-order term is explicitly left to investigator speculation and where the J=0 case performs poorly. This does not invalidate the method, but it makes the central claim conditional: M2 is effective when the forcing dictionary is physically informed, not when the nonlinearity is learned entirely from data. The CONDITIONAL verdict already reflects this, and no further adjustment is needed. A single ablation test using the PCC-based automatic dictionary selection would settle whether the prior can be supplied in a data-driven way or whether the reported skill depends on the manual choice of the correct nonlinear terms.","tokens_in":16257,"tokens_out":10932,"duration_ms":106399,"concrete_test":"Using the released GitHub code, rerun the Lorenz-96 F=16 case with the forcing dictionary selected automatically from Eq. (13): include all second-order monomials among variables whose PCC with X_dot_I exceeds 0.1 (as in Fig. 4b), and include no higher-order terms. Compare the resulting Λmax t_l and Eave with the manually chosen J=2 row of Table II (t_l ≈ 4.05, Eave ≈ 7.67). If the automatically selected dictionary reproduces those values, the prior-knowledge dependence is mitigated; if the horizon drops materially, the reported skill is tied to user-specified nonlinear terms and the central claim should remain conditional on that prior.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that M2 makes accurate data-driven forecasts for several Lyapunov timescales—requires that the user choose forcing functions f that capture the true nonlinearity. The paper is explicit about this: Sec. II.B (after Eq. 12) states that the highest-order term required for the forcing vector can be determined by the investigator's speculation or intuition, and Fig. 4(a) shows that for Lorenz-96 with F=16, using X_j^2 alone (J=0) collapses the prediction horizon below the reported 4.05/Λmax. The PCC criterion in Eq. (13) identifies which neighboring grid points enter the nonlinear terms, but it does not determine the monomial order or the functional form. For the K-S equation, the choice f_i=(u_i)^2 works because u u_x = (u^2)_x/2, but this is also equation knowledge. For a truly unknown system, a misspecified dictionary—for example, one lacking cross terms such as X_{j-1}X_{j+1}—will not deliver the reported horizons. This is not an internal inconsistency, but it narrows the headline claim: M2 is not a fully autonomous data-driven forecaster; its predictive skill is contingent on an external prior about the nonlinearity. The absence of robustness tests over different forcing dictionaries makes it difficult to judge how often that prior is satisfiable and how quickly the method degrades when the dictionary is wrong.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a data-driven spatiotemporal forecasting framework, M2, based on a finite-dimensional Koopman approximation. Vector-valued delay-embedded observables form the state, while nonlinear terms are modeled as exogenous forcing functions f in the linear state-space model u_{m+1} = A u_m + B f_m, with A and B estimated using dynamic mode decomposition with control (DMDc). The forcing dictionary is chosen using physical knowledge: u^2 for the Kuramoto-Sivashinsky equation, monomials X_j X_{j+k} for Lorenz-96, and Reynolds stresses for the lid-driven cavity flow. The method is evaluated on Lorenz-63, K-S equations with domain lengths L = 22 to 1600, Lorenz-96 with F = 4, 8, and 16, and 2D lid-driven cavity flow at Re = 20000 and 30000. Forecast horizons are reported in Lyapunov timescales for the chaotic systems and advective timescales for the cavity flow, with the central claim that M2 provides accurate forecasts for several Lyapunov timescales and modestly outperforms the reservoir computing approach of Pathak et al.","tokens_in":16605,"tokens_out":5358,"duration_ms":46086,"significance":"If the reported forecast horizons are representative, M2 is a useful contribution to short-term spatiotemporal forecasting of high-dimensional chaotic systems. The paper has clear strengths: evaluation on twenty independent test sets, direct comparison with a recent reservoir computing benchmark, physically interpretable forcing construction, and an eigenvalue analysis that gives a mechanistic picture of unstable modes balanced by nonlinear terms. The principal limitation is that the method's skill depends on a correctly specified forcing dictionary that is supplied by the user rather than learned from data, which makes the 'data-driven' framing incomplete. That limitation is acknowledged in the text but not quantified or stress-tested, and it materially affects the scope of the central claim.","major_comments":[{"comment":"The predictive skill of M2 depends critically on the user-supplied forcing dictionary, which is not learned from data. The paper states after Eq. (12) that the highest-order term required for constructing the forcing vector 'can be determined by the investigator's speculation or intuition,' and Fig. 4(a) shows that for Lorenz-96 with F = 16, omitting the J >= 2 cross terms (J = 0) drastically reduces the prediction horizon. For the K-S equation, the choice f_i = u_i^2 is motivated by the identity u u_x = (u^2)_x / 2, which is equation knowledge rather than a data-driven discovery. As written, the headline claim of a 'data-driven' forecasting framework is therefore overstated: the reported horizons are conditional on an external physical prior about the functional form of the nonlinearity. I recommend either adding a systematic study of dictionary misspecification (for example, random subsets of monomials or missing cross terms) or explicitly reframing the contribution as a physics-assisted Koopman forecasting method.","section":"Sec. II.B, Eq. (12), and Fig. 4(a)"},{"comment":"The comparison with the reservoir computing results of Pathak et al. and the statement that M2 'modestly outperforms' that approach are based on point estimates without uncertainty quantification. Although errors are averaged over 20 testing sets, no standard deviations, confidence intervals, or per-realization spreads are reported for t_l or E_ave. Because the best parameters (r, p, q) are selected by a validation search, the reported values are selected maxima and are susceptible to selection bias; the margin over Pathak et al. (roughly 8/Lambda_max versus 6/Lambda_max) may lie within the spread across realizations. Reporting the distribution of t_l and E_ave over the 20 test sets, and ideally over bootstrapped training sets, is necessary to support the comparative claim.","section":"Table I and Sec. III.B"},{"comment":"For the lid-driven cavity flow the averaged errors are high: E_ave = 16.6% at Re = 20000 and 16.8% at Re = 30000, with prediction horizons of 5.2 and 2.7 advective timescales. The abstract and conclusions describe the performance as 'accurate' and 'similar' to the chaotic test cases, but a 16-17% mean relative error is not in the same range as the 3-11% errors reported for the K-S and Lorenz systems, and the cavity results are not compared with a baseline method. The claim should either be qualified with a discussion of what 'accurate' means for this application or supported by an error-versus-time curve analogous to Fig. 2(b) and a comparison with M1 or another baseline.","section":"Figs. 12-14 and Sec. III.D"}],"minor_comments":[{"comment":"There are multiple grammatical and typographical errors, including 'introduce a data-driven method and shows' in the abstract and 'an reservoir computing' in the introduction; a careful proofread is needed.","section":"Abstract and Sec. I"},{"comment":"The dimensions and roles of the SVD factors in the DMDc formulas would benefit from a more explicit statement; in particular, the partition of U-tilde into U-tilde_1 and U-tilde_2 is only clear after rereading the derivation.","section":"Eq. (10) and surrounding text"},{"comment":"The caption says 'Similar to Fig. 13' but the intended cross-reference is clearly Fig. 12.","section":"Fig. 13 caption"},{"comment":"The sentence claiming that q*tau is approximately 0.2*tau_d for the cavity flow appears inconsistent with the stated values tau approximately 0.025*tau_d and q = 3, which give q*tau approximately 0.075*tau_d; please check the reported value.","section":"Sec. III.D, last paragraph"},{"comment":"There are numerous typographical errors (for example, 'choticity', 'whcih', 'unergo', 'compoenents', 'beahaviour', 'fucntion', 'fro cing', 'sudies') that should be corrected before publication.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a competent paper with promising results, but the central framing as a fully data-driven method is not supported by the manuscript as written. The main risk is that without a dictionary-learning procedure or a misspecification stress test, the reported forecast horizons are conditional on the authors' physical knowledge of each test problem. This issue is addressable by reframing the contribution as physics-assisted forecasting and adding targeted experiments, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on Khodkar et al. The headline: the method works—when you already know the nonlinearity. M2 is a delay-embedded DMDc where the nonlinear terms are supplied as exogenous forcings, and with the right dictionary it forecasts K-S, Lorenz-96, and Lorenz-63 for several Lyapunov timescales, and the cavity flow for a few advective times. The new bit is not the architecture—that's HAVOK/DMDc, and the authors say so—but the physics-based selection of the forcing terms. That's a legitimate, modest contribution.\n\nWhat's genuinely good: the tests are out-of-sample, averaged over 20 testing sets, training and testing are separated, and they report a comparison with reservoir computing where they do modestly better. They also flag the role of growing modes being suppressed by the nonlinear forcing, which is a nice physical interpretation. Code and data are on GitHub.\n\nThe soft spots are real, though. The central one is not hidden: the paper says the highest-order term in the forcing vector is chosen by the investigator's 'speculation or intuition.' For Lorenz-96, including only X_j^2 drops the horizon from 4.05 to below the useful range, so the skill is conditional on the dictionary. The PCC threshold only picks neighboring points; it doesn't tell you the monomial order. So 'data-driven' is doing less work than the abstract suggests. A robustness test over misspecified dictionaries would tell us how often the prior is satisfiable and how quickly the thing degrades. As is, we have one J=0 point.\n\nThe parameter search (r,p,q) on validation sets is standard but means the reported horizons are tuned. They don't give uncertainties on the error statistics, and the cavity flow errors sit around 16-17%, which is decent but not as clean as the dynamical-system cases. The text also has typos and one abstract grammar slip; not a substantive issue.\n\nWho's it for: people working on Koopman-based forecasting or surrogate models for flow control who can bring their own physics. As an autonomous discovery method, no. As a hybrid with a known model form, yes. The central claims hold up given the caveat.\n\nRecommendation: send to peer review. The editor should ask for a sensitivity analysis on forcing dictionaries and a clearer framing that this is a physics-informed surrogate, not a purely data-driven one. I'd engage with it.","headline":"A useful physics-informed hybrid for short-term forecasting of chaos, honestly labeled as data-driven only when the forcing dictionary is right; worth refereeing with a requested robustness analysis.","tokens_in":17113,"tokens_out":2546,"would_cite":false,"duration_ms":23167,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Koopman-based method that treats nonlinearities as external forcings predicts spatiotemporal chaos for several Lyapunov timescales.","keywords":["Koopman operator","data-driven forecasting","chaotic dynamics","Kuramoto-Sivashinsky equation","Lorenz-96 system","delay embedding","dynamic mode decomposition with control","exogenous forcing"],"falsifier":"Recompute the case labeled $J=0$ in Fig. 4(a) for Lorenz-96 with $F=8$: use the same training data and method parameters but construct the forcing vector from only $X_j^2$, with no neighboring-point products. The paper reports a short horizon; if the prediction horizon instead matches the full-forcing case, the claim that cross-term forcings are essential would be contradicted.","tokens_in":16070,"feed_emoji":"🌀","tokens_out":9410,"duration_ms":73270,"temperature":0.7,"pith_summary":"The paper proposes a data-driven method, called M2, for forecasting high-dimensional chaotic dynamics by approximating the Koopman operator on delay-embedded, vector-valued observables and treating the system's nonlinearity as an exogenous forcing term. The state update is linear in the augmented state $u_{m+1}=A u_m + B f_m$, where $f_m$ is constructed from physical knowledge of the governing equations, and $A$ and $B$ are learned jointly from data using dynamic mode decomposition with control. On the Kuramoto-Sivashinsky equation, Lorenz-96 system, and a 2D lid-driven cavity flow, the method yields accurate spatiotemporal forecasts for several Lyapunov timescales, with prediction horizons above $8/\\Lambda_{\\max}$ for moderately chaotic K-S systems and average relative errors below 7 percent. The performance modestly exceeds a reservoir computing baseline. The paper shows that a few growing linear modes are stabilized by the nonlinear forcing term, and that delay-embedding is essential, as setting the embedding dimension to one collapses the horizon.","feed_headline":"Forced linear model forecasts chaos for 8+ Lyapunov timescales","feed_subtitle":"Treating nonlinearities as external forcings extends prediction horizons in K-S, Lorenz-96, and high-Re cavity flows.","key_machinery":"The machinery is a delay-embedded state space: the observable is a window of $q$ consecutive snapshots of vector-valued measurements, arranged in a Hankel matrix, and the approximated Koopman operator acts on a rank-$r$ subspace of that matrix. The forcing extension appends a block of nonlinear functions $f$ (e.g., squared velocity components for K-S, neighboring-point products for Lorenz-96, Reynolds stresses for the cavity flow) to form an augmented data matrix, and DMDc learns the reduced linear map $A$ and forcing map $B$ by solving a least-squares problem. The key operational rule from the paper is that the optimal delay $q$ satisfies $q\\tau = O(\\tau_d)$, where $\\tau_d$ is the decorrelation time, and that the optimal truncation rank $r$ tracks the singular-value hard threshold, with the forcing rank $p$ set proportional to the ratio of forcing-vector length to state-vector length.","core_discovery":"The central claim is that a finite-dimensional linear approximation of the Koopman operator can be made to predict chaotic spatiotemporal dynamics accurately for several Lyapunov timescales, provided the nonlinearity is represented as an exogenous forcing term of known functional form. Using delay-embedded vector observables and the DMDc algorithm to learn the pair $(A,B)$, the model $u_{m+1}=A u_m + B f_m$ tracks the true trajectory across systems ranging from Lorenz-63 (more than 10 Lyapunov times with 2.10 percent average error) to the K-S equation with domain length up to 1600 and a 2D cavity flow at Reynolds numbers 20000 and 30000. The paper emphasizes that the linear operator $A$ of M2 generically possesses a few eigenvalues outside the unit circle, and that the forcing term $Bf$ suppresses the unbounded growth of these modes, analogous to the stabilizing role of nonlinear energy transfer in turbulence. Without the forcing term, the Hankel-DMD baseline (M1) produces predictions that decay to zero within about a Lyapunov time.","pith_inferences":["We infer that the method's reliance on user-specified forcing terms could be relaxed by combining its PCC-based neighbor selection with sparse regression over candidate monomials; the paper leaves the order of nonlinearity to 'investigator's speculation,' which is the main human input.","If the forced linear representation is as general as the K-S and Lorenz-96 results suggest, then the practical bottleneck for data-driven forecasting of high-dimensional chaos may be identifying the correct forcing subspace rather than fitting a nonlinear evolution operator; this hypothesis is testable by comparing M2 with a deep learning model on the same benchmarks.","The authors' speculation about compressing 3D turbulence with autoencoders before applying M2 is a natural extension; a concrete next step would be to train an autoencoder on isotropic turbulence and check whether the latent-space forced linear model retains the predictive horizon seen in the 2D cavity flow."],"forward_implications":["For K-S systems with domain length $L \\le 200$, predictions remain accurate for more than $8/\\Lambda_{\\max}$ with average error below 7 percent, which modestly exceeds the roughly $6/\\Lambda_{\\max}$ horizon of a reservoir computing baseline.","For the Lorenz-96 system with $F=8$, the prediction horizon is $8.16/\\Lambda_{\\max}$ with 6.82 percent average error, and for $F=4$ the predictions never diverge within the test window, indicating the method can sometimes recover quasiperiodic dynamics exactly.","For the 2D lid-driven cavity flow, the method predicts the vorticity field accurately for 5.2 advective timescales at $Re=20000$ and 2.7 advective timescales at $Re=30000$, corresponding to hundreds of DNS time steps.","Because the forcing terms are updated from the predicted state at each step, the trained model can be iterated forward in a closed loop without access to the true state."],"supporting_citations":[{"why":"Supplies the Hankel-DMD formulation that M1 is based on and that M2 extends by adding a forcing term.","marker":"[30]"},{"why":"DMDc algorithm used to determine the matrices A and B jointly from the augmented data.","marker":"[45]"},{"why":"HAVOK framework that motivates modeling nonlinear dynamics as a linear system with a forcing term; M2 replaces the singular-vector forcing with physics-informed terms.","marker":"[38]"},{"why":"Reservoir computing baseline whose K-S setup and prediction horizons are matched and modestly outperformed.","marker":"[18]"},{"why":"Delay-embedding theorem that justifies the use of time-delayed vector observables in the method.","marker":"[34]"},{"why":"Exact DMD formulation underlying the Hankel-DMD and DMDc computations.","marker":"[28]"}],"fun_headline_variants":["Forced linear model forecasts chaos for several Lyapunov times","Nonlinearity as external forcing extends chaos prediction horizon","Koopman with forcing: linear model tracks chaotic spatiotemporal dynamics","Model nonlinearity as forcing to predict chaos beyond one Lyapunov time","Forced Koopman model beats baseline in chaos forecasting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes the user can specify the functional form of the nonlinear forcing term; for Lorenz-96, including only $X_j^2$ yields short prediction horizons, and the paper states that the highest-order term needed can be determined by the investigator's speculation or intuition.","fun_headline_variants_meta":{"raw":{"variants":["Forced linear model forecasts chaos for several Lyapunov times","Nonlinearity as external forcing extends chaos prediction horizon","Koopman with forcing: linear model tracks chaotic spatiotemporal dynamics","Model nonlinearity as forcing to predict chaos beyond one Lyapunov time","Forced Koopman model beats baseline in chaos forecasting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2910,"prompt_tokens":884,"completion_tokens":2026,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":1940}},"tokens_in":500,"tokens_out":2026,"duration_ms":14131,"temperature":1.0,"reasoning_tokens":1940,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:02:29.106951+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the case labeled $J=0$ in Fig. 4(a) for Lorenz-96 with $F=8$: use the same training data and method parameters but construct the forcing vector from only $X_j^2$, with no neighboring-point products. The paper reports a short horizon; if the prediction horizon instead matches the full-forcing case, the claim that cross-term forcings are essential would be contradicted.","supporting_citations":[{"cited_title":"A data–driven approximation of the Koopman operator: Extending dynamic mode decomposition,","cited_arxiv_id":null,"evidence_quote":"Supplies the Hankel-DMD formulation that M1 is based on and that M2 extends by adding a forcing term."},{"cited_title":"Clustering of Series via Dynamic Mode Decomposition and the Matrix Pencil Method","cited_arxiv_id":"1802.09878","evidence_quote":"DMDc algorithm used to determine the matrices A and B jointly from the augmented data."},{"cited_title":"Nonlinear laplacian spectral analysis for time series with intermittency and low-frequency variability,","cited_arxiv_id":null,"evidence_quote":"HAVOK framework that motivates modeling nonlinear dynamics as a linear system with a forcing term; M2 replaces the singular-vector forcing with physics-informed terms."},{"cited_title":"Data-driven fore- casting of high-dimensional chaotic systems with long short-term memory networks,","cited_arxiv_id":null,"evidence_quote":"Reservoir computing baseline whose K-S setup and prediction horizons are matched and modestly outperformed."},{"cited_title":"Model reduction for ﬂow analysis and control,","cited_arxiv_id":null,"evidence_quote":"Delay-embedding theorem that justifies the use of time-delayed vector observables in the method."},{"cited_title":"Spectral analysis of nonlinear ﬂows,","cited_arxiv_id":null,"evidence_quote":"Exact DMD formulation underlying the Hankel-DMD and DMDc computations."}],"review_version":1}