{"id":"59cd647c-b8df-425f-b1c1-2a85f2d33f06","arxiv_id":"1908.07725","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Wiener projections unify the Koopman, Mori-Zwanzig, and Wiener-filtering views of data-driven model reduction and yield a heuristic derivation of NARMAX models.","lead":"The paper introduces a new projection operator, called the Wiener projection, that connects Koopman/Mori-Zwanzig model reduction theory to the widely used NARMAX time-series models. It demonstrates the framework on two hard test problems: the chaotic Kuramoto-Sivashinsky equation and a stochastically forced Burgers equation.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dissipative extension of the Wiener projection is asserted without proof, and the KS/Burgers examples lie outside the unitary-Koopman assumption used to derive the strong orthogonality (3.6b).","rationale":"The reader identified the weakest assumption correctly: the whole Wiener projection construction in Section 3.1 is built on unitarity of the Koopman operator, while the numerical demonstrations are for dissipative systems. I agree that this is the single most load-bearing issue, because the new object in the paper—PW with the strong orthogonality (3.6b)—is exactly what distinguishes the Wiener projection from an ordinary finite-rank MZ projection. If the extension to non-invertible F is only heuristic, the central claim is reduced to 'causal Wiener filtering is a useful way to motivate NARMAX,' which is much weaker than the advertised unification with MZ theory. The proposed EDMD-style check targets the operator identity (3.7) directly; it is not a check of the fitted NARMAX models, whose residuals are only approximately orthogonal due to the rational approximation. Thus the concern and test address the theoretical claim, not the quality of the numerics. The paper does contain independent support: the unitary case is derived carefully, the numerical implementation is described in detail, and the results are compared against full models. These justify conditional acceptance with the stated conditions. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":37425,"tokens_out":12668,"duration_ms":136384,"concrete_test":"On a long stationary trajectory of the 108-mode KS model, build the empirical dictionary V_L = span{Ψ(x_n),...,Ψ(x_{n-L})} with L ≈ 500. Use least-squares/EDMD to estimate the Koopman matrix M restricted to V_L, and compute the empirical operator norm of P_{V_L} M (I - P_{V_L}) on this subspace. The unitary theory predicts P M Q = 0 (Eq. 3.7); a value of this norm that is not small compared with the norm of M would show that the key identity fails for the dissipative KS system, so the strong orthogonality (3.6b) is not justified. If the norm is zero to within sampling error, the empirical conclusion survives, though a proof for non-unitary M would still be needed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 defines W = span(Ψ ∪ M^{-1}Ψ ∪ M^{-2}Ψ ∪ ...) and uses the identity (3.4) to conclude Q_W M^{-ℓ}P_W = 0, hence P_W M^ℓ Q_W = 0, and then M Q_W = Q_W M Q_W (Eq. 3.7). The step from Q_W M^{-1}P_W = 0 to P_W M Q_W = 0 uses (M^{-1})* = M, i.e., unitarity of M. This is exactly the assumption stated as 'assume F is invertible so that M is invertible and unitary.' For the KS time-δ map and the stochastically forced Burgers equation, F is not invertible (and for Burgers the state is augmented by an infinite forcing history). The single sentence in Section 3.1 asserting that the formalism 'can be safely applied to dissipative dynamical systems' argues only that one need not compute F^{-1} in practice; it does not supply the missing operator identity. Without Eq. (3.7), the Dyson formula does not collapse to (3.5b), the noise is not shown to be a deterministic shift of ξ_1 (Eq. 3.8d), and the strong orthogonality (3.6b)—the property that separates the Wiener projection from an ordinary finite-rank MZ projection—is not established for the demonstrated systems. The numerical successes could then be interpreted as good ARMA/NARMAX fits rather than instances of the claimed exact Wiener projection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a framework for data-driven model reduction that connects the Mori-Zwanzig formalism with Wiener filtering. It defines a 'Wiener projection' PW onto the span of present and past observables and, under the assumption that the Koopman operator M is invertible and unitary, proves an exact decomposition x_{n+1} = Σ Ψ(x_{n-k})·h_k + ξ_{n+1} with the strong orthogonality ⟨ξ_n, Ψ(x_m)⟩ = 0 for n > m. It then heuristically derives a NARMAX model from a rational approximation H(z) ≈ B(z)/A(z) of the filter transfer function, and tests the resulting reduced models on a Kuramoto-Sivashinsky PDE and a stochastically forced Burgers equation, reporting good short-time forecasting and long-time statistical agreement.","tokens_in":37786,"tokens_out":4608,"duration_ms":47509,"significance":"If fully established in the demonstrated settings, the paper would provide a useful conceptual bridge between Koopman operator theory, the Mori-Zwanzig projection formalism, and Wiener filtering / NARMAX data-driven modeling. The unitary-Koopman derivation in Section 3.1 is clean and elegant, and the numerical studies are carefully executed, with ensemble confidence intervals and comparisons against Galerkin truncations. The paper is honest in labeling the NARMAX step as heuristic and in discussing limitations in Section 6. However, the two numerical examples lie outside the assumptions under which the central orthogonality property is proved, so the theoretical contribution needs to be reconciled with the demonstrated scope of applicability.","major_comments":[{"comment":"The strong orthogonality result (3.6b) and the stationarity property ξ_n = M^{n-1}ξ_1 (Eq. (3.8d)) are derived by taking adjoints in Eq. (3.4), which uses (M^{-1})* = M, i.e., the unitarity of M. The paper explicitly assumes 'F is invertible so that M is invertible and unitary' at the start of Section 3.1, but then asserts in the same section that the formalism 'can be safely applied to dissipative dynamical systems' because computing F^{-1} is not needed. That assertion addresses numerical implementation, not the missing operator identity. For the Kuramoto-Sivashinsky and Burgers examples, M is not unitary, so Eq. (3.7) and the resulting strong orthogonality (3.6b) are not established. The paper should either prove the needed identities under weaker assumptions or explicitly state that the exactness claims hold only for invertible/unitary systems and that the numerical results are heuristic extensions.","section":"Section 3.1, Eqs. (3.6)-(3.8)"},{"comment":"The NARMAX derivation rests on the 'uncontrolled' approximation H(z) ≈ B(z)/A(z). No error bound or precise approximation space (e.g., a norm on transfer functions) is provided, and the orders p and r are selected by trial and error in Section 5.1. As written, the paper provides heuristic motivation rather than a derivation: the optimal Wiener filter in Eq. (3.6a) is replaced by a rational-ansatz model, and the difference between the two is not quantified. Since the claimed unification of NARMAX with Wiener filtering rests on this step, the paper should either supply a rigorous approximation result with a bound or explicitly reframe the contribution as showing consistency between the NARMAX ansatz and the Wiener projection, not a derivation.","section":"Section 3.2, Eq. (3.13)"},{"comment":"The comparison between linear and nonlinear regression is confounded by model order. The text states that no (p,r) pair produced stable models for both procedures, so the comparison uses p = r = 1 for nonlinear regression and p = 1, r = 0 for linear regression. The observed performance gap could therefore be due to the different lag structure rather than to the loss function. The claim that nonlinear regression is better than linear regression for the KS equation should be supported by a matched-order comparison (for example, with regularized estimation to stabilize the linear method) or by explicitly stating that the comparison is not designed to isolate the effect of the loss function.","section":"Section 5.1, Fig. 3"},{"comment":"The Burgers reduced model is fit and evaluated using the same forcing realization w_n that drives the full model. This correlates the full and reduced models during fitting (as the authors note) and gives the reduced model knowledge of the true future forcing path during the response-forecasting tests. The forecasting skill in Fig. 5(b) is therefore not an assessment of a standalone data-driven reduced model but of a reduced model with an oracle or data-assimilation-like input. To support the paper's 'data-driven' framing, the authors should report performance without shared forcing in the main text, or clearly separate the 'response prediction given forcing' setting from the pure model-reduction setting.","section":"Section 5.2, Eq. (5.7)"}],"minor_comments":[{"comment":"The phrase 'decaying meory condition' should read 'decaying memory condition'; the typo appears twice in this section.","section":"Section 4.1, paragraph 2"},{"comment":"'data-driven driven model reduction' should read 'data-driven model reduction'.","section":"Section 2.1, last paragraph"},{"comment":"The sentence 'Our construction here is related to the \"shift operator\" discussed in [42]' is vague; please specify which construction is related and what the precise connection is.","section":"Section 3.3, first paragraph"},{"comment":"The indexing conventions in B(z) and the recurrence (3.14b) are easy to misread because the shifts are nonstandard; a short worked example with explicit low-order polynomials would improve readability.","section":"Section 3.2, Eqs. (3.13)-(3.14)"},{"comment":"The symbol N is used both for the data length and as the upper limit of the sum; the index convention relative to the transient period in Section 4.1 should be stated explicitly.","section":"Equation (3.16)"},{"comment":"The code availability statement says the source code 'is being prepared for public release,' so the numerical results are not yet reproducible; please include a release URL or a code snapshot with the final version.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the gap between the unitary-Koopman theory and the dissipative examples. If the authors can supply an additional argument for the non-unitary case (or carefully qualify the claims as heuristic in those cases), the paper could become a solid contribution. The scope fits math.NA and nonlinear dynamics. I recommend major revision rather than rejection because the conceptual framework is valuable and the numerical results are plausible, but the load-bearing theoretical gap and the confounded model-comparison need to be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth reading, but go in knowing that the headline theoretical result is proven only for invertible, unitary Koopman operators. The Wiener projection—orthogonal projection onto the span of Ψ and its M^{-1}-shifts—is a genuinely nice idea. When F is invertible, the calculation is clean: the Dyson formula collapses, the noise ξ_n becomes a deterministic shift of ξ_1, and you get the strong orthogonality ⟨ξ_n, Ψ(x_m)⟩=0 for n>m. That's a real unification of Mori-Zwanzig, Wiener filtering, and NARMAX, and as far as I know the explicit rational-approximation route from MZ to NARMAX is new.\n\nThe soft spot is exactly where the stress-test note lands. Section 3.1 assumes F invertible so that M is unitary. The step from Q_W M^{-1}P_W=0 to P_W M Q_W=0 uses (M^{-1})* = M. For the KS time-δ map and the stochastic Burgers dynamics, F is not invertible; the one-sentence assurance that the formalism 'can be safely applied' because you don't need to compute F^{-1} in practice is not a proof. Without Eq. (3.7), the strong orthogonality (3.6b) is not established for the examples, and the numerical results could be interpreted as well-tuned ARMA/NARMAX models rather than instances of the exact Wiener projection. To the authors' credit, they call the NARMAX derivation 'heuristic' and explicitly flag the rational approximation H≈B/A as uncontrolled. But the gap is significant because the strong orthogonality is what separates the Wiener projection from an ordinary finite-rank MZ projection.\n\nWhat the paper does well: the Hilbert-space argument in the unitary case is clean, the cascade-form implementation with the decaying memory constraint is thoughtful, and the numerical study on KS and stochastic Burgers is detailed, with ensemble confidence intervals, energy spectra, ACFs, and a careful linear-vs-nonlinear regression comparison. The authors are honest about order selection by trial and error and about the approximation.\n\nWho is this for: anyone working on data-driven model reduction, Koopman operator methods, or MZ closures. It gives a useful common vocabulary and a concrete starting point for connecting time-series models to dynamical-systems theory. The examples are nontrivial and the results are plausible.\n\nMy recommendation: send it to review, but a referee should push on the dissipative extension. Either the authors prove something for non-invertible maps under reasonable assumptions, or they should explicitly restrict the theoretical claims to invertible systems and present the KS/Burgers results as a heuristic extension supported numerically. As written, it's a good but incomplete paper.","headline":"The Wiener projection cleanly links Koopman-Mori-Zwanzig and NARMAX, but the paper's strongest orthogonality result is proven only for invertible/unitary systems, and the dissipative examples step outside that assumption.","tokens_in":38271,"tokens_out":4618,"would_cite":true,"duration_ms":117695,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["37M10","60G10","93E11"],"pacs":[],"model":"deepseek-v4-flash","headline":"A Wiener projection onto the span of a resolved observable and all its pullbacks turns data-driven reduced models into optimal linear filters, with NARMAX as the rational approximation.","keywords":["model reduction","Wiener projection","Mori-Zwanzig formalism","Koopman operator","NARMAX","non-Markovian dynamics","Kuramoto-Sivashinsky equation","stochastic Burgers equation"],"falsifier":"Take a strongly dissipative one-dimensional map that is not invertible, run the paper's NARMAX fitting procedure with a fixed observable $\\Psi$, and measure the sample covariance between the residuals $\\xi_n$ and past observables $\\Psi(x_m)$ for $n>m$; if it is significantly nonzero, the strong orthogonality (3.6b) fails and the Wiener-projection derivation does not extend to non-invertible dynamics as claimed.","tokens_in":37230,"feed_emoji":"🎯","tokens_out":9122,"duration_ms":228773,"temperature":0.7,"pith_summary":"This paper argues that many data-driven reduced models of complex dynamical systems are really the same object: an optimal linear filter in a chosen space of observables, whose transfer function is then approximated by a ratio of polynomials. The authors define a Wiener projection onto the span of the observable map and all its pullbacks under the inverse dynamics, and show it makes the Mori-Zwanzig memory expansion exact with a strong orthogonality condition: the residual noise at each step is uncorrelated with every past observable. From this, the widely used NARMAX model follows as the rational-approximation step, giving such empirical models a dynamical-systems justification. If this picture is right, then fitting a reduced model is not an ad hoc regression but a principled approximation of an optimal filter, and classical Wiener filtering becomes an alternative foundation for model reduction alongside Mori-Zwanzig theory.","feed_headline":"Wiener projection unifies NARMAX and Koopman–Mori-Zwanzig","feed_subtitle":"A single projection shows NARMAX models are rational approximations of an optimal memory filter.","key_machinery":"The Wiener projection $P_W$ is the orthogonal projection onto $W = \\operatorname{span}(\\Psi \\cup M^{-1}\\Psi \\cup M^{-2}\\Psi \\cup \\cdots)$, the closed subspace generated by the vector of resolved observables $\\Psi(x)$ and all its preimages under the Koopman operator $M$. Under the assumption that the dynamics map $F$ is invertible, making $M$ unitary on $L^2(\\mu)$, the identity $M^{-\\ell}P_W = P_W M^{-\\ell}P_W$ collapses the Dyson expansion, leaving the exact one-step predictor with the strong orthogonality $\\langle \\xi_n, \\Psi(x_m)\\rangle = 0$ for $n>m$. The second piece is the rational-approximation ansatz $H(z)\\approx B(z)/A(z)$ for the $z$-transform of the filter coefficients, which turns the infinite convolution into the finite-order NARMAX recursion; the paper enforces the decaying-memory condition by factoring $A(z)$ into second-order blocks, a cascade form whose roots lie inside the unit disc.","core_discovery":"The central discovery is that the orthogonal projection $P_W$ onto $W = \\operatorname{span}(\\Psi \\cup M^{-1}\\Psi \\cup M^{-2}\\Psi \\cup \\cdots)$, where $\\Psi$ collects the resolved observables and $M$ is the Koopman operator, yields the exact decomposition $x_{n+1} = \\sum_{k\\ge 0} \\Psi(x_{n-k})\\cdot h_k + \\xi_{n+1}$ with $\\langle \\xi_n, \\Psi(x_m)\\rangle = 0$ for every $n>m$. This Wiener projection absorbs all memory effects into the subspace $W$, so the Dyson expansion collapses and the residual is orthogonal to the entire past observable history, not just a finite-dimensional subspace as in the usual Mori-Zwanzig projection. Because the coefficients $h_k$ form the impulse response of a causal Wiener filter, replacing its $z$-transform $H(z)$ by a rational function $B(z)/A(z)$ produces a NARMAX-type recursion; the paper calls this the derivation of NARMAX from the underlying dynamics. The result is demonstrated on the Kuramoto-Sivashinsky equation, where the five-mode reduced model reproduces short-time tracking, autocovariances, energy spectra, and energy cross-correlations, and on a stochastically forced viscous Burgers equation, where a nine-mode model forecasts the response of individual forcing realizations.","pith_inferences":["An editor-level test of the paper's asserted dissipative extension: compute the sample correlation between NARMAX residuals and past observables for a strongly non-invertible map; nonzero values would show the strong orthogonality (3.6b) is not automatic outside the unitary setting.","Because the observable map $\\Psi$ is left unspecified, the same construction suggests that learned observables (reservoir states, delay-coordinate maps) could be plugged into the Wiener projection, giving a principled way to add memory terms to machine-learning models of dynamics.","The rational-filter picture suggests a spectral criterion for choosing NARMAX orders: the needed number of poles and zeros should reflect the location of Koopman eigenvalues of the resolved observables relative to the unit circle.","The contrast between the Kuramoto-Sivashinsky and Burgers results suggests that the advantage of nonlinear least squares over linear regression grows with the memory depth of the unresolved dynamics; a quantitative criterion might come from comparing residual and signal power spectra."],"forward_implications":["NARMAX models acquire a dynamical-systems derivation: they are rational approximations of the Wiener filter associated with the Koopman–Mori-Zwanzig decomposition, giving a principled reason for their empirical success.","For stationary chaotic and randomly forced systems, the decomposition is exact in law before approximation, so the only modeling choices are the observable map $\\Psi$ and the order of the rational approximation.","The strong orthogonality $\\langle \\xi_n, \\Psi(x_m)\\rangle = 0$ for $n>m$ provides a least-squares optimality guarantee for reduced models that use only past observables, supporting the common practice of adding independent noise after fitting.","Classical Wiener filtering can serve as an alternative foundation for model reduction alongside Mori-Zwanzig when statistics are time-stationary.","Because the method works with the recent history of $\\Psi$, it is formally applicable to dissipative systems without computing the inverse map $F^{-1}$, as demonstrated on the Kuramoto-Sivashinsky and stochastic Burgers equations."],"supporting_citations":[{"why":"Supplies the discrete-time Mori-Zwanzig expansion that the Wiener projection specializes and simplifies.","marker":"[34]"},{"why":"Provides the Mori-Zwanzig formalism and the rational-transfer-function approximation tradition the paper builds on.","marker":"[6]"},{"why":"Earlier discrete data-driven MZ construction whose NARMAX estimation framework this paper compares with and extends.","marker":"[8]"},{"why":"Provides the Kuramoto-Sivashinsky reduced-model ansatz, data, and likelihood-based baseline that the new least-squares fits are compared against.","marker":"[53]"},{"why":"Defines the NARMAX family of time-series models that the paper derives from the underlying dynamics.","marker":"[18]"},{"why":"Supplies the Wiener filtering and spectral factorization background used in the optimal-filter interpretation and Appendix B.","marker":"[40]"},{"why":"Gives the ergodic-theory background for Koopman operators being unitary isometries when F is invertible.","marker":"[20]"}],"fun_headline_variants":["Wiener projection ties NARMAX to Koopman-Mori-Zwanzig","One projection: NARMAX as rational Wiener filter","Wiener projection reveals NARMAX from Koopman dynamics","Memory filter: Wiener projection unifies NARMAX and Mori-Zwanzig","Exact memory via Wiener projection: NARMAX explained"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire clean orthogonality and the NARMAX derivation assume the dynamics map $F$ is invertible, so that the Koopman operator $M$ is invertible and unitary; the paper's numerical successes on dissipative systems like Kuramoto-Sivashinsky rest on an asserted, not proved, extension to that case.","fun_headline_variants_meta":{"raw":{"variants":["Wiener projection ties NARMAX to Koopman-Mori-Zwanzig","One projection: NARMAX as rational Wiener filter","Wiener projection reveals NARMAX from Koopman dynamics","Memory filter: Wiener projection unifies NARMAX and Mori-Zwanzig","Exact memory via Wiener projection: NARMAX explained"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000664,"raw_usage":{"total_tokens":3086,"prompt_tokens":1050,"completion_tokens":2036,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":1944}},"tokens_in":666,"tokens_out":2036,"duration_ms":14295,"temperature":1.0,"reasoning_tokens":1944,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:58:41.781130+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a strongly dissipative one-dimensional map that is not invertible, run the paper's NARMAX fitting procedure with a fixed observable $\\Psi$, and measure the sample covariance between the residuals $\\xi_n$ and past observables $\\Psi(x_m)$ for $n>m$; if it is significantly nonzero, the strong orthogonality (3.6b) fails and the Wiener-projection derivation does not extend to non-invertible dynamics as claimed.","supporting_citations":[{"cited_title":"Darve, J","cited_arxiv_id":null,"evidence_quote":"Supplies the discrete-time Mori-Zwanzig expansion that the Wiener projection specializes and simplifies."},{"cited_title":"Zwanzig, Nonequilibrium Statistical Mechanics, Oxford, 2001","cited_arxiv_id":null,"evidence_quote":"Provides the Mori-Zwanzig formalism and the rational-transfer-function approximation tradition the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier discrete data-driven MZ construction whose NARMAX estimation framework this paper compares with and extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Kuramoto-Sivashinsky reduced-model ansatz, data, and likelihood-based baseline that the new least-squares fits are compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the NARMAX family of time-series models that the paper derives from the underlying dynamics."},{"cited_title":"Kailath, Lectures on Wiener and Kalman Filtering, Springer, 1981","cited_arxiv_id":null,"evidence_quote":"Supplies the Wiener filtering and spectral factorization background used in the optimal-filter interpretation and Appendix B."},{"cited_title":"Walters, An introduction to ergodic theory , Vol","cited_arxiv_id":null,"evidence_quote":"Gives the ergodic-theory background for Koopman operators being unitary isometries when F is invertible."}],"review_version":1}