{"id":"99926e41-5484-4901-9db2-79110e987dc1","arxiv_id":"2608.07189","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Autoregressive rollout error in latent-space wake models is dominated by a linear drift in the oscillation phase, and can be corrected offline with one fitted parameter per latent coordinate.","lead":"A machine-learning model that predicts swirling fluid behind cylinders goes wrong on long runs mostly because it keeps the right oscillation pattern but runs slightly too fast or too slow. The paper shows this timing error can be measured in a short window and corrected with one number per coordinate, removing most of the long-horizon error.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Hilbert phase estimator is never validated against a known injected drift; with the measured drift only ~0.02-0.03 rad and 4-14% end-effect sensitivity, a synthetic recovery test is needed before the 95-98% phase-share claim is secure.","rationale":"Reading the paper in good faith, the central diagnosis is credible: the error spectrum peaks at the shedding frequency, the amplitude envelope is nearly perfect, the phase difference grows smoothly, and the correction improves with horizon. The paper is unusually self-aware, and the Re=1000 breakdown and the compression-limited caveat are stated plainly. I do not fully share the reader's emphasis on the narrowband criterion at Re=1000 as the key risk: that is an explicitly excluded boundary case, and the reported 98-99% in-band energy together with small cross terms establish well-posedness for the four analysed configurations. The more load-bearing risk is internal to the measurement: the entire quantitative edifice rests on a Hilbert phase difference that is tiny (0.02-0.03 rad) and whose estimator is only coarsely characterised (4-14% drift-rate sensitivity at the calibration start). The paper never runs a positive control with a known injected phase drift. If the finite-length Hilbert phase has an edge bias that grows toward the end of the record, the linear phi(t) and the dominant phase share could be partly an artifact of the estimator. This is checkable with the released code and would either retire the concern or confirm the headline. The reader's requested alternative estimator would also address the issue, but the synthetic injection is the more direct and more stringent test. Because this unvalidated measurement underpins the 95-98% quantitative claim, I keep the verdict at CONDITIONAL; I partially agree with the reader, since we both locate risk in the Hilbert phase reliability, but I focus on the absence of a recovery control rather than on the narrowband boundary itself.","tokens_in":28933,"tokens_out":10054,"duration_ms":95023,"concrete_test":"Construct a synthetic control per case: start from the true latent trajectory z(t) and form a 'prediction' by applying a known constant phase drift, z_pred(t) = A(t) cos(theta(t) + omega_inj t), with omega_inj chosen from the observed range |omega_d| ~ 0.5-1.5e-3 rad per unit time, keeping A(t) and the mean exactly equal to the truth. Run the full pipeline: Hilbert decomposition, variance shares, calibration-window linear fit, and phase correction of Eq. (11), using identical N_cal and extrapolation windows. Check that (i) the estimated omega_d recovers omega_inj within the paper's reported 4-14% uncertainty, (ii) the phase share is approximately 100% with a negligible cross term, and (iii) the correction removes essentially all of the latent error. If these fail, the measured small drift and the 95-98% phase share are contaminated by Hilbert edge effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is quantitative: 95-98% of rollout-error variance is phase, the predicted amplitude is correct to 0.15%, and the phase difference phi(t) grows linearly to only ~0.02-0.03 rad. This entire decomposition is produced by the Hilbert analytic-signal representation of Sec. IIC2. The paper acknowledges finite-length end artifacts and reports that trimming 5-20 samples from the start of the calibration window changes the fitted drift rate by 4-14%, but this sensitivity is measured at the beginning of the calibration window, not at the end of the rollout where the extrapolation region lies and where Hilbert edge effects are largest. More seriously, no positive control is reported in which a known phase drift is injected into the true latent trajectory and the pipeline is asked to recover it. Because the measured phi(t) is only ~0.02 rad, roughly an order of magnitude above the quoted estimator uncertainty, a small edge bias that grows toward the end of the record could contribute substantially to the apparent linear drift and to the dominant phase share. Since the correction of Eq. (11) is also defined through the same Hilbert phase, the high reported efficacy could in part be removing an estimator artifact rather than a network defect. A synthetic phase-injection control would separate these possibilities. Without it, the quantitative headline (95-98%, 0.15%, linear drift) is not fully secured, even though the qualitative conclusion, that the error is a coherent oscillation at the shedding frequency and is correctable, is plausible and independently supported by the spectral peaks and comparison baselines.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies autoregressive rollout error in convolutional-autoencoder-LSTM reduced-order models of two-dimensional bluff-body wakes. Using a Hilbert analytic-signal decomposition, it reports that 95-98% of the rollout-error variance is phase error, that the predicted limit-cycle amplitude is correct to within 0.15%, and that the accumulated phase difference grows linearly to only 0.02-0.03 rad over the rollout. It then proposes a one-parameter-per-coordinate phase-realignment correction, fitted on a calibration window, which removes 72-83% of latent error in representative runs and 91-98% of the correctable field error, with an R^2-based diagnostic that predicts success (r=0.85 over 50 networks). The paper includes two geometries, several Reynolds numbers, seed ensembles, DMD and harmonic-model baselines, a documented breakdown case at Re=1000, and a non-stationary modulated-inflow test where periodic tiling fails but the phase correction retains about half of the error.","tokens_in":29178,"tokens_out":9574,"duration_ms":84864,"significance":"If the central diagnosis is correct, the paper reframes compounding rollout error in latent-space ROMs from unstructured noise to a single interpretable defect: a slow drift of the phase of an otherwise correctly learned limit cycle. This is a substantive and practically useful reframing, and the one-parameter, self-diagnosing correction is a natural consequence. The paper's strengths are its honesty and breadth: two shedding mechanisms, multiple Reynolds numbers and latent dimensions, 50- and 40-network seed ensembles, open code and case files, explicit acknowledgment of the Re=1000 breakdown, and a non-stationary test that separates coherence from exact periodicity. The main weakness is that the quantitative headline rests on a Hilbert phase estimator that has not been validated against a known injected phase drift, and several quantitative tables are single-run entries despite large documented seed spread.","major_comments":[{"comment":"The Hilbert analytic-signal decomposition is the sole estimator behind the quantitative headline (95-98% phase share, linear drift of 0.02-0.03 rad) and behind the correction Eq. (11), yet no positive control is reported in which a known phase drift is injected into a true or synthetic latent trajectory and recovered by the pipeline. The measured drift is only about an order of magnitude above the quoted estimator jitter, and the reported end-effect check in Sec. IIC2 only trims 5-20 samples from the start of the calibration window, not from the end of the rollout where the correction extrapolates and where Hilbert edge artifacts are largest. Because Eq. (11) is defined through the same Hilbert phase, the reported 72-83% latent and 91-98% field improvements could in part consist of removing an estimator artifact rather than a network defect. Please add a synthetic phase-injection study (known linear and, if feasible, slowly varying phase drift, with varying drift magnitude and record length) reporting the bias and variance of the recovered omega_d and of the phase share, together with a quantification of end-of-record edge effects on the extrapolated correction.","section":"Sec. IIC2, Sec. IIID, Sec. IIIE"},{"comment":"The seed study in Sec. IIIK shows that drift magnitude varies widely across random initializations (error-growth ratio mean 5.7, standard deviation 4.0 over 40 seeds), yet several quantitative comparisons that feed the headline claims are single-run entries: Table II (phase share and amplitude fidelity), Table VII (calibration-length dependence), Table VIII (linear-baseline error ratios), and Table IX (periodic-tiling comparison). The manuscript acknowledges this in Sec. IVD(g) and labels the entries representative, but the abstract and conclusions present '95-98%', '0.15%', and the correction percentages as general findings. Please replace or supplement these single-run entries with ensemble medians and dispersion (e.g., 5-95% ranges) over the existing seed ensembles, or explicitly rephrase the affected conclusions to the qualitative claim that is ensemble-supported. Without this, the reader cannot distinguish typical performance from a favorable run.","section":"Sec. IIIK and Tables II, VII, VIII, IX"}],"minor_comments":[{"comment":"The column headers contain the run-together tokens 'nphase' and 'n acc'; separate the network count n from the improvement column so the ensemble columns can be parsed unambiguously.","section":"Table III"},{"comment":"The y-axis of panel (b) is labeled only ' [rad]' with no variable name; label it as phi(t) [rad].","section":"Fig. 5(b)"},{"comment":"The statement that 'the decomposition returned a phase share above 95%' across seeds should state whether 95% is the observed minimum, a rounded lower bound, or a threshold, so that it matches the abstract's '95 to 98%' wording.","section":"Sec. IIIK"},{"comment":"The abstract's Reynolds-number range 'Re = 100 to 800' conflates the circular-cylinder range (300-800) with the square-cylinder case (Re=100); please state the ranges per geometry for clarity.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually transparent about its limitations and provides open code and case files. The central risk is the unvalidated Hilbert phase estimator and the single-run quantitative entries; if the requested synthetic phase-injection control confirms the estimator and the ensemble distributions support the headline percentages, I would support acceptance. The journal scope is appropriate for this work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid paper and its central claim holds. The rollout error of CAE-LSTM latent surrogates for 2D bluff-body wakes is not unstructured noise; it is a coherent oscillation at the shedding frequency, and the amplitude/phase decomposition puts 95-98% of the variance in phase, with amplitude good to 0.15%. That reframes what long-horizon error means for these models. The one-parameter phase-alignment correction is cheap, post-hoc, and self-diagnosing via the R2 of the drift fit (r=0.85 over 50 networks). The non-stationary inflow case, where periodic tiling fails by over an order of magnitude while the phase correction still removes about half the error, is a nice existence proof that coherence, not exact periodicity, is what matters.\n\nThe paper is also honest and well executed. It reports seed ensembles for the main efficacy numbers, tests linear baselines (DMD variants and a harmonic model) and shows they are clearly worse, and flags its own limitations: the Re=1000 breakdown, the 2D restriction, the compression-limited regime where the mechanism lives, and the fact that on the square cylinder periodic tiling of raw snapshots beats the full pipeline. It also concedes that refs 34 and 35 already reached the qualitative phase-error observation, so the novelty is the quantification, the spectral account, and the correction. The citation pattern is fair.\n\nThe soft spots are real but not fatal. The 95-98% phase share rests on the Hilbert analytic-signal phase, and the stress-test note lands: the measured drift is only ~0.02-0.03 rad, trimming 5-20 calibration samples changes the fitted rate by 4-14%, and there is no positive control where a known phase drift is injected and recovered. Since the correction uses the same Hilbert phase, part of the reported efficacy could conceivably be removing estimator edge bias rather than a genuine network defect. A synthetic phase-injection test would settle this cheaply, and the paper already points to band-pass of the signal or a geometric phase as alternative estimators but does not run them. I would want that check before treating the quantitative headline as fully secure. Also, several supporting tables (II, VII, VIII, IX) are single runs without error bars; they read as representative rather than definitive, and the paper says as much.\n\nThis is for anyone working on latent-space ROMs of oscillatory or convection-dominated flows, and it deserves serious refereeing despite the caveats. I would send it out.","headline":"Rollout error in these periodic-wake ROMs is mostly linear phase drift; the central claim holds, but the exact phase share needs a synthetic-control check before it is fully secure.","tokens_in":29789,"tokens_out":2541,"would_cite":true,"duration_ms":39751,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Latent rollout error in wake reduced-order models is accumulated phase drift, not unstructured noise.","keywords":["reduced-order models","rollout error","phase drift","bluff-body wake","cylinder wake","convolutional autoencoder","LSTM","amplitude-phase decomposition"],"falsifier":"A concrete test: take a trained latent-space ROM on a periodic wake and compute the cumulative phase difference φ(t) over the rollout. If the error spectrum is broadband rather than sharply peaked at the vortex-shedding frequency, or if the amplitude and phase contributions to the error variance do not partition the error to within a fraction of a percent (with a negligible cross term), the central claim fails. Applying the same decomposition to a genuinely three-dimensional wake or to a flow without a dominant frequency should produce a phase share well below 95% and a cross term of order unity.","tokens_in":28669,"feed_emoji":"🌊","tokens_out":3236,"duration_ms":29178,"temperature":0.7,"pith_summary":"Across circular-cylinder wakes at Re=300–800 and a square-cylinder wake at Re=100, the paper finds that 95–98% of the autoregressive rollout error of a convolutional-autoencoder-LSTM reduced-order model is phase error, not amplitude error. The network reproduces the limit-cycle amplitude within 0.15% but traverses the cycle at a slightly wrong rate. This phase drift accumulates linearly to only about 0.02–0.03 rad over the rollout yet accounts for the long-horizon error growth. Because the drift is linear, it can be corrected offline with one parameter per latent coordinate, removing 72–83% of latent error in representative runs and 91–98% of correctable field error. The paper also shows that the correction's diagnostic, the signal-to-noise ratio of the phase fit, predicts success before application across 50 networks.","feed_headline":"Wake-model rollout error is pure phase drift, not noise","feed_subtitle":"In cylinder wakes, 95–98% of the error comes from a slightly wrong clock; one parameter per latent coordinate fixes it offline.","key_machinery":"The central object is the amplitude–phase decomposition of each latent coordinate via the Hilbert analytic signal, which requires the coordinate to be narrowband (98–99% of spectral energy within ±10% of the shedding frequency). The load-bearing identity is e(t)=−2A sin(φ(t)/2) sin(2πf_s t+φ(t)/2), which shows that a correct-amplitude, phase-drifting prediction inevitably produces an error spectrum peaked at the fundamental. The correction itself is a one-parameter phase realignment, $ẍz^{{(d)}}$(t)= $Â^{{(d)}}$(t) cos($ẜḢ^{{(d)}}$(t)-ω_d t)+ $ẍz^{{(d)}}$, applied offline after the rollout without feedback into the network.","core_discovery":"The paper claims that the rollout error of latent-space reduced-order models of periodic bluff-body wakes is a coherent phase drift rather than unstructured compounding noise. The error's power spectrum peaks sharply at the vortex-shedding frequency of each flow, and an amplitude–phase decomposition attributes 97–98% of the error variance to the phase of the predicted oscillation and only 2.0–2.3% to its amplitude. The model learns the geometry of the attractor almost exactly and misjudges only the traversal rate. The accumulated phase error φ(t) grows linearly with time, so each latent coordinate is characterized by a single drift rate. The identity e(t)≈−Aφ(t)sin(2πf_s t) explains why so small a phase slip produces an error peak at the shedding frequency, and it is the algebraic basis for the offline, one-parameter phase-realignment correction that removes most of the correctable error.","pith_inferences":["The phase-drift mechanism, once identified on periodic cylinder wakes, may extend to any autoregressive surrogate of an oscillatory system (e.g., modal or spectral weather emulators) where a single dominant frequency exists; the paper's own harmonic-propagator result supports that the mechanism does not depend on the LSTM architecture.","The R² diagnostic could be used as an early-stopping or model-selection criterion during training, since it predicts correction efficacy without needing ground truth beyond the calibration window.","In genuinely three-dimensional or broadband flows, where the analytic-signal phase is not uniquely defined, the paper's own breakdown at Re=1000 suggests the claim would need a different phase estimator—such as a band-passed or geometric phase—before it can be tested.","A time-varying drift rate, estimated from a sliding-window phase derivative, is the natural extension and would likely recover more than half of the error on the modulated-inflow wake that the paper reports as a constant-rate residual."],"forward_implications":["A model can be excellent by every conventional metric—near-zero one-step validation error, correct amplitude, visually perfect reconstructions—and still fail on long horizons because of phase drift.","The error spectrum peaking at the shedding frequency is not a learned property of the network; it follows algebraically from the amplitude–phase decomposition once the amplitude is known to be correct.","Because the phase drift is linear, the correction improves with prediction horizon, unlike an error-fitted envelope, and the correction degrades far less than a trivial periodic-extension baseline when the flow is not exactly periodic.","The calibration-window R² of the phase fit is a prospective diagnostic that separates correction successes from failures (r=0.85) and prevents harmful application on poorly resolved drifts.","On a stationary limit cycle the correction only matches the trivial baseline of repeating the last shedding period; its genuine advantage appears on coherent but non-periodic flows, where it removes roughly half of the rollout error while tiling fails by an order of magnitude."],"supporting_citations":[{"why":"Supplies the two-stage convolutional-autoencoder LSTM construction whose rollout error is analyzed.","marker":"9"},{"why":"Documents the well-known autoregressive rollout instability and the need for an added stochastic term.","marker":"10"},{"why":"Defines the dynamic mode decomposition baseline against which the phase-drift mechanism and the correction are compared.","marker":"27"},{"why":"Provides the prior qualitative observation that phase errors accumulate fast in convolutional-autoencoder latent rollouts.","marker":"34"},{"why":"Reports the dominant error in autoregressive cylinder-wake rollouts as progressive phase shifts and positional drift.","marker":"35"},{"why":"Establishes the square-cylinder benchmark configuration and the reference Strouhal and drag values used for flow validation.","marker":"22"}],"fun_headline_variants":["Wake-model error is a drifting clock, not noise","95% of wake-model error is just phase drift","One-parameter fix for wake-model rollout drift","Latent wake models fail due to timing, not dynamics","Wake rollout error: phase slip, not random noise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Each latent coordinate must be narrowband around a single dominant frequency so that the Hilbert analytic signal gives a well-defined instantaneous phase; at Re=1000 or in genuinely broadband flows this condition fails and the 95–98% phase share becomes an artifact of the chosen phase definition.","fun_headline_variants_meta":{"raw":{"variants":["Wake-model error is a drifting clock, not noise","95% of wake-model error is just phase drift","One-parameter fix for wake-model rollout drift","Latent wake models fail due to timing, not dynamics","Wake rollout error: phase slip, not random noise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000311,"raw_usage":{"total_tokens":1785,"prompt_tokens":973,"completion_tokens":812,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":736}},"tokens_in":589,"tokens_out":812,"duration_ms":6490,"temperature":1.0,"reasoning_tokens":736,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:11:00.682029+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: take a trained latent-space ROM on a periodic wake and compute the cumulative phase difference φ(t) over the rollout. If the error spectrum is broadband rather than sharply peaked at the vortex-shedding frequency, or if the amplitude and phase contributions to the error variance do not partition the error to within a fraction of a percent (with a negligible cross term), the central claim fails. Applying the same decomposition to a genuinely three-dimensional wake or to a flow without a dominant frequency should produce a phase share well below 95% and a cross term of order unity.","supporting_citations":[{"cited_title":"and Noack, Bernd R","cited_arxiv_id":null,"evidence_quote":"Supplies the two-stage convolutional-autoencoder LSTM construction whose rollout error is analyzed."},{"cited_title":", title =","cited_arxiv_id":null,"evidence_quote":"Documents the well-known autoregressive rollout instability and the need for an added stochastic term."},{"cited_title":", title =","cited_arxiv_id":null,"evidence_quote":"Defines the dynamic mode decomposition baseline against which the phase-drift mechanism and the correction are compared."},{"cited_title":", title =","cited_arxiv_id":null,"evidence_quote":"Provides the prior qualitative observation that phase errors accumulate fast in convolutional-autoencoder latent rollouts."},{"cited_title":"and Durran, Dale R","cited_arxiv_id":null,"evidence_quote":"Reports the dominant error in autoregressive cylinder-wake rollouts as progressive phase shifts and positional drift."}],"review_version":1}