{"id":"51552944-0f28-4dcd-86c7-62b0565d803e","arxiv_id":"2411.13391","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A microtidal adaptation of NS_Tide that pairs quadratic river discharge with linear storm surge reconstructs Neretva water levels with about 2-3% residual variance and links the upstream S1 constituent growth to hydropower peaking.","lead":"The paper modifies an existing non-stationary tidal analysis tool so it can separate tides, storm surges, and river flow in a microtidal estuary, and tests it on Croatia's Neretva River. It finds that adding a storm-surge term and a quadratic river-flow term improves water level prediction, and that hydropower peaking shows up as an amplified 24-hour signal in upstream river sections.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training and validation periods are inconsistent (§3.1 vs §3.3), so the reported 2-3% out-of-sample residual variance may reflect data leakage rather than genuine predictive skill.","rationale":"This paper proposes a non-stationary harmonic model µNS_Tide and reports residual variance of 2-3% on a validation period, plus an attribution of S1 variability to hydropower peaking via STREAM simulations. I read both claims in good faith. The harmonic model is a plausible extension of NS_Tide, and the reported improvements are internally consistent if the validation is truly out-of-sample. However, the manuscript contains an internal contradiction about the training period: §3.1 says training starts January 2017, while §3.3 says January 2016. Since the validation period is June 2015–December 2016, the latter split would include a full year of the validation record in the training data. This is not a stylistic issue; it directly affects whether the 2-3% residual variance is a genuine prediction or a refit. The GAM-based model selection (§3.2) introduces a second, related leakage pathway if the GAM was fit on the full record, because the choice of quadratic discharge and linear surge terms would then be informed by the validation data. The STREAM-based power-peaking conclusion, while dependent on model assumptions, is less central to the paper's main contribution and rests on a previously calibrated model; it can be tested separately. The most efficient way to settle the concern is to check the actual split used in the analysis. If the split is clean, the paper's central claim stands and only minor text corrections are needed; if not, the validation statistics and model comparison require recomputation. This is why the appropriate verdict remains conditional on clarification, matching the reader's verdict.","tokens_in":24844,"tokens_out":5542,"duration_ms":53834,"concrete_test":"Inspect the code/scripts used to produce Fig. 4 (or, if unavailable, ask the authors) and determine the exact training/validation split. If the training set includes any month from June 2015 through December 2016, refit muNS_Tide and the three comparison models using only January 2017–December 2021 for training and recompute the validation residual variance, RMSE, and AIC for all five stations. If the residual variance rises above ~5% at any station or the model ranking changes, the out-of-sample performance claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 states the training period is January 2017 to December 2021 and validation is June 2015 to December 2016, but Section 3.3 states the training set is January 2016 to December 2021. These are irreconcilable. If the latter is used, the validation period overlaps with training for all of 2016, so the reported residual variance of 2-3% in Fig. 4 is not an out-of-sample measure and the comparison among NS_Tide variants is biased toward the model that best fits the validation data. If the former is used, the text in §3.3 is wrong but the validation is clean; the paper must clarify which split produced Fig. 4. Additionally, the GAM-based selection of quadratic/linear functional forms (§3.2) does not state whether it used the full record or only the training period; if the GAM was fitted to the validation period, the functional forms themselves are leakage. This concern is more fundamental than the STREAM-based power-peaking attribution because it threatens the paper's primary quantitative claim (2-3% residual variance and model ranking), whereas the STREAM experiment is a secondary supporting analysis based on a previously calibrated model.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a modification of the non-stationary harmonic analysis tool NS_Tide, called μNS_Tide, for microtidal estuaries. The stage term is expressed with quadratic river discharge and linear storm surge covariates, and the tidal-fluvial term with linear discharge and linear storm surge covariates (Eqs. 8–9). The model is applied to hourly water levels at five stations along the Neretva River estuary (Croatia) covering June 2015–December 2021, with a training/validation split and comparisons against the original NS_Tide, a quadratic-discharge variant (qNS_Tide), and a storm-surge variant (sNS_Tide). The authors report that μNS_Tide achieves about 2–3% residual variance during validation at all stations, outperforming the other formulations, and they use the STREAM numerical model in a two-simulation experiment to attribute the upstream amplification of the S1 tidal constituent to high-frequency discharge fluctuations from hydropower peaking.","tokens_in":25111,"tokens_out":6384,"duration_ms":63989,"significance":"If the validation design is sound, the paper offers a practical extension of NS_Tide that is better suited to microtidal environments, where storm surge rather than tidal range drives subtidal and tidal variability. The flexible user-defined predictor framework and the incorporation of recent uncertainty estimates are useful contributions. The power-peaking attribution for S1 is plausible and consistent with observations in other regulated estuaries. However, the primary quantitative claim—the 2–3% out-of-sample residual variance and the ranking of the four models—is compromised by an inconsistency in the stated training period between Sections 3.1 and 3.3, and by a potential circularity at the Usce station where the storm surge covariate is derived from the same water level series being modeled. These issues must be resolved before the central claims can be accepted.","major_comments":[{"comment":"The training period is stated as January 2017 to December 2021 in Section 3.1, but Section 3.3 states that the model parameters were determined using the training set January 2016 to December 2021. These two statements are mutually inconsistent. If the Section 3.3 split was actually used, the validation period (June 2015 to December 2016) overlaps with the training period for all of 2016, so the residual variances and model comparisons in Fig. 4 are not out-of-sample measures and are biased toward models that fit the validation data. If the Section 3.1 split was used, then Section 3.3 contains a factual error. Please clarify which split produced Fig. 4 and the other reported statistics, and confirm that all model choices (covariate lags, tidal constituent retention via the SNR criterion, and functional forms) were made using only the training period.","section":"§3.1 and §3.3"},{"comment":"The functional forms of the stage and tidal-fluvial terms are selected on the basis of GAM reconstructions, but the manuscript does not state whether the GAM was fitted to the training period only or to the full record. The supplementary figures (e.g., Fig. A.3 and B.2) show distributions of the full observation period, which suggests the full record may have been used. If the validation period informed the choice of quadratic versus linear forms, the subsequent validation is not independent. The authors should state explicitly that the GAM selection was performed on the training subset only, or rerun the selection on the training data and show that the chosen forms are stable across the two periods.","section":"§3.2"},{"comment":"At the Usce station (0 rkm), the storm surge covariate SS(t) is defined as a low-passed residual of the stationary harmonic analysis of the same observed water level series that the model then reconstructs. Because the stage term in μNS_Tide (Eq. 8) contains a linear term in SS, the model at Usce is partly using a filtered version of the target variable as a predictor, which can artificially inflate the fit and reduce the residual variance below a true out-of-sample value. The reported 2–3% residual variance at Usce is therefore not a genuine predictive metric. The authors should either compute SS from an independent coastal station (e.g., Ploce) and apply it at Usce, or quantify the fit at Usce with SS omitted from the stage term to assess the circularity.","section":"§3.2 and §4.5"},{"comment":"The attribution of the upstream S1 amplification to power peaking relies on the difference between STREAM simulation A (measured discharge) and simulation B (low-pass filtered discharge). This assumes that the low-pass filter removes only the hydropower peaking signal and no other physically relevant discharge variability, and that the STREAM model faithfully represents the tidal-fluvial interactions. The paper does not provide a sensitivity analysis with respect to the filter cutoff (0.03 cph) nor a quantitative validation of simulation A against the observed water levels over the study period (the cited calibration is from an earlier study). Please add such a validation and a discussion of how sensitive the S1 and other constituent differences are to the filtering choice, to strengthen the causal claim.","section":"§5.1"}],"minor_comments":[{"comment":"There is a typo in the text: 'esides' should read 'Besides'.","section":"§4.3"},{"comment":"There is a typo: 'emodel' should read 'The model'.","section":"§4.2"},{"comment":"The SNR criterion for retaining tidal constituents is described as 'SNR greater than 2 for at least 50% of the reconstructed time steps,' but it is not stated whether the SNR is computed over the training period only or over the full record. This should be clarified to ensure the constituent selection does not use validation-period information.","section":"§3.2"},{"comment":"The sentence 'we used μNS_Tide, but only the storm surge term, SS(t), was included as a predictor in Eqs. 8 and 9' should specify whether the discharge terms are omitted entirely or set to zero, and how this affects interpretation of the Usce reconstruction.","section":"§4.5"},{"comment":"The S1 amplitude at Gabela has a mean amplitude of 9.33 cm but an amplitude standard error of 7.32 cm, indicating substantial uncertainty. This should be acknowledged in the discussion of the S1 amplification, as the effect at that station is not tightly constrained.","section":"Table E.1"},{"comment":"The statement that the NS_Tide code and data are 'available upon request' is weaker than the standard for reproducibility. Consider depositing the code in a permanent repository (e.g., Zenodo) and providing a clear data-access statement.","section":"Data and code availability"}],"recommendation":"major_revision","confidential_remarks":"The reader's conditional verdict is appropriate. The central methodological idea is sound and the application is interesting, but the inconsistency in the stated training period between Sections 3.1 and 3.3 is load-bearing: if the January 2016 start is true, the reported out-of-sample validation is invalid because of overlap. The Usce circularity is a second serious issue that should be fixed, though it affects only the coastal station. If the authors can confirm that the Jan 2017–Dec 2021 training split was used and that all model-selection steps were confined to that period, the paper would be close to acceptable. The STREAM-based attribution is secondary and can be strengthened with a sensitivity analysis, but does not by itself justify rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: muNS_Tide is a legitimate extension of NS_Tide for microtidal systems, and the S1/power-peaking fingerprint is genuinely interesting. But the validation split is stated inconsistently across §3.1 and §3.3, and until that is resolved the headline 2–3% residual variance figure is not trustworthy.\n\nWhat is new and good: the predictor combination—linear storm surge plus quadratic river discharge in the stage, linear surge/discharge in the tidal-fluvial terms—is not in the cited NS_Tide literature. The constituent-by-constituent decomposition of tide-river and tide-surge interactions is a real step beyond species-level analyses. Using GAMs to guide functional form selection is pragmatic, and the model comparison against three baselines is well designed. On the intended split, the improvement is dramatic and the residual variance drops to 2–3%.\n\nThe soft spots are real. The training period is January 2017–December 2021 in §3.1 and January 2016–December 2021 in §3.3, with validation June 2015–December 2016 in both. If the §3.3 split is the one actually used, validation overlaps with training for all of 2016, so the out-of-sample claim collapses and the model ranking in Fig. 4 is biased. The paper must state which split produced Fig. 4 and re-run if necessary. Also, the GAM selection in §3.2 does not say whether functional forms were chosen on the full record or training only; if the former, the forms themselves leak information from the validation period.\n\nA smaller but worth-noting issue: at Usce, the storm surge covariate SS is a low-passed residual of the same water level series the model reconstructs, so part of the stage fit is a filtered version of the target. That is a form of circularity and should be acknowledged or handled differently. Upstream stations use SS as an external coastal predictor, so the concern is localized.\n\nCode and data are \"available upon request\" only, which limits reproducibility but is not fatal for a methods paper if the package is eventually released. The STREAM-based attribution of S1 amplification to power peaking relies on a previously calibrated model and a low-pass filtering of discharge; it is a reasonable secondary analysis but not as strong as the harmonic-model evidence itself.\n\nBottom line: the core modeling idea is sound and the paper deserves a serious referee. The right outcome is conditional acceptance after the split inconsistency is resolved and the GAM selection timing is clarified. I would bring it to the reading group as a case study in validation design for non-stationary tidal analysis.","headline":"Sensible extension of NS_Tide with a useful S1 diagnostic, but the validation split is stated inconsistently and the 2–3% out-of-sample claim needs verification before it can be trusted.","tokens_in":25715,"tokens_out":2515,"would_cite":true,"duration_ms":27945,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A microtidal-specific non-stationary harmonic model reproduces Neretva estuary water levels to within 2–3% residual variance and attributes an upstream S1 tidal signal to hydropower peaking.","keywords":["tidal dynamics","microtidal estuary","hydropower peaking","storm surge","tide-surge-river interaction","non-stationary harmonic analysis","NS_Tide","Neretva River"],"falsifier":"Run the same paired STREAM simulations during a period when upstream hydropower stations switch between peaking and steady operation with equal total daily volume; if the S1 amplitude at Gabela does not drop when the 24-hour discharge pulse is removed, the attribution fails. Equivalently, a second independent hydrodynamic model producing no S1 suppression under filtered inflow would also break the claim.","tokens_in":24574,"feed_emoji":"🌊","tokens_out":7259,"duration_ms":70947,"temperature":0.7,"pith_summary":"Microtidal estuaries are hard to forecast because storm surges can rival or exceed the tides. The paper tries to establish that a modified non-stationary harmonic model, muNS_Tide, which uses a storm-surge time series and quadratic-plus-linear river discharge as predictors, reproduces measured water levels along the Neretva estuary with residual variance around 2–3% at five stations, from the river mouth to the tidal-river limit. It also argues that the unusually strong S1 (24-hour) tidal-constituent signal upstream is not astronomical but is pumped into the river by hydropower peaking, using before-and-after filtering experiments with the STREAM numerical model. If true, the model gives a practical, site-flexible tool for decomposing water levels in microtidal estuaries and for attributing non-tidal variability to dam operations.","feed_headline":"Tidal model explains 97% of Neretva estuary water levels","feed_subtitle":"A non-stationary harmonic model separates tides, surge, and river flow, pinning a diurnal peak on dam operations.","key_machinery":"The carrying object is the muNS_Tide non-stationary harmonic model: it splits water level into a subtidal stage term S(t) and a tidal-fluvial term F(t), and lets both depend on user-chosen predictors. In this paper the stage uses $Q$ and $Q^2$ plus storm surge $SS(t)$, while each constituent's cosine and sine coefficients vary linearly with $Q$ and $SS(t)$. The identification of power peaking rests on the STREAM numerical model, a one-dimensional two-layer shallow-water model, whose paired simulations with full versus low-pass-filtered inflow isolate the effect of high-frequency dam operations.","core_discovery":"On the paper's own terms, the central discovery is that the original NS_Tide functional forms—built on river discharge to the 2/3 power and on coastal tidal range—fail when tidal range is smaller than storm-surge amplitude, leaving residual variance above 40% at downstream stations. Replacing the tidal-range predictor with a lagged storm-surge series, and using a quadratic discharge law in the stage term (Eq. 8) with linear discharge and surge modulations of each tidal constituent's amplitude and phase (Eq. 9), reduces validation residual variance to 2–3% at every station. The paper then uses the STREAM two-layer model to run paired simulations, one with measured discharge and one with only low-pass-filtered discharge, and finds that the S1 constituent's upstream growth, visible in observed data and in the unfiltered simulation, disappears when high-frequency discharge fluctuations are removed. That is the evidence behind the claim that power peaking amplifies S1 and modulates other constituents such as K2 and S2 in the tidal river.","pith_inferences":["Beyond the paper: in a regulated microtidal river, stationary harmonic analysis will misclassify dam-induced S1 energy as a real tidal constituent, so discharge filtering or non-stationary predictors should be standard before interpreting diurnal constituents there.","Related consequence: the success of a quadratic discharge law in the Neretva suggests that the classical 2/3 discharge exponent should not be assumed in other microtidal estuaries; exploratory GAM-style fits can pick the stage–discharge relationship before fitting the harmonic model.","Testable extension: if the S1 attribution holds, hydropower peaking should also raise sub-daily variance of water levels upstream of the salt wedge, potentially affecting salt-wedge intrusion forecasts; this can be checked by correlating peaking schedules with salinity records."],"forward_implications":["The muNS_Tide formulation can be applied to other microtidal estuaries where surge dominates tidal range, with GAM-based checks guiding the choice of predictors.","Tide-river interaction becomes larger than tide-surge interaction in the Neretva and peaks at the most upstream station, so predictions there need accurate discharge rather than just sea-level forcing.","Constituent-by-constituent decomposition allows significance testing of tide-river and tide-surge interactions; the paper finds tide-surge interaction significant for up to 10 constituents even though its amplitude contribution is small.","Removing the tidal-range term from NS_Tide prevents spurious oscillations and overfitting, improving predictability and AIC.","The significance tests and flexible predictor definitions make the method usable as a routine diagnostic for whether a diurnal signal in a regulated estuary is astronomical or operational."],"supporting_citations":[{"why":"Supplies the original NS_Tide non-stationary harmonic framework that the paper modifies.","marker":"Matte et al. (2013)"},{"why":"Gives the theoretical discharge and tidal-range functional forms whose exponents the paper tests and replaces for microtidal conditions.","marker":"Kukulka and Jay (2003a)"},{"why":"Provides the analytical GLS uncertainty estimation for temporally correlated noise used to compute constituent SNRs and standard errors.","marker":"Innocenti et al. (2022)"},{"why":"Introduces the STREAM two-layer model used to simulate the Neretva estuary and to isolate power-peaking effects.","marker":"Krvavica et al. (2017)"},{"why":"Provides the calibrated STREAM setup and the salt-wedge estuary background that the numerical experiment relies on.","marker":"Krvavica et al. (2021)"},{"why":"Documents S1 and P1 amplification by dam operations in a regulated estuary, serving as the comparison for interpreting the Neretva results.","marker":"Jay et al. (2015)"},{"why":"Reports similar S1 increases from hydroelectric operations in the Qiantang estuary, supporting the paper's attribution.","marker":"Zhou et al. (2024)"},{"why":"The GAM reconstructions that guide the choice of quadratic discharge and linear surge predictors in the new formulation.","marker":"Hastie and Tibshirani (1986); Wood (2017)"}],"fun_headline_variants":["New harmonic model explains 97% of Neretva levels","Dam peaking amplifies S1 tide in Neretva estuary","Model separates surge, river, and dam effects on tides","New non-stationary model for microtidal estuaries","Power peaking shapes tides in Neretva River estuary"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The power-peaking conclusion rests on the assumption that the numerical experiment that removes rapid discharge fluctuations strips out only the dam-operation signal, and that the model otherwise captures real tide-flow interactions, so the disappearance of the S1 peak can be credited solely to peaking.","fun_headline_variants_meta":{"raw":{"variants":["New harmonic model explains 97% of Neretva levels","Dam peaking amplifies S1 tide in Neretva estuary","Model separates surge, river, and dam effects on tides","New non-stationary model for microtidal estuaries","Power peaking shapes tides in Neretva River estuary"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000794,"raw_usage":{"total_tokens":3531,"prompt_tokens":1014,"completion_tokens":2517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":2432}},"tokens_in":630,"tokens_out":2517,"duration_ms":21759,"temperature":1.0,"reasoning_tokens":2432,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:27:51.410474+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same paired STREAM simulations during a period when upstream hydropower stations switch between peaking and steady operation with equal total daily volume; if the S1 amplitude at Gabela does not drop when the 24-hour discharge pulse is removed, the attribution fails. Equivalently, a second independent hydrodynamic model producing no S1 suppression under filtered inflow would also break the claim.","supporting_citations":[],"review_version":1}