{"id":"c9a6ce5a-55e0-455d-9f9b-b7711f8f772f","arxiv_id":"1909.04488","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A new 'state' method measures candidate shadowing times using order statistics of residual vectors, and in the rotating annulus model these times fall logarithmically with initial distance from truth.","lead":"This paper presents a faster statistical method for measuring how long model trajectories stay consistent with observations, and tests it on a simulated rotating annulus atmosphere experiment. The method works in high-dimensional systems and could help benchmark weather and climate forecast models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Spatial correlations in residual fields violate the IID order-statistic null in Eq. (3), so shadowing times for diverging trajectories may be systematically short; a block-bootstrap calibration check is needed.","rationale":"The reader's weakest assumption focused on IID residual components, and spatial correlation is indeed the central worry, but the reader's formulation slightly misses the mark: under the null hypothesis the residuals are IID by construction, so the false-rejection rate in Fig. 4 is correctly calibrated. The real issue is that the shadowing time is defined by rejection under the alternative, where the residual field is correlated; the analytic null distribution overstates the effective sample size and makes the test over-sensitive. This directly affects every τS reported and hence the empirical law Eq. (16) and its successful 2000 s prediction. The paper's internal validations (τD, visual divergence) are qualitative and do not quantify the correlation-induced bias. The proposed block-bootstrap check is a concrete, computationally feasible way to settle whether the analytic null is adequate. If the bias is small, the central claim stands; if large, the method needs a correlation-aware null or the times need re-interpretation. This does not change the reader's CONDITIONAL verdict: the method is promising and honestly presented, but the statistical foundation needs this check before the benchmark is used. Hence verdict_should_be = UNCHANGED.","tokens_in":19339,"tokens_out":24026,"duration_ms":240993,"concrete_test":"Recompute τS for the 256-candidate δ = 0.1 set using a null distribution for the median and 90th percentile built from a spatial block bootstrap of the residual vector (resample contiguous blocks of the MORALS grid, e.g., blocks of 4×4×4 grid points, to preserve local correlations). Compare the resulting distribution with Fig. 7 at p = 10^{-6}. If the median or interquartile range shifts by more than ~50 s, the IID order-statistic null is load-bearing and the analytic p-values (or the interpretation of τS) must be revised; if the shift is negligible, the concern is settled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The state method's critical values come from the IID order-statistic formula (Sect. 2.4.1, Eq. 3) with N = 24,192. When the null hypothesis is true (perfect initial condition), the residual vector is exactly IID observational noise, so the Type I error rate is correctly calibrated (Fig. 4). The load-bearing gap appears precisely when the method is used as intended: for diverging candidates, the residual field e_t∘r^{-1} is spatially correlated (large-scale, smooth error structures), so the effective number of independent components is much smaller than N. The sample median and 90th percentile then have substantially wider sampling distributions than the analytic IID formula predicts. Consequently, the test rejects earlier than it would with a correlation-aware null, and the measured τS (and the empirical scaling law Eq. 16) reflect not only the marginal distribution of the residuals but also their spatial correlation structure. The agreement with τD and visual divergence (Figs. 5, 6, 9) shows the method is not unreasonable, but it does not establish how much the neglected spatial correlation shortens the shadowing times. Since Eq. (16) is an empirical benchmark derived from these possibly biased times, this is the weakest link in the central claim that the state method measures 'reasonable' shadowing times.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a \"state\" method for measuring candidate trajectory shadowing times in high-dimensional models, building on the sequence method of Smith et al. (2010). At each verification time, the residual vector between model trajectory and observations is tested for consistency with the observational error distribution using the median and 90th percentile of the residual components; the sampling distributions of these order statistics are obtained analytically from the incomplete beta function under an IID assumption, avoiding Monte Carlo sampling. The method is applied to MORALS, a rotating-annulus model with N=24,192 state variables, in the perfect model scenario, with observations generated by adding IID Gaussian noise scaled by each grid point's natural variability. The authors calibrate the significance level using null trajectories (Type I error control), measure shadowing times for candidate trajectories started at eight distances δ from truth, and report an empirical scaling law τS ≈ 250(1−log10δ), which successfully predicts a shadowing time near 2000 s for δ=10^-7. The candidate shadowing times are compared with a simple distance metric τD and with visual divergence of the trajectories. The paper is explicitly a record of work submitted in 2010 and not accepted; the current version is posted to arXiv for citation and follow-up.","tokens_in":19577,"tokens_out":5167,"duration_ms":53790,"significance":"If the state method is valid, it provides a computationally cheap way to assign candidate shadowing times in intermediate- and high-dimensional systems, and the empirical scaling law would be a useful benchmark for future shadowing studies in the annulus and similar systems. The paper has genuine strengths: the analytic order-statistic result removes a sampling bottleneck; the Type I error calibration in Fig. 4 is careful and is checked against actual noise realizations; the out-of-sample prediction at δ=10^-7 is a real falsifiable check; and the agreement with an independent distance metric and visual divergence provides external corroboration. The central limitation is that the order-statistic formula in Eq. (3) assumes the N residual components are IID. That assumption is exact for a perfect initial condition, but it is violated for the diverging candidates the method is intended to measure, because residual fields are spatially coherent. Unless this is shown to be quantitatively harmless, the reported shadowing times and the scaling law Eq. (16) may be systematically too short.","major_comments":[{"comment":"The IID assumption is load-bearing in the derivation of F(r)(x)=IP(x)(r,N−r+1). Under the null hypothesis δ=0, the residual vector e_t∘r^{-1} is IID observational noise by construction, so the Type I calibration in Fig. 4 is meaningful. For a diverging candidate, however, e_t∘r^{-1} develops large-scale, spatially coherent structure (visible in Figs. 5b and 9), so the effective number of independent components is far smaller than N=24,192. The analytic confidence intervals for the median and 90th percentile are then much narrower than the true sampling distributions, causing premature rejection. The shadowing times in Fig. 7 and the scaling law Eq. (16) therefore reflect not only the marginal distribution of the residuals but also their spatial correlation structure. I request a block-bootstrap or spatial-decorrelation estimate of the effective sample size, with the order-statistic test repeated using that effective N, and a comparison of the resulting τS values with those reported.","section":"Section 2.4.1, Eq. (3)"},{"comment":"The insensitivity of τS to the significance level p is presented as evidence that the method is robust, but this does not address the IID concern. For p between 0.1 and 0.001, Fig. 4 shows that Type I errors dominate for a perfect trajectory, so near-constancy of τS in Fig. 7 is expected whenever divergence is fast relative to the critical-value differences in that range. The key check is whether rejection occurs because the residual marginal distribution has moved by more than the correlation-corrected sampling error, rather than simply because the IID critical intervals are too tight. Without that check, the agreement between τS and τD in Sect. 4.3 supports the qualitative claim that the method detects divergence, but not the specific numerical values entering Eq. (16).","section":"Section 4.2, Fig. 7"},{"comment":"The Šidák correction r=1−(1−p)^{2n} assumes independence of successive verification times and of the two percentile statistics. This assumption is valid for the null candidate, where additive noise is IID across components and over time, but it is not valid for the diverging candidates used to estimate τS, because the test statistics are serially and spatially correlated. The paper should state explicitly that Eq. (14) calibrates the null case only, and should provide a sensitivity analysis (for example, recomputing τS with the correlation-corrected effective N) to show that the conclusions are not an artifact of miscalibrated critical values.","section":"Section 4.1, Eq. (14)"}],"minor_comments":[{"comment":"The shadowing time definition is ambiguous if the test fails at one time and passes at a later time: the text says \"the longest time for which the null hypothesis is retained\" and also \"the final t at which the test is passed, counting upwards from zero.\" Please clarify the stopping rule used in the implementation.","section":"Section 3.4"},{"comment":"There is a typo in the phrase \"use the the Euclidean length\" in the description of the Smith et al. (2010) method.","section":"Section 2.3"},{"comment":"The notation \"e.10^-2\" and \"e.10^-1\" is nonstandard and is never defined; these presumably mean e·10^-2 and e·10^-1, but the notation should be introduced explicitly.","section":"Section 4.2"},{"comment":"The empirical fit τS≈250(1−log10δ) is presented without uncertainty estimates or goodness-of-fit measures; given that Eq. (17) is used for prediction, reporting the fitted slope and its uncertainty would strengthen the claim.","section":"Section 4.2, Eq. (16)"},{"comment":"The caption correctly notes that the order of grid points along the x-axis is unimportant, but the axis is otherwise unlabeled; adding a label such as \"grid point index\" would improve readability.","section":"Figure 9"}],"recommendation":"major_revision","confidential_remarks":"The authors are transparent that this manuscript was submitted to Physica D in 2010 and not accepted; I did not treat that history as dispositive. However, the lack of reproducibility artifacts (no code or data) and the reliance on a single parameter regime and a small number of noise realizations make the quantitative scaling law less secure than the paper's framing suggests. The main technical concern, the IID assumption in Eq. (3), is genuinely load-bearing and is not resolved by the existing comparisons. If the editor prefers to publish only fully refereed work, the provenance may weigh against publication; on scientific content alone, the paper merits a major revision rather than outright rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper does something genuinely useful. It provides a practical 'state' method for measuring candidate shadowing times in high-dimensional models, replacing the sampling-heavy sequence method of Smith et al. (2010) with an analytic order-statistic test on the residual vector components. The authors demonstrate it on a 24,192-dimensional MORALS rotating annulus model in the perfect model scenario, and they validate it three ways: visual divergence, a simple distance metric τD, and an out-of-sample prediction of a 2000 s shadowing time for δ=10^-7 that landed within a few percent. That prediction is real evidence the method is not just calibrated to noise.\n\nThe IID assumption in the order-statistic test (Sect. 2.4.1, Eq. 3) is the load-bearing gap. It is exact when the null is true—residuals are IID observational noise—and the Type I error calibration in Fig. 4 confirms that. But when the method is actually used, for diverging candidates, the residual field is spatially correlated and smooth, so the effective number of independent components is far smaller than N=24,192. The analytical percentiles then understate the sampling variability, which likely makes the test reject earlier. The measured shadowing times and the scaling law Eq. (16) therefore probably carry a correlation-induced bias. The agreement with τD and visual divergence shows the method is not unreasonable, but it does not tell us the magnitude of that bias. A block-bootstrap or a spatial-decorrelation test would settle this.\n\nTwo smaller soft spots: the scaling law is fit to eight points without uncertainty bounds, though the out-of-sample point helps; and the paper's account of which significance level to use is a bit dense, though not wrong. The authors' note that the paper was rejected from Physica D in 2010 and they no longer have time to work on it is honest and does not change the technical content.\n\nWho should read this: anyone working on shadowing in intermediate- or high-dimensional dynamical models, and people developing practical consistency tests for forecasts. It deserves a serious referee; the IID issue needs addressing before the method becomes a benchmark, but the core contribution is a real step forward.","headline":"A practical shadowing-time method for high-dimensional models, honestly demonstrated on the rotating annulus, with a load-bearing IID assumption that needs a correlation check.","tokens_in":20136,"tokens_out":2181,"would_cite":false,"duration_ms":21849,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A residual-vector consistency test measures how long high-dimensional model trajectories shadow noisy observations, and the measured times obey a simple logarithmic law.","keywords":["shadowing trajectories","rotating annulus","perfect model scenario","candidate shadowing time","order statistics","chaotic dynamics","high-dimensional forecasting","MORALS"],"falsifier":"Run the state method on perfect initial conditions ($\\delta = 0$) with observational noise that is spatially correlated across grid points instead of independent, and compare the observed fraction of false rejections with the Šidák prediction of Eq. (13); a systematic mismatch would show that the IID assumption is load-bearing and would require recalibrating the significance levels and shadowing times.","tokens_in":19108,"feed_emoji":"🌀","tokens_out":9680,"duration_ms":87669,"temperature":0.7,"pith_summary":"This paper sets out to make shadowing—the time a model trajectory stays consistent with a sequence of noisy observations—measurable in high-dimensional models, and to verify that the measurement behaves as physical intuition demands. The authors replace the earlier sequence-based test, which samples the noise distribution and collapses residuals to distances, with a \"state\" method that checks each residual vector directly: at every verification time, the median and 90th percentile of the scaled residual field must lie inside analytically computed confidence intervals for the same percentiles of the observational noise. Applied to a 24,192-variable rotating-annulus model in the perfect model scenario, the method yields candidate shadowing times that match visual divergence and a simple distance threshold, and that decrease as the initial distance from truth increases. The quantitative payoff is an empirical scaling law, $\\tau_S \\approx 250(1-\\log_{10}\\delta)$ seconds, which predicted a 2000 s shadowing time for $\\delta = 10^{-7}$ that the holdout experiment confirmed. A sympathetic reader would care because this supplies a benchmark for judging candidate trajectories in weather- and climate-scale models, where shadowing questions have been well posed but rarely answered.","feed_headline":"Shadowing time follows 250(1 − log δ) in a 24 000-dim model","feed_subtitle":"A residual-percentile test matches visual divergence and predicted a 2000 s shadow for δ=10⁻⁷.","key_machinery":"Load-bearing machinery is the order-statistic identity that makes the test computational: for IID components, the CDF of the $r$-th order statistic of a sample of size $N$ is the incomplete $\\beta$ function $I_{P(x)}(r, N-r+1)$, so confidence intervals for the median and 90th percentile of the observational noise can be computed by inverting it instead of drawing $M$ Monte Carlo samples. This identity is what lifts the method from low-dimensional toy systems to $N \\approx 24{,}000$. The second piece of machinery is the Šidák bound, $p < 1-(1-R/E)^{1/2n}$, which sets the per-test significance level so that fewer than $R$ of $E$ candidates suffer a false rejection over $n$ verification times; at the candidate counts used here it forces $p \\approx 10^{-5}$–$10^{-6}$. The sanity check is a simple threshold on the Euclidean distance $D(t)$, using the fact that the expected distance between truth and observations is $\\sigma\\sqrt{N}$, so divergence is declared when $D(t) > m\\sigma\\sqrt{N}$; the paper uses $m=2$ and finds $\\tau_D$ insensitive to $m$.","core_discovery":"On the paper's own terms, the central claim is that candidate shadowing times in a high-dimensional chaotic model can be obtained by testing, time step by time step, whether the distribution of the residual state vector is consistent with the known observational error distribution, and that the resulting times are reasonable. The procedure uses the perfect model scenario: MORALS, discretized on a staggered grid with $N = 24{,}192$ independent variables, is both system and model; observations are the true state plus independent Gaussian noise scaled by the 99\\% natural variability range at each grid point; candidates begin at fixed scaled distances $\\delta$ from truth. For each candidate, the shadowing time $\\tau_S$ is the longest prefix over which both the median and the 90th percentile of the scaled residual vector fall within confidence intervals obtained from the incomplete $\\beta$ function. The measured distributions agree with the time of visual divergence and with the distance-based divergence time $\\tau_D$, are largely insensitive to position on the attractor and to the significance level once Type I errors are controlled, and follow $\\tau_S \\approx 250(1-\\log_{10}\\delta)$ across eight perturbation sizes; a prediction at $\\delta = 10^{-7}$ gave a median shadowing time near 2000 s as expected.","pith_inferences":["If the logarithmic scaling survives in other chaotic parameter regimes of the annulus, it offers a cheap predictability diagnostic: the slope of 250 s per decade of initial distance could be compared across flow regimes and model resolutions without running full shadowing searches.","A testable extension is to replace the IID noise model with spatially correlated observational error; the incomplete-beta calibration would then be approximate, and the size of the resulting bias in $\\tau_S$ would quantify how much the method depends on that assumption.","The same per-state percentile consistency test could be applied to operational ensemble forecasts, declaring an ensemble member to have stopped shadowing when its median and 90th-percentile residual leave the noise-model intervals; this would give a state-vector-level version of useful forecast duration.","Because the test is applied per time step, small persistent outliers can go undetected, the paper's own caveat in Sect. 2.4.2; combining the state method with a whole-trajectory cumulative test would exploit both sensitivities."],"forward_implications":["In the perfect model scenario, candidate shadowing times measured from random-direction perturbations are lower bounds on the true shadowing time; candidates placed on the attractor or its stable manifold can shadow longer, but such initial conditions occur with probability zero under the paper's perturbation scheme.","The empirical law $\\tau_S \\approx 250(1-\\log_{10}\\delta)$ gives a benchmark for future shadowing searches: once a candidate's distance from truth is known, its expected shadowing time is fixed, and any candidate that beats this expectation is doing genuine work.","With an imperfect model or real observations the same measurement should produce upper bounds rather than lower bounds, because model and system attractors are typically disjoint, so model error—not initial-condition uncertainty—will limit shadowing.","The method is deliberately inapplicable to low-dimensional systems ($N = O(100)$ or below), since the 90th percentile of a three- or ten-component residual vector carries no statistical meaning; this defines the regime of applicability of the technique.","Type I errors accumulate across repeated verification times, so significance levels must be set by the Šidák bound; at practical candidate counts this forces $p \\approx 10^{-5}$–$10^{-6}$ rather than conventional 0.05."],"supporting_citations":[{"why":"Supplies the original \"sequence\" method and the definition of shadowing consistency that the paper modifies; the alternative \"state\" method is built to fix its sampling problems.","marker":"Smith et al. (2010)"},{"why":"Provides the order-statistic distribution used as the basis for replacing Monte Carlo sampling with the incomplete beta function.","marker":"David (1981)"},{"why":"Gives the identity linking the binomial sum to the incomplete beta function $I_{P(x)}(r, N-r+1)$, used to compute percentile confidence intervals analytically.","marker":"NIST (2010)"},{"why":"Source of the \"small annulus\" configuration and fluid-property parameterisation used to set up the MORALS simulations.","marker":"Hignett et al. (1985)"},{"why":"Documents the MORALS model equations that define the perfect model and its state space.","marker":"Farnell and Plumb (1976)"},{"why":"Supplies the Lyapunov-exponent estimation procedure and the comparison flow-regime values used to classify the wavenumber-3 chaotic regime.","marker":"Young and Read (2008)"},{"why":"States that the Euclidean norm of Gaussian vectors follows a chi distribution, used to set the expected observation distance $\\sigma\\sqrt{N}$ and the $\\tau_D$ threshold.","marker":"Cramér (1946)"},{"why":"Provides the asymptotic normal approximation for the chi distribution, used to justify the $m\\sigma\\sqrt{N}$ distance threshold in the sanity-check metric.","marker":"Stuart and Ord (1994)"}],"fun_headline_variants":["Residual tests reveal shadowing times in 24k-dim chaos","Empirical law predicts chaos shadowing times: 250(1-logδ)","Shadowing times in chaotic annulus: residual tests work","A 24k-dim model's shadowing times follow a simple law","Test candidate trajectories: measure shadowing time in chaos"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the components of the residual vector are independent and identically distributed draws from the observational error distribution; if spatial correlations among grid points or inhomogeneous scaling break that assumption, the analytic confidence intervals and every reported shadowing time would shift systematically.","fun_headline_variants_meta":{"raw":{"variants":["Residual tests reveal shadowing times in 24k-dim chaos","Empirical law predicts chaos shadowing times: 250(1-logδ)","Shadowing times in chaotic annulus: residual tests work","A 24k-dim model's shadowing times follow a simple law","Test candidate trajectories: measure shadowing time in chaos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00046,"raw_usage":{"total_tokens":2366,"prompt_tokens":1071,"completion_tokens":1295,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":687,"completion_tokens_details":{"reasoning_tokens":1206}},"tokens_in":687,"tokens_out":1295,"duration_ms":9971,"temperature":1.0,"reasoning_tokens":1206,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:10:43.705077+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the state method on perfect initial conditions ($\\delta = 0$) with observational noise that is spatially correlated across grid points instead of independent, and compare the observed fraction of false rejections with the Šidák prediction of Eq. (13); a systematic mismatch would show that the IID assumption is load-bearing and would require recalibrating the significance levels and shadowing times.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original \"sequence\" method and the definition of shadowing consistency that the paper modifies; the alternative \"state\" method is built to fix its sampling problems."},{"cited_title":"(1981), Order statistics , second ed., Wiley","cited_arxiv_id":null,"evidence_quote":"Provides the order-statistic distribution used as the basis for replacing Monte Carlo sampling with the incomplete beta function."},{"cited_title":"http://dlmf.nist.gov/","cited_arxiv_id":null,"evidence_quote":"Gives the identity linking the binomial sum to the incomplete beta function $I_{P(x)}(r, N-r+1)$, used to compute percentile confidence intervals analytically."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the \"small annulus\" configuration and fluid-property parameterisation used to set up the MORALS simulations."},{"cited_title":"Plumb (1976), Numerical integration of flow in a rotating annulus II: three dimensional model , Tech","cited_arxiv_id":null,"evidence_quote":"Documents the MORALS model equations that define the perfect model and its state space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Lyapunov-exponent estimation procedure and the comparison flow-regime values used to classify the wavenumber-3 chaotic regime."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the asymptotic normal approximation for the chi distribution, used to justify the $m\\sigma\\sqrt{N}$ distance threshold in the sanity-check metric."}],"review_version":1}