{"id":"85ec6ab3-7d9d-4906-bb3c-d88c24590fd8","arxiv_id":"2608.09692","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Rolling-origin evaluation windows misrepresent the atom mass of intermittent time series, a permutation control isolates coupling contributions at fixed CRPS, and an autoregressive hurdle beats a conditional flow on five of six datasets under run-length statistics.","lead":"Standard evaluation windows for generative time-series models can contain far fewer zero values than the datasets they come from, and on one benchmark this flipped which model looked best. The paper adds a cheap permutation control that isolates how much a model's temporal coupling actually contributes to its scores.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'beats the flow on five of six datasets' headline is metric-dependent: under the occurrence ACF (Table III) the flow ranks above the Hurdle-AR, so the unqualified claim is not robust.","rationale":"The evaluation-level findings (Sec. IV-A window mismatch; Sec. IV-B decorrelation control) are definitionally solid and independent of the model comparison. The model-level benchmark is the part that carries risk. The paper's own Table III contains the evidence for the concern: ACF, the only occurrence statistic in the 'independent' group that is computed from the binary process without run-length construction, places the flow above the hurdle, and the rank correlation between W1 and ACF is the lowest among the run-length family. The abstract's sentence 'an autoregressive hurdle beats a conditional flow on five of six datasets' is therefore not a robust summary; it depends on the choice of occurrence statistic. The paper is honest about the instability in Sec. V-A, which is why this is a concern about presentation and robustness rather than a claim of fabrication. The reader's weakest assumption (W1's choice is load-bearing) is exactly this point; I agree. No code or data artifacts are released, and the flow's CRPS row lacks seed error bars, but the metric-dependence alone would require the same conditional verdict: the paper should qualify the headline, report per-statistic win counts, and ship artifacts. Thus the reader's CONDITIONAL verdict is appropriate and unchanged.","tokens_in":8640,"tokens_out":6163,"duration_ms":54258,"concrete_test":"Apply the paper's own matched protocol (Sec. V) but set the occurrence autocorrelation error as the primary target statistic; count wins for Hurdle-AR vs Flow on the six datasets. If Hurdle-AR does not win on at least five datasets — or if a bootstrap 95% interval on mean rank in Table III puts Flow ahead of Hurdle-AR — then the 'beats the flow on five of six datasets' headline must be restricted to zero-run-length statistics and the paper's abstract/conclusion revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central model-level claim — repeated in the abstract, Section V, and conclusion — is that an autoregressive occurrence hurdle beats the conditional flow on five of six datasets, by up to 153x. The support is Table II's W1 zero-run-length distance, but Table III explicitly shows the model ordering is not stable across occurrence statistics. Under the occurrence autocorrelation (ACF), an independent second-order statistic, the flow's mean rank is 2.0 versus 2.2 for the hurdle; the W1–ACF rank correlation is only +0.55, and the two statistics that do not share the run-length construction (ACF and spectrum) agree least with each other. The paper reports this honestly in Sec. V-A, yet the abstract and conclusion still state 'beats' without the W1 qualifier. Because the model-level benchmark is one of the paper's three advertised contributions, the unqualified dominance claim is load-bearing: it is true only under the chosen zero-run-length lens. A secondary reproducibility issue (no code/data/config release, no error bars in the flow's CRPS row of Table II) makes the comparison impossible to audit, but the metric-dependence is the more direct threat to the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how generative time-series models are evaluated when the data carry a point mass at zero. It makes three main contributions. First, it shows that the standard rolling-origin protocol can produce evaluation windows whose zero-atom structure is very different from the dataset as a whole (e.g., covid_deaths is 42% zeros as a dataset but 13% in the evaluation windows, rideshare 47% versus 5%), and that this mismatch can reverse model rankings. Second, it proposes a permutation control under which the per-step predictive marginals, and therefore CRPS, are invariant by construction while the temporal coupling is destroyed, allowing one to measure how much a model's learned coupling contributes to a chosen statistic. Third, it benchmarks seven models on six datasets under a matched protocol, reporting that an autoregressive occurrence hurdle beats a conditional flow on five of six datasets under zero-run-length W1 distance, and that the flow's occurrence statistics are unstable across training seeds while the baselines are deterministic. The paper explicitly distinguishes evaluation-level claims, which it derives from definitions, from model-level empirical claims, and it reports the dependence of model ordering on the choice of occurrence statistic.","tokens_in":8800,"tokens_out":3482,"duration_ms":33957,"significance":"The evaluation-level findings are solid and useful. The CRPS invariance under per-step permutation is a correct consequence of CRPS being a function of the per-timestep marginals, and the paper verifies it exactly with a sorted-sample closed form rather than a random pairing estimator. The protocol-mismatch observation in Sec. IV-A is an important caution for the generative time-series benchmarking literature, and the paper honestly documents a case where the protocol reversed one of its own conclusions. The permutation control in Sec. IV-B is simple but genuinely diagnostic, and the paper is careful about its scope. The empirical benchmark is transparent about the threshold protocol and the exclusion of covid_deaths, and the paper explicitly reports that model ordering changes across occurrence statistics. However, the headline model-level claim, repeated in the abstract and conclusion, is metric-dependent and is not supported under all five reported occurrence statistics: under the occurrence autocorrelation in Table III the flow's mean rank is 2.0 versus 2.2 for the hurdle.","major_comments":[{"comment":"The claim that an autoregressive hurdle 'beats' a conditional flow on five of six datasets is load-bearing and is stated without qualification in the abstract and conclusion, but it holds only under the run-length family of statistics. Under the occurrence autocorrelation in Table III, the flow's mean rank is 2.0 versus 2.2 for Hurdle-AR, and the W1–ACF rank correlation is only +0.55. The paper honestly reports this in Sec. V-A, but the abstract and conclusion still present the dominance as an unqualified finding. Please either qualify the headline to the zero-run-length statistic or provide a substantive justification for why W1 should be the primary lens.","section":"Abstract; Sec. V-A; Table III; Conclusion"},{"comment":"The flow's CRPS row in Table II reports point estimates without standard deviations, even though the paper argues in Sec. V that a single-run number for a flow is not reportable on occurrence statistics at the margins these comparisons are decided by. If seed spread matters for the primary statistic, it should be reported for CRPS as well. In addition, no code, data, or configuration release is mentioned anywhere, which makes the empirical benchmark impossible to audit. Please provide per-seed results and an availability statement, or at minimum detailed hyperparameters and data splits.","section":"Table II (lower panel); Sec. V; reproducibility"},{"comment":"The claim that the flow's occurrence statistics vary by up to 62% across training seeds is supported only for the zero-run-length W1 statistic in Table II, where the flow is shown as mean ± s.d. over five seeds. The same 62% figure appears in the abstract and conclusion, and the text states 'the occurrence ACF gives the same ordering' without showing any per-seed or uncertainty values for the ACF or the other four statistics. Please provide the per-seed numbers (or a table of spreads) for all reported occurrence statistics before this stability claim can be verified.","section":"Sec. V; Table II; abstract"}],"minor_comments":[{"comment":"The notation 'π_h ∼ Unif(S_K) i.i.d.' is slightly compressed; please state explicitly that the permutations are independent across horizons h, since that independence is what makes the multiset of proposed values unchanged at each step.","section":"Eq. (4)"},{"comment":"The descriptors s and φ are defined in Eqs. (2) and (3), but H(Z|calendar) is not fully specified in the text; please state that it is the conditional entropy estimated with a cross-fitted logistic model on calendar harmonics.","section":"Table I"},{"comment":"The definition of 'spell quantiles' should be expanded: the text gives the set {50, 90, 99} but does not specify whether the mean relative error is computed before or after averaging across windows, or how ties and zero-length spells are handled.","section":"Sec. V-A"},{"comment":"The figure would be clearer if the caption explained the meaning of the negative covid value and the statement 'neither protocol recovers it'; currently the reader must infer the axis convention from the main text.","section":"Fig. 1"}],"recommendation":"major_revision","confidential_remarks":"The evaluation-level contributions are sound and the paper is admirably honest about the metric dependence of its own results. The main concern is that the abstract and conclusion overstate a model-level dominance result that is not robust across the reported occurrence statistics, and the lack of code/data release makes the empirical benchmark unauditable. With a qualified headline and release of artifacts, this would be a solid contribution; as written, the central benchmark claim needs revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — read this for the evaluation-level findings, not the model shootout. The permutation control is the real contribution: permuting sample indices per horizon step leaves CRPS fixed by construction while killing temporal coupling, so it measures exactly what a marginal score can and cannot see. The paper also documents that the standard rolling-origin protocol can score models on windows whose zero fraction bears no resemblance to the dataset (47% vs 5% on rideshare), and honestly reports that this mismatch inverted one of their own conclusions. That is carefully done, clearly written, and the evaluation-level claims follow from definitions.\n\nThe model benchmark is less solid. The 'autoregressive hurdle beats the flow on five of six datasets, up to 153x' is true only under the zero-run-length W1 statistic. Table III shows the ordering is not stable: under occurrence ACF the flow has better mean rank than the hurdle (2.0 vs 2.2). The paper says this honestly in Sec V-A, but the abstract and conclusion state 'beats' without the qualifier. That is a real soft spot, though not fatal — the body contains the caveat.\n\nReproducibility is the bigger problem. No code, data, or exact configs are released. The flow's CRPS row in Table II has no error bars, and while the baseline is deterministic, the flow has 5-seed variance. The empirical comparison is impossible to audit in full. The paper's own caution about seed instability makes this omission more pointed.\n\nThe negative result on the hybrid (semi-Markov substitution worsens the flow on five of six) is a nice falsification of the obvious fix, and the seed-spread finding is a useful warning. The literature coverage is appropriate, including the link to the Schaake shuffle.\n\nOverall: the measurement contributions deserve serious attention; the model-level headline needs a qualifier and the artifacts need to ship. A good referee could turn this into a solid paper. Yes, send it to peer review.","headline":"Solid evaluation-level contribution; the 'hurdle beats flow' headline is metric-dependent and needs a qualifier.","tokens_in":9356,"tokens_out":2211,"would_cite":true,"duration_ms":20886,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Standard evaluation windows can badly misrepresent zero-heavy time series, and a simple autoregressive occurrence model beats a conditional flow on five of six datasets.","keywords":["point mass","zero-inflated time series","generative time-series models","flow matching","rolling-origin evaluation","occurrence process","CRPS","run-length statistic"],"falsifier":"Score all seven models under the occurrence autocorrelation statistic on the same windows and seeds; the paper's own Table III already shows the flow (mean rank 2.0) above the autoregressive hurdle (2.2), so the five-of-six claim does not survive when ACF is the chosen occurrence statistic.","tokens_in":8367,"feed_emoji":"📊","tokens_out":9644,"duration_ms":80939,"temperature":0.7,"pith_summary":"Many time-series benchmarks concentrate a large share of probability mass on a single value, usually zero, yet the generative models being scored are continuous and cannot place exact mass on that value. This paper shows that the standard rolling-origin evaluation protocol can score a model on windows whose zero rate differs from the dataset's by dozens of percentage points, and that this artifact reversed one of the paper's own conclusions about which model is best. It introduces a permutation control that destroys temporal coupling while leaving every per-step marginal, and therefore CRPS, unchanged by construction, making the contribution of learned coupling to a chosen statistic directly measurable. On six atom-bearing datasets, an autoregressive occurrence hurdle beats a conditional flow on five of them under zero-run-length error, by up to a factor of 153, while the flow's occurrence statistics vary by up to 62% across training seeds. The paper argues that reporting the atom rate of the evaluation windows, comparing a model with its own decorrelated samples, and reporting seed spreads should become standard practice.","feed_headline":"Simple occurrence model beats flow on 5 of 6 zero-heavy datasets","feed_subtitle":"Evaluation windows can be 40 points off the dataset's zero rate, and changing the statistic flips the model ranking.","key_machinery":"The load-bearing object is the atom at zero: in all six datasets a single value carries between 31% and 94% of the probability mass, while the flow and diffusion samplers under study push forward an absolutely continuous law through an invertible map, so they cannot represent an exact atom (Eq. 1). The measurement mechanism is the decorrelation control (Eq. 4): permuting the sample index independently at each horizon step leaves the multiset of proposed values, and hence every per-step marginal and CRPS, completely unchanged, while destroying the trajectory coupling that carries run-length, autocorrelation and spectral statistics. Comparing a model's samples with their own decorrelated versions therefore measures exactly how much its learned coupling contributes to a chosen statistic. The paper's headline statistic is the Wasserstein distance between zero-run-length distributions ($W_1$), a run-length-family measure that the paper itself contrasts with four alternatives, including the occurrence autocorrelation under which the flow's rank improves.","core_discovery":"On data where a single value, typically zero, carries a large share of the probability mass, the paper finds that evaluation practice, not model capacity, is the main source of misleading conclusions. Because flow-matching and diffusion samplers move an absolutely continuous distribution through an invertible (or effectively non-singular) map, they cannot put exact probability on a point, yet six standard benchmarks put 31–94% of observations at zero. The paper shows that the rolling-origin protocol, which scores the last H steps of each series, can select windows whose zero rate is far from the dataset's (42.3% to 13.1% on one benchmark; 46.9% to 5.3% on another), and that this artifact inverted the paper's own ranking of the best occurrence model. It then isolates the contribution of temporal coupling by independently permuting the sample index at each horizon step, which leaves every per-step marginal, and hence CRPS, exactly unchanged while destroying the coupling, and it benchmarks seven models on identical windows. An autoregressive occurrence hurdle is the best occurrence model on five of six datasets under zero-run-length error, beating the conditional flow by up to a factor of 153, while the flow's occurrence statistics move by up to 62% across seeds and the ordering changes when a different occurrence statistic is used.","pith_inferences":["If window-atom mismatches of this size are common, previously published comparisons between continuous-state generative models on intermittent data may be ranking models on regimes the datasets barely contain; a series-wise split with window-atom reporting could change published conclusions beyond this paper's six datasets.","The decorrelation control is cheap, one permutation per horizon step, and could serve as a routine diagnostic for any generative model, including diffusion and autoregressive samplers, to test whether learned temporal dependence actually contributes to a reported statistic.","The success of the simple autoregressive hurdle suggests that for intermittent series the occurrence process is the bottleneck, and that explicit zero/non-zero modelling is likely to remain competitive even as continuous-state generators improve; the paper's negative hybrid result shows which explicit process is substituted matters.","Because the five occurrence statistics disagree on ordering, and the two independent ones agree least, benchmark conclusions about occurrence behaviour should be presented as a family of numbers rather than a single head-to-head, otherwise the chosen statistic can pick the winner."],"forward_implications":["The atom rate of the evaluation windows, not just of the dataset, must be reported; on two of the benchmarks the gap is 29 and 42 percentage points respectively, and it reversed one of the paper's own conclusions.","A model should be compared with its own decorrelated samples before crediting it with temporal structure: on one dataset the flow and its decorrelated version are within 12% on zero-run error, and on another the decorrelated version is better.","An autoregressive occurrence hurdle, a logistic classifier rolled out one step at a time, is the strongest occurrence model on five of six datasets, beating the conditional flow by up to a factor of 153 on zero-run error.","A single-seed comparison of occurrence statistics is not reportable for this flow: its run-length error spread over five seeds ranges from 7% to 62%, while every occurrence baseline is deterministic.","Model rankings under occurrence statistics depend on the statistic; the two statistics that do not share the run-length construction agree with each other least, so claims about the best occurrence model must name the statistic."],"supporting_citations":[{"why":"Supplies the conditional flow model that the benchmark compares against, the concrete instance of a continuous-state ODE-based sampler with an invertible map.","marker":"[1]"},{"why":"Defines the flow matching framework, whose absolute-continuity limitation motivates the paper's atom argument.","marker":"[12]"},{"why":"Introduces the occurrence/size decomposition of intermittent demand that motivates the hurdle and semi-Markov baselines.","marker":"[13]"},{"why":"Provides the chain-dependent occurrence process used as a baseline occurrence model.","marker":"[18]"},{"why":"Establishes that CRPS is a strictly proper but marginal scoring rule, grounding the claim that per-step marginals cannot see temporal coupling and supplying the energy score.","marker":"[24]"},{"why":"Supplies the variogram score, a dependence-sensitive alternative that the paper reports but which does not isolate coupling from marginals.","marker":"[26]"},{"why":"Describes the inverse operation to the paper's decorrelation control, used to impose rank dependence on postprocessed marginal samples.","marker":"[27]"},{"why":"Is the source of the benchmark time-series datasets on which the atom rates and evaluations are computed.","marker":"[30]"}],"fun_headline_variants":["Evaluation flaw flips time-series model ranking","Occurrence model beats flow on zero-heavy series","Rolling-origin bias skews time-series benchmarks","Zero-inflated data exposes evaluation artifact"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that zero-run-length error is the right primary lens for occurrence behaviour, since under the occurrence-autocorrelation statistic the flow ranks above the hurdle and the headline five-of-six comparison does not survive.","fun_headline_variants_meta":{"raw":{"variants":["Evaluation flaw flips time-series model ranking","Occurrence model beats flow on zero-heavy series","Rolling-origin bias skews time-series benchmarks","Zero-inflated data exposes evaluation artifact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1291,"prompt_tokens":1057,"completion_tokens":234,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":673,"completion_tokens_details":{"reasoning_tokens":175}},"tokens_in":673,"tokens_out":234,"duration_ms":2918,"temperature":1.0,"reasoning_tokens":175,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:36:13.922243+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score all seven models under the occurrence autocorrelation statistic on the same windows and seeds; the paper's own Table III already shows the flow (mean rank 2.0) above the autoregressive hurdle (2.2), so the five-of-six claim does not survive when ACF is the chosen occurrence statistic.","supporting_citations":[{"cited_title":"Flow matching with gaussian process priors for probabilistic time series forecasting,","cited_arxiv_id":null,"evidence_quote":"Supplies the conditional flow model that the benchmark compares against, the concrete instance of a continuous-state ODE-based sampler with an invertible map."},{"cited_title":"Forecasting and stock control for intermittent demands,","cited_arxiv_id":null,"evidence_quote":"Introduces the occurrence/size decomposition of intermittent demand that motivates the hurdle and semi-Markov baselines."},{"cited_title":"Precipitation as a chain-dependent process,","cited_arxiv_id":null,"evidence_quote":"Provides the chain-dependent occurrence process used as a baseline occurrence model."},{"cited_title":"Strictly proper scoring rules, prediction, and estimation,","cited_arxiv_id":null,"evidence_quote":"Establishes that CRPS is a strictly proper but marginal scoring rule, grounding the claim that per-step marginals cannot see temporal coupling and supplying the energy score."},{"cited_title":"Variogram-based proper scoring rules for probabilistic forecasts of multivariate quantities,","cited_arxiv_id":null,"evidence_quote":"Supplies the variogram score, a dependence-sensitive alternative that the paper reports but which does not isolate coupling from marginals."},{"cited_title":"The schaake shuffle: A method for reconstructing space–time variability in forecasted precipitation and temperature fields,","cited_arxiv_id":null,"evidence_quote":"Describes the inverse operation to the paper's decorrelation control, used to impose rank dependence on postprocessed marginal samples."}],"review_version":1}