{"id":"16fbf79b-0263-403d-9d3d-2ab47cb95545","arxiv_id":"2608.09684","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Using point-process registration and Wasserstein PCA, the authors find that earlier, not more extensive, COVID-19 restrictions are associated with flatter first-wave infection curves across US states.","lead":"This study separates the size of COVID-19 infection waves from their timing across US states, then relates both to the timing and overall level of government restrictions. States that imposed restrictions earlier showed flatter infection curves, while the cumulative amount of stringency showed no clear association with total cases.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The timing-flatness association may be an artifact of aligning every state's densities to its own case-defined onset; a common-calendar rerun is needed.","rationale":"The paper is reproducible, self-contained in its WTPCA derivation, and honest about endogeneity and the failure of the Cox process model (Section 4). Those are real strengths. However, the central quantitative claim rests on the sign of one regression coefficient linking a timing score to a flatness score, and both scores are derived from densities supported on state-specific windows anchored to the same case-count milestone. That creates a built-in coupling between the outcome (flatness) and the reference frame for 'earlier' restrictions. I regard this alignment issue as more load-bearing than the Cox-process misspecification emphasized by the reader: WTPCA scores are descriptive summaries of the observed densities and remain well-defined even if the generative point-process model is wrong, whereas the self-referential clock can manufacture precisely the association the paper reports. The proposed fixed-calendar rerun is immediate, uses the provided code, and would either corroborate or overturn the headline; hence the CONDITIONAL verdict remains appropriate (unchanged). Separately, the joint likelihood-ratio null in Section 4.2 is mis-specified for the conclusions it is said to support: it sets beta_3c = beta_2c = beta_1c = beta_11 = beta_12 = 0, omitting the budget coefficients on the case scores (beta_31, beta_32) and the Stringency PC2 coefficients on those scores (beta_21, beta_22). This should be corrected in a revision, and if the corrected null yields p<0.05 the null claims would need to be rephrased.","tokens_in":26339,"tokens_out":15417,"duration_ms":146762,"concrete_test":"Re-estimate the full pipeline using a common calendar window for all 50 states (e.g., March 1 to June 30, 2020) instead of the state-specific onset-defined windows, keeping the same smoothing bandwidth, WTPCA rank, and vector-on-vector regression specification. If the Stringency PC1 coefficient on Case PC2 remains negative with p<0.05, the alignment artifact cannot explain the headline; if it attenuates toward zero or changes sign, the reported association is an artifact of the case-defined clock.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 fixes each state's first-wave window as 21 days before the day cumulative cases reach 10 per million, so every infection density has the same cumulative mass at t=21 and all 'phase' scores are measured relative to a clock defined by the state's own case trajectory. The response Case PC2 (flatness) and the predictor Stringency PC1 (timing) are both computed on this self-referential clock. A state with a slow early epidemic reaches the 10-per-million threshold later in calendar time, tends to have a flatter density, and has more calendar days before the threshold in which restrictions can be enacted; mechanically, such a state receives a low Stringency PC1 ('early') and a high Case PC2 ('flat'). A fast-growing state tends to receive the opposite combination. The negative coefficient of Stringency PC1 on Case PC2 (Table 1, full Model: -0.35, SE 0.17; MANOVA p=0.015 in Table 2) may therefore reflect the growth rate used to define the window rather than a genuine association between policy timing and curve shape. The stability checks in Section 4.3 vary smoothing and window lengths but do not remove the onset-based alignment; the authors' endogeneity caveat does not address this mechanical coupling.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes daily COVID-19 infection counts in the 50 US states during the first wave, viewed as realizations of a point process with random time warping. Using Panaretos and Zemel's amplitude-phase separation, Wasserstein tangent-space PCA on the infection and stringency densities, and a vector-on-vector regression with total cases and infection PC scores as responses, the authors report that earlier implementation of restrictions is associated with flatter infection curves, while no aspect of stringency is significantly associated with total infection counts and the overall stringency budget is not significantly associated with the outcomes considered.","tokens_in":26618,"tokens_out":3271,"duration_ms":30714,"significance":"If the phase scores measure the intended temporal dynamics, the paper provides a credible and interpretable application of recent optimal-transport tools to a policy-relevant question, and it introduces a general distribution-on-distribution regression workflow with external covariates. The authors' care is visible in the MANOVA testing, regression diagnostics, added-variable plots, stability checks, and a self-contained consistency proof (Proposition 1). The manuscript is also reproducible via the provided repository. However, the central empirical claim rests on the assumption that the estimated warps and PC scores faithfully separate timing from intensity; this assumption is challenged both by the paper's own evidence that the Cox point-process model is not the right model for infection counts and by the onset-aligned window construction, which may induce a mechanical association between stringency timing and case-curve flatness.","major_comments":[{"comment":"The paper itself states, in the paragraph accompanying Figure 4, that the FPCA of the registered log-count curves shows a second eigenfunction explaining 11% of the variance, and concludes that 'the Cox point process model adopted, for example, by Gajardo and Müller is arguably not the right model for infection counts.' Since the phase scores used as regression inputs are estimated under exactly this Cox-process registration model, the misspecification directly bears on the validity of the headline association between Stringency PC1 and Case PC2. The authors need to show that the estimated warps are still consistent for the true warps under a more general model, or to re-estimate the phase component with a registration method that does not require the rank-one Cox structure. Without such evidence, the temporal interpretation of the phase scores in Table 1 is not established.","section":"Section 4, Figure 4"},{"comment":"Each state's first-wave window is defined as 21 days before the day cumulative cases reach 10 per million. This onset-based alignment creates a built-in coupling between the 'timing' of stringency and the 'flatness' of the case curve: a state with slower early growth reaches the threshold later in calendar time, has more pre-threshold calendar days available for enacting restrictions, and tends to have a flatter case density, while a fast-growing state has fewer such days and a more peaked density. The negative coefficient of Stringency PC1 on Case PC2 in Table 1 (full model: -0.35, SE 0.17; MANOVA p=0.015 in Table 2) may therefore reflect the growth rate used to define the window rather than a genuine policy association. The stability checks in Section 4.3 vary bandwidths and window lengths but do not remove the onset-based alignment; a common-calendar analysis, or a version using a fixed calendar window for all states, is needed to break this mechanical coupling.","section":"Section 2, Section 4.3"},{"comment":"The statement that 'the most important qualitative conclusions ... are quite robust and can be reached regardless of the specific choices, as long as the core first wave time window is included' is not substantiated by any reported results in the paper. The stability analysis is described only verbally; the repository is said to contain the scripts, but the paper itself reports no tables or figures showing, for example, the range of coefficients on Stringency PC1 across different bandwidths, window lengths, or control sets. Given that the central claim depends on the stability of this coefficient, the authors should present the actual stability-check results in the manuscript.","section":"Section 4.3, first paragraph"}],"minor_comments":[{"comment":"The table header includes a stray space in 'T able 1'; also, the significance codes and the formatting of the p-value column should be made consistent between the submodel and the full model.","section":"Table 1"},{"comment":"The phrase 'As an initial exploratory step, we perform FPCA on the smoothed, registered, and log-transformed infection count curves' should specify that the analysis is on the registered curves, since Figure 4 is also used to argue against the Cox-model assumption.","section":"Section 4, first paragraph"},{"comment":"The sentence beginning 'How many components to retain is largely settled in the case of the stringency index' is followed by a discussion of infection counts; for clarity, the authors should state explicitly that the decision to retain exactly two components for both variables is a modelling choice, not driven by a hard threshold.","section":"Section 4.1, last paragraph"},{"comment":"The score plot in Figure A.1 includes the District of Columbia, while the main analysis excludes it; the figure caption should note this difference to avoid confusion.","section":"Appendix A, Figure A.1"},{"comment":"The authors note that the Democratic vote share variable 'was ultimately dropped as unimportant,' but no diagnostic or test is shown to support this decision; a brief example of the sensitivity to including this control would be helpful.","section":"Section 2, paragraph on controls"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-executed application of sophisticated methodology, but the central finding is undercut by two intertwined threats: the paper's own admission that the point-process model is misspecified for infection counts, and the onset-aligned window that can create a mechanical correlation between the predictor and response scores. I would encourage the editor to request a common-calendar robustness check and a more detailed stability table; without those, the headline association may be an artifact of the pre-processing. The paper's contribution to distribution-on-distribution regression is valuable, but the applied claim needs stronger support before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a genuinely careful piece of applied statistics with a useful methodological kernel: Wasserstein tangent-space PCA scores plus total masses fed into a vector-on-vector regression is a clean, interpretable recipe for distribution-valued data, and the paper does the work of explaining the geometry and providing consistency proofs. The writing is clear, the limitations section is candid, and the code is available. Second, the headline applied result—that timing of restrictions, not cumulative stringency, is associated with flatter first-wave curves—is endangered by the way the time windows are defined.\n\nThe problem is the onset-based alignment. Each state's 120-day window starts 21 days before that state reaches 10 cumulative cases per million. So every state's infection density has the same cumulative mass at day 21, and all phase scores are measured on a clock defined by that state's own case trajectory. A slow, late-hit state will tend to have a flatter density and will also start its window later in calendar time, when national stringency levels were already high. That mechanically produces an early stringency score (mass concentrated early in the window) and a flat case curve. The negative coefficient on Stringency PC1 for Case PC2 (Table 1, full model: -0.35, SE 0.17; MANOVA p=0.015) may therefore reflect the growth rate used to build the window rather than a genuine policy effect. The stability analysis varies bandwidths and window lengths but never removes the onset-based clock; it cannot detect this coupling. The endogeneity caveat is real but does not address this mechanical confound.\n\nThe paper's own registered-curve FPCA (Figure 4) also shows an 11% second eigenfunction, which the authors honestly interpret as evidence that the Cox-process warping model is not right for these data. That admission undercuts the point-process framing, though the WTPCA on the raw densities could still stand on its own as a dimension-reduction tool.\n\nWhat does the paper do well? The methodological exposition is strong, the regression diagnostics are thorough, and the interpretation of the PC scores as time shift versus flatness is careful. It is a serious contribution to the FDA toolbox, even if the applied conclusion is shaky.\n\nThe paper deserves a serious referee, but the referee should require a common-calendar alignment (or at least a control for calendar onset date) before accepting the substantive claim. Without that, the key result remains unconvincing. Send it to review, but with major revision in mind.","headline":"Competent and honest paper with a useful WTPCA regression recipe, but the headline timing-flatness association is likely an artifact of the state-specific onset-based window and needs a common-calendar rerun before it can be taken seriously.","tokens_in":27114,"tokens_out":3300,"would_cite":true,"duration_ms":35283,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62R10","49Q22","62H25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Earlier COVID-19 restrictions, not stricter or more numerous ones, are the dimension that associates with flatter first-wave infection curves, and that signal lives in the timing component of the data.","keywords":["functional data analysis","optimal transport","point process registration","Wasserstein PCA","Oxford Stringency Index","phase variation","amplitude-phase separation","vector-on-vector regression"],"falsifier":"Re-estimate the regression after first removing the second eigenfunction of the registered log-count curves (the 11% mode the paper itself documents) or after allowing the latent process to have more than one amplitude dimension; if the negative association between restriction timing and infection flatness weakens, disappears, or changes sign, the headline result is an artifact of the time-warping assumption rather than a genuine policy signal.","tokens_in":26152,"feed_emoji":"📉","tokens_out":10901,"duration_ms":87759,"temperature":0.7,"pith_summary":"This paper asks which dimension of government response mattered for the shape of the first COVID-19 wave across the fifty US states. Treating each state's daily infection counts as a random point process, a stream of events arriving over time, it separates amplitude variation (how many cases arrived and how concentrated) from phase variation (when the wave arrived relative to a common local clock), and it does the same for the Oxford Stringency Index. These components, together with the total case count and the cumulative stringency budget, then enter a vector-on-vector regression. The central finding is that the timing of restrictions is the dimension that associates with the shape of the curve: states that placed their stringency earlier tend to have flatter infection curves, while no aspect of stringency is significantly associated with total infection counts and the overall stringency budget is not associated with any outcome considered. These are presented as associations, not causal effects, in a design the authors acknowledge has a feedback loop between cases and policy.","feed_headline":"Restriction timing, not strictness, tracks flatter COVID curves","feed_subtitle":"Across US states, earlier measures pair with slower spread; cumulative stringency shows no link.","key_machinery":"The engine of the analysis is canonical amplitude-phase separation for point processes (the procedure of [53]): each state's observed infection process is modelled as a random time warp of a latent point process, the warps are estimated as the optimal transport maps from each state's smoothed intensity to the empirical Frechet mean in the Wasserstein metric, and the transport maps themselves become the phase scores while the registered processes are the amplitude. On top of this, Wasserstein tangent-space PCA linearises the space of densities through their quantile functions, so the principal component scores carry two transparent readings: the first component is an overall time shift (when the wave, or the restrictions, happened), and the second contrasts lower against upper quantiles, that is, flatness versus spikiness of the curve. These scores, plus the total case count and the stringency budget, are the inputs to a vector-on-vector regression whose joint significance is assessed with the Pillai test.","core_discovery":"On the paper's own terms, the discovery is that phase variation, the temporal dynamics of the pandemic, usually treated as a nuisance to be registered away, is the carrier of the policy-relevant signal in COVID-19 infection data. Using the point-process registration of [53] and Wasserstein tangent-space PCA, the first principal component of the stringency curves acts as an overall time shift of restrictions, and the regression finds a significant negative association between this time shift and the flatness component of the infection curves: earlier stringency goes with flatter first-wave case curves, and later stringency with more spiked ones. In the model with control variables, higher population density is associated with higher total cases and higher GDP with earlier infection increases, but no aspect of stringency is significantly associated with total infection counts, and a likelihood-ratio test (p-value about 0.2) finds no joint evidence that the stringency budget or its timing contributes to the total-count response. The paper is explicit that these are associations in an observational design with a feedback loop between cases and restrictions, and that the point-process model is an idealisation the data only approximately satisfy.","pith_inferences":["A test the paper does not run: delete the 11% second eigenfunction from the registered curves before computing phase scores and re-fit the regression; a stable coefficient would show the timing-flatness result does not depend on the Cox-process assumption the paper itself rejects.","The results imply a simple, checkable surrogate, days from first local cases to first major restrictions, should reproduce the negative association with curve flatness in the same public data; if it does not, the Wasserstein phase score encodes something extra that the simple proxy misses.","Because restriction timing and infection timing are read from the same state-specific clock, part of the association may run from the epidemic to the policy (states hit earlier locked down earlier); the paper's observational design cannot separate that direction from the reverse, so the headline is best read as a description of co-movement between policy timing and curve shape."],"forward_implications":["If the association is correct, evaluations of pandemic policy should record when restrictions were imposed relative to the local epidemic clock, separately from how strict or how cumulative they were, because only the timing dimension shows a significant association.","The phase component of infection curves is not noise to be discarded: discarding it would throw away exactly the variation that carries the clearest policy relationship.","Analyses that summarise policy by average or cumulative stringency alone would find no association, and would wrongly conclude that non-pharmaceutical measures did not matter.","The WTPCA-plus-total-mass scheme gives applied researchers a template for putting naturally distribution-valued covariates (age, income, exposure) into ordinary multivariate regression while keeping a transport-based interpretation."],"supporting_citations":[{"why":"Supplies the canonical registration procedure that separates phase from amplitude in point processes; the paper's central machinery.","marker":"[53]"},{"why":"Provides the Oxford Stringency Index, the main predictor, treated as a measure on the same time window.","marker":"[30]"},{"why":"Source of the daily US state infection counts that form the response processes.","marker":"[67]"},{"why":"The local-polynomial density estimator used to smooth each state's stringency index into a density.","marker":"[11]"},{"why":"The first-wave window definition (120 days from 21 days before 10 cases per million) adopted for the time domain.","marker":"[10]"},{"why":"Underlies the transformation of densities into a Hilbert space that Wasserstein tangent-space PCA builds on.","marker":"[56]"},{"why":"Reference for the vector-on-vector regression estimator used to relate the scores.","marker":"[46]"},{"why":"Provides the MANOVA/Pillai test framework used to assess joint significance across the multivariate response.","marker":"[48]"}],"fun_headline_variants":["Earlier lockdowns flatten COVID curves, strictness doesn't","COVID curve shape tracks restriction timing, not severity","Timing of restrictions predicts flatter COVID curves, study finds","Flatter COVID waves tied to earlier restrictions, not stringency","Restriction timing beats strictness for COVID curve shape"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis rests on the assumption that every state's infection curve is one shared underlying pattern stretched and squeezed in time, so that after undoing those time changes the only differences left among states are in size; because the paper itself finds a second pattern of variation (11% of the variance) that this assumption cannot produce, the time-adjustment scores driving the main result could be measuring the wrong thing.","fun_headline_variants_meta":{"raw":{"variants":["Earlier lockdowns flatten COVID curves, strictness doesn't","COVID curve shape tracks restriction timing, not severity","Timing of restrictions predicts flatter COVID curves, study finds","Flatter COVID waves tied to earlier restrictions, not stringency","Restriction timing beats strictness for COVID curve shape"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00043,"raw_usage":{"total_tokens":2214,"prompt_tokens":983,"completion_tokens":1231,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":1151}},"tokens_in":599,"tokens_out":1231,"duration_ms":9447,"temperature":1.0,"reasoning_tokens":1151,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:45:07.060434+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the regression after first removing the second eigenfunction of the registered log-count curves (the 11% mode the paper itself documents) or after allowing the latent process to have more than one amplitude dimension; if the negative association between restriction timing and infection flatness weakens, disappears, or changes sign, the headline result is an artifact of the time-warping assumption rather than a genuine policy signal.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the canonical registration procedure that separates phase from amplitude in point processes; the paper's central machinery."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Oxford Stringency Index, the main predictor, treated as a measure on the same time window."},{"cited_title":"URL: https://github.com/nytimes/covid-19-data","cited_arxiv_id":null,"evidence_quote":"Source of the daily US state infection counts that form the response processes."},{"cited_title":"and Ma, X","cited_arxiv_id":null,"evidence_quote":"The local-polynomial density estimator used to smooth each state's stringency index into a density."},{"cited_title":"and Wang, J.-L","cited_arxiv_id":null,"evidence_quote":"The first-wave window definition (120 days from 21 days before 10 cases per million) adopted for the time domain."},{"cited_title":"and Müller, H.-G","cited_arxiv_id":null,"evidence_quote":"Underlies the transformation of densities into a Hilbert space that Wasserstein tangent-space PCA builds on."},{"cited_title":"and Bibby, J","cited_arxiv_id":null,"evidence_quote":"Reference for the vector-on-vector regression estimator used to relate the scores."},{"cited_title":"and Kelley, K","cited_arxiv_id":null,"evidence_quote":"Provides the MANOVA/Pillai test framework used to assess joint significance across the multivariate response."}],"review_version":1}