{"id":"3ed641ed-4706-4950-9f90-1814ac2efb46","arxiv_id":"2508.14596","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A statistics paper claims anytime-valid top-m variable screening with FCR-controlled post-screening inference, but the supplied manuscript body does not contain that work.","lead":"The abstract describes a statistical method for sequentially removing variables that cannot be among the top m, while keeping a high-probability guarantee that the true top-m set stays covered, plus post-screening intervals that control the false coverage rate. The body text supplied for review is an unrelated paper about optimizing turbulent pipe flow, so the claimed statistics content could not be examined.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Submission body is arXiv:2508.14593 (pipe-flow paper), not the arXiv:2508.14596 SCS/PSI manuscript; the central claims cannot be evaluated from the supplied text.","rationale":"The reader's verdict is UNVERDICTED with low confidence, and their weakest assumption correctly identifies that the abstract's guarantees require explicit assumptions and derivations that cannot be located because the body text is a different paper. My stress-test finds no additional load-bearing concern beyond this structural mismatch, but it is decisive: without the actual manuscript, no content-based evaluation of the central claim is possible. The reviewing rule requires treating the supplied text as in-scope evidence; doing so confirms that the text asserts no statistical methodology whatsoever. The only concrete check that could change this is retrieving the true arXiv:2508.14596 document. If the true document is retrieved and does contain the SCS/PSI method, then a fresh review would be required, but that is a new submission. I agree with the reader's judgment: the verdict should remain UNVERDICTED until the correct manuscript is supplied.","tokens_in":21644,"tokens_out":1072,"duration_ms":14013,"concrete_test":"Download the actual arXiv:2508.14596v1 record (metadata, abstract, and full body) from arXiv and mechanically compare the full text against the submitted pipe-flow manuscript. If the body is identical to arXiv:2508.14593v1, the mismatch is confirmed and the UNVERDICTED status stands because the claimed SCS/PSI content is entirely absent. If a corrected submission containing the SCS/PSI manuscript is provided, re-run the review on that text, treating the current annotation as withdrawn.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The paper's central claim—that Sequential Correct Screening provides an anytime-valid sequence of subsets containing the true top-m variables with high probability, and that the PSI procedure controls FCR whenever conducted—requires explicit theorem statements, regularity conditions, and derivations. The submitted full text is an entirely different manuscript (arXiv:2508.14593v1, 'Bayesian minimisation of energy consumption in turbulent pipe flow via unsteady driving') by different authors. No SCS/PSI definitions, assumptions, proofs, simulations, or data analysis appear in the supplied text. Consequently, there is no way to assess the correctness of the claimed anytime-validity or FCR control, no way to identify hidden assumptions, and no way to verify even the internal consistency of the statistical arguments. This is a load-bearing structural failure: the evidence necessary to evaluate the central claim is absent. The abstract alone is insufficient because it omits the mathematical model, the estimation procedure, the filtration/stopping-time details, and the precise conditions under which the guarantees hold. Nothing in the supplied text supports or contradicts the abstract; the only honest statement is that the submission, as received, is unverifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission (arXiv:2508.14596) presents an abstract claiming a new Sequential Correct Screening (SCS) methodology that returns a sequence of variable subsets containing the true top-m set with high probability at any stopping time, and a post-screening inference (PSI) procedure that controls the false coverage rate (FCR) whenever conducted. The abstract further states that theoretical guarantees, simulation studies, and a suicide-rate application are provided. The supplied full text is not that manuscript. It is arXiv:2508.14593, a Journal of Fluid Mechanics submission by Kranz, Morón, and Avila on Bayesian optimisation of energy consumption in turbulent pipe flow. The full text contains pipe-flow equations, DNS results, and optimisation tables; it contains no SCS algorithm, no PSI procedure, no theorem statements, no proofs, no simulation study of screening, and no suicide-rate data analysis. Consequently, the central claims of the paper are not present in the submitted manuscript and cannot be evaluated.","tokens_in":21857,"tokens_out":4050,"duration_ms":49615,"significance":"If the claimed results held, the paper would address a valuable problem: anytime-valid selection of top-m variables and post-selection FCR control at arbitrary stopping times are genuinely underdeveloped in the sequential testing literature. The specific target 'FCR whenever it is conducted' is a meaningful and nontrivial goal. However, because the submitted full text is unrelated to these claims, I cannot assess correctness, novelty, or assumptions. There are no machine-checked proofs, no reproducible code, and no derivations to credit. The manuscript as submitted therefore has no auditable statistical content beyond the abstract, and its significance cannot be established.","major_comments":[{"comment":"The entire submitted body is a different paper. Sections 1-5 and Appendices A-D describe direct numerical simulations of turbulent pipe flow and Bayesian optimisation of pulsatile driving waveforms, with equations (1.1)-(2.9) and Table 1; none of the SCS/PSI content from the abstract appears. There is no definition of top-m, no sequential algorithm, no filtration or stopping time, no theorem statement, and no proof of anytime validity or FCR control. The central claims are therefore entirely unsubstantiated in the supplied manuscript.","section":"Full text, all sections"},{"comment":"The abstract promises theoretical guarantees and a real-data application on suicide rates. The supplied text contains no regularity conditions under which the anytime-valid coverage holds, no derivation of the FCR bound, and no simulation or data section matching that description. I cannot identify any hidden assumption or circular step because the statistical argument itself is absent; this absence is load-bearing and prevents any substantive review of the methodology.","section":"Abstract vs. body"}],"minor_comments":[],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract (arXiv:2508.14596) and the supplied full text (arXiv:2508.14593) is a packaging defect, not a reviewable statistical argument. If the correct full text is resubmitted, it should receive a fresh review; the present submission cannot be revised into a verifiable manuscript within the current scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: the uploaded full text is not this paper. It's Kranz, Morón & Avila's turbulence pipe-flow manuscript (arXiv:2508.14593), so the statistical methodology claimed in the abstract—Sequential Correct Screening and post-screening inference with FCR control—has no accompanying derivations, assumptions, proofs, simulations, or data analysis in the submission. That's a load-bearing mismatch. No content-based verdict is possible.\n\nFrom the abstract alone, the proposal has promise. Top-m selection with anytime validity—where the reported subset contains the true top-m variables with high probability at every stopping time—is a genuinely useful property for sequential screening. The post-screening confidence intervals that control FCR whenever they are constructed also target a real gap; the abstract says FCR control after screening is \"largely overlooked,\" which is plausibly true for online/sequential settings. If the actual manuscript delivers those guarantees with honest regularity conditions, it's a contribution worth referee time.\n\nBut the abstract is all we have. The body text doesn't define SCS, the filtering rule, the estimators, the stopping times, or the probability space. It doesn't state the conditions under which \"with high probability, always\" holds. It doesn't show how the FCR argument avoids the usual traps—e.g., conditioning on selection, or treating screened sets as fixed. It provides no simulation results, and despite the abstract mentioning an application to suicide rates, no such data analysis appears anywhere. The reader's report marks everything UNVERDICTED, and I agree. The stress-test note is correct: the central claims are unverifiable from the supplied document.\n\nOne possibility is an administrative mix-up: the wrong file was uploaded. If so, the fix is simple. But as a submission, the evidence is absent, and we can't send a phantom to referees. The pipe-flow paper itself is unrelated; even its own merits don't bear on this abstract.\n\nFor whoever handles the desk: return the paper without external review, asking the authors to resubmit the actual SCS/PSI manuscript. If and when the correct text arrives, it deserves a serious referee—the abstract claims a real methodological advance in a subfield where anytime-valid FCR control is indeed underexplored. But that review has to wait until there is a there there.","headline":"The full text is a different paper; the abstract promises a useful top-m screening/PSI method, but I can't review what isn't there.","tokens_in":22362,"tokens_out":2181,"would_cite":false,"duration_ms":26710,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L10","62F25","62J15"],"pacs":[],"model":"deepseek-v4-flash","headline":"Sequential Correct Screening guarantees that its reported subsets contain the true top-m variables at every stopping time, and its intervals control the false coverage rate.","keywords":["top-m variable selection","sequential screening","anytime validity","false coverage rate","post-screening inference","confidence intervals","multiple testing","suicide rate application"],"falsifier":"Run repeated simulations of top-m screening in which the candidate estimates are serially dependent or heavy-tailed, stop the procedure adaptively (for example, as soon as the retained set stops shrinking), and measure the empirical fraction of runs in which the final subset misses the true top-m set, and the fraction of reported intervals that miss their targets. If either fraction exceeds the level the method promises — with the same procedure on independent, well-behaved data as a control — the anytime-coverage or false-coverage claim is refuted.","tokens_in":21508,"feed_emoji":"📊","tokens_out":23075,"duration_ms":213052,"temperature":0.7,"pith_summary":"Selecting the m variables with the largest population parameters from a larger candidate pool is a fundamental statistical task, and this paper proposes Sequential Correct Screening (SCS) to do it while data arrive in stages. The paper's defining claim is anytime validity: with high probability, every subset the procedure reports along the way — no matter when the analyst stops — contains the true top-m variables. On top of screening, the paper develops a post-screening inference (PSI) procedure that builds confidence intervals for the selected parameters and is designed to control the false coverage rate (FCR): the expected fraction of reported intervals that miss their target stays at its nominal level whenever inference is actually conducted. If both guarantees hold, an analyst can screen, stop at any convenient time, and still report a valid top-m set together with error-controlled intervals. The paper states theoretical guarantees for both procedures and supports them with simulation studies and an application to a real-world dataset on suicide rates.","feed_headline":"Keeps top-m coverage valid at every stop","feed_subtitle":"You can stop the screening anytime: the top-m set stays covered and intervals control false coverage.","key_machinery":"Two mechanisms carry the proposal. Sequential Correct Screening (SCS) is the screening engine: it examines candidate variables over successive rounds, discards any variable whose evidence rules it out of the top-m, and returns nested survivor subsets; its defining property is anytime validity, so the probability that the true top-m set is covered at every reported time stays at the nominal level regardless of the stopping rule used. The companion device is the post-screening inference (PSI) procedure, which constructs confidence intervals for the retained parameters and is engineered to control the false coverage rate (FCR) — the expected proportion of reported intervals that miss their true","core_discovery":"The paper's claim is that screening for the top-m variables can be made anytime-valid: SCS produces a sequence of variable subsets such that, with probability at least the nominal level, every subset in the sequence contains the true top-m variables simultaneously across all stopping times. This is stronger than a fixed-sample guarantee, because the user may stop after any number of screening rounds and the subset on the table is still covered. The companion PSI procedure constructs confidence intervals for the parameters of the variables that survive screening, and is designed to control the false coverage rate whenever it is conducted, meaning that across repeated selections the expected p","pith_inferences":["If the proof technique generalizes, SCS-style anytime-valid screening could extend to other structured selection tasks — top-k groups, clusters, or discovery rules — and to data streams that continue indefinitely; the paper itself does not construct those extensions.","The full text bundled with this submission is an unrelated manuscript on turbulent pipe flow, so the proofs and the exact regularity conditions of SCS and PSI could not be verified here; a reader should confirm what assumptions the proof places on the candidate estimators — for example independence, moment, or mixing conditions — before applying the method to serially dependent or heavy-tailed dat","A natural end-to-end extension the paper does not build is combining SCS with online multiple-testing corrections so that variables receive error-controlled significance labels as data accumulate, not just a covered top-m set and intervals."],"forward_implications":["An analyst who must stop early — because of budget, time, or a fixed deadline — can report the current subset with the same nominal coverage guarantee instead of needing a pre-specified sample size.","Reported intervals, taken as a package at the moment of stopping, carry a controlled false coverage rate, so the analyst does not need a separate correction for having selected the variables from the data.","The method targets the concrete task of picking the m largest population parameters, so it applies wherever practitioners choose the best features, treatments, or units from a larger pool, as in the paper's suicide-rate application.","The paper's simulation studies and data application indicate that the guarantees are attainable in finite samples, not only in the limit."],"supporting_citations":[],"fun_headline_variants":["Anytime-valid top-m screening with FCR-controlled intervals","Stop anytime, top-m still covered, FCR controlled","Screening sequence with anytime validity and post-selection inference","Top-m sets stay valid at every stop, intervals control FCR"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the statistics used to rank candidate variables behave as the proof requires, enough to make the anytime coverage and false-coverage guarantees hold, and the abstract does not state those conditions nor can they be located here, as the full text bundled with this submission is an unrelated manuscript on turbulent pipe flow.","fun_headline_variants_meta":{"raw":{"variants":["Anytime-valid top-m screening with FCR-controlled intervals","Stop anytime, top-m still covered, FCR controlled","Screening sequence with anytime validity and post-selection inference","Top-m sets stay valid at every stop, intervals control FCR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000167,"raw_usage":{"total_tokens":1050,"prompt_tokens":654,"completion_tokens":396,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":398,"completion_tokens_details":{"reasoning_tokens":325}},"tokens_in":398,"tokens_out":396,"duration_ms":5093,"temperature":1.0,"reasoning_tokens":325,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:25:05.336388+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run repeated simulations of top-m screening in which the candidate estimates are serially dependent or heavy-tailed, stop the procedure adaptively (for example, as soon as the retained set stops shrinking), and measure the empirical fraction of runs in which the final subset misses the true top-m set, and the fraction of reported intervals that miss their targets. If either fraction exceeds the level the method promises — with the same procedure on independent, well-behaved data as a control — the anytime-coverage or false-coverage claim is refuted.","supporting_citations":[],"review_version":1}