{"id":"e569e460-edb2-4914-b04e-78f7af86ac61","arxiv_id":"1908.06597","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Projection correlation can drive model-free feature screening with sure screening and rank consistency, and a two-step knockoff procedure controls FDR when the target level is at least one over the number of active features.","lead":"A statistics paper proposes two feature screening methods for ultra-high-dimensional data: PC-Screen ranks features by projection correlation without assuming a regression model, and PC-Knockoff adds a two-step knockoff procedure to control false discoveries. The paper provides theoretical guarantees and simulations showing gains over distance-correlation and SIS baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2 is a thresholding result, not a top-d result, so the probability of Algorithm 1's screening event E is not established and Theorem 5's high-probability claim is not supported as stated.","rationale":"The reader's conditional verdict and weakest-assumption analysis already identify the top-d/threshold gap, and I agree that it is real and load-bearing. I concentrate on it rather than on the approximate-knockoff issue because it is an internal inconsistency: Theorem 2, the only cited support for the screening event E, does not imply the event that the top d features contain the active set. The paper's own Section 3.4 statement requires the additional random condition that the d-th largest sample correlation is at most c3 n1^(-kappa), whose probability is not controlled. A concrete correlated-copy design shows the distinction is not merely cosmetic: with many inactive features that are near-deterministic functions of a strong active predictor, a fixed d can exclude a weak active feature even though a threshold rule would include it. The fix is straightforward, for example adding Condition 1(b) together with d at least s and invoking the rank-consistency Theorem 3, or making d adaptive to an estimated threshold, so the paper's central methodological idea remains sound. This does not change the reader's CONDITIONAL verdict; it sharpens the condition under which the paper should be accepted. The approximate-knockoff concern noted by the reader is real but already acknowledged by the authors in Remark 3 and Table 4, so it is a limitation rather than an inconsistency; my agreement is therefore partial rather than full.","tokens_in":25329,"tokens_out":10106,"duration_ms":112694,"concrete_test":"Simulate a design satisfying Condition 1(a) but not Condition 1(b): let n1 = 250, p = 5000, s = 10, with Y = beta1 X1 + ... + beta10 X10 + epsilon, beta1 large and beta10 = c n1^(-kappa) (the minimum active signal), and generate 4990 inactive features as X_k = X1 plus small independent noise, so they are conditionally inactive but have large marginal projection correlation with Y. Use Gaussian X in the second half so second-order knockoffs are exact, and run Algorithm 1 with d = 100. If Pr(all active features are in the top-d set) is substantially below 1 - O(s exp{-c n1^(1-2kappa)}), the top-d justification fails. An analytical companion check: attempt to prove a top-d analogue of Theorem 2; any such proof must add either Condition 1(b) or a bound on the number of features above the threshold, neither of which appears in Theorems 4-5.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing gap is in the screening step of Algorithm 1. The paper's justification for the event E = {all active features are selected in the top-d screening step} is Theorem 2, invoked in Section 3.4 and Remark 4. But Theorem 2 is a thresholding statement: it bounds Pr(A subset of the set of features whose sample projection correlation is at least delta), and it gives no control over the size of that set. Algorithm 1 selects exactly the top d features, not all features above a threshold. Under Condition 1(a) alone, arbitrarily many inactive features can have marginal projection correlations with y as large as or larger than active features (for example, inactive features that are nearly deterministic functions of active predictors), so the top-d set can miss active features even when the thresholded set contains them. Theorem 3 would provide a top-d guarantee, but it requires Condition 1(b), a uniform gap between active and inactive marginal signals, and that condition is not stated for Theorems 4 and 5. The sentence in Section 3.4 requiring the d-th largest sample correlation to be at most c3 n1^(-kappa) simply assumes the d-th order statistic is small; the probability of this event is not bounded under the paper's assumptions. Hence Pr(E) is not shown to tend to 1, and the unconditional form of Theorem 5, namely Pr(A subset of the final selected set) with high probability, is not established for Algorithm 1 as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two procedures for ultra-high-dimensional feature screening. PC-Screen ranks features by the sample projection correlation between each feature (or feature vector) and a possibly multivariate response, and screens either by threshold or by taking the top d features. The paper proves non-asymptotic exponential concentration inequalities for the sample projection correlation (Theorem 1) and uses them to establish sure screening (Theorem 2) under a minimum-signal-strength condition and rank consistency (Theorem 3) under a uniform signal gap. The second procedure, PC-Knockoff, first applies PC-Screen on one subsample to reduce the dimension to a moderate d, then constructs second-order knockoff features on the remaining subsample and selects features by the knockoff+ thresholding rule based on differences of sample projection correlations. Theorems 4 and 5 claim conditional FDR control and joint sure-screening-plus-FDR-control properties when exact knockoffs are available, with a phase transition at alpha = 1/s. The paper also reports extensive simulations and a real supermarket data analysis.","tokens_in":25533,"tokens_out":4928,"duration_ms":52920,"significance":"If the theoretical claims are fully established, the paper would make a useful contribution: PC-Screen is genuinely model-free, robust to heavy-tailed errors and multivariate responses, and the non-asymptotic concentration inequalities for projection correlation are of independent interest. The knockoff-based threshold selection idea for model-free screening is appealing and the conditional-FDR proof in Appendix A.1 is coherent. The numerical comparisons are extensive and show clear advantages for PC-Screen in the settings considered. However, two load-bearing gaps currently prevent the central claims from being accepted as stated: the connection between the thresholding theorem and the top-d screening step is missing, and the implemented second-order knockoffs are not covered by the FDR/sure-screening theorems.","major_comments":[{"comment":"The screening step of Algorithm 1 selects the top d features, but the only screening guarantee invoked, Theorem 2 (Eq. 2.6), is a thresholding result for the set {k: \\hat\\omega_k >= delta} and does not control the cardinality of that set or the event that the d-th largest sample correlation satisfies \\hat\\omega_{(d)} <= c3 n_1^{-\\kappa}. Under Condition 1(a) alone, many inactive features can have population projection correlations as large as or larger than the active ones, so the top-d set may miss active features even when the thresholded set contains them. Consequently Pr(E), the event that all active features are in the top-d set, is not shown to tend to 1 under the stated assumptions. This undermines the unconditional statement in Remark 5 and the claim that Theorems 4 and 5 apply to Algorithm 1 as written; the authors should either impose Condition 1(b) or another explicit gap condition on the order statistics in Theorems 4 and 5, or replace the top-d rule by a threshold rule with a size guarantee, or directly bound Pr(\\hat\\omega_{(d)} <= c3 n_1^{-\\kappa}).","section":"Section 3.4, Algorithm 1, Remark 4"},{"comment":"The FDR and sure-screening guarantees are proved only for exact knockoff features satisfying Condition 2, but Algorithm 1 constructs second-order knockoffs from an estimated covariance matrix via (3.1)-(3.3). The paper acknowledges in Remark 3 and in the discussion of Table 4 (Model 4.c) that these approximate knockoffs may not be close to exact ones, and indeed the reported empirical FDR in Model 4.c at alpha = 0.25 is 0.254, exceeding the nominal level. No theoretical result bounds the FDR inflation or the loss of screening power caused by the second-order approximation. Thus the abstract's blanket statement that the proposed two-step approach controls FDR is not established for the implemented procedure; the authors should state the guarantees only for exact knockoffs, or provide explicit conditions and a bound on the approximation error that yields an FDR correction.","section":"Section 3.2-3.4, Theorems 4 and 5"},{"comment":"The probability bound on the screening step is stated as 1 - O(s exp{c4 n1^{1-2kappa}}) and is described as following from Theorem 2, but the displayed expression has a sign error (the exponent should be -c4 n1^{1-2kappa}) and, more importantly, the bound holds only on the event that \\hat\\omega_{(d)} <= c3 n1^{-\\kappa}. The probability of this order-statistic event is not bounded under the assumptions of Theorems 4 and 5, so the combined probability bound in Remark 5 does not follow as written. This is the same root gap as the previous comment, but it directly affects the stated rate in the main text and should be corrected explicitly.","section":"Section 3.4, Remark 5 and the paragraph before Theorem 4"}],"minor_comments":[{"comment":"The definition 'x = 0.9x1 + 0.1x2' is ambiguous about whether x is a scalar or a vector and what covariance matrix the t2 component has; please clarify the data-generating scheme.","section":"Section 4.2, Model 4.c"},{"comment":"The word 'receptively' should be 'respectively'.","section":"Figure 1 caption"},{"comment":"The sentence 'DC-SIS and bcDC-SIS preform comparably' contains a typo; 'preform' should be 'perform'.","section":"Section 4.1, Example 2"},{"comment":"The expression '1-O(s exp{c4 n1^{1-2kappa}})' is missing a minus sign inside the exponent; it should read '1-O(s exp{-c4 n1^{1-2kappa}})'.","section":"Section 3.4, sentence before Theorem 4"},{"comment":"Algorithm 1 defines \\hat A(T_alpha) as {j: j in \\hat A1, \\hat W_j >= T_alpha}, while Eq. (3.7) defines \\hat A(T_alpha) over all p features; the notation should be reconciled so that the final selected set is unambiguously the intersection with \\hat A1.","section":"Algorithm 1 and Eq. (3.7)"}],"recommendation":"major_revision","confidential_remarks":"The top-d screening gap is the key obstacle: Theorem 2 cannot justify the probability of the screening event E for Algorithm 1, and the issue is fixable by adding a rank-consistency-type condition or by modifying the screening step. The second-order knockoff issue is acknowledged by the authors themselves, but the theoretical claims should be confined to exact knockoffs unless a formal approximation bound is supplied. I believe the paper is worth a major revision rather than rejection, provided the authors address these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The screening part is genuinely useful: the non-asymptotic concentration inequalities for empirical projection correlation are new, and PC-Screen delivers model-free, heavy-tail-robust screening with sure screening and rank consistency under weak conditions. The knockoff part has a real theory-implementation gap: the sure screening event for the top-d step is not actually proved, and the FDR guarantee is for exact knockoffs while the implementation uses second-order ones.\n\nWhat's new and good: the paper proves exponential-type concentration for sample projection correlation without moment conditions and uses it to derive screening and rank consistency results. The simulations are extensive, and PC-Screen clearly outperforms DC-SIS and competitors in heavy-tailed and multivariate-response settings. The two-step sample-splitting knockoff framework is a reasonable path to FDR control after screening, and the conditional FDR proof in Theorem 4 looks coherent. The self-citation to Zhu et al. (2017) is appropriate since projection correlation comes from there.\n\nSoft spots. The bigger one is the top-d issue. Theorem 2 is a thresholding result: it bounds the probability that all active features have sample projection correlation above delta. Algorithm 1 selects the top d features. Under Condition 1(a) alone, there is no control on how many inactive features have correlations comparable to the active ones, so the top-d set can miss active features even when the thresholded set contains them. Remark 4 essentially assumes the d-th largest sample correlation is small; the probability of that event is not bounded. Theorem 3 would give a top-d guarantee, but it requires Condition 1(b), which is not assumed for Theorems 4 and 5. This is a load-bearing gap in the advertised theory for Algorithm 1, though it is fixable: prove a top-d result under a gap condition, or switch Algorithm 1 to thresholding with a data-dependent cutoff. The second soft spot is approximate knockoffs. The FDR theorem requires exact knockoffs satisfying Condition 2; the implementation uses second-order knockoffs from an estimated covariance. The authors are honest about this, reporting FDR 0.254 at alpha 0.25 in Model 4.c, but it means the empirical procedure does not inherit the theorem's guarantee without additional conditions. That is a caveat, not a fatal flaw.\n\nWho this is for: researchers working on feature screening in ultra-high-dimensional, model-free, or heavy-tailed settings. The screening contribution alone is worth a serious referee; the knockoff part needs revision to close the top-d gap and to be precise about what approximate knockoffs imply.","headline":"PC-Screen is a solid model-free screening method with new concentration results, but the knockoff-based FDR theory doesn't actually cover the implemented top-d algorithm.","tokens_in":26140,"tokens_out":4097,"would_cite":true,"duration_ms":39558,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ranking features by projection correlation achieves sure screening with no model, and adding a knockoff threshold yields simultaneous FDR control and sure screening whenever the target FDR level is at least 1/s.","keywords":["feature screening","projection correlation","model-free screening","knockoff features","false discovery rate","sure screening","rank consistency","ultra-high dimensional data"],"falsifier":"Run Algorithm 1 at $\\alpha = 0.25$ with $s=10$ on the heavy-tailed mixture data of Model 4.c and check whether the empirical FDR exceeds $0.25$ as reported (0.254 in Table 4); or, separately, simulate Condition 1(a) with a known active set and check whether a top-$d$ screening step retains all active features with high probability—if it does not, the event $\\mathcal{E}$ on which Theorems 4 and 5 condition is not established for Algorithm 1.","tokens_in":25041,"feed_emoji":"🎯","tokens_out":16412,"duration_ms":140605,"temperature":0.7,"pith_summary":"PC-Screen ranks features by their projection correlation with the response and is claimed to achieve sure screening—every active feature is retained with probability tending to 1—and rank consistency under only a minimum-signal-strength condition, with no regression model, no sub-Gaussian assumption, and no limit on the response dimension. The paper also proposes PC-Knockoff, a two-step procedure that first screens to a moderate set and then uses knockoff features to choose a threshold with false discovery rate control. The main theoretical result is that if the nominal FDR level $\\alpha$ is at least $1/s$, where $s$ is the number of active features, PC-Knockoff controls FDR and keeps all active features simultaneously with high probability; below $1/s$ there is a phase transition and no such guarantee. These results give practitioners a model-free screening tool that works for heavy-tailed, nonlinear, and multivariate response data, and a principled data-adaptive way to set the screening threshold instead of picking a conservative cutoff by hand.","feed_headline":"Model-free screening with a guaranteed false-discovery cap","feed_subtitle":"Projection-correlation ranking keeps all active features while a knockoff threshold selects the cutoff adaptively.","key_machinery":"The argument is carried by two paired objects. The first is the squared projection correlation $\\omega_k = \\mathrm{PC}(X_k,\\mathbf y)^2$, whose sample version (2.4)–(2.5) is a triple-sum average of arccosine angles, and for which Theorem 1 provides a moment-free, dimension-free exponential deviation bound; this is what makes sure screening and rank consistency valid without model or tail assumptions. The second is the knockoff contrast $\\hat W_j = \\widehat{PC}(X_j,Y)^2-\\widehat{PC}(\\tilde X_j,Y)^2$, whose sign for inactive features is symmetric—exactly fair coin flips conditioned on the absolute values (Lemma 1)—and the knockoff+ threshold $T_\\alpha$ from (3.6). Lemma 1 turns the threshold-selection problem into a backward super-martingale, so the optional stopping theorem yields the FDR bound in Theorem 4, and the signal-separation argument in Theorem 5 gives the simultaneous sure-screening guarantee at $\\alpha\\ge 1/s$.","core_discovery":"The paper's central claim is that projection correlation—the dependence measure defined in (2.1)–(2.2) as an average over all unit projections of the squared covariance of indicator transforms—is the right engine for model-free screening in ultra-high dimensions. Theorem 1 establishes a non-asymptotic exponential concentration inequality for the empirical squared projection correlation with constants that do not depend on dimension or on any moment conditions. From it, Theorem 2 gives sure screening, $\\Pr(\\mathcal{A}\\subseteq\\hat{\\mathcal{A}}(\\delta))\\geq 1-O(s\\exp\\{-c_4 n^{1-2\\kappa}\\})$, when the smallest active signal exceeds $2c_3 n^{-\\kappa}$, and Theorem 3 gives rank consistency under a signal-gap condition. On the FDR side, the paper proves that the knockoff+ threshold (3.6) applied to $\\hat W_j=\\widehat{PC}(X_j,Y)^2-\\widehat{PC}(\\tilde X_j,Y)^2$ controls the false discovery rate conditionally on the screening event $\\mathcal{E}$ (Theorem 4), and that for $\\alpha\\geq 1/s$ it still retains every active feature with probability $1-O(n_2\\exp\\{-c_4 n_2^{1-2\\kappa}\\})$ (Theorem 5(i)); for $\\alpha<1/s$, the procedure either recovers the whole active set or returns an empty set, with no sure-screening guarantee.","pith_inferences":["A natural testable extension is to build a diagnostic that checks swap exchangeability of the constructed second-order knockoffs on the second subsample; the paper's Model 4.c result (empirical FDR 0.254 at $\\alpha = 0.25$) suggests the guarantee can degrade badly when that diagnostic fails.","The phase transition at $1/s$ can be inverted into a formal estimator of the active-set size $s$ by scanning $\\alpha$ and finding the largest level that yields an empty selection, as the paper sketches informally; a future analysis could attach confidence intervals to that estimate.","Because the $W$-statistic cancels spurious marginal signals of inactive features, screening on $\\hat W$ may tolerate strong marginal correlations between inactive and active features better than PC-Screen itself, potentially opening a path to factor-model or confounded settings—though the paper does not analyze that regime.","The sample-splitting scheme in Algorithm 1 suggests a general template: any marginal dependence measure with a dimension-free concentration inequality could replace projection correlation, and the FDR step would remain valid as long as the first-stage event holds, a direction the paper notes but does not develop."],"forward_implications":["Screening can be applied before any model is chosen: the same guarantees cover linear, nonlinear, additive, quantile, Poisson, and multivariate-response data, so model specification is no longer a prerequisite for dimension reduction.","With exact knockoffs, FDR control and sure screening are compatible exactly when the target level is not below $1/s$; below that, the procedure exhibits a hard phase transition and cannot promise both.","The screening threshold no longer needs to be fixed conservatively: Algorithm 1's knockoff step sets the cutoff data-adaptively while bounding false discoveries.","Because the concentration inequality is dimension-free and moment-free, the theoretical error rates do not degrade as $p$ grows or as tails become heavier, unlike distance-correlation screening whose rate carries an extra $\\eta$ term."],"supporting_citations":[{"why":"Supplies the definition of projection correlation, its zero-iff-independence property, and the sample estimator in (2.4) that PC-Screen ranks.","marker":"Zhu et al. (2017)"},{"why":"Supplies the knockoff constructions, the knockoff+ threshold (3.6), and the super-martingale proof strategy used for Theorem 4.","marker":"Barber and Candès (2015)"},{"why":"Provides the Model-X knockoff framework and the exact knockoff conditions (Condition 2) required by Theorems 4 and 5.","marker":"Candès et al. (2018)"},{"why":"Defines the sure independence screening property that PC-Screen extends and serves as the primary SIS baseline.","marker":"Fan and Lv (2008)"},{"why":"Provides the distance-correlation screening (DC-SIS) baseline and the theoretical rate comparison showing PC-Screen is faster.","marker":"Li et al. (2012)"},{"why":"Supplies the two-step sample-splitting framework and the data-recycling idea on which PC-Knockoff's structure and Remark 5 rely.","marker":"Barber and Candès (2019)"},{"why":"Supports the two-step FDR-control argument for knockoff-based large-scale inference, cited alongside Barber and Candès (2019).","marker":"Fan, Demirkaya, Li and Lv (2020)"}],"fun_headline_variants":["Projection-correlation screening with built-in FDR cap","Knockoff-powered FDR control without model assumptions","Model-free screening: sure selection and FDR cap","Heavy-tail-robust screening via projection correlation","Ultra-high-dim screening with guaranteed FDR control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dual premise that the knockoff features are exact (swap-exchangeable and conditionally independent of the response) and that the first-stage top-$d$ screen already contains every active feature is what makes both FDR control and sure screening true, and the implemented algorithm guarantees neither—the paper's own Model 4.c shows FDR inflation when second-order knockoffs fail to be exact.","fun_headline_variants_meta":{"raw":{"variants":["Projection-correlation screening with built-in FDR cap","Knockoff-powered FDR control without model assumptions","Model-free screening: sure selection and FDR cap","Heavy-tail-robust screening via projection correlation","Ultra-high-dim screening with guaranteed FDR control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000713,"raw_usage":{"total_tokens":3228,"prompt_tokens":985,"completion_tokens":2243,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":2166}},"tokens_in":601,"tokens_out":2243,"duration_ms":15621,"temperature":1.0,"reasoning_tokens":2166,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:41:20.977236+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Algorithm 1 at $\\alpha = 0.25$ with $s=10$ on the heavy-tailed mixture data of Model 4.c and check whether the empirical FDR exceeds $0.25$ as reported (0.254 in Table 4); or, separately, simulate Condition 1(a) with a known active set and check whether a top-$d$ screening step retains all active features with high probability—if it does not, the event $\\mathcal{E}$ on which Theorems 4 and 5 condition is not established for Algorithm 1.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the knockoff constructions, the knockoff+ threshold (3.6), and the super-martingale proof strategy used for Theorem 4."}],"review_version":1}