{"id":"29704368-cf63-46e5-b4f1-5f83750b07e5","arxiv_id":"2607.22440","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In high-dimensional binary GLMs, a stagewise multiple-testing selector is proved to recover the true sparse support and attain oracle post-selection inference, provided population dominance conditions hold.","lead":"This paper introduces a nonlinear boosting-with-multiple-testing (BMT) method that adds one covariate at a time to a binary-response model, using a family-wise threshold to stop. It claims the method recovers the true sparse set of predictors with probability approaching one and that post-selection estimates behave like an oracle estimator.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Exact-recovery theorem rests on A4, a stagewise population dominance condition that closely mirrors the conclusion; its satisfaction in the simulation and empirical settings is not demonstrated.","rationale":"The reader correctly isolates Assumption A4 as the load-bearing condition. My stress-test concurs: the exact-recovery and oracle results are direct consequences of A4, and the primitive conditions in Section 6 are themselves dominance/gap conditions rather than verifiable properties of the designs or data. This does not make the paper incorrect—the theorems are explicitly conditional—but it does mean the central claim's applicability is only as secure as A4. The paper's careful proof structure and the honest admission that HAC standard errors are not covered by Appendix D support a conditional rather than unconditional verdict. The proposed check is concrete and would settle whether A4 actually holds in the simulation designs where the paper claims success. I do not see an internal inconsistency or a fatal gap; the issue is the gap between high-level assumptions and their verification. Therefore the reader's CONDITIONAL verdict should remain unchanged.","tokens_in":44534,"tokens_out":8357,"duration_ms":105397,"concrete_test":"For a representative Monte Carlo design from Section 7 (e.g., k=4, T=300, p=400, VIF=2 with ω=0.75), compute the population noncentralities NC*_j(S) for every S⊆S0 and every candidate j from the known DGP, using a very long simulation (e.g., T=10^7) or numerical integration. Estimate E_T = max|W_T,j(S)−NC*_j(S)| at the same design. Check whether (13)–(14) hold with d_T at least, say, 4 times the estimated E_T. If A4 holds, the exact-recovery theorem is the right explanation for the reported MCC≈0.99; if A4 fails while BMT still recovers the model, the simulation success is not explained by Theorem 4 and the practical content of A4 is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—BMT selects exactly the true sparse set and yields oracle inference—rests on Assumption A4 (Eqs. 13–14): for every intermediate conditioning set S with |S∩S0|<k, every remaining signal's population Wald noncentrality NC*_j(S) must exceed every non-signal's NC*_m(S) by d_T, and also clear c_T by d_T, with d_T dominating the uniform approximation error E_T. This is essentially a population-level statement of the exact ordering BMT is supposed to discover. If a pseudo-signal ever has conditional strength comparable to a remaining signal, the algorithm can select it before the true signal and the oracle result fails. Section 6's primitive conditions (L1–L3, GM1, G1, G2) are themselves dominance or partial-correlation gap conditions, not facts verified for the data or the DGP. The proofs are careful and assumptions are stated, so the theorem is internally valid; the concern is that the central claim is not robust to plausible violations of A4, and no simulation or empirical check establishes A4 in the very settings where BMT is reported to succeed. Appendix D also verifies the uniform Wald approximation A3.4 only for information/one-period sandwich standard errors, not the HAC standard errors used in the empirical application, so the empirical illustration is not covered by the verified theory.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends the boosting-with-multiple-testing (BMT) framework to high-dimensional binary-response GLMs. Selection is forward-stagewise: at each stage, the BMT procedure adds the covariate with the largest conditional Wald statistic among those passing a family-wise multiple-testing threshold. The main theoretical results state that, under stated assumptions, BMT selects exactly the true sparse set with probability tending to one (Theorem 4), that the post-selection MLE is asymptotically equivalent to the oracle estimator (Theorem 5), and that uniform stochastic control of stagewise statistics holds under high-level approximation conditions (Theorems 1-2). The paper also contains a local- and global-regime analysis of dominance conditions, a comparison with binary LASSO, an extension to exponential-family GLMs, Monte Carlo evidence, and an empirical inflation application using FRED-MD data.","tokens_in":44850,"tokens_out":6615,"duration_ms":85737,"significance":"The theoretical apparatus is substantial: the proofs are detailed and transparently conditional on explicitly stated assumptions, and the paper is honest about the strength of its conditions. The extension of BMT to nonlinear GLMs with oracle post-selection inference is a useful contribution if the assumptions are credible. The comparison with binary LASSO irrepresentability (Appendix C), including a design where BMT fails and LASSO succeeds, is particularly valuable because it delineates the scope of the method. However, the central exact-recovery claim is heavily conditional on a stagewise population dominance condition that is close to the desired selection ordering; the paper does not verify that condition in the simulation or empirical settings, and the empirical application uses HAC standard errors for which the primitive verification of the key approximation assumption is explicitly not provided. These gaps limit the practical force of the theoretical claims.","major_comments":[{"comment":"The exact-recovery theorem is logically valid, but its content is almost entirely delegated to Assumption A4: at every intermediate set, every remaining signal must have population Wald noncentrality NC*_j(S) exceeding every remaining non-signal's NC*_m(S) by d_T that dominates the maximal approximation error E_T. This is essentially a population-level version of the ordering the algorithm is supposed to discover. The primitive conditions in Section 6 do not remove the concern: Assumption L3 (Eq. 38) is a partial-correlation gap of the same shape as A4, and Assumption GM1 (Eq. 40) is a variance-normalized coefficient dominance condition. The paper does not verify A4 in the Monte Carlo DGP or the empirical data and gives no diagnostic check. The severity of A4 is underscored by the paper's own Theorem C.2, which presents a simple Gaussian design with one inactive covariate in which BMT se","section":"§5.2, Assumption A4 (Eqs. 13-14); §6"},{"comment":"Appendix D (Theorem D.1 and the final paragraph) gives primitive sufficient conditions for Assumption A3.4 only for observed-information standard errors under the information equality and for one-period sandwich standard errors with serially uncorrelated scores. It explicitly states that kernel HAC estimators require a separate bandwidth-dependent argument and remain at the high-level A3.4 level. The empirical illustration in Section 8.1 (footnote 6) uses a kernel HAC covariance estimator with the Bartlett kernel and Newey-West bandwidth. Thus the verified theory does not cover the empirical implementation's standard errors. Please either extend the uniform HAC approximation with explicit bandwidth conditions, or change the application to a covered standard-error estimator; at minimum, state this coverage gap in the main text.","section":"Appendix D and §8.1"},{"comment":"The simulation threshold c_T = Phi^{-1}(1-0.05/(2aT)) with a=1,2 is of order sqrt(log T). Under the exponential-tail regime of Theorem D.1, E_T = O_p(sqrt(log(p∨T))), and Assumption A5 requires E_T/c_T -> 0 in probability. In the experiments p <= 400 and T <= 300, so log p and log T are comparable; no verification is provided that A5 holds for the implemented threshold. The favorable simulation outcomes suggest it may hold in these designs, but the theoretical link between the simulations and the exact-recovery theorem is not established. Please add a discussion or check of A5 for the chosen threshold, or use a threshold with provable slack.","section":"§7.1, Assumption A5"}],"minor_comments":[{"comment":"The estimate bse_j(S) is described verbally but not explicitly defined by an equation. A displayed definition of the standard-error estimator would improve clarity, especially because the theory distinguishes between information, sandwich, and HAC versions.","section":"§2.1"},{"comment":"Theorem 1 is essentially a restatement of Assumption A3.4 via the definition of E_T. This is fine, but the main text should make clear that the uniform approximation rate is an assumption except under the primitive conditions verified in Appendix D.","section":"§4, Theorem 1"},{"comment":"The summary statistics are computed over 20 design points but no Monte Carlo standard errors are reported for the medians or RMSE. Given the small numbers of replications for some extreme cases, a few standard errors or confidence intervals would be useful.","section":"§7, Tables 1-5"},{"comment":"The 85/15 sample split is described, but the total number of observations in the evaluation sample is not stated. Reporting T for training and evaluation periods would help interpret the out-of-sample metrics.","section":"§8.2"},{"comment":"Remark 1 correctly notes that Assumption G2 alone does not imply dominance. This is an important caveat and should be echoed in the concluding section, where 'verifiable via multiple primitive routes' could be read too optimistically.","section":"§6.2, Assumption G2"}],"recommendation":"major_revision","confidential_remarks":"The paper is well structured and the proofs are carefully hedged, but the exact-recovery theorem is conditional on a very strong population dominance assumption that the paper does not verify in its own empirical or simulation settings. The HAC standard-error mismatch in the empirical application is a concrete gap that should be fixed before publication. I do not see the issues as unfixable; they require either additional theory (uniform HAC approximation), a restricted empirical implementation, or a clearly stated boundary on when exact recovery is claimed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is exactly what it claims to be — a careful, nontrivial extension of the BMT framework to binary-response GLMs — and the headline is honest in the way that matters: exact recovery and oracle inference hold under the stated conditions, not unconditionally. The load-bearing condition, Assumption A4, is a stagewise population-dominance assumption that closely mirrors the selection ordering the algorithm is supposed to discover. The stress-test note lands on this correctly.\n\nWhat is actually new: the pseudo-true parameter treatment of misspecified stagewise submodels, the uniform control of stagewise Wald statistics along the path, exact-recovery theorems for single and multiple signals, and oracle post-selection inference. The comparison with binary LASSO in Appendix C is a genuinely useful addition — it gives explicit Gaussian designs showing that generalized irrepresentability and BMT stagewise dominance do not nest, and it states the directional stability condition under which a strict irrepresentability margin does imply one-signal-at-a-time BMT dominance. The proofs are detailed; the assumptions are stated openly; Appendix D makes clear what A3.4 needs.\n\nSoft spots, in proportion. First, A4 (and its global-regime variant GM1) requires that at every intermediate stage every remaining signal's population noncentrality exceed every non-signal's by a margin that also dominates the maximal sample error. That is, in population terms, the exact ordering BMT is supposed to learn. The theorem is internally valid, but the result's reach is narrower than the abstract suggests, and the simulations do not verify A4 for their DGP — they show the method works there, not that the condition holds. Moderate. Second, Appendix D verifies the uniform Wald approximation for information and one-period sandwich standard errors, not for the kernel HAC standard errors used in the empirical section. The authors disclose this, but it leaves the empirical illustration outside the verified theory. Moderate, and fixable. Third, no code or data bundle, and the out-of-sample evaluation has no uncertainty quantification. Minor.\n\nWho this is for: econometricians specializing in high-dimensional variable selection in nonlinear models, and applied macro-finance people who want parsimonious selection with explicit multiple-testing control. It earns serious refereeing; referees should press on how primitive A4 can be made and on closing the HAC gap. Send it out.","headline":"Careful, honest extension of BMT to binary-response GLMs — the exact-recovery theorem lives or dies with the population-dominance Assumption A4, but the paper says so and deserves a serious referee.","tokens_in":45320,"tokens_out":3368,"would_cite":true,"duration_ms":41551,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F07","62F12","62J12"],"pacs":[],"model":"deepseek-v4-flash","headline":"A stagewise boosting procedure can recover the exact sparse model in high-dimensional binary-response GLMs, and its post-selection estimator matches an oracle that knows the true signals in advance.","keywords":["variable selection","binary response","generalised linear models","boosting","multiple testing","oracle property","high dimensionality","Wald statistic"],"falsifier":"Construct a design like that in Appendix C.3 where an inactive covariate has a larger first-stage Wald statistic than any true signal (for example, a covariate with correlation 3/5 with two signals and -1/4 with a third). For a small true coefficient, BMT will select this covariate first with probability tending to one, so the selected set cannot be exactly the true support; simulating this design and recording the first selection would refute the exact-recovery claim.","tokens_in":44402,"feed_emoji":"🎯","tokens_out":3099,"duration_ms":37076,"temperature":0.7,"pith_summary":"The paper extends boosting-with-multiple-testing from linear to binary-response generalized linear models. It claims that by adding one covariate at a time based on the largest conditional Wald statistic and filtering through a family-wise threshold, the procedure asymptotically selects all true signals and no others. If this is right, post-selection maximum likelihood inference is as efficient as if the correct sparse model were known. The argument rests on a population measure of conditional strength, the Wald noncentrality, and on a dominance gap requiring true signals to stand out from pseudo-signals at every stage.","feed_headline":"Stagewise boosting proves exact model recovery for high-dim binary data","feed_subtitle":"A conditional Wald filter selects all signals and no noise, matching an oracle estimator for post-selection inference.","key_machinery":"The population Wald noncentrality NC*_j(S) = sqrt(T)|theta*_j(S)| / sqrt(V*_theta,theta,j(S)), the asymptotic mean of the stagewise Wald statistic for adding covariate j to the current model. The procedure's identification and selection are driven by this quantity: a uniform approximation result bounds the maximal deviation of sample Wald statistics from it, and a stagewise dominance condition (Assumption A4) requires every remaining true signal to exceed every non-signal in this measure by a margin that also dominates the approximation error. The multiple-testing threshold c_T, chosen with slack beyond the approximation error, acts as both screening device and stopping rule.","core_discovery":"The central claim is that the BMT algorithm—forward stagewise inclusion of the single most significant covariate, subject to a multiple-testing threshold that also serves as the stopping rule—achieves exact recovery: with probability tending to one, its selected set is exactly the union of the always-in controls and the true signal set. Conditional on that event, the post-BMT maximum likelihood estimator has the same first-order limiting distribution as an oracle estimator that knows the support in advance. The proof works by showing that the stagewise Wald statistics concentrate uniformly around their population expectations, and that the population ordering of these expectations is preserv","pith_inferences":["If the dominance margin fails only mildly, a natural extension is to derive finite-sample or high-probability bounds on the number of false selections under weaker separation, rather than only exact recovery as the paper states.","The one-at-a-time conditional updating resembles orthogonal matching pursuit and forward stepwise regression; the paper's population noncentrality is a nonlinear analogue of partial correlation, suggesting the machinery might transfer to other single-index models.","The theory requires the threshold to have slack beyond the uniform approximation error; a practical, data-driven calibration of this slack would be a testable extension.","In the local small-index regime, dominance reduces to a partial-correlation gap, implying practitioners could pre-screen designs by computing these gaps before deciding whether BMT is appropriate."],"forward_implications":["In sparse binary GLMs with a fixed number of signals, BMT selects all true covariates and no false ones with probability tending to one, for both logit and probit links.","The post-BMT (quasi-)MLE is asymptotically equivalent to the infeasible oracle estimator, so confidence intervals from the selected model are valid to first order.","The procedure does not rely on sparsity-inducing penalties or marginal screening; it is designed to resist pseudo-signals correlated with true signals.","The framework extends to general one-parameter exponential family GLMs, preserving exact recovery and oracle inference under suitable conditions.","Monte Carlo evidence indicates BMT yields smaller models and lower estimation error than OCMT and LASSO in the studied designs, and an inflation-forecasting application uses five predictors with competitive out-of-sample accuracy."],"fun_headline_variants":["Boosting with multiple testing nails exact support recovery","Exact signal selection in high-dim binary models via stagewise boosting","Stagewise boosting achieves oracle inference for high-dimensional GLMs","Boosting with multiple testing: exact recovery and oracle MLE","Nonlinear boosting for binary responses: exact support and oracle estimation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"At every stage before all true signals are selected, every remaining true signal must have a population conditional strength that exceeds every remaining non-signal's strength by a margin large enough to dominate sampling noise—essentially, the population ordering the algorithm is supposed to discover is already guaranteed by the assumptions.","fun_headline_variants_meta":{"raw":{"variants":["Boosting with multiple testing nails exact support recovery","Exact signal selection in high-dim binary models via stagewise boosting","Stagewise boosting achieves oracle inference for high-dimensional GLMs","Boosting with multiple testing: exact recovery and oracle MLE","Nonlinear boosting for binary responses: exact support and oracle estimation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000411,"raw_usage":{"total_tokens":1943,"prompt_tokens":697,"completion_tokens":1246,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":1163}},"tokens_in":441,"tokens_out":1246,"duration_ms":10278,"temperature":1.0,"reasoning_tokens":1163,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:44:15.316441+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a design like that in Appendix C.3 where an inactive covariate has a larger first-stage Wald statistic than any true signal (for example, a covariate with correlation 3/5 with two signals and -1/4 with a third). For a small true coefficient, BMT will select this covariate first with probability tending to one, so the selected set cannot be exactly the true support; simulating this design and recording the first selection would refute the exact-recovery claim.","supporting_citations":[],"review_version":1}