{"id":"177e7100-5e3c-48f9-93f0-958de4cce169","arxiv_id":"2411.17618","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A conditional posterior based on a variance-weighted projection yields asymptotically valid frequentist credible intervals for a binary treatment effect in high-dimensional logistic regression.","lead":"This paper develops a Bayesian method for estimating the effect of a binary treatment on a binary outcome while adjusting for many possible confounding variables. It introduces a variance-weighted projection that restores the orthogonality needed for valid credible intervals, with asymptotic frequentist guarantees.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption A5 assumes a parametric-rate uniform logistic approximation to the true propensity despite the paper's no-correct-model claim; misspecification biases the variance-weighted projection and can break Theorem 1's coverage.","rationale":"The reader's weakest-assumption analysis correctly identifies Assumption A5 as the most fragile load-bearing element. The theorem's proof explicitly routes the propensity-estimation error through Lemma 3 and the second-order term of Lemma 7, and both need the strict parametric uniform rate from A5. The paper tries to avoid a correct-model assumption by asserting a concentration condition, but that condition is not primitive and, for a misspecified logistic model, is generally false. Section 5.3's admission that the working model might not strictly apply, and Section 6.2's use of a LASSO shortcut outside the stated assumptions, make the abstract's generalization especially vulnerable. The reader's CONDITIONAL verdict is appropriate: the authors should either prove A5 under explicit conditions on the misspecification (e.g., a converging sieve approximation) or substantially weaken the claim that valid inference is obtained without a correct propensity model. A separate internal inconsistency also exists in the definition of sigma_n (Eq. 3.9 gives sqrt(-L'') while the proof in Appendix B treats sigma_n as the standard deviation, the inverse square root of the observed information). That inconsistency needs correction, but even with sigma_n fixed, A5 remains the principal obstacle to the claim of valid inference under general binary-covariate settings. Thus the stress-test does not change the reader's conditional verdict.","tokens_in":69,"tokens_out":10855,"duration_ms":138425,"concrete_test":"Simulate the Section 5.1 setting (n=400, d=500 or n=500, d=600) with a deliberately misspecified propensity model: generate X_i from P(X_i=1|Z_i)=Phi(Z_i^T gamma0) (probit) or from a logistic model with an additional quadratic term in Z_i that is not included in the working model, while keeping the outcome model logistic. For a fixed true theta0 (e.g., 0.5), run 1000 Monte Carlo replicates of the CB procedure and compute (i) the maximum over i of |logit(P(X_i=1|Z_i,gamma)) - logit(P(X_i=1|Z_i))| and (ii) the empirical 95% credible-interval coverage. If A5 holds, the max log-ratio should shrink at approximately sqrt(log(d∨n)/n) and CB coverage should stay near 0.95. If the log-ratio is bounded away from zero and coverage drops materially (e.g., below 0.90), A5 fails under misspecification and the central coverage claim is conditional on an unverified correct-specification assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central coverage claim in Theorem 1 and Corollary 1.1 rests on Assumption A5 (Eq. 3.5), which requires the logistic working model (γ-WM) to estimate the true propensity P(X_i=1|Z_i) with log-odds error uniformly bounded by M sqrt(log(d∨n)/n). The paper explicitly says it does not assume a correct model for P(X_i=1|Z_i), but A5 is effectively a correct-specification assumption at a parametric rate. Under misspecification, the posterior for γ concentrates at the pseudo-true value minimizing Kullback-Leibler divergence to the true propensity; the log-odds bias Z_i^T(γ*−γ_true_logit) is generically of constant order for at least some individuals, not o_p(1), so A5 fails. Lemma 3 uses A5 to control the T3 term in |h(Z_i)−h0(Z_i)|, and Lemma 7 uses the same rate to make the second-order remainder in the Taylor expansion negligible. If A5 fails, the estimated variance-weighted projection h(Z_i) is asymptotically biased, the orthogonality correction does not drive E[B_n] to zero after sqrt(n) scaling, and the posterior centers away from the oracle estimator, so the credible interval coverage in Corollary 1.1 is not attained. Section 5.3 concedes that the working model '(γ-WM) might not strictly apply,' and Section 6.2 uses a LASSO nuisance shortcut not covered by Assumption A4. The abstract's unqualified 'valid inference' claim is therefore not supported in settings where the propensity model is misspecified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new Bayesian method for inference on a low-dimensional parameter θ in a high-dimensional logistic regression model where the covariate of interest X is binary. The central construction is an orthogonal score based on a variance-weighted projection h0(Z) = E[X Var(Y|X,Z) | Z] / E[Var(Y|X,Z) | Z], which the authors argue is needed because the usual weighted-regression orthogonality condition fails for binary X. The method draws nuisance parameters (θ̃, β) and a propensity parameter γ from conditional posteriors, then forms a working model for θ using estimated h(Z) and φ. The main theoretical result (Theorem 1, Section 3.3) states that the marginal posterior of θ is asymptotically normal around the oracle MLE θ̂0 at the n^{-1/2} rate, and Corollary 1.1 concludes that Bayesian credible intervals have correct frequentist coverage. The paper includes simulations and a real-data analysis of the CKD dataset, plus an extension to categorical X and a general form of the variance-weighted projection.","tokens_in":40307,"tokens_out":6433,"duration_ms":59567,"significance":"If the theorem and its coverage claim were fully established, the paper would provide a valuable Bayesian alternative to existing frequentist de-biasing methods for logistic regression with a binary treatment in high dimensions. The variance-weighted projection is a natural and well-motivated construction, and the use of multiple nuisance samples from posterior distributions is an interesting twist on integrated-likelihood ideas. The simulation study is thorough, covers a range of signal strengths and two sample sizes, and compares well against several existing methods. The paper also identifies a genuine technical obstacle (the failure of continuous-X orthogonality conditions when X is binary) and offers a concrete score that addresses it. The categorical extension and the general expression in Section 8 broaden the potential applicability. The central claim, however, rests on a set of assumptions whose realism and internal consistency need careful scrutiny, and the proof as written leaves several gaps that a revision would need to close.","major_comments":[{"comment":"Assumption A5 is effectively a correct-specification assumption at a parametric rate, despite the text's claim that no correct model for P(X_i=1|Z_i) is assumed. The assumption requires the logistic working model (γ-WM) to estimate the true propensity with uniform log-odds error M√(log(d∨n)/n). If the true propensity is not exactly logistic, the posterior for γ concentrates at a pseudo-true value, and the log-odds bias Z_i^T(γ* − γ_true_logit) is generically of constant order for at least some i, violating A5. This is not a minor technicality: Lemma 3 uses A5 to control the T3 term in |h(Z_i)−h0(Z_i)|, and Lemma 7 uses the same rate to make the second-order remainder in the score expansion negligible. If A5 fails, the estimated h(Z) is asymptotically biased, the orthogonality correction does not drive E[B_n] to zero after √n scaling, and the posterior no longer centers at θ0. Consequently, the coverage guarantee in Corollary 1.1 is not supported in settings where the propensity model is misspecified. The paper's own Section 5.3 concedes that the working model '(γ-WM) might not strictly apply' in the real-data experiment, so the discrepancy between the theory and the claimed empirical validity is concrete. A revision should either (i) state explicitly that the theory assumes a correctly specified logistic propensity model and soften the 'valid inference' claim accordingly, or (ii) develop a doubly robust version that permits misspecification of one of the nuisance models.","section":"Section 3.2, Assumption A5 (Eq. 3.5)"},{"comment":"The definition of σ_n is inconsistent with its use. Eq. (3.9) defines σ_n = √( −∂² L_n/∂θ² |_{θ=θ̂0} ), which is O(√n) and is the square root of the observed Fisher information, not the standard deviation of θ̂0. The correct standard deviation is 1/√(−∂² L_n/∂θ²). The proof then uses (σ_n²)^{-1} = L''_n(θ̂0|h0,φ0), which has the wrong sign and is also inconsistent with (3.9). This matters because Theorem 1 and Corollary 1.1 rely on σ_n as the scaling in the normal approximation; if σ_n is mis-defined, the limiting distribution statement is not well posed. The text in Section 3.3 even says 'σn = O(√n)' immediately after defining σ_n as the standard deviation, which confirms the error. A revision must correct the definition of σ_n and ensure that all subsequent expressions (B.16), (B.20), and (B.25) are consistent with that corrected definition.","section":"Section 3.3, Eq. (3.9) and proof in Appendix B (around B.16)"},{"comment":"The proof of Lemma 7 invokes Lemma 2 from Belloni et al. (2013) to control the empirical process deviations in (B.10), but the observations (Y_i, X_i, Z_i) are not independent conditional on the estimated nuisance quantities h(Z) and φ, because h(Z_i) and φ_i are constructed from the same posterior sample (θ̃,β). Lemma 2 as stated applies to independent observations; the manuscript does not verify the required conditions (e.g., uniform entropy or bracketing numbers) for the dependent case. Moreover, the statement 'By Neyman orthogonality, E[∂_t b_n(η0)[η−η0]] = 0' is only true when h0 is the true variance-weighted projection; if A5 fails (see major comment above), this expectation is not zero and the claimed √n|E[b_n − b_n0]| → 0 bound in (B.9) does not follow. The proof of Lemma 7 is therefore incomplete as written.","section":"Lemma 7 (Appendix B, Eq. B.9–B.10)"},{"comment":"The proof establishes the conclusion only on a high-probability set E and then conditions on a further high-probability set B. The passage from (B.14) shows that the marginal posterior is close to the conditional posterior on B, but the stochastic convergence results used to prove (B.25) and (B.27) are unconditional statements (over the data), not statements conditional on Y, X, Z. To prove that the posterior probability P(θ ∈ (θ̂0 + aσ_n, θ̂0 + bσ_n) | Y,X,Z) converges in probability to the normal CDF, one needs to show that the ratio of integrals converges uniformly over the data with high probability. The manuscript does not provide a rigorous justification connecting the unconditional convergence of the score terms to the conditional posterior probability. This gap is load-bearing because the theorem's conclusion is exactly a convergence in probability of a function of the data.","section":"Theorem 1 proof (Appendix B, first paragraph and B.40)"}],"minor_comments":[{"comment":"The sentence 'We use the notations h0(Z) and h(Z) to represent the vectors (h0Z1),...,h0(Zn)) and (h(Z1),...,h(Zn)), respectively' contains missing parentheses and is hard to read; it should be edited for clarity.","section":"Section 2.1, paragraph after Eq. (2.9)"},{"comment":"The phrase 'Connections to Causal Inference' contains a typo in the body text: 'casual inference' appears in the first sentence; it should be 'causal inference'.","section":"Section 1.3, title"},{"comment":"The sentence 'The theorem states that the posterior concentrates around θ̂0 which is the oracle estimator (3.8) at n−1/2 rate because σn = O(√n)' is misleading: if σ_n were the standard deviation it would be O(n^{-1/2}); the sentence becomes correct only if σ_n is re-defined as the inverse standard deviation. This is a direct consequence of the Eq. (3.9) error and should be fixed.","section":"Section 3.3, text before Theorem 1"},{"comment":"The paper states that in the real-data analysis 'we obtain the nuisance estimates via LASSO' without noting that this violates Assumption A4, which is about posterior concentration of the spike-and-slab method. A sentence explaining that this is a heuristic used for computational convenience, not covered by the theory, would be appropriate.","section":"Section 6.2, paragraph 1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a real gap in the Bayesian high-dimensional inference literature and the proposal is attractive, but the theoretical guarantee as stated is not yet solid: the propensity-model assumption A5 is effectively a correct-specification assumption, the definition of σ_n is inconsistent, and the proof of the main theorem relies on unverified empirical-process bounds for dependent data. These issues are fixable in a revision, but they are substantial. The simulations are extensive and generally support the method's practical promise, though the real-data analyses are not backed by the theory. I would recommend major revision rather than rejection, because the core idea and the simulation evidence are valuable and the gaps appear bridgeable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the variance-weighted projection in (2.9) is a real idea, and the paper is right that the continuous-X orthogonality trick fails when X is binary. The trouble is that Theorem 1 and Corollary 1.1 lean on Assumption A5, which is closer to assuming a correct propensity model than the paper admits. That's the main soft spot, and it's load-bearing.\n\nWhat's new: h0(Z_i) = E[X_i Var(Y_i|X_i,Z_i)|Z_i] / E[Var(Y_i|X_i,Z_i)|Z_i] is a novel construction, and the derivation from the conditional moment equation (8.2) is clean. Combining this with the conditional-posterior framework is a reasonable step forward. The simulations are reasonably thorough; CB's coverage is close to nominal across the grid and competitive with the oracle and LSW. That's real evidence the method works when the propensity model is correctly specified.\n\nSoft spots, in rough order of severity. First, A5 requires the estimated log odds of the propensity to be uniformly within M sqrt(log(d∨n)/n) of the truth. If the logistic working model for X|Z is misspecified, the posterior for γ concentrates at the pseudo-true value, and the log-odds bias is generically of constant order for at least some individuals, so A5 fails. The paper says it doesn't assume a correct model, but then imposes exactly the kind of parametric-rate closeness that a correct model would deliver. Section 5.3 concedes the working model 'might not strictly apply,' and Section 6.2 uses a LASSO shortcut not covered by Assumption A4, so the empirical sections don't rescue the abstract's unqualified 'valid inference' claim. Second, the definition of sigma_n in (3.9) has the opposite sign from the proof in (B.16). That's fixable, but as written the proof is internally inconsistent. Third, Lemma 7's proof applies a weak law to terms that are dependent through the estimated nuisance functions; a rigorous treatment needs to condition on a high-probability set more explicitly. The overall strategy of proving the theorem on a high-probability set is standard and fine.\n\nFor whom: this is for a reader working on Bayesian inference for low-dimensional parameters in high-dimensional GLMs, especially with binary/categorical treatments. The projection idea deserves attention; the theorem as stated doesn't fully support the coverage claim. I'd send this to peer review, with a strong request to fix A5 (either prove it under explicit conditions or weaken the claims), correct the sigma_n and dependence issues, and provide code and data. With those fixes it could be a solid contribution.","headline":"A genuinely new variance-weighted projection for binary treatments, but the coverage theorem rests on an assumption that quietly requires a correct propensity model.","tokens_in":40940,"tokens_out":4520,"would_cite":false,"duration_ms":39705,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J12","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that a variance-weighted projection score yields Bayesian credible intervals with correct frequentist coverage for a binary treatment effect in high-dimensional logistic regression.","keywords":["high-dimensional logistic regression","Bayesian inference","Neyman orthogonality","variance-weighted projection","credible intervals","binary treatment","propensity score","conditional posterior"],"falsifier":"Run a simulation in which Y follows the assumed logistic model but the true probability that X equals 1 given Z is a nonlinear function, such as a probit or a logistic with interaction terms, that the linear logistic working model cannot represent; then compute the empirical coverage of the nominal 95% credible intervals over many replications. If coverage drops below nominal as n grows, the uniform propensity-concentration assumption fails and the theorem's conclusion is false.","tokens_in":39724,"feed_emoji":"🎯","tokens_out":8007,"duration_ms":67226,"temperature":0.7,"pith_summary":"High-dimensional logistic regression with a binary treatment and many confounders makes inference on the log-odds ratio hard because the usual orthogonality trick of regressing the treatment on the confounders breaks down. The paper constructs an orthogonal score that replaces the binary treatment with its variance-weighted projection onto the confounders, then builds a Bayesian conditional posterior for the treatment effect around that score. The main theorem states that this posterior is asymptotically normal at the oracle $n^{-1}$/2 rate, so the resulting credible intervals have correct frequentist coverage even when the number of confounders grows faster than the sample size. If true, this gives practitioners a Bayesian interval for treatment effects that avoids the bias of selection-based methods and does not require nonzero nuisance signals to be strong.","feed_headline":"Bayesian intervals for binary treatment effects hit nominal coverage","feed_subtitle":"Restores orthogonality that fails for binary treatments, enabling valid high-dimensional Bayesian intervals.","key_machinery":"The machinery is the variance-weighted projection h0(Z_i) = E[X_i Var(Y_i | X_i, Z_i) | Z_i] / E[Var(Y_i | X_i, Z_i) | Z_i], estimated from posterior samples of the nuisance parameters and a logistic working model for the propensity score. Subtracting this projection in the score (X - h0(Z))^T (Y - μ) forces the expected derivative with respect to the nuisance parameter β to vanish, giving Neyman orthogonality. The conditional posterior for θ then uses the working model with linear predictor (X_i - h(Z_i))θ + φ_i, which coincides exactly with the original logistic model at the oracle values; this is what lets the slower nuisance concentration rate be converted into the optimal $n^{-1}$/2 posterior rate for θ.","core_discovery":"The paper's central claim is that valid frequentist inference for the log-odds ratio of a binary covariate in a high-dimensional logistic model can be obtained from a Bayesian posterior built on an orthogonal score. Because the standard weighted regression of the binary treatment X on the nuisance covariates Z fails—the required moment condition would force a linear function to equal a bounded nonlinear ratio—the paper instead defines h0(Z_i) as the conditional expectation of X_i weighted by outcome variances, and uses the score (X - h0(Z))^T (Y - μ). Plugging posterior samples of the nuisance parameters into h and a reparametrized φ yields a conditional posterior for the treatment effect θ. Theorem 1 shows this marginal posterior converges in probability to a normal distribution centered at the oracle MLE θ̂0 with scale σ_n, and Corollary 1.1 concludes that the credible intervals have correct frequentist coverage.","pith_inferences":["The same conditional-posterior construction should transfer to other high-dimensional generalized linear models with discrete exposures, with h0 defined by the same variance-weighted conditional expectation.","Because h0 is defined through conditional expectations, a nonparametric or doubly robust estimator of h0 could replace the logistic propensity working model and relax the paper's concentration assumption on the propensity score.","A practical user should treat the logistic propensity working model as a genuine assumption: in real datasets where the treatment depends nonlinearly on covariates, coverage should be checked empirically rather than taken for granted."],"forward_implications":["Credible intervals from the conditional posterior achieve nominal frequentist coverage for θ at the oracle n^-1/2 rate.","The method remains valid when the number of covariates grows sub-exponentially in n and does not require nonzero nuisance coefficients to be bounded away from zero.","The construction extends to a K-category treatment by applying the variance-weighted projection separately to each dummy variable.","Averaging over posterior samples of the nuisance parameters captures multiple orthogonality instances, in contrast to methods that rely on a single debiased estimate."],"supporting_citations":[{"why":"Supplies the Neyman-orthogonality and double-selection framework, plus the empirical-process Lemma 2 used to control score differences.","marker":"Belloni et al. (2013)"},{"why":"Introduces the conditional Bayesian posterior approach for high-dimensional linear models that this paper extends to logistic regression.","marker":"Wu et al. (2023)"},{"why":"Previous conditional-posterior method for high-dimensional logistic regression with continuous covariate of interest, whose moment conditions fail for binary X.","marker":"Ojha and Narisetty (2023)"},{"why":"Provides the orthogonality principle on which the proposed orthogonal score is based.","marker":"Neyman (1959, 1979)"},{"why":"Supplies the spike-and-slab Gibbs sampler whose posterior concentration rate underlies Assumption A4 and the implemented priors.","marker":"Narisetty et al. (2019)"},{"why":"Posterior concentration theorem used to show that the oracle-likelihood posterior concentrates around the oracle MLE at n^-1/2 rate.","marker":"Walker (1969)"},{"why":"Provides the asymptotic posterior normality results used in the proof of Theorem 1.","marker":"Schervish (2012)"},{"why":"Establishes consistency and asymptotic normality of the one-dimensional oracle MLE θ̂0 used as the center of the posterior.","marker":"Fahrmeir and Kaufmann (1985)"}],"fun_headline_variants":["Variance-weighted projection enables valid Bayesian intervals for binary treatment","New Bayesian method for valid inference on binary treatment in high dimensions","Bayesian inference for binary treatments in high-dimensional logistic regression","Projection-based Bayesian inference gives valid intervals for binary treatments","Binary treatment gets valid Bayesian intervals via variance-weighted projection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the logistic working model for the probability that the binary treatment equals 1 given the covariates estimates the true probability well enough that its error vanishes uniformly as the sample size grows, even though the paper does not assume that working model is correct; if this premise fails, the projection that removes nuisance bias is itself biased and the credible intervals can lose coverage.","fun_headline_variants_meta":{"raw":{"variants":["Variance-weighted projection enables valid Bayesian intervals for binary treatment","New Bayesian method for valid inference on binary treatment in high dimensions","Bayesian inference for binary treatments in high-dimensional logistic regression","Projection-based Bayesian inference gives valid intervals for binary treatments","Binary treatment gets valid Bayesian intervals via variance-weighted projection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001451,"raw_usage":{"total_tokens":5791,"prompt_tokens":843,"completion_tokens":4948,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":4867}},"tokens_in":459,"tokens_out":4948,"duration_ms":39160,"temperature":1.0,"reasoning_tokens":4867,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:55:23.409629+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a simulation in which Y follows the assumed logistic model but the true probability that X equals 1 given Z is a nonlinear function, such as a probit or a logistic with interaction terms, that the linear logistic working model cannot represent; then compute the empirical coverage of the nominal 95% credible intervals over many replications. If coverage drops below nominal as n grows, the uniform propensity-concentration assumption fails and the theorem's conclusion is false.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Neyman-orthogonality and double-selection framework, plus the empirical-process Lemma 2 used to control score differences."},{"cited_title":"Skinny Gibbs: A Cons istent and Scalable Gibbs Sampler for Model Selection,","cited_arxiv_id":null,"evidence_quote":"Supplies the spike-and-slab Gibbs sampler whose posterior concentration rate underlies Assumption A4 and the implemented priors."}],"review_version":1}