{"id":"dcaf2302-ba3a-4a50-9c98-182e6af8026c","arxiv_id":"2501.11181","paper_version":5,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper provides closed-form sample size and power formulas for observational causal inference via IPW variance decomposition, requiring only two additional user-specified parameters for confounder-treatment overlap (via Bhattacharyya coefficient) and confounder-outcome association (via R-squared–","lead":"This paper derives analytical formulas for sample size and power in observational causal studies by decomposing the variance of an inverse probability weighting estimator into propensity score, outcome, and correlation parts. Researchers planning such studies might use the two proposed confounding parameters and the associated R package to size their work more accurately than with randomized-trial formulas alone.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Unique identifiability of PS distribution from BC + pi requires an extra parametric family on law of e(X)","rationale":"The reader's weakest_assumption correctly flags the parametric PS + semiparametric outcome models as the point where identifiability is asserted. The load-bearing risk is that the asserted uniqueness embeds a third modeling choice (the family for the PS law) not counted among the two sensitivity parameters; this is internal to the argument rather than an external consensus issue. If the full derivations show the family is canonical and results are insensitive, the claim can be accepted; otherwise the sufficiency statement requires the extra qualification. This moves the verdict from UNVERDICTED to CONDITIONAL pending verification of the test.","tokens_in":1758,"tokens_out":478,"duration_ms":42295,"concrete_test":"From the methods/appendix, extract the exact parametric family and mapping used to obtain the distribution of e(X) from (BC, pi). Simulate data from an alternative family (e.g., mixture of betas or kernel density) calibrated to the same BC and pi but different variance of e(X); recompute the exact Var of the IPW estimator on the simulated sample and compare to the paper's closed-form expression—if the ratio differs by >15% for any of the numerical examples in the paper, the two-parameter claim does not hold.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim asserts that two parameters (BC for confounder-treatment strength, R²-bounded sensitivity for confounder-outcome) suffice to determine minimal n for IPW-based ATE power. The abstract states this works because a parametric PS model plus BC + treatment proportion yields a uniquely identifiable PS distribution (needed for the propensity component of Var(IPW)). However, BC equals a single functional E[sqrt(e(X)(1-e(X)))] (normalized by pi), which does not determine the full distribution of the weights 1/e(X) and 1/(1-e(X)) without an additional low-dimensional parametric restriction on the law of e(X) itself. The paper explicitly avoids assumptions on the multivariate covariate distribution, so this restriction cannot be derived from the covariates; it must be imposed separately on the PS distribution. Misspecification of that family directly alters the computed variance and thus the reported minimal sample size.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript develops analytical formulas for sample size and power calculations for estimating the ATE via IPW in observational studies. It decomposes the IPW variance into propensity-score distribution, potential-outcome distribution, and their correlation components. The central claim is that, beyond the usual RCT inputs, only two additional parameters suffice: the Bhattacharyya coefficient (measuring confounder-treatment association and, with treatment proportion, yielding a uniquely identifiable PS distribution under a parametric PS model) and a sensitivity parameter for confounder-outcome association bounded by the R² from regressing the outcome on covariates. The procedure uses a parametric PS model and semiparametric restricted-mean outcome model without distributional assumptions on the multivariate covariates; an R package and online calculator are provided.","tokens_in":1957,"tokens_out":598,"duration_ms":47687,"significance":"If the identifiability and bounding arguments hold, the work supplies a practical, low-input framework for power analysis in observational causal inference, which is a frequent practical need. The explicit variance decomposition and software release are strengths that would aid reproducibility and adoption. The avoidance of full covariate-distribution assumptions is a positive feature relative to simulation-based alternatives.","major_comments":[{"comment":"Abstract (final paragraph) and the section deriving the PS distribution: the claim that the Bhattacharyya coefficient plus treatment proportion 'leads to a uniquely identifiable' PS distribution under the parametric PS model is load-bearing for the two-parameter sufficiency result. Because the BC equals a single functional E[sqrt(e(X)(1-e(X)))] (normalized by π), uniqueness of the full law of the weights 1/e(X) and 1/(1-e(X)) requires an explicit statement of the low-dimensional parametric family imposed on the distribution of e(X) itself; the manuscript's statement that no assumptions are made on the multivariate law of X leaves open whether this family is an additional modeling choice or is derived.","section":"Abstract and PS-distribution derivation"},{"comment":"Variance-decomposition section (around the IPW variance formula): the correlation term between the PS weights and the potential outcomes must be shown to be either bounded or eliminated by the two parameters without introducing further user inputs; otherwise the reduction to exactly two extra parameters does not follow from the decomposition alone.","section":"Variance decomposition"}],"minor_comments":[{"comment":"The abstract states reliance on 'a parametric propensity score model' but does not name the family (e.g., logistic, beta, etc.); adding this detail would improve immediate readability.","section":"Abstract"},{"comment":"Figure captions or the software section could include a small numerical example showing how the two parameters translate into a concrete minimal n, to illustrate the formulas.","section":"Software and examples"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and constructive major comments. Both points identify areas where additional explicit statements and derivations would strengthen the manuscript. We agree that clarifications are warranted and will revise accordingly.","responses":[{"response":"We agree that the uniqueness result requires an explicit statement of the parametric family on the distribution of e(X). The manuscript already states reliance on a parametric propensity score model; under this model the BC together with the treatment proportion π uniquely determines the parameters of the induced distribution of e(X) (and hence the law of the IPW weights). To remove any ambiguity, we will revise the relevant section and abstract to name the specific low-dimensional parametric family for e(X) (derived directly from the parametric PS model) and to clarify that no further distributional assumptions on the multivariate law of X are introduced beyond those already declared.","revision_made":"yes","referee_comment":"[Abstract and PS-distribution derivation] Abstract (final paragraph) and the section deriving the PS distribution: the claim that the Bhattacharyya coefficient plus treatment proportion 'leads to a uniquely identifiable' PS distribution under the parametric PS model is load-bearing for the two-parameter sufficiency result. Because the BC equals a single functional E[sqrt(e(X)(1-e(X)))] (normalized by π), uniqueness of the full law of the weights 1/e(X) and 1/(1-e(X)) requires an explicit statement of the low-dimensional parametric family imposed on the distribution of e(X) itself; the manuscript's statement that no assumptions are made on the multivariate law of X leaves open whether this family is an additional modeling choice or is derived."},{"response":"The referee correctly notes that the correlation term must be controlled by the two parameters. The R²-bounded sensitivity parameter for the confounder-outcome association is intended to bound the feasible range of this correlation (via its effect on the covariance between weights and potential outcomes) without additional user-specified inputs. We will add an explicit bounding argument or derivation in the variance-decomposition section demonstrating that the correlation is indeed governed by these two quantities alone, thereby confirming that exactly two extra parameters suffice.","revision_made":"yes","referee_comment":"[Variance decomposition] Variance-decomposition section (around the IPW variance formula): the correlation term between the PS weights and the potential outcomes must be shown to be either bounded or eliminated by the two parameters without introducing further user inputs; otherwise the reduction to exactly two extra parameters does not follow from the decomposition alone."}],"tokens_in":1481,"tokens_out":537,"duration_ms":41618,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that this paper gives explicit formulas and software to size observational studies for IPW-based ATE estimation by adding only two user-specified parameters beyond the usual RCT inputs: a Bhattacharyya coefficient to capture confounder-treatment strength and an R-squared-bounded sensitivity parameter for the confounder-outcome link. The variance is decomposed into propensity, outcome, and correlation pieces, and the claim is that these two numbers plus treatment proportion suffice once a parametric PS model is fixed. They also release an R package and online calculator, which is the kind of deliverable that actually gets used. That is the concrete advance over standard power literature. The approach is upfront about relying on a parametric propensity score model and a semiparametric restricted-mean outcome model, and it avoids needing the full multivariate covariate distribution. The BC-plus-pi step is presented as yielding a unique PS distribution under that parametric restriction, which is a reasonable way to make the problem tractable. The R2 bound for the outcome component is a sensible, conservative choice that stays within observable quantities. The main limitation is exactly the one flagged in the stress test: the parametric family on the PS distribution is doing the heavy lifting for identifiability, since BC alone is only one functional and does not pin down the full weight distribution without it. If the chosen family is misspecified, the computed variance and therefore the minimal n will be off. The abstract does not report numerical checks or simulation studies that would show how sensitive the results are to that choice, so the practical robustness remains to be verified. The derivations themselves are not visible here, which keeps the soundness assessment provisional. This is aimed at biostatisticians and epidemiologists who plan observational studies and want a structured alternative to purely ad-hoc sensitivity analyses. It is worth sending to peer review because it fills a genuine applied gap with usable formulas and code; referees can check the derivations and ask for validation simulations without the work being fundamentally broken.","headline":"Two extra parameters (BC for overlap, R2-bound sensitivity) plus a parametric PS model let you power IPW observational studies, but the identifiability claim depends on that model choice.","tokens_in":2484,"tokens_out":474,"would_cite":false,"duration_ms":24077,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"IPW variance decomposition and Bhattacharyya-based overlap identifiability lie outside RS forcing chain","alignment":"orthogonal","rationale":"The paper's central machinery (parametric logit PS + Beta approximation, Bhattacharyya coefficient ϕ yielding unique (a,b) via Prop. 1-2, Lyapunov CLT reduction of X'β, R²-bounded ρ sensitivity in variance formula (17)) is classical frequentist causal-inference methodology. It never invokes reciprocal cost J, ratio symmetry, φ-ladder, 8-tick periodicity, or parameter-free constant derivation. RS modules (Cost/FunctionalEquation.lean: washburn_uniqueness_aczel; Foundation/DimensionForcing.lean: D=3; Foundation/AbsoluteFloorClosure.lean: reality_from_one_distinction) therefore supply no relevant theorems; the statistical construction is domain-orthogonal.","tokens_in":62587,"confidence":"high","tokens_out":197,"duration_ms":8756,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"To calculate the minimal sample size for an observational causal study, it suffices to know two parameters quantifying the confounder-treatment and confounder-outcome associations in addition to standard randomized trial inputs.","keywords":["sample size calculation","power analysis","observational studies","causal inference","inverse probability weighting","Bhattacharyya coefficient","propensity score","confounding"],"falsifier":"In a dataset with known confounder strengths, compute the actual variance of the IPW estimator directly and compare it to the variance predicted by the formula that uses only the two proposed parameters; a systematic mismatch would falsify the claim that these two suffice.","tokens_in":2624,"feed_emoji":"","tokens_out":659,"duration_ms":39772,"temperature":0.7,"pith_summary":"This paper develops analytical formulas for sample size and power calculations in observational studies using inverse probability weighting for the average treatment effect. It decomposes the variance into propensity score distribution, potential outcome distribution, and their correlation. The key finding is that only two additional parameters are needed beyond those for randomized trials: the Bhattacharyya coefficient for covariate overlap and a sensitivity parameter bounded by the R-squared of the outcome regression. This makes power analysis practical for observational data without requiring full knowledge of multivariate covariate distributions. Sympathetic readers would care because designing studies with adequate power is essential, and observational studies often face challenges in specifying all components.","feed_headline":"Two parameters suffice to size observational causal studies","feed_subtitle":"Power analysis for average treatment effects needs only confounder-treatment overlap via Bhattacharyya coefficient and outcome association R","key_machinery":"Variance decomposition of the inverse probability weighting estimator for average treatment effect, with propensity score distribution identified from Bhattacharyya coefficient plus treatment proportion and outcome correlation bounded by R-squared of outcome on covariates.","core_discovery":"By analyzing the variance of an inverse probability weighting estimator of the average treatment effect, we decompose the power calculation into three components: propensity score distribution, potential outcome distribution, and their correlation. We show that to determine the minimal sample size of an observational study, in addition to the standard inputs in the power calculation of randomized trials, it is sufficient to have two parameters, which quantify the strength of the confounder-treatment and the confounder-outcome association, respectively. For the former, we propose using the Bhattacharyya coefficient, which measures the covariate overlap and, together with the treatment比例,leads","pith_inferences":["Pilot data could be used to estimate the Bhattacharyya coefficient for study planning.","The two-parameter approach may combine with existing sensitivity analysis techniques in causal inference.","Similar variance decompositions could be derived for estimators other than inverse probability weighting.","Empirical checks comparing predicted versus observed power in completed studies would test practical accuracy."],"forward_implications":["Minimal sample size follows from standard randomized trial inputs plus the two parameters.","Propensity score distribution is uniquely identifiable from the Bhattacharyya coefficient and treatment proportion.","The sensitivity parameter for the outcome association is bounded by the R-squared statistic without needing full covariate distributional assumptions.","The procedure applies under a parametric propensity score model and semiparametric restricted mean outcome model.","An R package and online calculator implement the formulas."],"fun_headline_variants":[],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The parametric propensity score model is correctly specified and the outcome association can be bounded using only the R-squared statistic from regressing outcome on covariates.","fun_headline_variants_meta":{"error":"Expecting value: line 2 column 9069 (char 9070)"},"cache_creation_input_tokens":0},"created_at":"2026-05-23T04:54:36.769351+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"In a dataset with known confounder strengths, compute the actual variance of the IPW estimator directly and compare it to the variance predicted by the formula that uses only the two proposed parameters; a systematic mismatch would falsify the claim that these two suffice.","supporting_citations":[],"review_version":1}