{"id":"0621e8e3-687b-4b37-960a-4cb5b377b887","arxiv_id":"1908.07352","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For general matched observational studies, sensitivity analysis for the weak null cannot be simultaneously sharp for the sharp null; two procedures, one valid for heterogeneous effects but conservative and one sharp under constant effects but valid under extra conditions, implement the two sides.","lead":"This statistics paper develops sensitivity analyses for matched observational studies when the interest is in the average treatment effect and individual effects vary. It proves that in general matched designs, no single sensitivity analysis can be simultaneously valid for average effects and sharp for constant effects, then offers two workable procedures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's weak-null expectation bound for \\tilde K appears false: a concrete Γ=2 construction gives E(\\tilde K)>0, so the paper's universally valid weak-null sensitivity analysis lacks support.","rationale":"The reader's conditional verdict focuses on the D-bar procedure's covariance conditions in §7.2 and on the omitted proof of Proposition 5. My reading found a more direct and more serious problem: Theorem 2, the paper's advertised procedure valid over the entirety of the weak null, appears to be contradicted by an explicit finite-sample construction under the paper's own sensitivity model. The construction uses only two strata, potential outcomes with overall average effect zero, and unmeasured confounders in {0,1}, so it is squarely inside the theorem's hypotheses. The computed expectation of \\tildeK is strictly positive rather than bounded above by zero. This is not a matter of disagreement with current consensus or of missing regularity conditions; it is an internal inconsistency in the central methodological claim. The paper does include a constructive impossibility theorem and a simulation study, and those may retain independent value, but the abstract's promise of 'a sensitivity analysis valid for the weak null' is unsupported if Theorem 2 fails. Because the counterexample is algebraic and small, a single exact computation settles the issue; no code or machine-checked proof is needed. Given that the central guaranteed-valid procedure fails as stated, I would move the verdict from CONDITIONAL to REJECT.","tokens_in":29219,"tokens_out":43843,"duration_ms":459738,"concrete_test":"Evaluate the exact expectation in the displayed two-stratum design: compute E(\\tildeK)=(5/103)(-200/7)+(199/103)(40400/50399). If this number is positive, approximately 0.162, then Theorem 2 as stated is false; the same check also undermines Theorem 5 for \\tildeK because its expectation bound is the starting point of the validity argument. Repeating the two-stratum construction across many blocks keeps the sample average treatment effect zero and preserves the positive expectation, so the failure is not a finite-B artifact.","verdict_should_be":"REJECT","load_bearing_attack":"The most load-bearing concern is not the untestable covariance condition in §7.2 but Theorem 2 itself. Under the paper's own model, take Γ=2 and B=2. Block 1 has n1=3, rC=(0,0,0), rT=(-100,0,0); block 2 has n2=100, rC=(0,...,0), rT=(100,0,...,0). Then Σ n_i \\barτ_i = 0, so H_N^(0) holds. Take u1=(0,1,1) and u2=(1,0,...,0); thus p1=(1/5,2/5,2/5) and p2=(2/101,1/101,...,1/101), which satisfies (1) at Γ=2. With \\tildeκ1=5, Γ_{n1}=5/2, \\tildeκ2=199, Γ_{n2}=398/101, we have E(\\tildeD1)=0.2·(-100)·(1+3/7)=-200/7, and E(\\tildeD2)=(2/101)·100·(1-297/499)=40400/50399>0. Hence E(\\tildeK)=(5/103)(-200/7)+(199/103)(40400/50399)≈0.162>0, contradicting the claimed E_u(\\tildeK)≤0. Replacing Γ by the stratum-dependent Γ_{n_i} in (13), while weighting by \\tildeκ_{Γ n_i}, destroys the cancellation that Proposition 2 relies on; the appendix supplies no proof of Theorem 2.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops sensitivity analyses for Neyman's weak null hypothesis about the sample average treatment effect in matched observational studies, allowing for arbitrary effect heterogeneity. It proves an impossibility result (Theorem 1) for general matched designs, proposes two test statistics: \\tilde K, claimed to be valid for the entire weak null under Rosenbaum's sensitivity model, and \\bar D, which is asymptotically sharp under constant effects but is claimed valid for the weak null only under additional covariance conditions. The paper also gives a conservative variance estimator (Proposition 5), a binary-outcome integer-programming bound, simulations, and a data example. The central claim is that \\tilde K provides a valid sensitivity analysis for the weak null over the entirety of Rosenbaum's model. I find this central claim to be false: Theorem 2 is contradicted by an explicit finite-population construction, and the appendix does not contain the promised proof of Theorem 2.","tokens_in":29551,"tokens_out":11715,"duration_ms":110241,"significance":"The paper's topic is important and the impossibility result in Theorem 1 is a genuine conceptual contribution, as is the connection between the proposed statistics and inverse probability weighting. The simulation study is carefully designed and the binary-outcome integer-programming extension is interesting. However, the headline methodological guarantee---that \\tilde K is valid for the weak null under the full Rosenbaum model---is the load-bearing result of the paper, and it is false as stated. Because the counterexample gives positive expectation for a data-generating process satisfying the paper's own model and the weak null, the claimed universal validity of the \\tilde K-based sensitivity analysis is unsupported. This is not a local or cosmetic flaw; it undermines one of the two main procedures and the abstract's claim of a sensitivity analysis valid for the entirety of the weak null.","major_comments":[{"comment":"Theorem 2's assertion that E_u(\\tilde K_Γ^{(τ0)}|F,Z) ≤ 0 under (1) at Γ and H_N^{(τ0)} is false. Counterexample: take τ0=0, Γ=2, B=2, n1=3, n2=100. In block 1 set rC=(0,0,0) and rT=(-100,0,0); in block 2 set rC=(0,...,0) and rT=(100,0,...,0). Then n1·\\barτ1 + n2·\\barτ2 = 3·(-100/3)+100·1 = 0, so the weak null H_N^{(0)} holds. Take hidden covariates u1=(0,1,1) and u2=(1,0,...,0), giving assignment probabilities p1=(1/5,2/5,2/5) and p2=(2/101,1/101,...,1/101); these are of the form exp(γu)/Σexp(γu) with γ=log2, hence admissible under the model stated in §2.2. With \\tildeκ1=5, Γ_{n1}=5/2, \\tildeκ2=199, and Γ_{n2}=398/101, direct calculation gives E(\\tildeD1)=-200/7 and E(\\tildeD2)=40400/50399≈0.802. Therefore E(\\tildeK)=(1/103)(5·(-200/7)+199·(40400/50399))≈0.162>0, contradicting Theorem 2. Repeating the two-block construction M times gives the same positive per-observation expectation with B=2M, so the failure is not an artifact of small B. The problem is that replacing Γ by the stratum-dependent Γ_{n_i} in (13) while weighting by \\tildeκ_{Γn_i} destroys the cancellation on which Proposition 2 relies.","section":"§5.3, Theorem 2"},{"comment":"The text states that Theorem 2 is proved in the appendix, but Appendix C contains no proof of Theorem 2. It proves Propositions 1-4 and Theorem 3, but the proof of Theorem 2 is absent. In light of the counterexample above, this is not merely an omitted detail: the claimed theorem is false, and no proof can be supplied without changing the statement or the statistic.","section":"Appendix C"},{"comment":"The validity of the \\bar D-based procedure for the weak null rests on conditions (a) and (b), which are untestable from the observed data and are not shown to hold under any substantive primitive condition on the data-generating process. The paper's own simulation setting (h) in Table 2 violates these conditions and produces Type I error 0.138 at nominal 0.10 for continuous outcomes. This is acknowledged in the text, but it means that the abstract's characterization of the restrictions as 'benign' is supported only by the choice of simulation settings, not by a formal argument. Since \\bar D is the less conservative procedure recommended for practice, this limitation is load-bearing for the paper's practical claims.","section":"§7.2, Theorem 4"}],"minor_comments":[{"comment":"In the first paragraph of §8, 'the onyl feasible value' should read 'the only feasible value'.","section":"§8"},{"comment":"In the heading of Appendix B.2, 'continous' should be 'continuous'.","section":"Appendix B.2"},{"comment":"The constraint on Z is written with reversed indices: ∑_{n_i}^{i=1} Z_{ij}=1 should be ∑_{j=1}^{n_i} Z_{ij}=1.","section":"§2.1"},{"comment":"In the proof of Proposition 2, equation (20) and the surrounding display use ∑_{i=1}^{n_i} where the summation index should be j; please correct all such index typos.","section":"Appendix C.3"},{"comment":"The sentence 'the discussion has focused on the the extent' contains a duplicated article; it should read 'on the extent'.","section":"§7.1"}],"recommendation":"reject","confidential_remarks":"The central theoretical guarantee of the paper is false: a simple finite-population construction with Γ=2 and two blocks gives positive expectation for \\tilde K under the weak null, and the promised proof of Theorem 2 is absent from the appendix. Because the \\tilde K procedure is the only one claimed to be valid over the entirety of the weak null, the paper's main conclusion cannot stand without a substantially different statistic or a substantially different theorem. The error is load-bearing and not a local fix, so I recommend rejection. The impossibility result in Theorem 1 and the IPW connection are interesting and may be worth salvaging in future work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about this paper before spending time on it. First, the impossibility result (Theorem 1) is a genuine contribution: it shows that in general matched designs, any sensitivity analysis that is tight under Fisher's sharp null cannot also be valid over the whole weak null. That result is new, the proof is constructive, and it will matter for how people teach and design sensitivity analyses. Second—and this is the larger story—the paper's advertised remedy, the procedure based on \\tilde K in §5.3, does not actually work. Theorem 2, which claims E_u(\\tilde K) ≤ 0 under the weak null, is false as stated.\n\nI checked the stress-test construction by hand. At Γ = 2, take two blocks: block 1 with n = 3, rC = (0,0,0), rT = (-100,0,0); block 2 with n = 100, rC all zero, rT = (100,0,...,0). The average treatment effect is zero. Let u1 = (0,1,1) and u2 = (1,0,...,0), giving assignment probabilities (1/5,2/5,2/5) and (2/101,1/101,...,1/101), both within Rosenbaum's model at Γ = 2. Then E(\\tilde D_1) = -200/7 and E(\\tilde D_2) = 40400/50399. Weighting by \\tilde κ = 5 and 199 and dividing by N = 103 gives E(\\tilde K) ≈ 0.162 > 0. That directly contradicts Theorem 2. The appendix contains no proof of Theorem 2, so this is not a case of a fixable typo in an otherwise verified argument.\n\nWhat the paper does well, beyond Theorem 1, is the IPW interpretation of D as a worst-case inverse probability weighted estimator, and the binary-outcome integer program in §8. Both are solid and worth keeping. The simulations are honest, including the setting (h) where \\bar D fails. The omission of Proposition 5's proof is a minor issue since it is delegated to prior published work; the bigger omission is Theorem 2's proof.\n\nThe \\bar D procedure may still be salvageable as a valid method under the covariance conditions in §7.2, and the binary IP method looks defensible. But the claim of a universally valid weak-null sensitivity analysis for general matched designs, which is the abstract's main selling point, is unsupported. A serious referee should engage with the paper because the impossibility theorem and the binary section are important, but the revision path is not cosmetic: Theorem 2 needs either a corrected statement with additional assumptions or removal from the paper.\n\nFor a reading group, maybe, if people want to discuss what happens when a central theorem fails. I would not cite it as it stands.","headline":"The paper's headline claim—a sensitivity analysis valid for the entire weak null in general matched designs—is undermined by a concrete counterexample to Theorem 2; the rest of the paper still contains real ideas worth refereeing.","tokens_in":30031,"tokens_out":12126,"would_cite":false,"duration_ms":112684,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62F40"],"pacs":[],"model":"deepseek-v4-flash","headline":"In matched observational studies, no one sensitivity analysis can be simultaneously sharp for constant effects and valid for average effects.","keywords":["sensitivity analysis","weak null hypothesis","matched observational studies","heterogeneous treatment effects","sample average treatment effect","hidden bias","randomization inference","binary outcomes"],"falsifier":"Run the paper's continuous-outcome simulation setting (h), where the sign of the stratum-level effect dictates whether individual effects are right- or left-skewed, at $\\Gamma=5$, $B=500$, and nominal level $0.10$; the paper reports Type I error $0.138$ for the $\\bar D_{\\Gamma}^{(\\tau_0)}$ procedure. Any repeated run that keeps the error at $0.10$ or below under that same generative model would show the stated covariance condition is not actually necessary.","tokens_in":28968,"feed_emoji":"📊","tokens_out":12477,"duration_ms":116143,"temperature":0.7,"pith_summary":"This paper asks whether a single sensitivity analysis for matched observational studies can be both sharp under the constant-effects null (no individual treatment effect) and valid under the weak null (sample average treatment effect zero). It establishes that, for matched sets larger than pairs, the answer is no over a broad class of test statistics: any procedure whose worst-case expectation is tight under constant effects generally fails to control the worst-case expectation when effects are heterogeneous and average to zero (Theorem 1). The paper then develops two workable procedures: one is always valid for the weak null but conservative under constant effects, the other is asymptotically sharp under constant effects and valid for the weak null only under an extra covariance condition on hidden bias and stratum-level effects. The practical payoff is that researchers using any optimal without-replacement matching can assess robustness to hidden bias without assuming effect heterogeneity away, as long as they choose which null to prioritize.","feed_headline":"One sensitivity analysis cannot serve both constant and average effects","feed_subtitle":"This paper proves the trade-off and offers two procedures, one for each goal.","key_machinery":"The load-bearing object is the stratum-level statistic $D_{\\Gamma i}^{(\\tau_0)} = \\hat\\tau_i - \\tau_0 - \\frac{\\Gamma-1}{\\Gamma+1}|\\hat\\tau_i - \\tau_0|$: the treated-minus-control mean difference in matched set $i$, minus the worst-case bias that a paired design would allow under the sensitivity model. Weighted across strata, this statistic is proportional to a worst-case inverse-probability-weighted estimator under an interval restriction on assignment probabilities, and this connection is what lets the paper bound its expectation over the composite weak null. A conservative variance estimator is obtained by regressing the stratum contributions on a fixed design matrix and using the residual variance, which overestimates the true variance in expectation. Replacing $\\Gamma$ with the larger effective value $\\Gamma_{n_i}$ yields $\\tilde K_{\\Gamma}^{(\\tau_0)}$, the version valid over the entire weak null; the contrast between the two versions is the paper's main instrument.","core_discovery":"The central discovery is an incompatibility theorem. For any nondecreasing, nonconstant function $h_{\\Gamma n_i}$ of the stratum-wise treated-minus-control mean difference $\\hat\\tau_i - \\tau_0$, there exist stratum sizes, hidden-bias levels $\\Gamma$, and potential outcomes satisfying the weak null $\\bar\\tau = \\tau_0$ such that the separable algorithm's worst-case expectation under the sharp null fails to bound the worst-case expectation under the weak null. Thus, in general matched designs with $n_i \\ge 3$, sensitivity analysis for the sample average treatment effect cannot be unified with sensitivity analysis for constant effects. The paper proves this constructively, then shows that the statistic $D_{\\Gamma i}^{(\\tau_0)} = \\hat\\tau_i - \\tau_0 - \\frac{\\Gamma-1}{\\Gamma+1}|\\hat\\tau_i-\\tau_0|$ bounds the worst-case expectation under the weak null when weighted appropriately, and that a modified version $\\tilde K$ is valid over the whole weak null under the standard sensitivity model but is strictly conservative under the sharp null. A second procedure based on weighted $D_{\\Gamma i}^{(\\tau_0)}$ is asymptotically sharp under constant effects and valid for the weak null under additional covariance conditions; the paper's own simulations show one plausible generative model violates those conditions and the procedure's Type I error exceeds nominal levels.","pith_inferences":["The same incompatibility should be expected for test statistics outside the monotone-mean class, such as rank-based or M-estimators of stratum effects, because the driving mechanism is unequal weighting of heterogeneous effects by worst-case selection probabilities.","Corollary 1 suggests a concrete diagnostic: estimate each stratum's mean difference $\\hat\\tau_i$, then check whether the covariance between the estimated probability of $\\hat\\tau_i \\ge \\bar\\tau_i$ and $n_i(\\bar\\tau_i - \\tau_0)$ is negative; the paper does not propose this as a routine, but it follows directly from its necessary condition.","If the target estimand were a weighted average treatment effect rather than the unweighted sample average, tuning the weights $\\kappa_{\\Gamma i}$ in the interval-restriction construction could reduce the conservativeness of the weak-null-valid procedure, an extension the paper does not pursue.","The data example shows the choice between procedures can change the qualitative conclusion about hidden-bias robustness (changepoint $\\Gamma$ 1.29 versus 1.52 in the 1:5 matched lead study), so practitioners need guidance on when condition (a)/(b) is plausible."],"forward_implications":["In matched designs with sets of size three or more, a sensitivity analysis that is tight under constant effects cannot guarantee weak-null Type I error; Theorem 1 says failures can be constructed for every monotone function of the stratum mean differences.","The procedure $\\tilde K_{\\Gamma}^{(\\tau_0)}$ gives guaranteed asymptotic control of the weak null under the standard sensitivity model, but it is strictly conservative when effects are constant unless the design is paired or $\\Gamma=1$, so reported robustness to hidden bias can drop sharply.","The procedure $\\bar D_{\\Gamma}^{(\\tau_0)}$ is asymptotically sharp under constant effects and, when either condition (a) or (b) holds, controls the weak null; its practical validity therefore rests on an unverifiable covariance condition.","For binary outcomes, the worst-case expectation over the whole weak null can be computed exactly as an integer program, yielding a valid sensitivity analysis that avoids the conservativeness of $\\tilde K$; in the paper's simulations it solves in under half a second per iteration."],"supporting_citations":[{"why":"It supplies the asymptotically separable algorithm that defines worst-case expectations under the sharp null, which Theorem 1 shows cannot double as weak-null bounds.","marker":"Gastwirth et al. (2000)"},{"why":"It introduces the D-statistic and shows it controls the weak null in paired designs; the present paper extends and contrasts that result for general matched designs.","marker":"Fogarty (2019)"},{"why":"It reduces the worst-case hidden-bias search to binary vectors, a reduction used throughout the expectation bounds and proofs.","marker":"Rosenbaum and Krieger (1990)"},{"why":"It provides the integer-program formulation that computes worst-case expectations for binary outcomes under the weak null.","marker":"Fogarty et al. (2017)"},{"why":"It defines the weak null hypothesis that the paper tests, in contrast to Fisher's sharp null.","marker":"Neyman (1935)"},{"why":"It documents the impossibility of consistent variance estimation under effect heterogeneity, motivating the conservative standard-error construction.","marker":"Ding (2017)"},{"why":"It supplies the variance estimators for finely stratified designs that Proposition 5 extends to the sensitivity-analysis setting.","marker":"Fogarty (2018)"},{"why":"It establishes inverse-probability weighting, the interpretative lens connecting the D-statistic to a worst-case IPW estimator.","marker":"Horvitz and Thompson (1952)"},{"why":"It formulates the odds-ratio sensitivity model (1) that governs all of the paper's hidden-bias calculations.","marker":"Rosenbaum (2002, Chapter 4)"}],"fun_headline_variants":["Sensitivity analysis faces a trade-off: average vs. constant effects","Matched studies: no single sensitivity test for both nulls","Proving a sensitivity trade-off in matched observational studies","Weak null vs. sharp null: a sensitivity analysis clash","Two procedures needed for average and constant effects"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recommended procedure that is sharp under constant effects controls Type I error for the weak null only if, in the population of matched sets, larger stratum-level average effects do not tend to come with lower chances that the stratum's observed mean difference reaches its own average, for either the true hidden bias or the worst-case one; this condition is unverifiable and one of the paper's own simulation settings (h) violates it.","fun_headline_variants_meta":{"raw":{"variants":["Sensitivity analysis faces a trade-off: average vs. constant effects","Matched studies: no single sensitivity test for both nulls","Proving a sensitivity trade-off in matched observational studies","Weak null vs. sharp null: a sensitivity analysis clash","Two procedures needed for average and constant effects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000415,"raw_usage":{"total_tokens":2171,"prompt_tokens":999,"completion_tokens":1172,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":1092}},"tokens_in":615,"tokens_out":1172,"duration_ms":8740,"temperature":1.0,"reasoning_tokens":1092,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:20:23.228060+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's continuous-outcome simulation setting (h), where the sign of the stratum-level effect dictates whether individual effects are right- or left-skewed, at $\\Gamma=5$, $B=500$, and nominal level $0.10$; the paper reports Type I error $0.138$ for the $\\bar D_{\\Gamma}^{(\\tau_0)}$ procedure. Any repeated run that keeps the error at $0.10$ or below under that same generative model would show the stated covariance condition is not actually necessary.","supporting_citations":[{"cited_title":"L., Krieger, A","cited_arxiv_id":null,"evidence_quote":"It supplies the asymptotically separable algorithm that defines worst-case expectations under the sharp null, which Theorem 1 shows cannot double as weak-null bounds."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It reduces the worst-case hidden-bias search to binary vectors, a reduction used throughout the expectation bounds and proofs."},{"cited_title":"B., Shi, P., Mikkelsen, M","cited_arxiv_id":null,"evidence_quote":"It provides the integer-program formulation that computes worst-case expectations for binary outcomes under the weak null."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines the weak null hypothesis that the paper tests, in contrast to Fisher's sharp null."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It documents the impossibility of consistent variance estimation under effect heterogeneity, motivating the conservative standard-error construction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the variance estimators for finely stratified designs that Proposition 5 extends to the sensitivity-analysis setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It formulates the odds-ratio sensitivity model (1) that governs all of the paper's hidden-bias calculations."}],"review_version":1}