Pith. sign in

REVIEW 2 major objections 1 minor 4 references

Powerful Multivariate Sensitivity Analysis via Sample Splitting in an Observational Study of the Effects of Poverty on Cardiovascular Disease Risk Factors

T0 review · 2 major / 1 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read Splitting the sample into planning and analysis parts lets researchers select an optimal linear combination of outcomes for sensitivity analysis while preserving the same asymptotic power as full-sample methods.

desk verdict Sample splitting to pick linear combinations while preserving asymptotic power is the core idea, but the dependence from selection is the spot that needs the derivation to hold up. read the letter →

arxiv 2606.04416 v1 pith:JEVCWE6Y submitted 2026-06-03 stat.ME math.STstat.APstat.TH

classification stat.MEmath.STstat.APstat.TH
keywords samplesplittingmultivariatesensitivityanalysislinearcombinationsunmeasuredconfoundingobservationalstudiesasymptoticpowerfinite-samplecardiovascularriskfactors
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

When testing causal effects on several outcomes at once, a well-chosen linear combination can reduce how much the result depends on possible unmeasured biases. Searching over all possible combinations normally triggers multiple-testing corrections that lower power. This paper introduces a split-sample procedure: one portion of the data is used to pick the combination that appears least sensitive, and the remaining portion is used for the actual test. The authors characterize the exact set of combinations for which the split-sample test matches the asymptotic power of any full-sample competitor. Simulations show the split improves power in moderate-sized samples, and the method is applied to data on poverty and cardiovascular risk factors in children and adolescents.

What carries the argument

Sample splitting into planning and analysis subsamples to select an optimal linear combination without incurring a multiple-testing penalty, together with a characterization of the combinations that preserve asymptotic power equivalence.

What would settle it

A Monte Carlo experiment in which the linear combination chosen from the planning sample produces a test whose asymptotic power falls below that of the corresponding full-sample procedure under the paper's sensitivity model and study design.

Watch

Extended reading notes

Core claim

Splitting the sample into a planning portion for selecting an optimal linear combination of scored outcomes and an analysis portion for inference yields the same asymptotic power as full-sample alternatives for a characterized class of combinations; finite-sample simulations confirm higher power, and the procedure applied to poverty effects on child cardiovascular risk factors identifies adverse impacts on body composition, physical activity, and tobacco exposure.

Load-bearing premise

The selection step performed on the planning sample must not create dependence that breaks the claimed asymptotic power equivalence with full-sample methods.

Editorial extensions

If this is right

  • The procedure can be used whenever multiple outcomes are available and a linear combination is expected to reduce sensitivity to unmeasured confounding.
  • Extensive simulations demonstrate strictly higher finite-sample power than full-sample competitors that search over combinations.
  • In the poverty application the method detects effects on body-composition, activity, and tobacco-exposure outcomes, with the tobacco finding showing greater robustness to hidden bias.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same splitting idea could be examined in settings where outcomes are measured at multiple time points rather than cross-sectionally.
  • If the characterization of power-preserving combinations can be stated in terms of the eigenvalues of the outcome covariance matrix, it may be possible to compute the admissible set without enumerating all directions.
  • Applied researchers could test the practical value of the method by comparing split-sample results against a pre-registered analysis that fixes the linear combination in advance.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper proposes splitting an observational study sample into a planning sample (to select an optimal linear combination of outcomes minimizing sensitivity to unmeasured bias) and an analysis sample (for inference on the global null). It claims a novel characterization of the linear combinations for which this split-sample procedure achieves the same asymptotic power as full-sample alternatives (e.g., Scheffé projections), reports simulation evidence of finite-sample power gains, and applies the method to poverty effects on multiple CVD risk factors in youth, finding adverse associations with body composition, physical activity, and tobacco exposure (with varying robustness).

Significance. If the characterization is valid under the induced dependence, the method would allow data-driven selection of outcome combinations without multiple-testing penalties eroding power, a practical advance for multivariate sensitivity analysis in observational epidemiology. The simulation studies and real-data application would strengthen the case for adoption if the asymptotic result holds.

major comments (2)
  1. [Abstract / characterization section] Abstract and the section presenting the novel characterization: the claimed asymptotic power equivalence must be derived under the dependence between the planning-sample selection indicator and the analysis-sample outcomes (induced by the observational design and sensitivity model). If the derivation treats the selected combination as fixed or assumes independence between samples, the guarantee fails precisely for the data-dependent combinations the procedure is designed to use; the manuscript must exhibit the limiting distribution of the test statistic accounting for this selection step.
  2. [Simulation studies] Simulation design (results section): without explicit details on how the planning-sample selection is implemented and whether the finite-sample power gains persist when the dependence structure matches the observational setting, it is unclear whether the reported enhancements support the central claim or merely reflect an idealized independent-split scenario.
minor comments (1)
  1. [Abstract] The abstract states the characterization exists but does not list the precise conditions (e.g., on the sensitivity model or outcome scoring); adding a one-sentence summary of those conditions would improve readability.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the careful reading and constructive comments on our manuscript. We address each major comment below, providing clarifications on the asymptotic characterization and agreeing to enhance the simulation details.

read point-by-point responses
  1. Referee: [Abstract / characterization section] Abstract and the section presenting the novel characterization: the claimed asymptotic power equivalence must be derived under the dependence between the planning-sample selection indicator and the analysis-sample outcomes (induced by the observational design and sensitivity model). If the derivation treats the selected combination as fixed or assumes independence between samples, the guarantee fails precisely for the data-dependent combinations the procedure is designed to use; the manuscript must exhibit the limiting distribution of the test statistic accounting for this selection step.

    Authors: We appreciate the referee raising this point regarding the limiting distribution. However, because the sample is randomly partitioned into planning and analysis subsamples, these two portions are independent by construction under standard assumptions. The selection of the linear combination is a measurable function of the planning sample alone and is therefore independent of the analysis-sample outcomes. The test statistic in the analysis sample thus has the same limiting distribution conditional on the selected combination as it would if the combination were fixed in advance. Our novel characterization identifies the linear combinations for which this conditional asymptotic power equals that of full-sample procedures such as Scheffé projections. We will revise the characterization section to explicitly state this independence argument and display the relevant conditional limiting distribution. revision: yes

  2. Referee: [Simulation studies] Simulation design (results section): without explicit details on how the planning-sample selection is implemented and whether the finite-sample power gains persist when the dependence structure matches the observational setting, it is unclear whether the reported enhancements support the central claim or merely reflect an idealized independent-split scenario.

    Authors: We agree that greater transparency on the simulation implementation is warranted. In the revised manuscript we will supply explicit pseudocode describing how the planning-sample selection is performed and will add simulations that embed the full dependence structure implied by the observational design and sensitivity model. These additional results will confirm that the reported finite-sample power gains continue to hold under the relevant dependence. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: asymptotic power characterization derived independently of fitted inputs or self-citations

full rationale

The paper derives a novel characterization of linear combinations preserving asymptotic power under sample splitting, based on the observational design and sensitivity model. This is presented as a mathematical result evaluated via simulations, with no quoted reduction showing the characterization equals a fitted quantity or prior self-citation by construction. The approach is self-contained against external benchmarks like full-sample methods, and the reader's assessment of score 2.0 aligns with absence of load-bearing circular steps.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract alone supplies no explicit free parameters, axioms, or invented entities; the method appears to rest on standard asymptotic arguments for sample splitting and sensitivity models common to the field.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Powerful Multivariate Sensitivity Analysis via Sample Splitting in an Observational Study of the Effects of Poverty on Cardiovascular Disease Risk Factors." pith.science (2026). https://pith.science/paper/JEVCWE6Y

@misc{pith2026260604416,
  author       = {Pith},
  title        = {Pith review of: Powerful Multivariate Sensitivity Analysis via Sample Splitting in an Observational Study of the Effects of Poverty on Cardiovascular Disease Risk Factors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JEVCWE6Y}},
  note         = {Machine review of arXiv:2606.04416}
}
read the original abstract

When assessing the causal effect of an exposure on two or more outcomes in an observational study, a linear combination of outcomes may lessen the sensitivity of a test of the global null hypothesis to potential unmeasured biases. While all linear combinations of scored outcomes can be considered using Scheffe projections or constrained variants thereof, finding the combination that minimizes sensitivity to unmeasured biases requires corrections for multiple testing, which can erode power, especially when many outcomes are of interest. To mitigate this issue, we propose splitting the sample into a planning sample to identify an optimal linear combination and an analysis sample to conduct inference. We provide a novel characterization of the set of linear combinations for which this approach is guaranteed to achieve the same asymptotic power as full-sample alternatives and conduct extensive simulation studies that demonstrate enhanced power in finite samples. Finally, we apply our method to investigate the effects of poverty on the emergence of cardiovascular disease risk factors in children and adolescents. We discover adverse consequences on outcomes related to body composition, physical activity, and tobacco exposure. Although the impact of poverty on elevated tobacco exposure shows some robustness to unmeasured confounding, the other findings remain sensitive to potential biases.

Figures

Figures reproduced from arXiv: 2606.04416 by the authors.

Figure 1
Figure 1. Power comparisons between the method of Cohen et al. (2020) (dashed, purple), the sample splitting approach (solid, blue), and the oracle test (dotted, beige) as Γ increases with I = 300. The left column has K = 5; the center has K = 15; and the right has K = 25. The first row has ρ = 0 and the second row has ρ = 0.2. outperforms the full-sample comparator, it does so by a smaller margin compared to when the outcome… view at source ↗
Figure 2
Figure 2. Power comparisons between the method of Cohen et al. (2020) (dashed, purple), the sample splitting approach (solid, blue), and the oracle test (dotted, beige) as Γ increases with I = 1000. The left column has K = 5; the center has K = 15; and the right has K = 25. The first row has ρ = 0 and the second row has ρ = 0.2. Appendix C. Additional Simulations We emulate the setting described in Section 5 of the main text,… view at source ↗
Figure 3
Figure 3. Love plot displaying the absolute standardized mean dif￾ferences across covariates. D.3. Creating a robust nutrition composite. We constructed a harmonized Healthy Eating Index–2015 (HEI-2015) measure from NHANES 1999–2016 Day 1 24-hour dietary recalls. Individual food records were linked to their correspond￾ing food-pattern equivalency files (MPED for 1999–2004; FPED for 2005–2016) at the respondent–line level usin… view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 1.00 [PITH_FULL_IMAGE:figures/full_fig_p076_4.png]
Figure 5
Figure 5. Figure 5: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 1.05 [PITH_FULL_IMAGE:figures/full_fig_p077_5.png]
Figure 6
Figure 6. Figure 6: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 1.15 [PITH_FULL_IMAGE:figures/full_fig_p078_6.png]
Figure 7
Figure 7. Figure 7: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 1.25 [PITH_FULL_IMAGE:figures/full_fig_p079_7.png]
Figure 8
Figure 8. Figure 8: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 1.50 [PITH_FULL_IMAGE:figures/full_fig_p080_8.png]
Figure 9
Figure 9. Figure 9: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 1.75 [PITH_FULL_IMAGE:figures/full_fig_p081_9.png]
Figure 10
Figure 10. Figure 10: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 2.00 [PITH_FULL_IMAGE:figures/full_fig_p082_10.png]
Figure 11
Figure 11. Figure 11: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 2.25 [PITH_FULL_IMAGE:figures/full_fig_p083_11.png]
Figure 12
Figure 12. Figure 12: Rejected hypotheses and p-values for our method (top row) and the full sample comparator of Cohen et al. (2020) (bottom row) at Γ = 2.50 [PITH_FULL_IMAGE:figures/full_fig_p084_12.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

4 extracted references

  1. [1]

    K., Bullard, K

    Ali, M. K., Bullard, K. M., Beckles, G. L., Stevens, M. R., Barker, L., Venkat Narayan, K. and Imperatore, G. (2011), ‘Household income and cardio- vascular disease risks in US children and young adults: Analyses from NHANES 1999–2008’, Diabetes Care34(9), 1998–2004. Aliprantis, C. D. and Border, K. C. (2006), Infinite Dimensional Analysis: A Hitchhiker’s...

  2. [2]

    Fogarty, C. B. and Small, D. S. (2016), ‘Sensitivity analysis for multiple compar- isons in matched observational studies through quadratically constrained linear programming’, Journal of the American Statistical Association111(516), 1820–

  3. [3]

    L., Krieger, A

    Gastwirth, J. L., Krieger, A. M. and Rosenbaum, P. R. (2000), ‘Asymptotic sepa- rability in sensitivity analysis’, Journal of the Royal Statistical Society Series B 86 SENSITIVITY ANALYSIS FOR OBSER V ATIONAL STUDIES WITH MANY OUTCOMES 62(3), 545–555. Goeman, J. J. and Finos, L. (2012), ‘The inheritance procedure: Multiple testing of tree-structured hypot...

  4. [4]

    and Gabriel, K

    Marcus, R., Eric, P. and Gabriel, K. R. (1976), ‘On closed testing procedures with special reference to ordered analysis of variance’, Biometrika63(3), 655–660. Rosenbaum, P. R. (2004), ‘Design sensitivity in observational studies’, Biometrika 91(1), 153–164. Rosenbaum, P. R. (2007), ‘Sensitivity analysis form-estimates, tests, and confidence intervals in...

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.