Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

A causal framework for evaluating the total effect of strategies aiming to expand screening and to improve outcomes

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that the total effect of a cluster-level strategy to expand screening and improve outcomes can be estimated without bias by a two-stage targeted estimator, even though the focus population is only observed among those…

desk verdict A competent and useful Two-Stage TMLE for total effects in cluster trials with screening-mediated outcomes; the untestable MAR assumption needs a sensitivity analysis before the OPAL application. read the letter →

arxiv 2506.06267 v4 pith:F2I6W4LG submitted 2025-06-06 stat.ME

classification stat.ME
keywords counterfactualstrataeffectsclusterrandomizedtrialsgroupmeasurementmediationmissingdatascreeningtargetedminimumloss-basedestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Health strategies that aim both to expand screening and to improve outcomes create a problem: the people the strategy is meant to help—say, persons at risk of HIV—are only identifiable after they screen. Standard analyses condition on screening or on observed eligibility, which blocks the mediating effect of screening and biases the total effect toward the null. The paper defines the target as a Counterfactual Strata Effect, $\Psi^* = E[Y^{c*}(1) - Y^{c*}(0)]$, where $Y^{c*}(a)$ is the counterfactual probability of the outcome under arm $a$ among the underlying focus population. It then shows that, under a conditional missing-at-random assumption, each cluster's endpoint can be written as a ratio and estimated by Two-Stage TMLE. Simulations with 20 to 70 clusters show negligible bias and confidence-interval coverage near the nominal 95% rate, while comparator estimators remain biased.

What carries the argument

The carrier of the argument is the cluster-level endpoint $Y^c = P(Y_2=1)/E[E(Y_1 \mid \Delta=1, W)]$, a ratio of the observed outcome proportion to the adjusted prevalence of the underlying focus population. This ratio identity re-expresses the counterfactual probability $\Pr(Y_2=1 \mid Y_1^*=1)$ as a joint probability divided by a missing-data-adjusted denominator, converting the multilevel-mediation-missing data problem into a denominator-estimation problem within each cluster. Two-Stage TMLE is the machinery that operationalizes the identity: Stage 1 fully stratifies on cluster and uses TMLE with Super Learner to estimate the denominator, and Stage 2 applies cluster-level TMLE with Adaptive Pre-specification to select the covariate adjustment that maximizes empirical efficiency. The unadjusted estimator is always included as a candidate, so the method defaults to it when adjustment does not help.

What would settle it

The claim stands or falls on whether a validation substudy that screens everyone in a subset of clusters reproduces the Stage-1 denominator estimate $E[E(Y_1 \mid \Delta=1, W)]$; if the complete-screening prevalence differs materially, the conditional independence assumption is violated and the total-effect estimate is not trustworthy.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that the total effect of a cluster-randomized strategy to expand screening and improve outcomes is identifiable and estimable despite the multilevel-mediation-missing data problem. The causal estimand is a Counterfactual Strata Effect, $\Psi^* = E[Y^{c*}(1)-Y^{c*}(0)]$ with $Y^{c*}(a) = \Pr(Y_2(a)=1 \mid Y_1^*=1)$, so the outcome is relevant only for persons in an underlying focus population whose membership is missing for those who do not screen. Under the assumption that, within each cluster, underlying focus-population status is conditionally independent of screening given baseline covariates ($Y_1^* \perp \Delta \mid W$) and positivity holds, the cluster-level endpoint is identified as $Y^c = P(Y_2=1)/E[E(Y_1 \mid \Delta=1, W)]$. The paper proposes estimating $Y^c$ separately in each cluster by TMLE in Stage 1, then comparing arms with a cluster-level TMLE that uses Adaptive Pre-specification in Stage 2. The simulation evidence is presented as showing negligible bias and near-95% coverage for $J = 20, 30, 50, 70$ clusters, with the approach maintaining type-I error under the null.

Load-bearing premise

The load-bearing premise is that, after accounting for the measured baseline covariates, whether a person is screened does not depend on whether they actually belong to the focus population; if people who suspect they are at risk screen at different rates, the effect estimate is biased.

Editorial extensions

If this is right

  • If the paper is right, cluster randomized trials of screening-plus-outcome strategies can report unbiased total-effect estimates without requiring universal screening or conditioning on being measured.
  • The comparison shows that restricting analyses to screened participants or to observed eligible participants biases the effect toward the null, and in the eligible-only case can even reverse the sign.
  • The method operates with as few as 20 clusters, which is within the range of many real cluster randomized trials, so it can serve as a practical primary-analysis tool.
  • Because the numerator $P(Y_2=1)$ is directly observed, the approach extends to settings with complete outcome measurement and is not restricted to binary outcomes.
  • The framework also extends to time-varying covariates and to strategies that change the composition of the focus population, covering a more realistic version of the motivating OPAL trial.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the practical success of the estimator hinges on whether the covariate set $W$ captures the drivers of both screening and focus-population status; in the HIV venue example, venue-level characteristics such as alcohol type and availability of on-site rooms are therefore load-bearing, not optional.
  • My inference: the same ratio identity could be used as a diagnostic—comparing the adjusted denominator estimate $E[E(Y_1\mid \Delta=1,W)]$ with a complete-screening prevalence in a validation substudy would probe the missing-at-random assumption.
  • My inference: for trial design, the method creates an incentive to measure rich baseline covariates even when the primary analysis is unadjusted at the cluster level, because Stage 2's adaptive selection can only exploit covariates that were collected.
  • My inference: adapting the approach to observational group-level strategies would require adding cluster-level confounding adjustment in Stage 2, which the paper lists as future work but does not develop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a causal framework and an estimator for the total effect of a cluster-level strategy that aims to expand screening and improve outcomes, motivated by the OPAL trial of PrEP uptake in Kenya and Uganda. The causal estimand is a Counterfactual Strata Effect, Psi = E[Y^{c*}(1) - Y^{c*}(0)], with Y^{c*}(a) = Pr(Y2(a)=1 | Y1*=1), where Y1* is the underlying focus-population status. For identification, the authors stratify on cluster, express the cluster-level endpoint as the ratio P(Y2=1)/P(Y1*=1), and identify the denominator as E[E(Y1 | Delta=1, W)] under a conditional missing-at-random assumption and positivity. Stage 1 estimates each cluster's endpoint by the empirical outcome proportion divided by a TMLE of the denominator; Stage 2 applies a cluster-level TMLE with Adaptive Pre-specification to estimate the average effect. A simulation study with J=20,30,50,70 clusters reports negligible bias and near-nominal confidence interval coverage for the proposed estimator, while several comparator approaches are substantially biased.

Significance. If the identification assumptions hold, the paper makes a useful and timely contribution to the analysis of cluster randomized trials in which screening is both a mediator and a measurement mechanism. The derivation in Section 3.4 is clean and correct under the stated assumptions, and the paper clearly explains why standard approaches that condition on screening or on observed eligibility are biased for the total effect. The simulation study is well designed, includes null simulations demonstrating Type I error control, and the code is publicly available, which are notable strengths. The proposed Two-Stage TMLE is doubly robust within clusters and offers a practical template for the OPAL trial and similar settings.

major comments (3)
  1. [Section 3.4 (denominator identification)] The consistency of the entire estimator rests on the untestable assumption Y1* is independent of Delta given W within each cluster. The paper states this assumption (Section 3.1: 'there is no effect of underlying status Y1* on screening Delta') and says it is 'needed for identification' in Section 3.4, but it provides no sensitivity analysis or simulation under a violation. The simulation DGP in Section 4.1 enforces the assumption by generating U_Y1* and U_Delta independently. Because individuals who suspect they are at higher HIV risk may be more likely to attend screening, violations are plausible in the OPAL setting. I request a sensitivity analysis with an unmeasured confounder affecting both Y1* and Delta (e.g., with specified odds ratios), or partial-identification bounds, so that the practical impact of this assumption on the reported estimate can be assessed.
  2. [Section 5 (extension allowing exposure to affect focus population)] In the extension, the paper defines Y^{c*}(a) = Pr(Y2(a)=1 | Y1*(a)=1) while allowing A_c to affect Y1*. When Y1*(1) and Y1*(0) differ, the two conditional probabilities condition on different populations, so the contrast Psi is not an effect on a fixed focus population and may be driven by compositional changes in who belongs to the focus population. This is the standard problem of conditioning on a post-exposure variable. The distinction from Principal Strata Effects in Section 3.2 does not resolve the issue because the paper does not define the latent strata on which the extension's estimand is based. Please clarify the target of inference in the extension, or restrict the claim to the setting where A_c has no effect on Y1*.
  3. [Section 3.6 and Discussion (asymptotic linearity of the ratio estimator)] The paper asserts that the ratio-based Two-Stage TMLE is asymptotically linear and refers to Balzer et al. [26] for conditions, but it does not explicitly verify how the ratio form of the Stage 1 endpoint satisfies those conditions. The remainder term for a ratio involves products of errors in the numerator and denominator, and the stated condition J/min_j(N_j) -> 0 is not derived for this specific estimator. I ask the authors to state the exact theorem they invoke, verify its conditions for the ratio parameter, or explicitly label the efficiency and asymptotic-linearity claims as heuristic and supported by the simulation evidence rather than by a formal proof.
minor comments (5)
  1. [Table 1] The abbreviation 'Pt' in the column 'Pt (95% CI)' is not defined in the table caption; please define it or spell out 'average point estimate and 95% confidence interval'. Also, the header 'sigmaCoverage' is missing a space and should read 'sigma Coverage'.
  2. [Section 3.5, Equation (1)] The estimator is written as 'Psi_hat_unad j' with a cluster subscript j, but it is a single global estimator across all clusters; the subscript appears to be a typo and should be removed.
  3. [Section 4.2] In the description of the comparators, 'we used the MRStdCRT Rpackage' should be 'R package' for consistency; similar fix applies to the mention of 'ltmle'.
  4. [Section 4.1] In the data-generating process for Delta(a_c), the terms '-.4 a_c W3' and '+.4(1-a_c) W3' are correct but could be simplified for readability, as they combine to -0.4 W3 under a_c=1 and +0.4 W3 under a_c=0.
  5. [Abstract and Section 3.1] The abstract contains 'and/ or' with nonstandard spacing; this should be 'and/or'. More substantively, the exclusion restriction 'no effect of underlying status on screening' in Section 3.1 is not exactly the same as the statistically stated conditional independence Y1* perpendicular Delta | W; the paper should clarify that the MAR assumption also rules out unmeasured common causes of Y1* and Delta after conditioning on W.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the causal estimand is defined independently, the statistical estimand is derived from stated assumptions, and the simulations validate the estimator against independently generated ground truth.

full rationale

The paper's derivation chain is self-contained at the level where circularity is assessed. The causal estimand Ψ* is defined from counterfactual outcomes under each trial arm (Section 3.2), and the statistical estimand Y^c = P(Y2=1)/E[E(Y1|Δ=1,W)] is derived from it via the stated missing-at-random assumption Y1*⊥Δ|W and positivity (Section 3.4), rather than by plugging in the estimator itself. The numerator simplification P(Y2=1,Y1*=1)=P(Y2=1) follows from the structural equation Y1=Δ×Y1* and the determinism that Y2=1 implies Y1=1, so it is a definitional identity, not a fitted input. Stage 2 targets Ψ* using cluster-level TMLE with Adaptive Pre-specification, and the unadjusted estimator is included as a default, so no parameter is fitted to the target effect and then renamed a prediction. The simulation study generates counterfactual full data independently of the estimator and computes the true Ψ* from a separate population of 5000 clusters, making the reported bias and coverage results external validation rather than a tautology. The cited prior work [24-29] is used for the Counterfactual Strata Effects framework and for asymptotic regularity conditions; these are self-citations with author overlap, but they are not the sole support for the identification displayed in the paper, and no uniqueness theorem is invoked to forbid alternative approaches. The MAR assumption is untestable and clinically strong, but an untestable assumption is a correctness and robustness limitation, not circularity. No step reduces by construction to its own inputs, and no fitted value is renamed as a prediction.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central claim rests on standard causal identification assumptions: randomization, conditional independence of underlying status and screening given covariates (MAR), positivity, and structural assumptions that the outcome requires screening and focus-population membership. No new entities or fitted constants are introduced; simulation coefficients are inputs, not parameters of the estimator.

assumptions (6)
  • domain assumption Trial arm A^c is randomized and independent of potential outcomes and covariates (unconfoundedness).
    Section 3.1 states 'the trial arm A^c is completely randomized; this assumption can easily be relaxed to reflect covariate-dependent randomization schemes.'
  • domain assumption No effect of underlying status Y1* on screening Delta; equivalently Y1* is independent of Delta given W within clusters (conditional missing-at-random).
    Section 3.1 states 'there is no effect of underlying status Y1* on screening Delta; this assumption is needed for identification'; Section 3.4 formalizes Y1* independent of Delta given W.
  • standard math Positivity for screening: P(Delta=1|W=w)>0 for all possible w.
    Section 3.4 states 'we require that the probability of being measured within covariate values be bounded away from 0'.
  • domain assumption Outcome Y2 is only possible when Y1=1 and Delta=1; outcome is completely measured.
    Section 3 states 'having the outcome (Y2=1) is only possible if the individual participates in screening and they are in the population of interest (Y1=1)' and 'we assume complete measurement of the outcome'.
  • domain assumption No effect of trial arm A^c on underlying focus population Y1* in the main analysis.
    Section 3.1 states 'there is no impact of the trial arm A^c on underlying status Y1*'; this is relaxed in Section 5.
  • domain assumption Within-cluster dependence is weak enough and J/min_j(N_j) tends to 0 for the Central Limit Theorem to apply in asymptotic inference.
    Section 3.6 lists conditions for asymptotic linearity, quoted from Balzer et al. [26].

how reviews work

0 comments
Cite this review

Pith. "Pith review of A causal framework for evaluating the total effect of strategies aiming to expand screening and to improve outcomes." pith.science (2026). https://pith.science/paper/F2I6W4LG

@misc{pith2026250606267,
  author       = {Pith},
  title        = {Pith review of: A causal framework for evaluating the total effect of strategies aiming to expand screening and to improve outcomes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F2I6W4LG}},
  note         = {Machine review of arXiv:2506.06267}
}
read the original abstract

For many health conditions, there are highly efficacious treatment and prevention products. Maximizing their impact requires strategies that improve the reach of health screening in order to establish who could benefit. For example, HIV prevention strategies aim to expand risk screening and to improve uptake of pre-exposure prophylaxis (PrEP) among those experiencing risk. Often, these strategies induce changes at the group-level (e.g., health clinics or communities) and are evaluated through cluster randomized trials. This scenario creates a complex, multilevel-mediation-missing data problem for the following reasons. First, the strategy is delivered at the cluster-level, while health screening and outcomes are at the individual-level. Second, the strategy improves health outcomes directly and indirectly through improved health screening. Third, everyone has an underlying status, which is only observed among those screened. To formally define the total effect in such settings, we use Counterfactual Strata Effects: causal estimands where the outcome is only relevant for a group whose membership is subject to missingness and/ or impacted by the exposure of interest. To identify and estimate the corresponding statistical estimand, we propose a novel extension of Two-Stage targeted minimum loss-based estimation (TMLE). Simulations demonstrate the practical performance of our approach as well as the limitations of existing approaches.

Figures

Figures reproduced from arXiv: 2506.06267 by the authors.

Figure 1
Figure 1. Simplified causal graph for the OPAL trial to increase HIV risk screening and PrEP use among [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. A simplified hierarchical causal graph to illustrate the data generating process in a CRT where [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. A simplified hierarchical causal graph to illustrate the data generating process where the cluster [PITH_FULL_IMAGE:figures/full_fig_p018_3.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Causal Inference with Missing Exposures and Missing Outcomes

    stat.ME 2025-06 unverdicted novelty 5.0 of 10

    The paper introduces counterfactual strata effects to identify causal estimands under missing exposures, missing baseline outcomes, and missing follow-up outcomes, demonstrated on the SEARCH-TB study with TMLE and Sup...

Reference graph

Works this paper leans on

74 extracted references · 54 canonical work pages · cited by 1 Pith paper

  1. [26]

    Two-Stage TMLE to reduce bias and improve efficiency in cluster randomized trials.Biostatistics, 24(2):502–517, April 2023

    Laura B Balzer, Mark Van Der Laan, James Ayieko, Moses Kamya, Gabriel Chamie, Joshua Schwab, Diane V Havlir, and Maya L Petersen. Two-Stage TMLE to reduce bias and improve efficiency in cluster randomized trials.Biostatistics, 24(2):502–517, April 2023. ISSN 1465-4644, 1468-4357. doi: 10.1093/bi ostatistics/kxab043. URLhttps://academic.oup.com/biostatisti...

  2. [1]

    The Urgency of Now: AIDS at a Crossroads — global AIDS report, 2024

    UNAIDS. The Urgency of Now: AIDS at a Crossroads — global AIDS report, 2024. URLhttps: //www.unaids.org/en/resources/documents/2024/global-aids-update-2024

  3. [2]

    Creating hiv prevention cascades: Operational guidance on a tool for monitoring programmes,

    UNAIDS. Creating hiv prevention cascades: Operational guidance on a tool for monitoring programmes,

  4. [3]

    The HIV test-and-treat cascade – AIDS

    UNAIDS. The HIV test-and-treat cascade – AIDS. Technical report, 2020. URLhttps://aids2020.u naids.org/chapter/chapter-2-2020-commitments/the-hiv-test-and-treat-cascade/

  5. [4]

    AIDS, crisis and the power to transform: UNAIDS global AIDS update, 2025

    UNAIDS. AIDS, crisis and the power to transform: UNAIDS global AIDS update, 2025. URLhttps: //www.unaids.org/en/resources/documents/2025/2025-global-aids-update. Licence: CC BY-NC-SA 3.0 IGO

  6. [5]

    Maryam Shahmanesh, T. Nondumiso Mthiyane, Carina Herbsst, Melissa Neuman, Oluwafemi Adeagbo, Paul Mee, Natsayi Chimbindi, Theresa Smit, Nonhlanhla Okesola, Guy Harling, Nuala McGrath, Lor- raine Sherr, Janet Seeley, Hasina Subedar, Cheryl Johnson, Karin Hatzold, Fern Terris-Prestholt, Frances M. Cowan, and Elizabeth Lucy Corbett. Effect of peer-distribute...

  7. [6]

    Kakande, James Ayieko, Helen Sunday, Edith Biira, Marilyn Nyabuti, George Agengo, Jane Kabami, Colette Aoko, Hellen N

    Elijah R. Kakande, James Ayieko, Helen Sunday, Edith Biira, Marilyn Nyabuti, George Agengo, Jane Kabami, Colette Aoko, Hellen N. Atuhaire, Norton Sang, Asiphas Owaranganise, Janice Litunya, Er- ick W. Mugoma, Gabriel Chamie, James Peng, John Schrom, Melanie C. Bacon, Moses R. Kamya, Diane V. Havlir, Maya L. Petersen, Laura B. Balzer, and SEARCH Study Team...

  8. [7]

    Hickey, Asiphas Owaraganise, Sabina Ogachi, Norton Sang, Erick M

    Matthew D. Hickey, Asiphas Owaraganise, Sabina Ogachi, Norton Sang, Erick M. Wafula, Jane Kabami, Nicole Sutter, Jennifer Temple, Anthony Muiru, Gabriel Chamie, Elijah Kakande, Maya L. Petersen, Laura B. Balzer, Diane V. Havlir, Moses R. Kamya, and James Ayieko. Community health worker–facilitated telehealth for moderate–severe hypertension care in Kenya ...

Show all 74 references
  1. [8]

    Ortblad, Peter Mogere, Victor Omollo, Alexandra P

    Katrina F. Ortblad, Peter Mogere, Victor Omollo, Alexandra P. Kuo, Magdaline Asewe, Stephen Gakuo, Stephanie Roche, Mary Mugambi, Melissa Latigo Mugambi, Andy Stergachis, Josephine Odoyo, Eliz- abeth A. Bukusi, Kenneth Ngure, and Jared M. Baeten. Stand-alone model for delivery...

  2. [9]

    Jane Kabami, Laura B. Balzer, Mucunguzi Atukunda, Elizabeth Arinaitwe, Gerald Mutungi, Brian Twinamatsiko, Ronald Aine Mwesigye, Michael Ayebare, Alan Asiimwe, Cecilia Akatukwasa, Joan Nangendo, Starley B. Shade, Edwin D. Charlebois, Emmy Okello, Saidi Kapiga, Heiner Grosskurt...

  3. [10]

    Greenland, J

    S. Greenland, J. Pearl, and J. M. Robins. Causal diagrams for epidemiologic research.Epidemiology (Cambridge, Mass.), 10(1):37–48, January 1999. ISSN 1044-3983

  4. [11]

    Petersen, Sandra E

    Maya L. Petersen, Sandra E. Sinisi, and Mark J. van der Laan. Estimation of direct causal effects. Epidemiology (Cambridge, Mass.), 17(3):276–284, May 2006. ISSN 1044-3983. doi: 10.1097/01.ede.000 0208475.99429.2d

  5. [12]

    Philip Dawid, and Sara Geneletti

    Vanessa Didelez, A. Philip Dawid, and Sara Geneletti. Direct and indirect effects of sequential treat- ments. InProceedings of the Twenty-Second Conference on Uncertainty in Artificial Intelligence, UAI’06, pages 138–146, Arlington, Virginia, USA, July 2006. AUAI Press. ISBN 9...

  6. [13]

    MacKinnon, Amanda J

    David P. MacKinnon, Amanda J. Fairchild, and Matthew S. Fritz. Mediation Analysis.Annual review of psychology, 58:593, 2007. ISSN 0066-4308. doi: 10.1146/annurev.psych.58.110405.085542. URL https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2819368/

  7. [14]

    J. M. Robins and S. Greenland. Identifiability and exchangeability for direct and indirect effects. Epidemiology (Cambridge, Mass.), 3(2):143–155, March 1992. ISSN 1044-3983. doi: 10.1097/00001648 -199203000-00013

  8. [15]

    Direct and indirect effects

    Judea Pearl. Direct and indirect effects. InProceedings of the Seventeenth conference on Uncertainty in artificial intelligence, UAI’01, pages 411–420, San Francisco, CA, USA, August 2001. Morgan Kaufmann Publishers Inc. ISBN 978-1-55860-800-9. REFERENCES 23

  9. [16]

    Austin, Douglas S

    Peter C. Austin, Douglas S. Lee, and Jason P. Fine. Introduction to the Analysis of Survival Data in the Presence of Competing Risks.Circulation, 133(6):601–609, February 2016. ISSN 1524-4539. doi: 10.1161/CIRCULATIONAHA.115.017719

  10. [17]

    Young, Mats J

    Jessica G. Young, Mats J. Stensrud, Eric J. Tchetgen Tchetgen, and Miguel A. Hern´ an. A causal framework for classical statistical estimands in failure-time settings with competing events.Statistics in Medicine, 39(8):1199–1236, April 2020. ISSN 1097-0258. doi: 10.1002/sim.8471

  11. [18]

    Rudolph, Oleg Sofrygin, Wenjing Zheng, and Mark J

    Kara E. Rudolph, Oleg Sofrygin, Wenjing Zheng, and Mark J. van der Laan. Robust and Flexible Estimation of Stochastic Mediation Effects: A Proposed Method and Example in a Randomized Trial Setting.Epidemiologic methods, 7(1):20170007, 2018. ISSN 2194-9263. doi: 10.1515/em-2017...

  12. [19]

    VanderWeele and Eric J

    Tyler J. VanderWeele and Eric J. Tchetgen Tchetgen. Mediation analysis with time varying exposures and mediators.Journal of the Royal Statistical Society. Series B, Statistical Methodology, 79(3):917–938, June 2017. ISSN 1369-7412. doi: 10.1111/rssb.12194

  13. [20]

    Stensrud, Miguel A

    Mats J. Stensrud, Miguel A. Hern´ an, Eric J Tchetgen Tchetgen, James M. Robins, Vanessa Didelez, and Jessica G. Young. A generalized theory of separable effects in competing event settings.Lifetime Data Analysis, 27(4):588–631, 2021. ISSN 1380-7870. doi: 10.1007/s10985-021-09...

  14. [21]

    Kara E Rudolph, Dana E Goin, Diana Paksarian, Rebecca Crowder, Kathleen R Merikangas, and Eliz- abeth A Stuart. Causal Mediation Analysis With Observational Data: Considerations and Illustration Examining Mechanisms Linking Neighborhood Poverty to Adolescent Substance Use.Amer...

  15. [22]

    Cambridge University Press, Cambridge, 2 edition, 2009

    Judea Pearl.Causality. Cambridge University Press, Cambridge, 2 edition, 2009. ISBN 978-0-521- 89560-6. doi: 10.1017/CBO9780511803161. URLhttps://www.cambridge.org/core/books/causali ty/B0046844FAE10CBF274D4ACBDAEB5F5B

  16. [23]

    Laura B Balzer, Wenjing Zheng, Mark J van der Laan, and Maya L Petersen. A new approach to hierarchical data analysis: Targeted maximum likelihood estimation for the causal effect of a cluster- level exposure.Statistical methods in medical research, 28(6):1761–1780, June 2019....

  17. [24]

    Laura Balzer, Joshua Schwab, Mark van der Laan, and Maya Petersen. Evaluation of Progress Towards the UNAIDS 90-90-90 HIV Care Cascade: A Description of Statistical Methods Used in an Interim Analysis of the Intervention Communities in the SEARCH Study.U.C. Berkeley Division o...

  18. [25]

    Balzer, James Ayieko, Dalsone Kwarisiima, Gabriel Chamie, Edwin D

    Laura B. Balzer, James Ayieko, Dalsone Kwarisiima, Gabriel Chamie, Edwin D. Charlebois, Joshua Schwab, Mark J. van der Laan, Moses R. Kamya, Diane V. Havlir, and Maya L. Petersen. Far from MCAR: Obtaining population-level estimates of HIV viral suppression.Epidemiology (Cambri...

  19. [27]

    Blurring cluster randomized trials and observational studies: Two-Stage TMLE for subsampling, missingness, and few independent units.Biostatistics, page kxad015, August 2023

    Joshua R Nugent, Carina Marquez, Edwin D Charlebois, Rachel Abbott, and Laura B Balzer. Blurring cluster randomized trials and observational studies: Two-Stage TMLE for subsampling, missingness, and few independent units.Biostatistics, page kxad015, August 2023. ISSN 1465-4644...

  20. [28]

    PhD thesis, UC Berkeley, 2024

    Shalika Gupta.Mechanism and Mediation: Counterfactual Strata Effects for Perinatal Epidemiology. PhD thesis, UC Berkeley, 2024. URLhttps://escholarship.org/uc/item/5m34j2p6

  21. [29]

    The Causal Roadmap in the Age of AI: From All Wheel Drive to Formula 1, April 2024

    Maya Petersen. The Causal Roadmap in the Age of AI: From All Wheel Drive to Formula 1, April 2024

  22. [30]

    Noah Kiwanuka, Ali Ssetaala, Ismail Ssekandi, Annet Nalutaaya, Paul Kato Kitandwe, Julius Ssempiira, Bernard Ssentalo Bagaya, Apolo Balyegisawa, Pontiano Kaleebu, Judith Hahn, Christina Lindan, and Nelson Kaulukusi Sewankambo. Population attributable fraction of incident HIV i...

  23. [31]

    Nyabuti, Maya L

    Marilyn N. Nyabuti, Maya L. Petersen, Elizabeth A. Bukusi, Moses R. Kamya, Florence Mwangwa, Jane Kabami, Norton Sang, Edwin D. Charlebois, Laura B. Balzer, Joshua D. Schwab, Carol S. Camlin, REFERENCES 25 Douglas Black, Tamara D. Clark, Gabriel Chamie, Diane V. Havlir, and Ja...

  24. [32]

    Predicting harmful alcohol use preva- lence in Sub-Saharan Africa between 2015 and 2019: Evidence from population-based HIV impact assessment.PloS One, 19(10):e0301735, 2024

    Mtumbi Goma, Wingston Felix Ng’ambi, and Cosmas Zyambo. Predicting harmful alcohol use preva- lence in Sub-Saharan Africa between 2015 and 2019: Evidence from population-based HIV impact assessment.PloS One, 19(10):e0301735, 2024. ISSN 1932-6203. doi: 10.1371/journal.pone.0301735

  25. [33]

    Alcohol consumption and high risk sexual behaviour among female sex workers in Uganda.African journal of AIDS research: AJAR, 13(2):145–151, 2014

    Martin Mbonye, Rwamahe Rutakumwa, Helen Weiss, and Janet Seeley. Alcohol consumption and high risk sexual behaviour among female sex workers in Uganda.African journal of AIDS research: AJAR, 13(2):145–151, 2014. ISSN 1727-9445. doi: 10.2989/16085906.2014.927779

  26. [34]

    Watt, Laurie Abler, Donald Skinner, Seth C

    Jennifer Velloza, Melissa H. Watt, Laurie Abler, Donald Skinner, Seth C. Kalichman, Alexis C. Dennis, and Kathleen J. Sikkema. HIV-risk behaviors and social support among men and women attending alcohol-serving venues in South Africa: Implications for HIV prevention.AIDS and b...

  27. [35]

    Kalichman, Ofer Harel, Jacqueline Mthembu, Michael P

    Demetria Cain, Valerie Pare, Seth C. Kalichman, Ofer Harel, Jacqueline Mthembu, Michael P. Carey, Kate B. Carey, Vuyelwa Mehlomakulu, Leickness C. Simbayi, and Kelvin Mwaba. HIV risks associated with patronizing alcohol serving establishments in South African Townships, Cape T...

  28. [36]

    Nakato, Jaquiline A

    Janice Litunya, Brian Beesiga, Joy Z. Nakato, Jaquiline A. Agola, Kara Marson, Wafula E. Mugoma, Jennifer Temple, Carol S. Camlin, Starley B. Shade, Judith A. Hahn, Elijah Kakande, Jane Kabami, Maya L. Petersen, Diane V. Havlir, Moses R. Kamya, Laura B. Balzer, James Ayieko, a...

  29. [37]

    Nakato, Jaquiline A

    Brian Beesiga, Janice Litunya, Joy Z. Nakato, Jaquiline A. Agola, Kara Marson, Wafula E. Mugoma, Pius Agaba, Jennifer Temple, Carol S. Camlin, Starley B. Shade, Sarah E. Woolf-King, Judith A. Hahn, Elijah Kakande, Jane Kabami, Maya L. Petersen, Diane V. Havlir, Laura B. Balzer...

  30. [38]

    Petersen, Mark J

    Alejandra Benitez, Maya L. Petersen, Mark J. van der Laan, Nicole Santos, Elizabeth Butrick, Dilys Walker, Rakesh Ghosh, Phelgona Otieno, Peter Waiswa, and Laura B. Balzer. Defining and estimating effects in cluster randomized trials: A methods comparison.Statistics in Medicin...

  31. [39]

    Bukusi, Craig R

    Maya Petersen, Laura Balzer, Dalsone Kwarsiima, Norton Sang, Gabriel Chamie, James Ayieko, Jane Kabami, Asiphas Owaraganise, Teri Liegler, Florence Mwangwa, Kevin Kadede, Vivek Jain, Albert Plenty, Lillian Brown, Geoff Lavoy, Joshua Schwab, Douglas Black, Mark van der Laan, El...

  32. [40]

    Havlir, Laura B

    Diane V. Havlir, Laura B. Balzer, Edwin D. Charlebois, Tamara D. Clark, Dalsone Kwarisiima, James Ayieko, Jane Kabami, Norton Sang, Teri Liegler, Gabriel Chamie, Carol S. Camlin, Vivek Jain, Kevin Kadede, Mucunguzi Atukunda, Theodore Ruel, Starley B. Shade, Emmanuel Ssemmondo,...

  33. [41]

    Incident Tuberculosis Infection Is Associ- ated With Alcohol Use in Adults in Rural Uganda.Clinical Infectious Diseases, page ciae304, June

    Rachel Abbott, Kirsten Landsiedel, Mucunguzi Atukunda, Sarah B Puryear, Gabriel Chamie, Judith A Hahn, Florence Mwangwa, Elijah Kakande, Maya L Petersen, Diane V Havlir, Edwin Charlebois, Laura B Balzer, Moses R Kamya, and Carina Marquez. Incident Tuberculosis Infection Is Ass...

  34. [42]

    Carina Marquez, Mucunguzi Atukunda, Joshua Nugent, Edwin D Charlebois, Gabriel Chamie, Florence Mwangwa, Emmanuel Ssemmondo, Joel Kironde, Jane Kabami, Asiphas Owaraganise, Elijah Kakande, REFERENCES 27 Bob Ssekaynzi, Rachel Abbott, James Ayieko, Theodore Ruel, Dalsone Kwariis...

  35. [43]

    Leonardo Grilli and Fabrizia Mealli. Nonparametric Bounds on the Causal Effect of University Studies on Job Opportunities Using Principal Stratification.Journal of Educational and Behavioral Statistics, 33(1):111–130, March 2008. ISSN 1076-9986. doi: 10.3102/1076998607302627. ...

  36. [44]

    Frangakis and Donald B

    Constantine E. Frangakis and Donald B. Rubin. Principal Stratification in Causal Inference.Biometrics, 58(1):21–29, March 2002. ISSN 0006-341X. URLhttps://www.ncbi.nlm.nih.gov/pmc/articles/PM C4137767/

  37. [45]

    Estimation of separable direct and indirect effects in continuous time.Biometrics, 79(1):127–139, March 2023

    Torben Martinussen and Mats Julius Stensrud. Estimation of separable direct and indirect effects in continuous time.Biometrics, 79(1):127–139, March 2023. ISSN 1541-0420. doi: 10.1111/biom.13559

  38. [46]

    Young, P ˚ al C

    Matias Janvin, Jessica G. Young, P ˚ al C. Ryalen, and Mats J. Stensrud. Causal inference with recurrent and competing events.Lifetime Data Analysis, 30(1):59–118, January 2024. ISSN 1572-9249. doi: 10.1007/s10985-023-09594-8. URLhttps://doi.org/10.1007/s10985-023-09594-8

  39. [47]

    DONALD B. RUBIN. Inference and missing data.Biometrika, 63(3):581–592, December 1976. ISSN 0006-3444. doi: 10.1093/biomet/63.3.581. URLhttps://doi.org/10.1093/biomet/63.3.581

  40. [48]

    James Robins. A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect.Mathematical Modelling, 7(9): 1393–1512, January 1986. ISSN 0270-0255. doi: 10.1016/0270- 0255(86)90088-6. URLhtt...

  41. [49]

    Van Der Laan and Sherri Rose.Targeted Learning: Causal Inference for Observational and Experimental Data

    Mark J. Van Der Laan and Sherri Rose.Targeted Learning: Causal Inference for Observational and Experimental Data. Springer Series in Statistics. Springer, New York, NY, 2011. ISBN 978-1-4419-9781- 4 978-1-4419-9782-1. doi: 10.1007/978-1-4419-9782-1. URLhttps://link.springer.co...

  42. [50]

    van der Laan, Eric C

    Mark J. van der Laan, Eric C. Polley, and Alan E. Hubbard. Super learner.Statistical Applications in Genetics and Molecular Biology, 6:Article25, 2007. ISSN 1544-6115. doi: 10.2202/1544-6115.1309. REFERENCES 28

  43. [51]

    Gail, Steven D

    Mitchell H. Gail, Steven D. Mark, Raymond J. Carroll, Sylvan B. Green, and David Pee. Design considerations for studies of intervention effects on recurrence.Statistics in Medicine, 15:123–135, 1996

  44. [52]

    Tsiatis, Marie Davidian, Min Zhang, and Xiaomin Lu

    Anastasios A. Tsiatis, Marie Davidian, Min Zhang, and Xiaomin Lu. Covariate adjustment for two- sample treatment comparisons in randomized clinical trials: A principled yet flexible approach.Statistics in Medicine, 27(23):4658–4677, Oct 2008. ISSN 0277-6715. doi: 10.1002/sim.3...

  45. [53]

    R. A. Fisher.Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh, 4th, revised and enlarged edition, 1932. URLhttps://archive.org/details/statisticalmetho00fish. Biological Monographs and Manuals

  46. [54]

    K. L. Moore and M. J. van der Laan. Covariate adjustment in randomized trials with binary outcomes: Targeted maximum likelihood estimation.Statistics in medicine, 28(1):39–64, January 2009. ISSN 0277-

  47. [55]

    van der Laan

    Michael Rosenblum and Mark J. van der Laan. Simple, Efficient Estimators of Treatment Effects in Randomized Trials Using Generalized Linear Models to Leverage Baseline Variables.The International Journal of Biostatistics, 6(1):13, April 2010. ISSN 1557-4679. doi: 10.2202/1557-...

  48. [56]

    Balzer, Mark J

    Laura B. Balzer, Mark J. van der Laan, and Maya L. Petersen. Adaptive Pre-specification in Randomized Trials With and Without Pair-Matching.Statistics in medicine, 35(25):4528–4545, November 2016. ISSN 0277-6715. doi: 10.1002/sim.7023. URLhttps://www.ncbi.nlm.nih.gov/pmc/artic...

  49. [57]

    Adaptive selection of the optimal strategy to improve precision and power in randomized trials.Biometrics, 80(1):ujad034, March 2024

    Laura B Balzer, Erica Cai, Lucas Godoy Garraza, and Pracheta Amaranath. Adaptive selection of the optimal strategy to improve precision and power in randomized trials.Biometrics, 80(1):ujad034, March 2024. ISSN 0006-341X. doi: 10.1093/biomtc/ujad034. URLhttps://doi.org/10.1093...

  50. [58]

    Balzer, Mark J

    Laura B. Balzer, Mark J. van der Laan, and Maya L. Petersen. Machine learning to optimize precision in the analysis of randomized trials: A journey in pre-specified, yet data-adaptive learning.Clinical Trials, (In Press), 2026. URLhttps://arxiv.org/abs/2512.13610

  51. [59]

    Rubin and Mark J

    Daniel B. Rubin and Mark J. van der Laan. Empirical efficiency maximization: improved locally efficient covariate adjustment in randomized experiments and survival analysis.The International Journal of Biostatistics, 4(1):Article 5, 2008. ISSN 1557-4679. REFERENCES 29

  52. [60]

    Stitelman and Mark J

    Ori M. Stitelman and Mark J. van der Laan. Collaborative targeted maximum likelihood for time to event data.The International Journal of Biostatistics, 6(1):Article 21, 2010. ISSN 1557-4679. doi: 10.2202/1557-4679.1249

  53. [61]

    A. W. Van Der Vaart.Asymptotic Statistics. Cambridge University Press, 1 edition, October 1998. ISBN 978-0-511-80225-6 978-0-521-49603-2 978-0-521-78450-4. doi: 10.1017/CBO9780511802256. URL https://www.cambridge.org/core/product/identifier/9780511802256/type/book

  54. [62]

    Hayes and Lawrence H

    Richard J. Hayes and Lawrence H. Moulton.Cluster Randomised Trials. Chapman and Hall/CRC, New York, January 2009. ISBN 978-0-429-14205-5. doi: 10.1201/9781584888178

  55. [63]

    Fiero, Shuang Huang, Eyal Oren, and Melanie L

    Mallorie H. Fiero, Shuang Huang, Eyal Oren, and Melanie L. Bell. Statistical analysis and handling of missing data in cluster randomized trials: a systematic review.Trials, 17:72, February 2016. ISSN 1745-6215. doi: 10.1186/s13063-016-1201-z. URLhttps://pmc.ncbi.nlm.nih.gov/ar...

  56. [64]

    Schnitzer, Mark J

    Mireille E. Schnitzer, Mark J. van der Laan, Erica E. M. Moodie, and Robert W. Platt. EFFECT OF BREASTFEEDING ON GASTROINTESTINAL INFECTION IN INF ANTS: A TARGETED MAX- IMUM LIKELIHOOD APPROACH FOR CLUSTERED LONGITUDINAL DATA.The Annals of Applied Statistics, 8(2):703–725, Jun...

  57. [65]

    KUNG-YEE LIANG and SCOTT L. ZEGER. Longitudinal data analysis using generalized linear models.Biometrika, 73(1):13–22, April 1986. ISSN 0006-3444. doi: 10.1093/biomet/73.1.13. URL https://doi.org/10.1093/biomet/73.1.13

  58. [66]

    Kahan, and Bingkai Wang

    Fan Li, Jiaqi Tong, Xi Fang, Chao Cheng, Brennan C. Kahan, and Bingkai Wang. Model-Robust Standardization in Cluster-Randomized Trials.Statistics in Medicine, 44(20-22):e70270, 2025. ISSN 1097-0258. doi: 10.1002/sim.70270. URLhttps://onlinelibrary.wiley.com/doi/abs/10.1002/si ...

  59. [67]

    ltmle: Longi- tudinal Targeted Maximum Likelihood Estimation, April 2023

    Joshua Schwab, Samuel Lendle, Maya Petersen, Mark van der Laan, and Susan Gruber. ltmle: Longi- tudinal Targeted Maximum Likelihood Estimation, April 2023. URLhttps://cran.r-project.org /web/packages/ltmle/index.html

  60. [68]

    Joy Z Nakato and Laura B. Balzer. Adaptive pooling to minimize bias and maximize power in cluster- randomized trials, 2025. Presentation at the Soceity for Epidemiological Research (SER) Conference 2025, Boston, MA. REFERENCES 30

  61. [69]

    van der Vaart.Asymptotic Statistics

    Aad W. van der Vaart.Asymptotic Statistics. Cambridge University Press, Cambridge, 1998

  62. [70]

    Screened

    Anastasios A. Tsiatis.Semiparametric Theory and Missing Data. Springer, New York, 2006. 31 Appendix A: The Screened and Eligible estimators for Stage 1 Recall that our Stage 1 causal estimand is the counterfactual probability of the outcome among the underlying target populati...

  63. [2021]

    Accessed 2025-01-23

    URLhttps://www.unaids.org/sites/default/files/media_asset/JC3038_creating-hiv -prevention-cascades_en.pdf. Accessed 2025-01-23

  64. [2023]

    doi: 10.1002/sim.9813

    ISSN 1097-0258. doi: 10.1002/sim.9813. URLhttps://onlinelibrary.wiley.com/doi/abs/ 10.1002/sim.9813. eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/sim.9813

  65. [2024]

    doi: 10.1093/cid/ciae304

    ISSN 1058-4838. doi: 10.1093/cid/ciae304. URLhttps://doi.org/10.1093/cid/ciae304

  66. [6715]

    URLhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC2857590/

    doi: 10.1002/sim.3445. URLhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC2857590/

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.