REVIEW 4 major objections 5 minor 33 references
Jointly modeling multiple endpoints for efficient treatment effect estimation in randomized controlled trials
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A one-factor structural equation model over primary and secondary endpoints can estimate a trial's primary treatment effect at least as efficiently as the difference in means when the model is correct, and model averaging keeps it…
desk verdict Solid, limited-scope efficiency method; headline application overstates robustness of the 27% gain. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the one-factor structural equation model of Equation (2), in which a single latent factor $\eta_i \mid A_i \sim N(\gamma A_i, 1)$ drives all endpoint means and correlations: given $\eta_i$, the endpoints are independent, and the treatment effect on endpoint $p$ is $\gamma\lambda_p$. This proportionality constraint is what turns secondary endpoints into information about the primary endpoint: the MLE $\hat\tau_{\mathrm{SEM}} = \hat\gamma\hat\lambda_1$ can be more precise than the saturated difference in means because it also uses the secondary mean differences and the covariance structure. Guarding against the constraint being wrong is the second piece of machinery: model averaging with BIC weights (and with Super Learner cross-validated weights) between the SEM and the saturated estimator, which is consistent in both regimes and achieves the SEM variance when the SEM holds.
What would settle it
Run the paper's Simulation B misspecification scenario with a null primary effect, nonzero secondary effects, and endpoint correlation incompatible with the one-factor SEM; record the Super Learner estimator's coverage and its difference from the saturated estimate over 1000 replicates. If coverage falls below nominal whenever the SEM is misspecified, the robustness claim is falsified; if the estimator remains close to the saturated estimate, the claimed bias protection holds.
Extended reading notes
Core claim
The central claim is that a correctly specified one-factor structural equation model (SEM) can estimate the average treatment effect on the primary endpoint $\tau_1 = E(Y_{i,1} \mid A_i=1) - E(Y_{i,1} \mid A_i=0)$ more efficiently than the saturated difference-in-means estimator by borrowing strength from secondary endpoints. Under the model, endpoints are Gaussian given a latent variable $\eta_i \sim N(\gamma A_i, 1)$, so $E(Y_{i,p} \mid A_i=1)-E(Y_{i,p} \mid A_i=0)=\gamma\lambda_p$; the proportionality of the endpoint effects through the loadings $\lambda_p$ is what lets the secondary endpoints inform $\tau_1$. Theorem 1 establishes $A\mathrm{Var}(\hat\tau_{\mathrm{SEM}}) \le A\mathrm{Var}(\hat\tau_{\mathrm{Saturated}})$ when the SEM holds, and Theorem 2 shows the BIC- and Super-Learner-averaged estimators are consistent and attain the SEM variance under the SEM. Simulations show efficiency gains translate into power gains, with the Super Learner estimator providing the best bias-efficiency trade-off when the SEM is misspecified; the application to the VLNC trial estimates a 10 percentage-point increase in abstinence among Black PWS with a standard error about 27 percent smaller than the saturated estimate.
Load-bearing premise
The load-bearing premise is that one latent factor explains all of the treatment's effect on every endpoint, so the effect on each endpoint is proportional to the same loading and endpoints are uncorrelated once that factor is conditioned on; when this is false, the SEM estimate is biased and model averaging only partly corrects the bias.
Editorial extensions
If this is right
- When the one-factor SEM holds, trials can gain power on a scientifically meaningful primary endpoint without enlarging the sample, because the same observations now inform the primary endpoint through correlations with secondary endpoints.
- The BIC- and Super-Learner-weighted estimators inherit the SEM's variance when the SEM is correct and collapse to the unbiased saturated estimator when it is not, so the primary estimand is unchanged throughout.
- Because the method targets the primary endpoint's average treatment effect directly, it avoids the interpretability cost of composite endpoints and win ratios.
- In the VLNC application, the estimated standard error reduction of about 27 percent corresponds to an effective sample size roughly twice the observed one for the abstinence endpoint among Black PWS.
Reading between the lines
- A natural extension is to average over several candidate SEMs, each using a different subset of secondary endpoints, with Super Learner weights; the paper notes this as a possibility, and it would make endpoint selection data-adaptive.
- The same borrowing logic could be applied within pre-specified subgroups of a randomized trial, which would directly address underpowered subgroup analyses; the motivating example already does this implicitly for Black PWS.
- Because the gain depends on the proportionality of endpoint effects, an investigator could screen candidate secondary endpoints using historical trial data and a likelihood-ratio test of the SEM restriction before committing to them in a new trial.
- Targeting estimands like risk ratios or odds ratios is feasible in principle from the fitted primary endpoint model, but the paper notes such functionals are more sensitive to the distributional assumptions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a method for estimating the treatment effect on a primary endpoint in randomized controlled trials by jointly modeling primary and secondary endpoints. The authors introduce a one-factor structural equation model (SEM) that imposes a mean--covariance relationship across endpoints, and a saturated model; they then combine the two estimators via model averaging based on BIC and Super Learner. The main theoretical results are Theorem 1, stating that the SEM estimator is asymptotically no less efficient than the saturated difference-in-means estimator when the SEM is correctly specified, and Theorem 2, stating that the model-averaged estimators are consistent and achieve the SEM variance when the SEM holds. The method is illustrated by estimating the effect of very low nicotine content cigarettes on abstinence among Black people who smoke, where the authors report a 27% reduction in standard error relative to the saturated estimator.
Significance. If the claims are appropriately qualified, this is a useful contribution to the literature on improving efficiency in randomized trials by leveraging secondary endpoints. The manuscript provides detailed M-estimation proofs for the efficiency result, comprehensive simulation studies covering Gaussian and binary endpoints, and an honest discussion of the assumptions underlying the SEM. The application to a real tobacco regulatory science trial is relevant and demonstrates the potential practical value of the approach. The paper also clearly positions itself relative to existing methods such as composite endpoints and win ratios. However, the central 'robustness' claim is overstated relative to the finite-sample evidence, and the proof of Theorem 1 has a gap at the null treatment effect. With revisions to qualify the claims and complete the proof, the method would be a solid contribution.
major comments (4)
- [Abstract, Section 5, Theorem 2] The abstract states the estimator 'is robust to model misspecification via model averaging,' but this claim is not supported by the paper's own finite-sample results. Theorem 2 is asymptotic: under misspecification, the BIC and Super Learner weights converge to the saturated model only as n grows. Simulations B and C (Figures 3 and 5) show that at n = 250, both tau_BIC and tau_SL are biased with degraded coverage when the SEM is misspecified. In the application (Section 4), n = 99 and the estimated weights are omega_BIC = 0.983 and omega_SL = 0.939, so the estimators are effectively the SEM rather than the robust fallback. This is a load-bearing mismatch between the abstract's claim and the actual guarantees; the robustness claim should be qualified as asymptotic, or the abstract should be revised.
- [Section 4, Figure 6] The reported 27% standard error reduction is conditional on the SEM being correctly specified, but the paper's own simulations show that the SEM-based estimators can be substantially biased under misspecification. The point estimate moves from 13 percentage points (saturated) to 10 percentage points (all secondary-endpoint estimators), a 23% relative shift that is consistent with the bias mechanism documented in Simulation B rather than with pure variance reduction. Given the small subgroup sample (n = 99), this shift should be explicitly discussed as a sensitivity concern, and the efficiency gain should not be presented without acknowledging the potential for bias.
- [Appendix C, proof of Theorem 1] The proof of Theorem 1 uses the reparameterization theta_0 = gamma^{-2}, which is not defined when gamma = 0, i.e., under the null treatment effect on all endpoints. Since Theorem 1 is stated for all correctly specified SEMs, including the global null, the proof is incomplete for the null case. The authors should either handle gamma = 0 separately or use a parameterization (e.g., based on lambda and the covariance diag(theta) + lambda lambda^T) that is regular at gamma = 0.
- [Appendix D, proof of Theorem 2 for tau_SL] The proof for tau_SL asserts that, when the SEM is correctly specified, 'it must be the case that omega_SL -> 1' to match the oracle estimator's performance, but this implication is not fully justified for the target parameter tau_1, which is a contrast of conditional means. The Super Learner oracle property is stated for the prediction loss for E(Y_1|A), and the argument that this transfers to the treatment effect contrast is only sketched. Please provide a formal proof or a precise reference for the transfer of the oracle property to the contrast in this setting.
minor comments (5)
- [Section 3.1, first paragraph] The sentence 'which in turn manipulates both Cov(Yi,1, Yi,3|Ai) and Cov(Yi,2, Yi,3|Ai) as well' should read 'determines' rather than 'manipulates'.
- [Section 3.2 and Figure captions] In the Figure 3, 4, and 5 captions, 'structural equation' should be 'structural equation model' (e.g., 'the structural equation is correctly specified').
- [Equation (7)] The display of equation (7) is malformed; the expression 'arg min_{omega in [0,1]} nP i=1 ...' should be typeset correctly, and the subscript 'nP' should be 'sum_{i=1}^n'.
- [Section 4] The text reports '20,000 bootstrapped iterations performed for each imputed dataset'; consider using 'bootstrap resamples' or 'bootstrap replicates' for clarity.
- [General] No code or data availability statement is provided. For a methodological paper, making simulation and application code available would improve reproducibility and reader confidence.
Circularity Check
No material circularity: efficiency claims follow from the stated SEM assumptions, and the self-citations are contextual rather than load-bearing.
full rationale
After walking the derivation chain, I find no step in which a claimed result is equivalent to its input by construction. Theorem 1 compares the SEM MLE, derived from the explicit one-factor model in Equation (2), with the saturated difference-in-means estimator; the Appendix C proof computes both asymptotic variances from estimating equations and shows a positive-semidefinite difference under the stated SEM assumption. This is a model-based efficiency comparison, not a definitional identity. Theorem 2 relies on the BIC's model-selection consistency and the Super Learner oracle property cited to van der Laan et al. (2007); the model-averaged estimator in Equation (5) is a convex combination, so its consistency under either model follows from the weight convergence rather than being assumed. The application's 27% standard-error reduction is an in-sample estimate of precision from the fitted SEM, not a prediction of a quantity that was itself used as a fitted input. The self-citations (Wolf et al. 2024a, 2024b) are motivational and contextual; neither enters the proofs nor is used to forbid alternative estimators. The paper's own Section 5 limitation statement, 'The consistency and efficiency gains of our initial estimator are only guaranteed when the one-factor SEM is correctly specified,' is an honest statement of scope. The finite-sample concern that, in the n=99 application, the BIC and Super Learner weights put 98.3% and 93.9% mass on the SEM, so asymptotic robustness may not yet operate, is a performance limitation documented in the paper's own simulations, not circular reasoning.
Assumptions & free parameters
assumptions (4)
- ad hoc to paper One-factor SEM: Y_i|η_i ~ N(ν+λη_i, diag(θ)), η_i|A_i ~ N(γA_i, 1), so the treatment effect on each endpoint is γλ_p and residual covariances are zero after conditioning on η.
- domain assumption Multivariate Gaussian endpoints for the theoretical results.
- domain assumption BIC and Super Learner weights select the correct model asymptotically (oracle property).
- domain assumption Missing data in the application are handled by multiple imputation, which requires missing-at-random assumptions.
invented entities (1)
-
Latent variable η
Cite this review
Pith. "Pith review of Jointly modeling multiple endpoints for efficient treatment effect estimation in randomized controlled trials." pith.science (2026). https://pith.science/paper/NDWRFAJG
@misc{pith2026250603393,
author = {Pith},
title = {Pith review of: Jointly modeling multiple endpoints for efficient treatment effect estimation in randomized controlled trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/NDWRFAJG}},
note = {Machine review of arXiv:2506.03393}
}
read the original abstract
Randomized controlled trials are the gold standard for evaluating the efficacy of an intervention. However, there is often a trade-off between selecting the most scientifically relevant primary endpoint versus a less relevant, but more powerful, endpoint. For example, in the context of tobacco regulatory science many trials evaluate cigarettes per day as the primary endpoint instead of abstinence from smoking due to limited power. Additionally, it is often of interest to consider subgroup analyses to answer additional questions; such analyses are rarely adequately powered. In practice, trials often collect multiple endpoints. Heuristically, if multiple endpoints demonstrate a similar treatment effect we would be more confident in the results of this trial. However, there is limited research on leveraging information from secondary endpoints besides using composite endpoints which can be difficult to interpret. In this paper, we develop an estimator for the treatment effect on the primary endpoint based on a joint model for primary and secondary efficacy endpoints. This estimator gains efficiency over the standard treatment effect estimator when the model is correctly specified but is robust to model misspecification via model averaging. We illustrate our approach by estimating the effect of very low nicotine content cigarettes on the proportion of Black people who smoke who achieve abstinence and find our approach reduces the standard error by 27%.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
A neutralizing monoclonal antibody for hospitalized patients with COVID -19
ACTIV-3/TICO LY-CoV555 Study Group (2021). A neutralizing monoclonal antibody for hospitalized patients with COVID -19. New England Journal of Medicine , 384(10):905--914
work page 2021
-
[2]
Bakal, J. A., Roe, M. T., Ohman, E. M., Goodman, S. G., Fox, K. A., Zheng, Y., Westerhout, C. M., Hochman, J. S., Lokhnygina, Y., Brown, E. B., and Armstrong, P. W. (2015). Applying novel methods to assess clinical outcomes: Insights from the TRILOGY ACS trial. European Heart Journal , 36(6):385--392
work page 2015
-
[3]
Beran, T. N. and Violato, C. (2010). Structural equation modeling in medical research: A primer. BMC Research Notes , 3(1):267
work page 2010
-
[4]
Carroll, R. J. and Ruppert, D. (1982). A comparison between maximum likelihood and generalized least squares in a heteroscedastic linear model. Journal of the American Statistical Association , 77(380):878--882
work page 1982
-
[5]
Chen, C., Han, P., and He, F. (2022). Improving main analysis by borrowing information from auxiliary data. Statistics in Medicine , 41(3):567--579
work page 2022
-
[6]
Chen, C., Wang, M., and Chen, S. (2023). An efficient data integration scheme for synthesizing information from multiple secondary datasets for the parameter inference of the main analysis. Biometrics , 79(4):2947--2960
work page 2023
-
[7]
Donny, E. C., Denlinger, R. L., Tidey, J. W., Koopmeiners, J. S., Benowitz, N. L., Vandrey, R. G., al’Absi, M., Carmella, S. G., Cinciripini, P. M., Dermody, S. S., Drobes, D. J., Hecht, S. S., Jensen, J., Lane, T., Le, C. T., McClernon, F. J., Montoya, I. D., Murphy, S. E., Robinson, J. D., Stitzer, M. L., Strasser, A. A., Tindle, H., and Hatsukami, D. K...
work page 2015
-
[8]
W., Heels-Ansdell, D., Montori, V
Ferreira-González, I., Permanyer-Miralda, G., Domingo-Salvany, A., Busse, J. W., Heels-Ansdell, D., Montori, V. M., Akl, E. A., Bryant, D. M., Alonso-Coello, P., Alonso, J., Worster, A., Upadhye, S., Jaeschke, R., Schünemann, H. J., Pacheco-Huergo, V., Wu, P., Mills, E. J., and Guyatt, G. H. (2007). Problems with use of composite end points in cardiovascu...
work page 2007
Show all 33 references
-
[9]
Freemantle, N., Calvert, M., Wood, J., Eastaugh, J., and Griffin, C. (2003). Composite outcomes in randomized trials: Greater precision but with greater uncertainty? JAMA , 289(19):2554--2559
2003
-
[10]
and Pocock, S
Frison, L. and Pocock, S. J. (1992). Repeated measures in clinical trials: Analysis using mean summary statistics and its implications for design. Statistics in Medicine , 11(13):1685--1704
1992
-
[11]
K., Jensen, J
Hatsukami, D. K., Jensen, J. A., Carroll, D. M., Luo, X., Strayer, L. G., Cao, Q., Hecht, S. S., Murphy, S. E., Carmella, S. G., Denlinger-Apte, R. L., Colby, S., Strasser, A. A., McClernon, F. J., Tidey, J., Benowitz, N. L., and Donny, E. C. (2024). Reduced nicotine in cigare...
2024
-
[12]
K., Luo, X., Jensen, J
Hatsukami, D. K., Luo, X., Jensen, J. A., al’Absi, M., Allen, S. S., Carmella, S. G., Chen, M., Cinciripini, P. M., Denlinger-Apte, R., Drobes, D. J., Koopmeiners, J. S., Lane, T., Le, C. T., Leischow, S., Luo, K., McClernon, F. J., Murphy, S. E., Paiano, V., Robinson, J. D., ...
2018
-
[13]
Hjort, N. L. and Claeskens, G. (2003). Frequentist model average estimators. Journal of the American Statistical Association , 98(464):879--899
2003
-
[14]
P., Carlin, B
Hobbs, B. P., Carlin, B. P., Mandrekar, S. J., and Sargent, D. J. (2011). Hierarchical commensurate and power prior models for adaptive incorporation of historical information in clinical trials. Biometrics , 67(3):1047--1056
2011
-
[15]
Ibrahim, J. G. and Chen, M.-H. (2000). Power prior distributions for regression models. Statistical Science , 15(1):46--60
2000
-
[16]
Jöreskog, K. G. (1970). A general method for estimating a linear structural equation system. ETS Research Bulletin Series , 1970(2):i--41
1970
-
[17]
Jöreskog, K. G. and Goldberger, A. S. (1975). Estimation of a model with multiple indicators and multiple causes of a single latent variable. Journal of the American Statistical Association , 70(351):631--639
1975
-
[18]
M., Koopmeiners, J
Kaizer, A. M., Koopmeiners, J. S., and Hobbs, B. P. (2018). Bayesian hierarchical modeling based on multisource exchangeability. Biostatistics , 19(2):169--184
2018
-
[19]
and Asparouhov, T
Muthén, B. and Asparouhov, T. (2012). Bayesian structural equation modeling: A more flexible representation of substantive theory. Psychological Methods , 17:313--335
2012
-
[20]
J., Ariti, C
Pocock, S. J., Ariti, C. A., Collier, T. J., and Wang, D. (2012). The win ratio: A new approach to the analysis of composite endpoints in clinical trials based on clinical priorities. European Heart Journal , 33(2):176--182
2012
-
[21]
J., Assmann, S
Pocock, S. J., Assmann, S. E., Enos, L. E., and Kasten, L. E. (2002). Subgroup analysis, covariate adjustment and baseline comparisons in clinical trial reporting: Current practice and problems. Statistics in Medicine , 21(19):2917--2930
2002
-
[22]
J., Ferreira, J
Pocock, S. J., Ferreira, J. P., Collier, T. J., Angermann, C. E., Biegus, J., Collins, S. P., Kosiborod, M., Nassif, M. E., Ponikowski, P., Psotka, M. A., Teerlink, J. R., Tromp, J., Gregson, J., Blatchford, J. P., Zeller, C., and Voors, A. A. (2023). The win ratio method in h...
2023
-
[23]
Pocock, S. J. and Simon, R. (1975). Sequential treatment assignment with balancing for prognostic factors in the controlled clinical trial. Biometrics , 31(1):103--115
1975
-
[24]
N., Nordwall, J., Babiker, A
Polizzotto, M. N., Nordwall, J., Babiker, A. G., Phillips, A., Vock, D. M., Eriobu, N., and Lane, H. C. (2022). Hyperimmune immunoglobulin for hospitalised patients with COVID -19 ( ITAC ): A double-blind, placebo-controlled, phase 3, randomised trial. The Lancet , 399(10324):530--540
2022
-
[25]
and Heumann, C
Schomaker, M. and Heumann, C. (2018). Bootstrap inference when using multiple imputation. Statistics in Medicine , 37(14):2252
2018
-
[26]
Senn, S. J. (1989). Covariate imbalance and random allocation in clinical trials. Statistics in Medicine , 8(4):467--475
1989
-
[27]
Taves, D. R. (1974). Minimization: A new method of assigning patients to treatment and control groups. Clinical Pharmacology & Therapeutics , 15(5):443--453
1974
-
[28]
and Detsky, A
Tomlinson, G. and Detsky, A. S. (2010). Composite end points in randomized trials: There is no free lunch. JAMA , 303(3):267--268
2010
-
[29]
FDA Center for Drug Evaluation and Research and U.S
U.S. FDA Center for Drug Evaluation and Research and U.S. FDA Center for Biologics Evaluation and Research (2021). E9( R1 ) Statistical principles for clinical trials: addendum: Estimands and sensitivity analysis in clinical trials. Technical report
2021
-
[30]
J., Polley, E
van der Laan, M. J., Polley, E. C., and Hubbard, A. E. (2007). Super learner. Statistical Applications in Genetics and Molecular Biology , 6(1):Article 25
2007
-
[31]
M., Tessier, K
White, C. M., Tessier, K. M., Koopmeiners, J. S., Denlinger-Apte, R. L., Cobb, C. O., Lane, T., Campos, C. L., Spangler, J. G., Hatsukami, D. K., Strasser, A. A., and Donny, E. C. (2022). Preliminary evidence on cigarette nicotine reduction with concurrent access to an e-cigar...
2022
-
[32]
M., Koopmeiners, J
Wolf, J. M., Koopmeiners, J. S., and Vock, D. M. (2024a). Commentary on Chen et al. (2022): The need for continued methodological research on leveraging information in secondary endpoints for more efficient RCTs . Contemporary Clinical Trials , 145:Article 107664
2024
-
[33]
M., Vock, D
Wolf, J. M., Vock, D. M., Luo, X., Hatsukami, D. K., McClernon, F. J., and Koopmeiners, J. S. (2024b). Leveraging information from secondary endpoints to enhance dynamic borrowing across subpopulations. Biometrics , 80(4):ujae118
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.