REVIEW 3 major objections 4 minor 19 references
Absolute Risk Prediction for Cannabis Use Disorder in Adolescence and Early Adulthood Using Bayesian Machine Learning
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Five simple factors forecast cannabis use disorder risk
desk verdict First absolute-risk model for CUD with real external validation, but the delinquency predictor is averaged over post-prediction waves, so the reported AUC and E/O overstate prospective accuracy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the cause-specific absolute-risk integral from standard competing-risk theory, computed by plugging in a Cox model for the cannabis-use-disorder hazard and an external all-cause mortality hazard. The Cox coefficients are regularized with a Bayesian lasso prior, the baseline hazard is modeled with M-splines, a smooth piecewise-polynomial basis, under a Dirichlet prior, survey weights enter the likelihood, and posterior draws from Markov chain Monte Carlo are averaged to produce predicted risks. The longitudinal predictor delinquency is summarized by its mean across waves, which the authors show performs at least as well as a model-based random intercept and is easier for clinicians to use.
What would settle it
Take the trained model and recompute five-year AUC and E/O on the validation sets after replacing the delinquency mean with values measured strictly before the prediction age, for example before first cannabis use. If the AUC drops below about 0.6 or the E/O values move away from 1, the reported prospective accuracy is inflated by look-ahead.
Extended reading notes
Core claim
The paper's central claim is that the absolute risk of cannabis use disorder for an adolescent or young adult who already uses cannabis can be modeled as a function of five risk factors through a Bayesian Cox proportional-hazards model, with the competing risk of death folded in from national life tables. In the final model, male sex, higher delinquency, higher neuroticism, higher openness, and lower conscientiousness each raise the estimated hazard, and the model returns a risk estimate for any user-specified age interval. The authors report that for five-year risk following first cannabis use, the AUC is 0.68 in cross-validation, 0.64 on a held-out national sample, and 0.75 on an external cohort, with expected-versus-observed ratios of 0.95, 0.98, and 1.00. They also present a six-factor variant that adds welfare status for settings where that information exists.
Load-bearing premise
The model assumes that a young person's delinquency average is available at the moment of prediction, yet the study's predictor construction only required values to be measured before the disorder or censoring, not before the prediction time — so the delinquency score used in some predictions may include behavior that happened after the prediction was supposed to be made.
Editorial extensions
If this is right
- A clinician who knows five easy-to-obtain facts about a young cannabis user can state a concrete five-year probability of developing cannabis use disorder, not just a relative-risk ranking.
- The model can be recalibrated to a new population using only that population's cannabis-use-disorder prevalence, so it can be ported to other countries or eras of cannabis policy.
- Because the risk is absolute and time-bounded, it can be compared directly with the competing risk of death and with risks of other outcomes, supporting decisions about intervention intensity.
- The five-factor form makes a paper-and-pencil or short-app risk score feasible for primary-care and school-based settings.
Reading between the lines
- If the look-ahead concern I noted is real, the reported AUC and E/O values are upper bounds on prospective accuracy; a fair out-of-time test with baseline delinquency only would likely show somewhat lower discrimination.
- The model treats delinquency as fixed at prediction; a dynamic version that updates delinquency and personality measures over follow-up could improve prediction but would require new methodology for time-varying covariates in absolute-risk settings.
- Post-legalization cohorts may have different base prevalence and cannabis potency, so absolute risks should be recalibrated before clinical use even though the ranking of individuals may persist.
- The same pipeline — Bayesian lasso, M-spline baseline hazard, life-table competing risk — could be turned into absolute-risk models for alcohol or opioid use disorders if suitable longitudinal cohorts exist.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a Bayesian machine-learning model for absolute risk of cannabis use disorder (CUD) in adolescents and young adults who use cannabis. Training uses the Add Health cohort, with CUD hazard modeled by a Cox proportional hazards model with an M-spline baseline and a lasso prior, and mortality from non-CUD causes treated as a competing risk using U.S. life-table rates. The final Model 1 contains five predictors: biological sex, delinquency, conscientiousness, neuroticism, and openness. Prediction performance is assessed by AUC and E/O in 5-fold CV, on an Add Health holdout, and on the external CHDS cohort, with CHDS recalibration via a logistic intercept update. The authors report AUCs of 0.68, 0.64, and 0.75 and E/O values of 0.95, 0.98, and 1, and conclude that the model is well calibrated and clinically useful.
Significance. If the reported performance were unbiased, this would be a useful contribution: an externally validated, parsimonious absolute-risk model for CUD, with a principled competing-risk framework and an M-spline baseline, would fill a genuine clinical gap. The paper also demonstrates a reproducible workflow for incorporating a longitudinal predictor as a summary measure in absolute-risk prediction. However, the main performance claim is undermined by a temporal-ordering flaw in the delinquency predictor, which can incorporate post-prediction information and therefore inflates the reported AUC and E/O values. The CHDS calibration claim is also partly circular because the reported E/O near 1 is obtained after fitting a recalibration intercept to CHDS outcomes. These issues are load-bearing for the central claim of good prospective discrimination and calibration.
major comments (3)
- [Supplement S5, step 1; Section 3.1, Table 1] The longitudinal delinquency predictor is not guaranteed to be measured before the time at which a prediction is made. Supplement S5 restricts predictor values to those measured "strictly before the time of CUD onset or censoring," not before the prediction age. Because Model 1 uses the mean of delinquency over waves, a prediction made at the age of first cannabis use or at age 16 can include delinquency measurements from later waves for any individual whose CUD onset or censoring occurs after that age, and for CUD-free individuals censoring is at wave IV. Since delinquency is the strongest predictor in the model (posterior mean HR 19.89, 95% CrI 12.06-32.79 in Table 1), this look-ahead likely inflates the reported AUC values (0.68/0.64/0.75) and the apparently good E/O calibration. The same construction rule was applied in the 5-fold CV and in both validation datasets, so the reported prospective performance is not credible as stated. The model building and validation should be redone with delinquency summarized only from waves available at the prediction time, followed by a re-estimation of all performance metrics.
- [Section 4.2 and Supplement S6] The E/O value of 1 reported for the CHDS validation is computed after logistic recalibration in which the intercept is estimated from the CHDS CUD outcomes themselves. This is not an independent calibration check: updating the intercept on the validation data will pull the overall E/O toward 1 by construction. The abstract and Section 4.2.2 should clearly separate the model's original calibration from the post-recalibration calibration, and the original (pre-recalibration) E/O should be reported. Without this separation, the calibration claim for the model in a new population is overstated.
- [Table 1 and Section 3.1] For Model 2, Table 1 reports a hazard ratio for welfare of 1.03 with 95% credible interval (0.85, 1.23), which includes 1. The text in Section 3.1 states that "higher risk of CUD is associated with receipt of welfare," but this association is not supported by the reported posterior interval. The wording should be corrected to indicate that the evidence for a welfare effect is inconclusive, or the model should be refit with a different specification if the authors wish to claim an effect.
minor comments (4)
- [Abstract and Section 3.2] The abstract refers to an AUC of 0.68 for the "training dataset," but this value is from 5-fold cross-validation, not from refitting the model to the full training data; the wording should be changed to "5-fold cross-validation" for accuracy.
- [Section 2.1 and Section 4] The paper describes both validation datasets as "independent," but the Add Health test set is a holdout from the same study and cohort used for training. It would be clearer to call this an internal validation set and reserve "external" for the CHDS data.
- [Supplement S2] The M-spline specification is not fully reported: the degree, number of basis terms L, and knot locations are described generically but their actual values in the fitted model are not given. Reporting these values is needed for reproducibility.
- [Section 4.2.2 and Supplement Table S6] The statement that "all E/O values were close to 1 (for CHDS) after recalibration" is difficult to reconcile with Supplement Table S6, which shows E/O values such as 4.48 for the first risk quartile and 1.66 for females in 5-year predictions at first cannabis use. The text should clarify whether Table S6 is pre- or post-recalibration and should not imply uniform calibration across subgroups.
Circularity Check
CHDS E/O validation reduces to the recalibration intercept; training and Add Health validation remain independent.
-
fitted input called prediction
[Section 4.2.1, Section 4.2.2, and Supplement Section S6]
"Specifically, we applied logistic re-calibration to update the intercept in the model for log odds of CUD risk to reflect the high prevalence in the CHDS validation data. ... All E/O values were close to 1 (for CHDS) after recalibration. These results indicate good predictive performance of the model."
Supplement S6 fits a logistic regression on the CHDS validation data with CUD status as the response and the logit of Model 1 predicted risks as the predictor, then uses the estimated intercept to update the risks and evaluates the model on those updated risks. The intercept is estimated from the same CHDS outcomes used to compute E/O, so the post-recalibration E/O is a byproduct of that fit rather than an independent test of the original model's calibration. Reporting this E/O as evidence of good predictive performance converts a fitted parameter into a validation result. The AUC is unaffected by the intercept update and remains genuine, which is why the circularity is only partial.
full rationale
The central derivation is not circular. The model is trained on Add Health data using a standard survival likelihood with lasso priors, and the 5-fold CV plus Add Health holdout validation evaluate predictions against held-out observations, so those AUC and E/O values are independent of the fitted parameters. The external CHDS validation contributes a genuine AUC comparison. The only circular element is the reported CHDS E/O calibration: the paper explicitly recalibrates the intercept on CHDS outcomes and then uses the recalibrated E/O, close to 1, as evidence of predictive performance. Equivalently, that calibration claim is manufactured by the recalibration step. A separate concern about delinquency being averaged over waves measured after the prediction time is a look-ahead and validity issue rather than a circularity, so it does not affect this score. Overall circularity score 4 reflects one fitted-input-called-prediction in an otherwise self-contained validation chain.
Assumptions & free parameters
free parameters (5)
- Univariate screening p-value threshold =
0.25
- Multivariate screening criteria =
p < 0.25 and at least 80% complete data
- Variable selection thresholds =
Credible interval levels and scaled neighborhood thresholds varied
- M-spline degree and knot locations =
Not specified in main text
- CHDS recalibration intercept =
Not reported numerically
assumptions (6)
- domain assumption Censoring is independent of CUD onset time (Supplement S3.1).
- domain assumption Mortality from non-CUD causes is independent of the model's risk factors and equals the all-cause US life table hazard, treated as known without error (Supplement S2.2).
- domain assumption The proportional hazards assumption holds (main text Section 3.1).
- domain assumption Self-reported ages of first cannabis use and CUD onset, and Wave IV CUD status, are accurate (Section 2.1).
- domain assumption The weighted likelihood with Add Health survey weights is a valid estimator for the target population (Section 2.4 and Supplement S3.1).
- standard math Only one event (CUD or non-CUD death) can occur at a given time, so the overall hazard is the sum of cause-specific hazards (Section 2.2).
Cite this review
Pith. "Pith review of Absolute Risk Prediction for Cannabis Use Disorder in Adolescence and Early Adulthood Using Bayesian Machine Learning." pith.science (2026). https://pith.science/paper/7464EZWL
@misc{pith2026250109156,
author = {Pith},
title = {Pith review of: Absolute Risk Prediction for Cannabis Use Disorder in Adolescence and Early Adulthood Using Bayesian Machine Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/7464EZWL}},
note = {Machine review of arXiv:2501.09156}
}
read the original abstract
Introduction: Substance use disorders (SUDs) have emerged as a pressing public health concern in the United States, with adolescent substance use often leading to SUDs in adulthood. Effective strategies are needed to stem this progression. To help fulfill this need, we developed a novel absolute risk prediction model for cannabis use disorder (CUD) for adolescents or young adults who use cannabis. Methods: We trained a Bayesian machine learning model that provides a personalized CUD absolute risk for adolescents or young adults who use cannabis with data from the National Longitudinal Study of Adolescent to Adult Health. Model performance was assessed using 5-fold cross-validation (CV) with area under the curve (AUC) and ratio of the expected to observed number of cases (E/O). Independent validation of the final model was conducted using two datasets. Results: The proposed model has five risk factors: biological sex, delinquency, and scores on personality traits of conscientiousness, neuroticism, and openness. For predicting CUD risk within five years of first cannabis use, AUC values for the training dataset and two validation datasets were 0.68, 0.64, and 0.75, respectively, and E/O values were 0.95, 0.98, and 1, respectively. This indicates good discrimination and calibration performance of the model. Discussion and Conclusion: The proposed model can aid clinicians in assessing the risk of developing CUD among adolescents and young adults who use cannabis, enabling clinically appropriate interventions.
Figures
Reference graph
Works this paper leans on
-
[1]
Fitting Linear Mixed-Effects Models using lme4
Douglas Bates, Martin M¨ achler, Ben Bolker, and Steve Walker. Fitting linear mixed-effects models using lme4. arXiv preprint arXiv:1406.5823 , 2014
work page Pith review arXiv 2014
-
[2]
D. R. Cox. Regression models and life-tables. Journal of the Royal Statistical Society Series B: Statistical Methodology, 34(2):187–202, 1972
work page 1972
-
[3]
J. O. Ramsay. Monotone regression splines in action. Statistical Science, 3(4):425 – 441, 1988
work page 1988
-
[4]
John P. Klein and Melvin L. Moeschberger. Survival analysis: techniques for censored and truncated data . Springer New York, NY, 2003
work page 2003
-
[5]
National Academies of Sciences, Medicine Division, Board on Population Health, Public Health Practice, Committee on the Health Effects of Marijuana, An Evidence Review, and Research Agenda. The health effects of cannabis and cannabinoids: the current state of evidence and recommendations for research . National Academies Press, Washington, DC, 2017
work page 2017
-
[6]
Underlying cause of death 1999-2020 on cdc wonder online database, 2021
CDC. Underlying cause of death 1999-2020 on cdc wonder online database, 2021
work page 1999
-
[7]
United States life tables, 2020
Elizabeth Arias and Jiaquan Xu. United States life tables, 2020. National Vital Statistics Reports, 71(1), 2022
work page 2020
-
[8]
Ruth M. Pfeiffer and Mitchell H. Gail. Absolute risk: methods and applications in clinical management and public health. Chapman and Hall/CRC, 1st edition, 2017
work page 2017
Show all 19 references
-
[9]
Shrinkage priors for Bayesian penalized regression
Sara Van Erp, Daniel Oberski, and Joris Mulder. Shrinkage priors for Bayesian penalized regression. Journal of Mathematical Psychology, 89:31–50, 2019
2019
-
[10]
A constructive definition of Dirichlet priors
Jayaram Sethuraman. A constructive definition of Dirichlet priors. Statistica Sinica, 4(2):639–650, 1994
1994
-
[11]
R: A language and environment for statistical computing, 2023
R Core Team. R: A language and environment for statistical computing, 2023
2023
-
[12]
Shape-restricted regression splines with R package splines2
Wenjie Wang and Jun Yan. Shape-restricted regression splines with R package splines2. Journal of Data Science, 19(3):498–517, 2021
2021
-
[13]
RStan: the R interface to Stan, 2023
Stan Development Team. RStan: the R interface to Stan, 2023. R package version 2.32.3
2023
-
[14]
R. M. D. S. Rajapaksha, F. Filbey, S. Biswas, and P. Choudhary. A Bayesian learning model to predict the risk for cannabis use disorder. Drug and Alcohol Dependence, 236:109476, 2022
2022
-
[15]
Regression modeling strategies for improved prognostic prediction
Frank E Harrell Jr, Kerry L Lee, Robert M Califf, David B Pryor, and Robert A Rosati. Regression modeling strategies for improved prognostic prediction. Statistics in Medicine , 3(2):143–152, 1984
1984
-
[16]
The Bayesian elastic net
Qing Li and Nan Lin. The Bayesian elastic net. Bayesian Analysis, 5(1):151 – 170, 2010
2010
-
[17]
Ruberu, Rajapaksha Mudalige Dhanushka S
Thanthirige Lakshika M. Ruberu, Rajapaksha Mudalige Dhanushka S. Rajapaksha, Mary M. Heitzeg, Ryan Klaus, Joseph M. Boden, Swati Biswas, and Pankaj Choudhary. Validation of a Bayesian learning model to predict the risk for cannabis use disorder. Addictive Behaviors, 146:107799...
2023
-
[18]
Steyerberg, Gerard J
Ewout W. Steyerberg, Gerard J. J. M. Borsboom, Hans C. Van Houwelingen, Marinus J. C. Eijkemans, and J. Dik F. Habbema. Validation and updating of predictive logistic regression models: a study on sample size and shrinkage. Statistics in Medicine , 23(16):2567–2586, 2004
2004
-
[19]
K. J. M. Janssen, K. G. M. Moons, C. J. Kalkman, D. E. Grobbee, and Y Vergouwe. Updating methods improved the performance of a clinical prediction model in new patients. Journal of clinical epidemiology , 61 (1):76–86, 2008. 13
2008
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.