Pith. sign in

REVIEW 3 major objections 4 minor 19 references

Absolute Risk Prediction for Cannabis Use Disorder in Adolescence and Early Adulthood Using Bayesian Machine Learning

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Five simple factors forecast cannabis use disorder risk

desk verdict First absolute-risk model for CUD with real external validation, but the delinquency predictor is averaged over post-prediction waves, so the reported AUC and E/O overstate prospective accuracy. read the letter →

arxiv 2501.09156 v2 pith:7464EZWL submitted 2025-01-15 stat.AP stat.ML

classification stat.APstat.ML MSC 62N0162F1562P10
keywords cannabisusedisorderabsoluteriskpredictionBayesianlassoCoxproportionalhazardscompetingriskslongitudinalpredictorsexternalvalidation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to give clinicians a tool that converts five easily collected facts about a young cannabis user — biological sex, a delinquency measure, and three personality-trait scores — into a personalized probability of developing cannabis use disorder within a chosen time window, such as five years. No absolute-risk model of this kind exists for substance use disorders. The authors argue that the model discriminates reasonably well, with area-under-the-curve values around 0.68 to 0.75 across training and two validation sets, and is well calibrated, with expected-to-observed ratios near 1. If these results hold, a brief questionnaire could flag high-risk adolescents for early intervention.

What carries the argument

The load-bearing object is the cause-specific absolute-risk integral from standard competing-risk theory, computed by plugging in a Cox model for the cannabis-use-disorder hazard and an external all-cause mortality hazard. The Cox coefficients are regularized with a Bayesian lasso prior, the baseline hazard is modeled with M-splines, a smooth piecewise-polynomial basis, under a Dirichlet prior, survey weights enter the likelihood, and posterior draws from Markov chain Monte Carlo are averaged to produce predicted risks. The longitudinal predictor delinquency is summarized by its mean across waves, which the authors show performs at least as well as a model-based random intercept and is easier for clinicians to use.

What would settle it

Take the trained model and recompute five-year AUC and E/O on the validation sets after replacing the delinquency mean with values measured strictly before the prediction age, for example before first cannabis use. If the AUC drops below about 0.6 or the E/O values move away from 1, the reported prospective accuracy is inflated by look-ahead.

Watch

Extended reading notes

Core claim

The paper's central claim is that the absolute risk of cannabis use disorder for an adolescent or young adult who already uses cannabis can be modeled as a function of five risk factors through a Bayesian Cox proportional-hazards model, with the competing risk of death folded in from national life tables. In the final model, male sex, higher delinquency, higher neuroticism, higher openness, and lower conscientiousness each raise the estimated hazard, and the model returns a risk estimate for any user-specified age interval. The authors report that for five-year risk following first cannabis use, the AUC is 0.68 in cross-validation, 0.64 on a held-out national sample, and 0.75 on an external cohort, with expected-versus-observed ratios of 0.95, 0.98, and 1.00. They also present a six-factor variant that adds welfare status for settings where that information exists.

Load-bearing premise

The model assumes that a young person's delinquency average is available at the moment of prediction, yet the study's predictor construction only required values to be measured before the disorder or censoring, not before the prediction time — so the delinquency score used in some predictions may include behavior that happened after the prediction was supposed to be made.

Editorial extensions

If this is right

  • A clinician who knows five easy-to-obtain facts about a young cannabis user can state a concrete five-year probability of developing cannabis use disorder, not just a relative-risk ranking.
  • The model can be recalibrated to a new population using only that population's cannabis-use-disorder prevalence, so it can be ported to other countries or eras of cannabis policy.
  • Because the risk is absolute and time-bounded, it can be compared directly with the competing risk of death and with risks of other outcomes, supporting decisions about intervention intensity.
  • The five-factor form makes a paper-and-pencil or short-app risk score feasible for primary-care and school-based settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the look-ahead concern I noted is real, the reported AUC and E/O values are upper bounds on prospective accuracy; a fair out-of-time test with baseline delinquency only would likely show somewhat lower discrimination.
  • The model treats delinquency as fixed at prediction; a dynamic version that updates delinquency and personality measures over follow-up could improve prediction but would require new methodology for time-varying covariates in absolute-risk settings.
  • Post-legalization cohorts may have different base prevalence and cannabis potency, so absolute risks should be recalibrated before clinical use even though the ranking of individuals may persist.
  • The same pipeline — Bayesian lasso, M-spline baseline hazard, life-table competing risk — could be turned into absolute-risk models for alcohol or opioid use disorders if suitable longitudinal cohorts exist.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper develops a Bayesian machine-learning model for absolute risk of cannabis use disorder (CUD) in adolescents and young adults who use cannabis. Training uses the Add Health cohort, with CUD hazard modeled by a Cox proportional hazards model with an M-spline baseline and a lasso prior, and mortality from non-CUD causes treated as a competing risk using U.S. life-table rates. The final Model 1 contains five predictors: biological sex, delinquency, conscientiousness, neuroticism, and openness. Prediction performance is assessed by AUC and E/O in 5-fold CV, on an Add Health holdout, and on the external CHDS cohort, with CHDS recalibration via a logistic intercept update. The authors report AUCs of 0.68, 0.64, and 0.75 and E/O values of 0.95, 0.98, and 1, and conclude that the model is well calibrated and clinically useful.

Significance. If the reported performance were unbiased, this would be a useful contribution: an externally validated, parsimonious absolute-risk model for CUD, with a principled competing-risk framework and an M-spline baseline, would fill a genuine clinical gap. The paper also demonstrates a reproducible workflow for incorporating a longitudinal predictor as a summary measure in absolute-risk prediction. However, the main performance claim is undermined by a temporal-ordering flaw in the delinquency predictor, which can incorporate post-prediction information and therefore inflates the reported AUC and E/O values. The CHDS calibration claim is also partly circular because the reported E/O near 1 is obtained after fitting a recalibration intercept to CHDS outcomes. These issues are load-bearing for the central claim of good prospective discrimination and calibration.

major comments (3)
  1. [Supplement S5, step 1; Section 3.1, Table 1] The longitudinal delinquency predictor is not guaranteed to be measured before the time at which a prediction is made. Supplement S5 restricts predictor values to those measured "strictly before the time of CUD onset or censoring," not before the prediction age. Because Model 1 uses the mean of delinquency over waves, a prediction made at the age of first cannabis use or at age 16 can include delinquency measurements from later waves for any individual whose CUD onset or censoring occurs after that age, and for CUD-free individuals censoring is at wave IV. Since delinquency is the strongest predictor in the model (posterior mean HR 19.89, 95% CrI 12.06-32.79 in Table 1), this look-ahead likely inflates the reported AUC values (0.68/0.64/0.75) and the apparently good E/O calibration. The same construction rule was applied in the 5-fold CV and in both validation datasets, so the reported prospective performance is not credible as stated. The model building and validation should be redone with delinquency summarized only from waves available at the prediction time, followed by a re-estimation of all performance metrics.
  2. [Section 4.2 and Supplement S6] The E/O value of 1 reported for the CHDS validation is computed after logistic recalibration in which the intercept is estimated from the CHDS CUD outcomes themselves. This is not an independent calibration check: updating the intercept on the validation data will pull the overall E/O toward 1 by construction. The abstract and Section 4.2.2 should clearly separate the model's original calibration from the post-recalibration calibration, and the original (pre-recalibration) E/O should be reported. Without this separation, the calibration claim for the model in a new population is overstated.
  3. [Table 1 and Section 3.1] For Model 2, Table 1 reports a hazard ratio for welfare of 1.03 with 95% credible interval (0.85, 1.23), which includes 1. The text in Section 3.1 states that "higher risk of CUD is associated with receipt of welfare," but this association is not supported by the reported posterior interval. The wording should be corrected to indicate that the evidence for a welfare effect is inconclusive, or the model should be refit with a different specification if the authors wish to claim an effect.
minor comments (4)
  1. [Abstract and Section 3.2] The abstract refers to an AUC of 0.68 for the "training dataset," but this value is from 5-fold cross-validation, not from refitting the model to the full training data; the wording should be changed to "5-fold cross-validation" for accuracy.
  2. [Section 2.1 and Section 4] The paper describes both validation datasets as "independent," but the Add Health test set is a holdout from the same study and cohort used for training. It would be clearer to call this an internal validation set and reserve "external" for the CHDS data.
  3. [Supplement S2] The M-spline specification is not fully reported: the degree, number of basis terms L, and knot locations are described generically but their actual values in the fitted model are not given. Reporting these values is needed for reproducibility.
  4. [Section 4.2.2 and Supplement Table S6] The statement that "all E/O values were close to 1 (for CHDS) after recalibration" is difficult to reconcile with Supplement Table S6, which shows E/O values such as 4.48 for the first risk quartile and 1.66 for females in 5-year predictions at first cannabis use. The text should clarify whether Table S6 is pre- or post-recalibration and should not imply uniform calibration across subgroups.

Circularity Check

1 steps flagged · score 4.0 of 10

CHDS E/O validation reduces to the recalibration intercept; training and Add Health validation remain independent.

  1. fitted input called prediction [Section 4.2.1, Section 4.2.2, and Supplement Section S6]
    "Specifically, we applied logistic re-calibration to update the intercept in the model for log odds of CUD risk to reflect the high prevalence in the CHDS validation data. ... All E/O values were close to 1 (for CHDS) after recalibration. These results indicate good predictive performance of the model."

    Supplement S6 fits a logistic regression on the CHDS validation data with CUD status as the response and the logit of Model 1 predicted risks as the predictor, then uses the estimated intercept to update the risks and evaluates the model on those updated risks. The intercept is estimated from the same CHDS outcomes used to compute E/O, so the post-recalibration E/O is a byproduct of that fit rather than an independent test of the original model's calibration. Reporting this E/O as evidence of good predictive performance converts a fitted parameter into a validation result. The AUC is unaffected by the intercept update and remains genuine, which is why the circularity is only partial.

full rationale

The central derivation is not circular. The model is trained on Add Health data using a standard survival likelihood with lasso priors, and the 5-fold CV plus Add Health holdout validation evaluate predictions against held-out observations, so those AUC and E/O values are independent of the fitted parameters. The external CHDS validation contributes a genuine AUC comparison. The only circular element is the reported CHDS E/O calibration: the paper explicitly recalibrates the intercept on CHDS outcomes and then uses the recalibrated E/O, close to 1, as evidence of predictive performance. Equivalently, that calibration claim is manufactured by the recalibration step. A separate concern about delinquency being averaged over waves measured after the prediction time is a look-ahead and validity issue rather than a circularity, so it does not affect this score. Overall circularity score 4 reflects one fitted-input-called-prediction in an otherwise self-contained validation chain.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The model is a standard statistical prediction exercise: model coefficients are estimated on training data, which is expected. The ledger lists the hand-chosen screening and selection thresholds, the smoothing and prior choices, and the external recalibration intercept that the central calibration claim partly rests on. No new particles, forces, or entities are introduced.

free parameters (5)
  • Univariate screening p-value threshold = 0.25
    Hand-chosen in model building step 1; determines which of 45 candidate predictors move forward.
  • Multivariate screening criteria = p < 0.25 and at least 80% complete data
    Hand-chosen in model building step 2; produces the 21 predictors that enter Bayesian model fitting.
  • Variable selection thresholds = Credible interval levels and scaled neighborhood thresholds varied
    Different thresholds generated six competing models; final Model 1 depends on these choices.
  • M-spline degree and knot locations = Not specified in main text
    Baseline hazard flexibility depends on these choices; absence of details hinders replication.
  • CHDS recalibration intercept = Not reported numerically
    Fitted to CUD status in the external validation data; the near-1 E/O after recalibration reflects this fit.
assumptions (6)
  • domain assumption Censoring is independent of CUD onset time (Supplement S3.1).
    Wave IV participation could be related to substance use; informative censoring would bias the estimated CUD hazard and absolute risks.
  • domain assumption Mortality from non-CUD causes is independent of the model's risk factors and equals the all-cause US life table hazard, treated as known without error (Supplement S2.2).
    Used to compute absolute risk; violation would shift absolute risk estimates, though competing mortality is small for this age group.
  • domain assumption The proportional hazards assumption holds (main text Section 3.1).
    Only a Schoenfeld residual test p-value summary is given; the entire hazard ratio model relies on PH.
  • domain assumption Self-reported ages of first cannabis use and CUD onset, and Wave IV CUD status, are accurate (Section 2.1).
    Time-to-event outcome and left truncation age are built from retrospective self-report; recall error directly affects the estimated hazard.
  • domain assumption The weighted likelihood with Add Health survey weights is a valid estimator for the target population (Section 2.4 and Supplement S3.1).
    Survey weights are entered as powers in the likelihood; this design-based approximation is standard but relies on the weights being correct and non-informative within strata.
  • standard math Only one event (CUD or non-CUD death) can occur at a given time, so the overall hazard is the sum of cause-specific hazards (Section 2.2).
    This is the standard competing-risk formulation underlying equation (1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Absolute Risk Prediction for Cannabis Use Disorder in Adolescence and Early Adulthood Using Bayesian Machine Learning." pith.science (2026). https://pith.science/paper/7464EZWL

@misc{pith2026250109156,
  author       = {Pith},
  title        = {Pith review of: Absolute Risk Prediction for Cannabis Use Disorder in Adolescence and Early Adulthood Using Bayesian Machine Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7464EZWL}},
  note         = {Machine review of arXiv:2501.09156}
}
read the original abstract

Introduction: Substance use disorders (SUDs) have emerged as a pressing public health concern in the United States, with adolescent substance use often leading to SUDs in adulthood. Effective strategies are needed to stem this progression. To help fulfill this need, we developed a novel absolute risk prediction model for cannabis use disorder (CUD) for adolescents or young adults who use cannabis. Methods: We trained a Bayesian machine learning model that provides a personalized CUD absolute risk for adolescents or young adults who use cannabis with data from the National Longitudinal Study of Adolescent to Adult Health. Model performance was assessed using 5-fold cross-validation (CV) with area under the curve (AUC) and ratio of the expected to observed number of cases (E/O). Independent validation of the final model was conducted using two datasets. Results: The proposed model has five risk factors: biological sex, delinquency, and scores on personality traits of conscientiousness, neuroticism, and openness. For predicting CUD risk within five years of first cannabis use, AUC values for the training dataset and two validation datasets were 0.68, 0.64, and 0.75, respectively, and E/O values were 0.95, 0.98, and 1, respectively. This indicates good discrimination and calibration performance of the model. Discussion and Conclusion: The proposed model can aid clinicians in assessing the risk of developing CUD among adolescents and young adults who use cannabis, enabling clinically appropriate interventions.

Figures

Figures reproduced from arXiv: 2501.09156 by the authors.

Figure 1
Figure 1. CUD hazards. 19 [PITH_FULL_IMAGE:figures/full_fig_p019_1.png] view at source ↗
Figure 2
Figure 2. Validation on Add Health test data: AUC and E/O trajectories for CUD risk prediction within a certain [PITH_FULL_IMAGE:figures/full_fig_p020_2.png] view at source ↗
Figure 3
Figure 3. AUC and E/O trajectories of Model 1 for CUD risk prediction within a certain number of years made at [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 19 canonical work pages

  1. [1]

    Fitting Linear Mixed-Effects Models using lme4

    Douglas Bates, Martin M¨ achler, Ben Bolker, and Steve Walker. Fitting linear mixed-effects models using lme4. arXiv preprint arXiv:1406.5823 , 2014

  2. [2]

    D. R. Cox. Regression models and life-tables. Journal of the Royal Statistical Society Series B: Statistical Methodology, 34(2):187–202, 1972

  3. [3]

    J. O. Ramsay. Monotone regression splines in action. Statistical Science, 3(4):425 – 441, 1988

  4. [4]

    Klein and Melvin L

    John P. Klein and Melvin L. Moeschberger. Survival analysis: techniques for censored and truncated data . Springer New York, NY, 2003

  5. [5]

    The health effects of cannabis and cannabinoids: the current state of evidence and recommendations for research

    National Academies of Sciences, Medicine Division, Board on Population Health, Public Health Practice, Committee on the Health Effects of Marijuana, An Evidence Review, and Research Agenda. The health effects of cannabis and cannabinoids: the current state of evidence and recommendations for research . National Academies Press, Washington, DC, 2017

  6. [6]

    Underlying cause of death 1999-2020 on cdc wonder online database, 2021

    CDC. Underlying cause of death 1999-2020 on cdc wonder online database, 2021

  7. [7]

    United States life tables, 2020

    Elizabeth Arias and Jiaquan Xu. United States life tables, 2020. National Vital Statistics Reports, 71(1), 2022

  8. [8]

    Pfeiffer and Mitchell H

    Ruth M. Pfeiffer and Mitchell H. Gail. Absolute risk: methods and applications in clinical management and public health. Chapman and Hall/CRC, 1st edition, 2017

Show all 19 references
  1. [9]

    Shrinkage priors for Bayesian penalized regression

    Sara Van Erp, Daniel Oberski, and Joris Mulder. Shrinkage priors for Bayesian penalized regression. Journal of Mathematical Psychology, 89:31–50, 2019

  2. [10]

    A constructive definition of Dirichlet priors

    Jayaram Sethuraman. A constructive definition of Dirichlet priors. Statistica Sinica, 4(2):639–650, 1994

  3. [11]

    R: A language and environment for statistical computing, 2023

    R Core Team. R: A language and environment for statistical computing, 2023

  4. [12]

    Shape-restricted regression splines with R package splines2

    Wenjie Wang and Jun Yan. Shape-restricted regression splines with R package splines2. Journal of Data Science, 19(3):498–517, 2021

  5. [13]

    RStan: the R interface to Stan, 2023

    Stan Development Team. RStan: the R interface to Stan, 2023. R package version 2.32.3

  6. [14]

    R. M. D. S. Rajapaksha, F. Filbey, S. Biswas, and P. Choudhary. A Bayesian learning model to predict the risk for cannabis use disorder. Drug and Alcohol Dependence, 236:109476, 2022

  7. [15]

    Regression modeling strategies for improved prognostic prediction

    Frank E Harrell Jr, Kerry L Lee, Robert M Califf, David B Pryor, and Robert A Rosati. Regression modeling strategies for improved prognostic prediction. Statistics in Medicine , 3(2):143–152, 1984

  8. [16]

    The Bayesian elastic net

    Qing Li and Nan Lin. The Bayesian elastic net. Bayesian Analysis, 5(1):151 – 170, 2010

  9. [17]

    Ruberu, Rajapaksha Mudalige Dhanushka S

    Thanthirige Lakshika M. Ruberu, Rajapaksha Mudalige Dhanushka S. Rajapaksha, Mary M. Heitzeg, Ryan Klaus, Joseph M. Boden, Swati Biswas, and Pankaj Choudhary. Validation of a Bayesian learning model to predict the risk for cannabis use disorder. Addictive Behaviors, 146:107799...

  10. [18]

    Steyerberg, Gerard J

    Ewout W. Steyerberg, Gerard J. J. M. Borsboom, Hans C. Van Houwelingen, Marinus J. C. Eijkemans, and J. Dik F. Habbema. Validation and updating of predictive logistic regression models: a study on sample size and shrinkage. Statistics in Medicine , 23(16):2567–2586, 2004

  11. [19]

    K. J. M. Janssen, K. G. M. Moons, C. J. Kalkman, D. E. Grobbee, and Y Vergouwe. Updating methods improved the performance of a clinical prediction model in new patients. Journal of clinical epidemiology , 61 (1):76–86, 2008. 13

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.