{"id":"0d167686-abb8-46c5-bea8-f0df1a73672a","arxiv_id":"2501.09156","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A five-factor absolute risk model for cannabis use disorder in young cannabis users shows moderate discrimination (AUC 0.64 to 0.75) and near-expected calibration in external samples.","lead":"A research team built a risk calculator that estimates the chance an adolescent or young adult who uses cannabis will develop cannabis use disorder within a set number of years, using five factors: sex, delinquency, and three personality scores. The model was built on a large US longitudinal study, checked on two other datasets, and could help clinicians decide who needs early intervention.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Delinquency summary can include post-prediction waves; because delinquency is the strongest predictor (HR ~19.9 vs. 1.3 for sex), reported AUC and E/O likely overstate prospective accuracy.","rationale":"The reader's weakest assumption correctly identifies the temporal look-ahead in longitudinal predictor construction. Supplement S5 explicitly limits predictor values to those before CUD onset or censoring, not before the prediction time. Because the final model uses the delinquency mean over all waves before onset/censoring, and because delinquency is by far the strongest risk factor, this is a load-bearing flaw for the central claim of a clinically usable absolute risk model. The concern is internal to the paper's own protocol rather than a disagreement with external consensus, and it is testable with the existing data and code. The reader's CONDITIONAL verdict already reflects this concern, so no verdict change is needed; the added value here is a precise experiment that would settle the magnitude of the inflation. I did not find a more serious objection: the CHDS recalibration issue is real but secondary, and the model comparison and Bayesian estimation are otherwise competently described.","tokens_in":20444,"tokens_out":2363,"duration_ms":25907,"concrete_test":"Recompute Model 1's 5-year-after-first-use predictions in the 5-fold CV and in the Add Health validation set, replacing the delinquency average with the average taken only over waves at or before the prediction age (for age-of-first-use predictions, restrict to wave I; for age-16 predictions, restrict to waves before age 16). Keep the same CV folds and Bayesian estimation procedure. If the truncated-predictor AUC drops by more than 0.03 or E/O moves outside 0.8–1.2, the look-ahead is material and the reported prospective performance is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The final Model 1 uses the mean of the longitudinal delinquency scale over waves as a predictor (Section 3.1, Table 1). Supplement S5, step 1 states that predictor values were restricted to those measured 'strictly before the time of CUD onset or censoring' — not before the time at which the prediction is made. For a prediction made at the age of first cannabis use, or at a specific age such as 16, delinquency measurements taken in later waves (e.g., wave III or wave IV) are therefore available to the model whenever CUD onset or censoring occurs after that age. For CUD-free individuals, censoring is at wave IV, so the delinquency average can include measurements taken many years after the stated prediction time. This is look-ahead: a clinician at the prediction time would not have those future delinquency values. Since delinquency is the strongest risk factor in the model (posterior mean HR = 19.89, 95% CrI 12.06–32.79), the reported AUC values of 0.68/0.64/0.75 and E/O values near 1 are not credible as prospective performance measures. The problem affects the 5-fold CV and both validation datasets, because the same construction rule was applied throughout. The abstract's claim that the model provides risk 'within five years of first cannabis use' using these five factors is therefore not supported unless delinquency is redefined to use only information available at the prediction time.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Bayesian machine-learning model for absolute risk of cannabis use disorder (CUD) in adolescents and young adults who use cannabis. Training uses the Add Health cohort, with CUD hazard modeled by a Cox proportional hazards model with an M-spline baseline and a lasso prior, and mortality from non-CUD causes treated as a competing risk using U.S. life-table rates. The final Model 1 contains five predictors: biological sex, delinquency, conscientiousness, neuroticism, and openness. Prediction performance is assessed by AUC and E/O in 5-fold CV, on an Add Health holdout, and on the external CHDS cohort, with CHDS recalibration via a logistic intercept update. The authors report AUCs of 0.68, 0.64, and 0.75 and E/O values of 0.95, 0.98, and 1, and conclude that the model is well calibrated and clinically useful.","tokens_in":20689,"tokens_out":3645,"duration_ms":37050,"significance":"If the reported performance were unbiased, this would be a useful contribution: an externally validated, parsimonious absolute-risk model for CUD, with a principled competing-risk framework and an M-spline baseline, would fill a genuine clinical gap. The paper also demonstrates a reproducible workflow for incorporating a longitudinal predictor as a summary measure in absolute-risk prediction. However, the main performance claim is undermined by a temporal-ordering flaw in the delinquency predictor, which can incorporate post-prediction information and therefore inflates the reported AUC and E/O values. The CHDS calibration claim is also partly circular because the reported E/O near 1 is obtained after fitting a recalibration intercept to CHDS outcomes. These issues are load-bearing for the central claim of good prospective discrimination and calibration.","major_comments":[{"comment":"The longitudinal delinquency predictor is not guaranteed to be measured before the time at which a prediction is made. Supplement S5 restricts predictor values to those measured \"strictly before the time of CUD onset or censoring,\" not before the prediction age. Because Model 1 uses the mean of delinquency over waves, a prediction made at the age of first cannabis use or at age 16 can include delinquency measurements from later waves for any individual whose CUD onset or censoring occurs after that age, and for CUD-free individuals censoring is at wave IV. Since delinquency is the strongest predictor in the model (posterior mean HR 19.89, 95% CrI 12.06-32.79 in Table 1), this look-ahead likely inflates the reported AUC values (0.68/0.64/0.75) and the apparently good E/O calibration. The same construction rule was applied in the 5-fold CV and in both validation datasets, so the reported prospective performance is not credible as stated. The model building and validation should be redone with delinquency summarized only from waves available at the prediction time, followed by a re-estimation of all performance metrics.","section":"Supplement S5, step 1; Section 3.1, Table 1"},{"comment":"The E/O value of 1 reported for the CHDS validation is computed after logistic recalibration in which the intercept is estimated from the CHDS CUD outcomes themselves. This is not an independent calibration check: updating the intercept on the validation data will pull the overall E/O toward 1 by construction. The abstract and Section 4.2.2 should clearly separate the model's original calibration from the post-recalibration calibration, and the original (pre-recalibration) E/O should be reported. Without this separation, the calibration claim for the model in a new population is overstated.","section":"Section 4.2 and Supplement S6"},{"comment":"For Model 2, Table 1 reports a hazard ratio for welfare of 1.03 with 95% credible interval (0.85, 1.23), which includes 1. The text in Section 3.1 states that \"higher risk of CUD is associated with receipt of welfare,\" but this association is not supported by the reported posterior interval. The wording should be corrected to indicate that the evidence for a welfare effect is inconclusive, or the model should be refit with a different specification if the authors wish to claim an effect.","section":"Table 1 and Section 3.1"}],"minor_comments":[{"comment":"The abstract refers to an AUC of 0.68 for the \"training dataset,\" but this value is from 5-fold cross-validation, not from refitting the model to the full training data; the wording should be changed to \"5-fold cross-validation\" for accuracy.","section":"Abstract and Section 3.2"},{"comment":"The paper describes both validation datasets as \"independent,\" but the Add Health test set is a holdout from the same study and cohort used for training. It would be clearer to call this an internal validation set and reserve \"external\" for the CHDS data.","section":"Section 2.1 and Section 4"},{"comment":"The M-spline specification is not fully reported: the degree, number of basis terms L, and knot locations are described generically but their actual values in the fitted model are not given. Reporting these values is needed for reproducibility.","section":"Supplement S2"},{"comment":"The statement that \"all E/O values were close to 1 (for CHDS) after recalibration\" is difficult to reconcile with Supplement Table S6, which shows E/O values such as 4.48 for the first risk quartile and 1.66 for females in 5-year predictions at first cannabis use. The text should clarify whether Table S6 is pre- or post-recalibration and should not imply uniform calibration across subgroups.","section":"Section 4.2.2 and Supplement Table S6"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is the first absolute risk model for CUD I know of, externally validated on a New Zealand cohort, and it is a clean application of Bayesian Cox with M-splines. But the reported prospective performance is not credible as stated. The supplement restricts longitudinal predictor values to those measured before CUD onset or censoring, not before the time of prediction. For a prediction made at age of first cannabis use, a CUD-free person's delinquency score is averaged over waves that can occur years later. Since delinquency has HR ~20, that look-ahead likely inflates both AUC and E/O. This affects the cross-validation and both validation sets. I don't see a way around it other than redefining delinquency to include only waves measured at or before the prediction time and re-running the analysis.\n\nWhat's genuinely new: the absolute risk estimand with a defined time horizon and competing mortality, and the parsimonious five-factor model. The external CHDS validation is real, although the E/O values there are reported after a recalibration intercept is fitted to the same data, so only the AUC carries independent weight. The Add Health holdout subjects are a non-random subset (no survey weights), which weakens generalizability claims. No code or data are released, which is common for restricted cohorts but still limits reproducibility.\n\nThe math itself is standard and appears correctly specified; the Bayesian machinery and M-spline baseline are appropriate. The writing is clear and the comparison with Rajapaksha et al. is fair. The paper would be useful to methodologists working on temporal validation of risk models, and to addiction epidemiology readers, but only after the look-ahead is fixed.\n\nRecommendation: deserves serious peer review, not desk reject. A reviewer should push for a revised model using only pre-prediction predictors, un-recalibrated CHDS E/O, and ideally code or a data-sharing plan. If the authors can show the delinquency look-ahead makes little difference, the headline numbers may stand; right now they are upper bounds on prospective accuracy.","headline":"First absolute-risk model for CUD with real external validation, but the delinquency predictor is averaged over post-prediction waves, so the reported AUC and E/O overstate prospective accuracy.","tokens_in":21305,"tokens_out":2760,"would_cite":false,"duration_ms":29683,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N01","62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Five simple factors forecast cannabis use disorder risk","keywords":["cannabis use disorder","absolute risk prediction","Bayesian lasso","Cox proportional hazards","competing risks","longitudinal predictors","external validation"],"falsifier":"Take the trained model and recompute five-year AUC and E/O on the validation sets after replacing the delinquency mean with values measured strictly before the prediction age, for example before first cannabis use. If the AUC drops below about 0.6 or the E/O values move away from 1, the reported prospective accuracy is inflated by look-ahead.","tokens_in":20194,"feed_emoji":"🧠","tokens_out":6559,"duration_ms":62862,"temperature":0.7,"pith_summary":"This paper aims to give clinicians a tool that converts five easily collected facts about a young cannabis user — biological sex, a delinquency measure, and three personality-trait scores — into a personalized probability of developing cannabis use disorder within a chosen time window, such as five years. No absolute-risk model of this kind exists for substance use disorders. The authors argue that the model discriminates reasonably well, with area-under-the-curve values around 0.68 to 0.75 across training and two validation sets, and is well calibrated, with expected-to-observed ratios near 1. If these results hold, a brief questionnaire could flag high-risk adolescents for early intervention.","feed_headline":"Five simple factors forecast cannabis use disorder risk","feed_subtitle":"A five-factor model gives personalized five-year CUD risk, validated on two independent cohorts.","key_machinery":"The load-bearing object is the cause-specific absolute-risk integral from standard competing-risk theory, computed by plugging in a Cox model for the cannabis-use-disorder hazard and an external all-cause mortality hazard. The Cox coefficients are regularized with a Bayesian lasso prior, the baseline hazard is modeled with M-splines, a smooth piecewise-polynomial basis, under a Dirichlet prior, survey weights enter the likelihood, and posterior draws from Markov chain Monte Carlo are averaged to produce predicted risks. The longitudinal predictor delinquency is summarized by its mean across waves, which the authors show performs at least as well as a model-based random intercept and is easier for clinicians to use.","core_discovery":"The paper's central claim is that the absolute risk of cannabis use disorder for an adolescent or young adult who already uses cannabis can be modeled as a function of five risk factors through a Bayesian Cox proportional-hazards model, with the competing risk of death folded in from national life tables. In the final model, male sex, higher delinquency, higher neuroticism, higher openness, and lower conscientiousness each raise the estimated hazard, and the model returns a risk estimate for any user-specified age interval. The authors report that for five-year risk following first cannabis use, the AUC is 0.68 in cross-validation, 0.64 on a held-out national sample, and 0.75 on an external cohort, with expected-versus-observed ratios of 0.95, 0.98, and 1.00. They also present a six-factor variant that adds welfare status for settings where that information exists.","pith_inferences":["If the look-ahead concern I noted is real, the reported AUC and E/O values are upper bounds on prospective accuracy; a fair out-of-time test with baseline delinquency only would likely show somewhat lower discrimination.","The model treats delinquency as fixed at prediction; a dynamic version that updates delinquency and personality measures over follow-up could improve prediction but would require new methodology for time-varying covariates in absolute-risk settings.","Post-legalization cohorts may have different base prevalence and cannabis potency, so absolute risks should be recalibrated before clinical use even though the ranking of individuals may persist.","The same pipeline — Bayesian lasso, M-spline baseline hazard, life-table competing risk — could be turned into absolute-risk models for alcohol or opioid use disorders if suitable longitudinal cohorts exist."],"forward_implications":["A clinician who knows five easy-to-obtain facts about a young cannabis user can state a concrete five-year probability of developing cannabis use disorder, not just a relative-risk ranking.","The model can be recalibrated to a new population using only that population's cannabis-use-disorder prevalence, so it can be ported to other countries or eras of cannabis policy.","Because the risk is absolute and time-bounded, it can be compared directly with the competing risk of death and with risks of other outcomes, supporting decisions about intervention intensity.","The five-factor form makes a paper-and-pencil or short-app risk score feasible for primary-care and school-based settings."],"supporting_citations":[{"why":"The earlier Bayesian logistic CUD risk model this work extends, from which the risk-factor definitions and comparison tables are drawn.","marker":"[16]"},{"why":"Supplies the national longitudinal cohort used for training and for one of the two validation datasets.","marker":"[31]"},{"why":"Supplies the all-cause mortality hazard used as the competing risk in the absolute-risk integral.","marker":"[32]"},{"why":"Supplies the external birth cohort used for independent validation of the final model.","marker":"[33]"},{"why":"Provides the external validation sample and the recalibration procedure used to adapt the model to a higher-prevalence population.","marker":"[48]"},{"why":"Gives the absolute-risk formula with competing risks and the E/O calibration metric used throughout the evaluation.","marker":"[6]"},{"why":"The Cox proportional-hazards model on which the cannabis-use-disorder hazard is built.","marker":"[36]"},{"why":"Provides the M-spline basis used to model the baseline hazard flexibly.","marker":"[38]"},{"why":"Supplies the random-intercept summary for longitudinal predictors that the final model simplifies to a mean.","marker":"[37]"},{"why":"Supplies the scaled-neighborhood variable-selection rule used to choose among competing models.","marker":"[42]"}],"fun_headline_variants":["Bayesian model gives personalized cannabis disorder risk for teens","Five simple traits forecast 5-year cannabis use disorder risk","New tool predicts teen cannabis disorder risk using five factors","Bayesian ML flags teen cannabis use disorder risk with five inputs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes that a young person's delinquency average is available at the moment of prediction, yet the study's predictor construction only required values to be measured before the disorder or censoring, not before the prediction time — so the delinquency score used in some predictions may include behavior that happened after the prediction was supposed to be made.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian model gives personalized cannabis disorder risk for teens","Five simple traits forecast 5-year cannabis use disorder risk","New tool predicts teen cannabis disorder risk using five factors","Bayesian ML flags teen cannabis use disorder risk with five inputs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3288,"prompt_tokens":995,"completion_tokens":2293,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":2227}},"tokens_in":611,"tokens_out":2293,"duration_ms":18128,"temperature":1.0,"reasoning_tokens":2227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:11:30.321925+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the trained model and recompute five-year AUC and E/O on the validation sets after replacing the delinquency mean with values measured strictly before the prediction age, for example before first cannabis use. If the AUC drops below about 0.6 or the E/O values move away from 1, the reported prospective accuracy is inflated by look-ahead.","supporting_citations":[{"cited_title":"The Bayesian elastic net","cited_arxiv_id":null,"evidence_quote":"The earlier Bayesian logistic CUD risk model this work extends, from which the risk-factor definitions and comparison tables are drawn."},{"cited_title":"Underlying cause of death 1999-2020 on cdc wonder online database, 2021","cited_arxiv_id":null,"evidence_quote":"Gives the absolute-risk formula with competing risks and the E/O calibration metric used throughout the evaluation."}],"review_version":1}