{"id":"f2dac7cb-0d07-42d3-ad11-a4f5936ca580","arxiv_id":"2505.15974","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A 12-week RCT found that college students using a smartwatch-based stress app had significantly fewer app-detected stress moments than controls, but no significant between-group differences on standard mental health questionnaires.","lead":"A 12-week randomized trial tested a smartwatch-and-app stress intervention, mHELP, in 117 college students with moderate anxiety. The treatment group showed a significantly larger drop in app-detected Moments of Stress than controls, but no significant between-group differences on standard anxiety, depression, or stress questionnaires.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on an unvalidated, intervention-contaminated outcome: MS counts user-confirmed alerts from the same app that delivers the treatment. Without external validation or a confirmation-free analysis, the reported group-by-time difference cannot be attributed to reduced stress.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing threat: MS is generated by the same app that delivers the intervention and is never externally validated. I agree with that assessment. This is not a statistical nitpick about the Gamma distribution or the unreported shift constant; the decisive problem is that the primary outcome is endogenous to the intervention. The treatment group's app prompts users to confirm detected stress and then guides them through coping exercises, so the act of measurement is part of the intervention. Without a validation study of the detector, or an analysis of pre-confirmation detections or raw physiological data, the central claim that mHELP reduces stress moments cannot be distinguished from the claim that it changes confirmation behavior or app engagement. The trial itself is a real deployed RCT with standard secondary questionnaires, adequate stated power, and a plausible design, which is why rejection should be based on the specific measurement problem rather than on the overall study concept. The proposed concrete test would directly separate the intervention effect on stress from the intervention effect on outcome generation, and it is feasible because the paper states that algorithm-detected events are separately logged in a stress detection column. If that test shows the group-by-time interaction survives, the central claim would be substantially strengthened; if not, the manuscript's headline conclusion is unsupported.","tokens_in":16381,"tokens_out":3622,"duration_ms":37770,"concrete_test":"Re-run the primary GLMM (Formula 2) using only raw algorithm-detected stress events from the app's designated stress detection column, before user confirmation, with the same fixed effects, random effects, and shift constant. If the treatment-by-time interaction is no longer significant or is substantially attenuated, the reported effect depends on user confirmation and engagement behavior rather than on an independently measurable stress response.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The primary outcome, Moments of Stress (MS), is not an independent measure of stress. In the Measures and Metrics section, MS events are defined as algorithm-detected stress moments that are subsequently approved by the user, with self-reported events also aggregated into the daily score. The treatment arm uses the same app that generates the alerts and also receives real-time prompts to confirm detected stress and then perform breathing or focus exercises. This creates a direct pathway for the intervention to mechanically change the outcome: prompts can alter heart rate and attentional state, which can change future algorithm detections, and the confirmation requirement means differential engagement with the app can change the recorded MS count even if underlying stress is unchanged. The manuscript reports no sensitivity, specificity, calibration, or external validation of the proprietary ML detector, and the raw watch HR and accelerometer data are not analyzed as an independent benchmark. Secondary self-report outcomes show no significant group-by-time interaction, which is consistent with the concern that MS reflects app responsiveness rather than a validated stress construct. The key treatment-by-time coefficient (Beta = -6.80, SE = 0.70) therefore cannot be interpreted as evidence of stress reduction until the outcome measure itself is validated or analyzed in a form that does not depend on user confirmation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a 12-week randomized controlled trial (n = 117 completers) comparing a full mHELP intervention to a logging-only control among college students with moderate anxiety. The primary outcome, \"Moments of Stress\" (MS), is defined as daily counts of stress events detected by a proprietary machine-learning algorithm embedded in the same app and confirmed by the user. Using a Gamma Generalized Linear Mixed Model with a log link on a shifted version of MS, the authors report a significantly steeper decline in MS for the treatment group than for the control group. Secondary questionnaire outcomes (GAD-7, PHQ-8, PSS) show no significant group-by-time interactions, although within-group declines for GAD-7 and PSS are interpreted as clinically meaningful. The paper concludes that the mHELP intervention reduces acute stress moments in real-world settings.","tokens_in":16557,"tokens_out":4052,"duration_ms":41931,"significance":"If the primary outcome were valid and independent of the intervention, this would be a valuable contribution to the mHealth and wearable-sensing literature: it reports a longitudinal, naturalistic RCT with passive sensing, a plausible intervention mechanism, and standardized secondary measures. The paper also gives credit for transparently discussing limitations and using a pre-registered-style power analysis. However, the central claim rests entirely on the MS outcome, which is generated by the same system being evaluated and is never validated against an independent stress measure. The secondary, independently validated outcomes do not show between-group effects, so the manuscript as it stands does not support its headline conclusion. The contribution is therefore currently more of an engineering demonstration than a validated efficacy result.","major_comments":[{"comment":"The primary outcome MS is not an independent or validated measure of stress. MS events are defined as moments detected by the app's proprietary machine-learning algorithm and subsequently approved by the user, with self-reported events also aggregated into the daily score. No sensitivity, specificity, calibration, or external validation of the detector is reported, and raw heart-rate or accelerometer data are not analyzed as an independent benchmark. Moreover, the treatment arm uses the same app that generates the alerts and additionally receives real-time prompts to confirm stress and perform breathing or focus exercises; this creates a direct pathway for the intervention to mechanically change the recorded outcome through altered detection and confirmation behavior, independent of any true change in stress. Because the manuscript's central claim is the treatment-by-time effect on MS, this lack of outcome validation and intervention-contamination is load-bearing. The authors should either validate MS against an independent physiological or clinical criterion, or report an analysis that does not depend on user confirmation (e.g., algorithm-only events, raw HR/accelerometer features); without such analysis, the reported p<0.001 cannot be interpreted as evidence of stress reduction.","section":"Measures and Metrics"},{"comment":"The Gamma GLMM is applied to a zero-inflated daily count after adding an unreported \"small constant\" to create MSshift. A Gamma distribution with log link is a model for a continuous, strictly positive response with constant coefficient of variation, not for a discrete count with many zeros; the arbitrary shift constant can materially change the fitted coefficients, yet its value is not reported. The reported unstandardized interaction coefficient of -6.80 for Treatment*Time is on the log scale and is implausibly large for a 12-week study unless time is scaled in a very unusual way; this suggests either a different scaling than described or a misspecification. The authors should report the exact shifted variable, the value of the constant, the time unit, fit diagnostics, and sensitivity analyses using alternatives such as negative binomial, hurdle, or zero-inflated models. Without these, the numerical magnitude and even the direction of the central effect are not reproducible.","section":"Data Analytics and Results (MS)"},{"comment":"The analysis does not establish that the randomization produced comparable groups or that missing data are handled appropriately. The treatment arm (n=86) is much larger than the control arm (n=31), which is expected under 3:1 allocation, but no baseline table reports MS, GAD-7, PHQ-8, or PSS by condition, and no test of baseline imbalance is provided. The text itself notes that the treatment condition had a lower baseline level of the log-transformed outcome, so the group comparison may partly reflect baseline differences rather than intervention effects. In addition, the Results sections report varying N values (125, 126) despite 117 completers, with no statement about missing data or the number of observations in the GLMMs. The authors should report baseline descriptive statistics and tests, the analysis sample, and a missing-data sensitivity analysis; without this, the robustness of the group-by-time interaction is unclear.","section":"Procedure and Data Analysis"}],"minor_comments":[{"comment":"The text refers to the \"binary nature of the data\" after describing MS as a daily aggregate count; these are inconsistent, and the authors should clarify whether the analysis unit is the individual stress event or the daily frequency.","section":"Measures and Metrics"},{"comment":"The simple-slope values reported in the text (treatment beta_std = -0.127, control beta_std = -0.026) do not match the standardized fixed-effect coefficients in Table 1 (beta_std = -0.10 for the interaction and -0.03 for Time); the notation should be made consistent so readers can follow which quantity is being reported.","section":"Data Analytics and Results (MS)"},{"comment":"The model description says random intercepts and slopes for Time were included, but Formula 1 specifies a random slope for condition (u1j cij) and the subsequent variance components are described as \"treatment slope\"; this discrepancy should be corrected.","section":"Data Analytics and Results (MS)"},{"comment":"The clinically meaningful change claims for GAD-7 and PSS are based on within-group modeled changes despite non-significant interaction terms; these should be labeled as exploratory post hoc observations, not as efficacy evidence.","section":"Discussion"},{"comment":"Reference [40] appears to duplicate reference [18]; please consolidate or distinguish them appropriately.","section":"References"}],"recommendation":"reject","confidential_remarks":"The underlying dataset and intervention are potentially valuable, and the secondary questionnaire results are reported honestly. However, the primary outcome is a proprietary, user-confirmed detector that is inseparable from the treatment itself, and no validation or confirmation-free analysis is supplied. This is not a fixable presentation issue; it requires a fundamentally different analysis or external validation of the outcome measure. If the authors can provide such validation and re-run the analysis, a resubmission could be considered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know up front: the paper's central claim is that the mHELP app reduces real-time stress moments, but \"Moments of Stress\" (MS) is defined as events the app's proprietary algorithm detects and the user then confirms. The treatment group is using that same app, with prompts to confirm and then do breathing exercises. That means the intervention can mechanically change the outcome—prompts alter heart rate and attention, and differential engagement changes how many alerts get confirmed—even if underlying stress is unchanged. The paper reports no sensitivity, specificity, calibration, or any external benchmark for the detector. On reading the full text, the stress-test concern holds up: the outcome is not an independent measure of stress.\n\nWhat the paper does well: it is a genuinely naturalistic 12-week RCT with a reasonable sample size, a clear design, and standard secondary questionnaires. The power analysis is reported, the statistical models are described in enough detail to be criticized, and the null secondary outcomes are in the table even though the discussion downplays them. The writing is clear and the trial itself looks adequately powered for its stated design.\n\nThe soft spots are substantial. Besides the contaminated outcome, the treatment group starts at a lower MS baseline, which the authors themselves note, making the steeper decline hard to attribute to the intervention. The Gamma GLMM is fitted to a zero-inflated count after adding an unreported \"small constant,\" and no sensitivity analysis is given for that choice. The secondary outcomes show no significant group-by-time interaction, yet the abstract and discussion highlight \"clinically meaningful\" within-group declines in GAD-7 and PSS. That is overclaiming from non-significant interactions, and the simple slopes on those models are post hoc. The self-report data are the one independent check, and they do not back up the primary claim.\n\nWho this is for: researchers in digital mental health, wearable sensing, and human factors. It is a useful case study in how outcome definition can sink an otherwise sensible trial. It deserves a serious referee rather than a desk rejection—the question is important and the trial was a real investment—but the current manuscript should not be accepted without a much stronger defense of the MS metric: external validation, a confirmation-free analysis, baseline balance checks, and pre-registration. My honest recommendation is to send it to peer review with reviewers instructed to focus on those points, and to be prepared for the central claim to fall apart unless the authors can show the outcome measures stress rather than app engagement.","headline":"A well-run 12-week RCT whose headline result is undermined because the primary outcome is generated by the same app that delivers the intervention, and no validation of that outcome is provided.","tokens_in":17182,"tokens_out":1770,"would_cite":false,"duration_ms":18247,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports a 12-week randomized trial in which college students using a smartwatch-and-machine-learning stress app showed a significantly steeper decline in detected, user-confirmed stress moments than controls who only logged…","keywords":["Machine Learning","Mobile Health","Psychological Stress","Wearable Devices","College Students","Randomized Controlled Trial","Stress Detection","mHealth Intervention"],"falsifier":"Re-analyze the trial's logged data (or run a new trial) separating algorithm alerts that the user confirmed from alerts the algorithm raised regardless of confirmation, and check the detector against a standardized stressor such as the Trier Social Stress Test in the same student population; if the treatment group's advantage disappears for unconfirmed alerts, or the detector's agreement with the stressor is poor, the headline reduction would be an artifact of confirmation behavior rather than a reduction in stress.","tokens_in":16111,"feed_emoji":"⌚","tokens_out":12794,"duration_ms":93987,"temperature":0.7,"pith_summary":"This paper reports a 12-week randomized controlled trial testing whether a smartwatch-based mobile health app, mHELP, can reduce acute stress in college students. The authors claim that students randomized to the full intervention—real-time machine-learning stress detection on an Apple Watch, in-the-moment coping prompts, and optional telehealth counseling—showed a significantly steeper decline in daily Moments of Stress (algorithm-detected, user-confirmed stress events) than a control group that used the app only for logging stress and taking weekly surveys. The trial found this objective moment-level reduction even though weekly questionnaires (GAD-7, PHQ-8, PSS) showed no statistically significant between-group differences; the treatment group did show clinically meaningful improvements in anxiety and perceived stress. The paper argues that wearable-enabled mHealth can reduce acute stress in naturalistic student life and that chronic depression may need longer or more targeted interventions.","feed_headline":"Trial: stress app cuts detected stress moments in students","feed_subtitle":"A 12-week RCT: full mHELP users' confirmed stress events fell faster than those of logging-only controls.","key_machinery":"The active mechanism is the mHELP app paired with an Apple Watch (Series 4/5): heart-rate and accelerometer data are sampled at 1 Hz, and a proprietary machine-learning model flags stress moments in real time; the user confirms each flag, and the app immediately offers coping tools such as breathing and focus exercises, with links to telehealth counseling. The statistical machinery that produces the headline result is a generalized linear mixed model with a Gamma distribution and log link, random intercepts and slopes per participant, and a treatment-by-time interaction, which estimates whether the rate of decline in daily stress moments differs between the two groups.","core_discovery":"Over 12 weeks, 117 college students with at least moderate anxiety (GAD-7 score of 7 or higher) were randomized 3:1 to the full mHELP intervention or to a logging-only control. The primary outcome, Moments of Stress, declined in both groups, but the decline was significantly steeper for the treatment group (treatment simple slope $\\beta_{\\mathrm{Std}} = -0.127\\,(0.013)$, $t(81) = -9.74$, $p < .001$; control $\\beta_{\\mathrm{Std}} = -0.026\\,(0.008)$, $t(81) = -3.29$, $p < .001$; treatment-by-time interaction $p < .001$). Weekly self-reported anxiety (GAD-7) and perceived stress (PSS) also declined, with the treatment group crossing clinically meaningful thresholds (about a 5.25-point GAD-7 drop and just over a 28% PSS drop), though the between-group interaction terms were not significant; depression (PHQ-8) did not improve. The authors conclude that the intervention reduced acute, moment-level stress responses in a real-world student setting, and that the divergence between real-time and weekly measures reflects a difference between acute stress and longer-term psychological state.","pith_inferences":["Editorial inference: the paper's primary outcome conflates the algorithm's detection with the user's confirmation; separating raw alerts from confirmations in future analyses would reveal whether the intervention lowers physiological stress events or changes users' willingness to label moments as stressful.","Editorial inference: if the proprietary detector were validated against a standardized stressor and against unconfirmed alerts, Moments of Stress could become a reusable outcome for other just-in-time mHealth interventions aimed at panic, cravings, or pain episodes.","Editorial inference: because the control arm was an active logging condition rather than a no-contact condition, the between-group comparison isolates the added value of the intervention; a wait-list control would likely show a larger effect, so the paper's estimate may be conservative."],"forward_implications":["A wearable-enabled mHealth app can produce a measurable reduction in moment-level acute stress over a semester, not just in self-reported symptoms.","The logging-only control group also showed downward trends in anxiety and stress, suggesting that weekly self-monitoring alone may carry some benefit and makes between-group differences harder to detect.","Depressive symptoms, which started near the mild range, did not respond within 12 weeks; testing the intervention longer or with depression-specific content is needed before concluding it works for depression.","Because the treatment-control separation grows with time (significant interaction), longer deployments should show larger acute-stress benefits if the effect is real."],"supporting_citations":[{"why":"Scoping review establishing that wearable stress-management interventions can mitigate acute stress, the prior evidence this trial's MS finding aligns with.","marker":"[16]"},{"why":"Meta-analysis of digital mental health interventions for university students that supplies expected effect sizes and supports the DMHI comparison context.","marker":"[17]"},{"why":"Evaluation of physiological stress-detection model reproducibility; the lab-detection baseline that the naturalistic ML detection extends.","marker":"[23]"},{"why":"GAD-7 instrument used as the weekly anxiety outcome and clinical-threshold benchmark.","marker":"[28]"},{"why":"PHQ-8 instrument used as the weekly depression outcome.","marker":"[29]"},{"why":"Psychometric review of the Perceived Stress Scale used as the weekly stress outcome.","marker":"[30]"},{"why":"Smartwatch-based hyperarousal event detection work that the mHELP stress-detection algorithm builds on.","marker":"[32]"},{"why":"Provides the PSS minimal clinically significant change threshold used to claim the treatment group's stress decline was meaningful.","marker":"[38]"},{"why":"Establishes the GAD-7 minimal clinically important difference used to interpret the treatment group's modeled anxiety decline.","marker":"[39]"}],"fun_headline_variants":["App cuts stress moments in students, but not depression","Wearable app reduces acute stress events in 12-week trial","Stress-detecting app lowers moment-level stress in RCT","mHELP app: fewer stress episodes, unchanged depression scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central finding rests on the assumption that the app's stress alerts, after the user confirms them, are a true measure of stress rather than a measure of how much the user engages with the app; the paper reports no check of the detector against a known stress test.","fun_headline_variants_meta":{"raw":{"variants":["App cuts stress moments in students, but not depression","Wearable app reduces acute stress events in 12-week trial","Stress-detecting app lowers moment-level stress in RCT","mHELP app: fewer stress episodes, unchanged depression scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000495,"raw_usage":{"total_tokens":2514,"prompt_tokens":1113,"completion_tokens":1401,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":729,"completion_tokens_details":{"reasoning_tokens":1333}},"tokens_in":729,"tokens_out":1401,"duration_ms":10431,"temperature":1.0,"reasoning_tokens":1333,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:09:34.618065+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the trial's logged data (or run a new trial) separating algorithm alerts that the user confirmed from alerts the algorithm raised regardless of confirmation, and check the detector against a standardized stressor such as the Trier Social Stress Test in the same student population; if the treatment group's advantage disappears for unconfirmed alerts, or the detector's agreement with the stressor is poor, the headline reduction would be an artifact of confirmation behavior rather than a reduction in stress.","supporting_citations":[{"cited_title":"Review of the Psychometric Evidence of the Perceived Stress Scale,","cited_arxiv_id":null,"evidence_quote":"Psychometric review of the Perceived Stress Scale used as the weekly stress outcome."},{"cited_title":"Wearables for Stress Management: A Scoping Review,","cited_arxiv_id":null,"evidence_quote":"Scoping review establishing that wearable stress-management interventions can mitigate acute stress, the prior evidence this trial's MS finding aligns with."},{"cited_title":"Posttraumatic stress disorder hyperarousal event detection using smartwatch physiological and activity data,","cited_arxiv_id":null,"evidence_quote":"Smartwatch-based hyperarousal event detection work that the mHELP stress-detection algorithm builds on."},{"cited_title":"Cross-cultural adaptation and validation of the Danish consensus version of the 10-item Perceived Stress Scale,","cited_arxiv_id":null,"evidence_quote":"Provides the PSS minimal clinically significant change threshold used to claim the treatment group's stress decline was meaningful."},{"cited_title":"Sensitivity to change and minimal clinically important difference of the 7-item Generalized Anxiety Disorder Questionnaire (GAD-7),","cited_arxiv_id":null,"evidence_quote":"Establishes the GAD-7 minimal clinically important difference used to interpret the treatment group's modeled anxiety decline."}],"review_version":1}