{"id":"d9055024-7f11-42c1-9534-afbc138fde1f","arxiv_id":"2506.15834","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A multi-objective function combining predicted EMA response likelihood and emotion-model uncertainty could select delivery times that favor responsive moments and less common emotion scores, based on offline analysis of two wearable datasets.","lead":"Researchers propose a smart EMA trigger that sends mobile surveys when a person is likely to respond and when the emotion model is uncertain, aiming to capture rarer emotional states. Offline tests on two wearable datasets suggest it would favor responsive moments and less common positive-affect scores, but the simulation's improvement numbers rely on the models' own predictions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ3's simulated receptivity gains are nearly tautological: the Smart Trigger maximizes the same R(t) used to compute response rates, so the reported 0.93 vs 0.82 improvement does not yet establish real-world benefit.","rationale":"The reader's weakest assumption identified exactly this circularity: using predicted responses from the same models that the trigger optimizes. My analysis confirms this is the most load-bearing concern because it directly undermines the paper's quantitative claim that the Smart Trigger improves receptivity and emotional diversity. RQ1 and RQ2 provide label-grounded evidence that J is higher during responsive moments and extreme positive-affect scores, so the core feasibility idea is plausible and the paper should not be rejected outright. However, the RQ3 numbers are structurally inflated and should not be cited as evidence until validated on real held-out EMA responses or reframed as a model-internal diagnostic. The unreported weights are a secondary but important reproducibility gap that compounds the problem. Because my concern reinforces the reader's conditional verdict rather than changing it, I leave the verdict as CONDITIONAL (UNCHANGED).","tokens_in":19849,"tokens_out":4604,"duration_ms":50971,"concrete_test":"Hold out a random subset of participants (or time points) for which actual EMA response labels and PA scores exist. Re-run the RQ3 simulation by selecting delivery times with J (Eq. 1), but compute each trigger's response rate using the held-out real response labels rather than the receptivity model's predicted R(t). If the Smart Trigger's real-label response rate does not significantly exceed the random trigger's, the RQ3 improvement is an artifact of circular evaluation. As a secondary check, report the exact w_u and w_r values used and recompute RQ1-RQ3 under a grid of weights to assess sensitivity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central feasibility claim rests on RQ1 and RQ2, which use actual EMA labels and are reasonably supportive. The load-bearing problem is RQ3 (§3.3.6, §4.2.3). In the simulation, the Smart Trigger selects the time point maximizing J = w_u·U(t)^2 + w_r·R(t)^2 (Eq. 1), where R(t) is the receptivity model's predicted probability of response. The paper then 'utilize[s] our receptivity and emotion recognition models to obtain predicted values for responsiveness (response or non-response) and emotion (PA score) at each selected time point' and calculates response rates from those predicted values. Because the trigger explicitly maximizes R(t), the selected time points will have higher predicted R(t) than random time points by construction. The reported gains (ADRD: 0.93 vs 0.82; Healthy: 0.94 vs 0.87) are therefore largely a tautological artifact of optimizing and then evaluating on the same model output, not evidence about actual EMA response behavior. The same circularity inflates the within-participant PA variance results, since the trigger maximizes U(t), the model's uncertainty, which naturally leads to more extreme predicted PA values. The paper itself acknowledges offline evaluation is speculative (§5.3.1), but the issue is stronger than speculation: the RQ3 numbers are structurally guaranteed to favor the Smart Trigger. Additionally, the weights w_u and w_r in Eq. 1 are never reported, so J is not fully defined and even RQ1/RQ2 cannot be reproduced or checked for sensitivity. Without either real-label validation of RQ3 or the actual weights, the headline quantitative improvements are not yet trustworthy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a multi-objective function J = max_{t in T} w_u U(t)^2 + w_r R(t)^2 for deciding when to deliver ecological momentary assessments (EMAs), where R(t) is a receptivity model's predicted response probability and U(t) is an emotion recognition model's predictive uncertainty. The intended effect is to prompt when participants are likely to respond and when their emotional state is underrepresented. The authors evaluate the idea offline on two datasets (73 ADRD caregivers and 45 healthy participants) through three research questions: RQ1 tests whether J is higher at times with actual EMA responses using mixed-effects models; RQ2 tests whether J is associated with participant-specific absolute z-scores of positive affect; RQ3 simulates a 'Smart Trigger' versus a random trigger and compares predicted response rates and predicted PA distributions. The paper concludes that the multi-objective function would improve receptivity and capture a broader emotional range.","tokens_in":20175,"tokens_out":4071,"duration_ms":47213,"significance":"The paper addresses a real problem: ML-based EMA scheduling that maximizes receptivity alone can bias collected emotions, because receptivity is correlated with affect. The proposed balance of receptivity and uncertainty is sensible, and the authors provide a genuinely useful offline evaluation framework for RQ1 and RQ2 using actual EMA response labels and PA labels. Modeling comparisons across algorithms and two populations add credibility. If the RQ3 simulation were replaced by an evaluation using actual labels, the contribution would be relevant to mHealth and affective computing. However, the current RQ3 results cannot be taken as evidence of improved compliance or broader emotion capture.","major_comments":[{"comment":"In RQ3, the Smart Trigger selects the time point that maximizes J = w_u U(t)^2 + w_r R(t)^2 (Eq. 1), and the response rate is then computed from the same receptivity model's predicted response values at that time point, not from actual EMA responses. Because J is increasing in R(t), the selected time point will have a higher predicted response probability than a random time point by construction, so the reported gains (0.93 vs 0.82 for ADRD; 0.94 vs 0.87 for Healthy) are largely a tautological artifact. Similarly, the PA-distribution and within-participant variance comparisons rely on emotion predictions at the selected times, so they do not establish that real emotion reports would be more diverse. The offline limitation is acknowledged in §5.3.1, but this is not merely speculation: the comparison is structurally guaranteed to favor the Smart Trigger. Please validate the simulation against actual EMA labels (e.g., limiting to prompt occasions where a real response outcome exists) or reframe RQ3 explicitly as 'expected behavior under the fitted models' with appropriate caveats and a sensitivity analysis.","section":"§3.3.6, §4.2.3"},{"comment":"The weights w_u and w_r are never reported, even though every result in RQ1–RQ3 depends on J constructed with specific values. Without these values, the analyses cannot be reproduced and the results cannot be checked for sensitivity to the weighting. Please report the weights used in all experiments, and ideally a sensitivity analysis over a range of weights (the paper already suggests adaptive weights in §5.2).","section":"§3.1.4, Eq. (1)"}],"minor_comments":[{"comment":"Please state the threshold used to convert the receptivity model's predicted probability into binary response/non-response for the simulated response rate.","section":"§3.3.6"},{"comment":"Define U(t) explicitly; the text says the variance across 200 stochastic forward passes is used, but U(t) is introduced only in Eq. 1 without a formal definition.","section":"§3.1.3, §3.1.4"},{"comment":"The repeated-measures ANOVA is non-significant for the Healthy study; please report the effect size rather than attributing the null result to sample size.","section":"§4.2.1"},{"comment":"The caption says 'PS' where 'PA' is meant.","section":"Figure 5"},{"comment":"No data or code availability statement is included; given the reproducibility concerns about the weights, an explicit statement is needed.","section":"§7 and reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for the journal; the RQ3 circularity is the key obstacle. I would be willing to review a revision that reworks RQ3 around actual labels and reports the weights used in the objective function."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2506.15834.\n\nFirst, the core idea is worth taking seriously. Scheduling EMAs on predicted receptivity alone biases the collected labels toward the emotional states that make people compliant. Adding model uncertainty as a second objective is a sensible, simple fix, and nobody has applied exactly that combination to EMA delivery before. The offline evaluation with two datasets (91 caregivers, 45 healthy controls) is a reasonable first test.\n\nSecond, the paper has genuine label-grounded evidence for the idea: RQ1 and RQ2 use actual EMA responses and reported positive affect, and mixed-effects models show that J is significantly higher during responsive moments and at PA values far from a participant's mean. Those analyses are the paper's real contribution. The neural network with MC dropout for uncertainty is standard but fine, and the model comparisons are thorough.\n\nThe problem is RQ3. The Smart Trigger selects times by maximizing J, which includes R(t), and then the 'response rate' is computed from the same model's predicted response probabilities at those times. It is structurally guaranteed to beat a random trigger. The same circularity inflates the PA variance results: U(t) is part of J, and the variance is computed from predicted PA values. The paper's limitations section calls the results speculative, but that is too generous—these numbers are not just unvalidated, they are an artifact of the evaluation procedure. A holdout of real responses (or at least evaluation on a separate model) is needed before reporting 0.93 vs 0.82 as a benefit.\n\nThere are two smaller issues. The weights w_u and w_r in Eq. 1 are never reported, so J is not fully defined and even the good RQ1/RQ2 results can't be reproduced or sensitivity-checked. And the ADRD RQ2 coefficient is 0.08 with 95% CI [0.095, 0.11]—the lower bound exceeds the estimate; likely a typo, but it should be fixed.\n\nBottom line: a conditional accept. The feasibility claim partly holds up (RQ1/RQ2), but the headline improvements don't yet. A serious referee should require reporting the weights and redoing RQ3 with actual response labels, or at least rephrasing it as a simulation bound. If that gets done, this becomes a solid contribution. Worth bringing to a reading group for the methodology discussion, but not for the RQ3 numbers.","headline":"Sensible idea and solid label-grounded feasibility evidence in RQ1/RQ2, but RQ3's headline gains are circular—computed from the same models the trigger optimizes—so treat those numbers as artifacts until validated on real responses.","tokens_in":20758,"tokens_out":2996,"would_cite":true,"duration_ms":33770,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EMA scheduling can be recast as maximizing a weighted combination of predicted response likelihood and emotion-prediction uncertainty, and offline evidence on two datasets indicates this raises receptivity while capturing less-common…","keywords":["ecological momentary assessment","EMA compliance","receptivity prediction","model uncertainty","emotion recognition","multi-objective optimization","mobile health","Monte Carlo dropout"],"falsifier":"Run a randomized field experiment comparing the Smart Trigger with a random trigger for one to two weeks per participant, tracking actual EMA response rates and the variance of the positive-affect scores actually reported; the central claim collapses if the measured compliance difference is not significant or if the smart-trigger group does not show greater within-person emotion variance, and it is also weakened if participants' actual responses at smart-selected times are systematically lower than the model predicted.","tokens_in":19638,"feed_emoji":"📱","tokens_out":7779,"duration_ms":84906,"temperature":0.7,"pith_summary":"Mobile health studies rely on Ecological Momentary Assessments (EMAs), short in-the-moment surveys, and low response rates starve emotion-recognition models of ground truth. Machine-learned triggers that fire only when a response is likely can nudge sampling toward emotions that happen to co-occur with responsiveness, narrowing the emotional range collected. This paper proposes a multi-objective score $J = \\max_{t} w_u U(t)^2 + w_r R(t)^2$ that combines predicted response likelihood $R(t)$ with uncertainty $U(t)$ of an emotion-prediction model, so prompts are sent when people are both likely to answer and experiencing emotions the model rarely sees. In offline evaluation on 91 spousal caregivers of people with Alzheimer's disease and related dementias (73 with usable labels) and 45 healthy participants, the score is significantly higher at actual responses, higher when reported positive affect deviates from a participant's typical range, and a simulated trigger outperforms random scheduling on compliance and within-person emotion variance. A reader should care because this is a modular way to improve the quality of the subjective data that downstream emotion models depend on, without changing survey content or incentives.","feed_headline":"Dual-rule trigger lifts simulated EMA response rates to 0.93","feed_subtitle":"Timing surveys by response likelihood plus emotion-model uncertainty also broadens the emotions captured.","key_machinery":"The central object is the multi-objective function $J=\\max_{t\\in T} w_u U(t)^2 + w_r R(t)^2$, a weighted sum of squared outputs from two learned models: a binary receptivity classifier (neural network for ADRD, random forest for Healthy) and a regression neural network with Monte Carlo dropout whose prediction variance serves as the uncertainty term $U(t)$. The paper's key design choice is using uncertainty rather than the emotion prediction itself: because positive-affect scores cluster near the participant mean, high-uncertainty moments are assumed to coincide with underrepresented emotional states, counteracting the bias a purely receptivity-driven trigger would introduce. The evaluation machinery is semi-personalized cross-validation, which trains on other participants first and then adds the target participant's earlier days, mimicking how a deployed system would adapt; the offline trigger evaluation compares the time point that maximizes $J$ against a randomly chosen time point within each three-hour window.","core_discovery":"The paper's claim is that EMA delivery can be treated as a maximization over candidate times of a weighted combination of receptivity and emotion-prediction uncertainty. The receptivity model outputs the probability $R(t)$ that the participant answers a prompt at time $t$; the emotion model, a dropout-based neural network regression over positive affect (PA), outputs a mean prediction and variance $U(t)$ from 200 stochastic forward passes. The function $J = \\max_{t\\in T} w_u U(t)^2 + w_r R(t)^2$ then selects the moment that balances the two goals. In both datasets, mixed-effects models show $J$ is a significant positive predictor of actual response labels (RQ1) and of the absolute participant-specific z-score of PA (RQ2), meaning higher $J$ aligns with reported emotions farther from the participant's mean. The offline simulation (RQ3) predicts average receptivity rates of 0.93 vs 0.82 (ADRD) and 0.94 vs 0.87 (Healthy) for Smart Trigger vs random, with statistically larger within-participant variance in predicted emotion scores, indicating the trigger would collect a broader emotional range while improving compliance.","pith_inferences":["Editorial inference: the same uncertainty-seeking objective could generalize to other skewed constructs such as stress or negative affect, or to intervention delivery, provided the prediction model is accurate enough; the paper reports that negative affect was too skewed to model well, so this extension is conditional on model quality.","Editorial inference: the RQ3 comparison is an in-silico upper bound because the Smart Trigger selects times using the receptivity model and is then evaluated on those same predicted labels; a prospective randomized trial is needed to confirm the rate gap.","Editorial inference: a third term measuring the gap between predicted and participant-average emotion could be added to the function to counter social-desirability bias, since the paper identifies that bias as a reason negative affect lacks high-intensity reports.","Editorial inference: because the function is evaluated only through model outputs, it can be re-derived for just-in-time adaptive interventions where the trigger must decide within minutes rather than across a scheduled window."],"forward_implications":["Deployed in an EMA app, the rule would schedule prompts without changing survey length, frequency, or incentives, sidestepping the compliance trade-offs of those alternatives.","The weights $w_u$ and $w_r$ let researchers favor compliance in less-adherent populations or emotional coverage in emotion-model training studies, and the paper suggests adaptively tuning them per participant.","If the offline response-rate estimates hold prospectively, a study using the Smart Trigger could expect receptivity gains on the order of 0.82 to 0.93 (caregivers) and 0.87 to 0.94 (healthy adults) relative to random timing.","Capturing more extreme positive-affect scores would give emotion-recognition models training labels at the edges of the distribution, where they currently have little supervision.","Because the function is modular, new constructs could be added as additional weighted terms, extending the same scheduling logic to other target variables or populations."],"supporting_citations":[{"why":"Reports the link between emotional state and EMA receptivity and warns that response-likelihood triggers bias collected emotions; defines the problem the paper addresses.","marker":"[22]"},{"why":"Supplies the cost-sensitive active-learning formulation that the multi-objective function adapts, combining uncertainty with likelihood.","marker":"[24]"},{"why":"Existing machine-learned receptivity model for just-in-time interventions; the approach the paper extends by adding emotion uncertainty.","marker":"[30]"},{"why":"Shows contextual cues can indicate when to deliver EMAs; a comparison baseline for ML-based EMA scheduling.","marker":"[31]"},{"why":"State-of-receptivity model for mHealth interventions whose reported F1 the receptivity model outperforms; supplies algorithm and feature context.","marker":"[23]"},{"why":"Identifies emotional state as one of the components shaping perceived interruption burden, motivating receptivity modeling.","marker":"[14]"},{"why":"Documents prompt-level links between emotional state and EMA compliance, supporting the premise that emotion-aware triggers are needed.","marker":"[34]"}],"fun_headline_variants":["Dual-goal EMA timing lifts response rates to 0.93","Uncertainty-guided surveys boost replies and emotion range","Smart trigger improves EMA compliance and captures rarer states","Offline study: optimal EMA times heighten response and variety"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the receptivity model's predicted response labels and the emotion model's predicted scores at unobserved time points are faithful stand-ins for what participants would actually do and feel; if those predictions drift from reality, the simulated compliance gains and emotion-coverage gains are artifacts of evaluating the trigger on the same model that chose the trigger's times.","fun_headline_variants_meta":{"raw":{"variants":["Dual-goal EMA timing lifts response rates to 0.93","Uncertainty-guided surveys boost replies and emotion range","Smart trigger improves EMA compliance and captures rarer states","Offline study: optimal EMA times heighten response and variety"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1397,"prompt_tokens":1062,"completion_tokens":335,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":678,"completion_tokens_details":{"reasoning_tokens":266}},"tokens_in":678,"tokens_out":335,"duration_ms":4645,"temperature":1.0,"reasoning_tokens":266,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:50:54.600179+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a randomized field experiment comparing the Smart Trigger with a random trigger for one to two weeks per participant, tracking actual EMA response rates and the variance of the positive-affect scores actually reported; the central claim collapses if the measured compliance difference is not significant or if the smart-trigger group does not show greater within-person emotion variance, and it is also weakened if participants' actual responses at smart-selected times are systematically lower than the model predicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports the link between emotional state and EMA receptivity and warns that response-likelihood triggers bias collected emotions; defines the problem the paper addresses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the cost-sensitive active-learning formulation that the multi-objective function adapts, combining uncertainty with likelihood."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Existing machine-learned receptivity model for just-in-time interventions; the approach the paper extends by adding emotion uncertainty."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows contextual cues can indicate when to deliver EMAs; a comparison baseline for ML-based EMA scheduling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"State-of-receptivity model for mHealth interventions whose reported F1 the receptivity model outperforms; supplies algorithm and feature context."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents prompt-level links between emotional state and EMA compliance, supporting the premise that emotion-aware triggers are needed."}],"review_version":1}