{"id":"e04d25b5-730b-41a4-bd2a-ecd87ac7cb69","arxiv_id":"2501.09530","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"In an eight-month Singapore deployment, smartwatch JITAI nudges were associated with self-reported increases in comfort-related behavior changes over the first three weeks, though causality is not established.","lead":"Researchers tested a smartwatch app that sends nudges about heat and noise to 103 people in Singapore. Participants reported more helpfulness and more behavior changes over the first three weeks, but the study had no control group.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported behavioral 'increases' may be an artifact of changing response-category definitions across weeks; the abstract's effect-size ranges are not tied to a fixed metric.","rationale":"I read the paper in good faith as a feasibility and deployment study, and I credit the authors for releasing data and code and for acknowledging the lack of a control group and objective outcome measures in Sections 4.7 and 4.8. The reader's identified weakest assumption, that observed increases might be caused by ambient variation or app familiarity rather than by the messages, is real. However, the more load-bearing problem is more basic: the reported increases may not even be measured consistently. Because the combined response categories differ across weeks, the apparent upward trends could be produced solely by counting a larger set of answer options in later weeks. This is an internal consistency issue that must be settled before any causal or attributional claim can be evaluated. The reader did not flag this changing-category problem, so I disagree with the identification of the weakest assumption. My proposed recomputation is straightforward and can be run immediately from the provided repository. I keep the verdict CONDITIONAL rather than moving to REJECT because a fixed-definition analysis might still support a positive effect; but the condition for acceptance must now include reproducing the headline percentages from a single consistent outcome definition, not merely adding a control arm.","tokens_in":18148,"tokens_out":14237,"duration_ms":139021,"concrete_test":"Using the open GitHub dataset, reconstruct the weekly contingency tables for the three behavioral questions in Figure 9 and recompute every week with one fixed outcome definition, e.g., the share answering 'Always', 'Often', or 'Sometimes'; also report the disaggregated Always/Often/Sometimes counts. Then compute Week 1-to-Week 3 percentage-point changes per phase and pooled. If the fixed-definition changes do not reproduce the abstract's ranges, or if Phase 2 thermal adjustment is still negative, the central claim is an artifact of changing response buckets and must be corrected or withdrawn.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central effect claims are not computable as stated. In Section 3.1, the combined response buckets change between weekly surveys. For Phase 1 earphone use, Week 1 combines 'Always and Sometimes', Week 2 combines 'Often and Sometimes', and Week 3 combines 'Always, Often and Sometimes'. For Phase 1 location changes and thermal adjustments, Week 1 and Week 2 use 'Often and Sometimes' while Week 3 adds 'Always'. Adding response categories in later weeks mechanically inflates the combined percentage even if no participant changed behavior. The abstract and conclusion ranges (4-11%, 2-17%, 3-13%) are therefore not tied to a fixed behavioral outcome definition. Moreover, Phase 2 thermal adjustments decline from 31% to 29% across Weeks 1-3 in the text, while the abstract claims a 3-13% increase. If the open dataset can produce these ranges from a fixed response category, that calculation is not shown in the paper; if it cannot, the headline effect sizes must be revised or removed.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a Just-in-Time Adaptive Intervention (JITAI) framework built on the Cozie Apple smartwatch platform, deployed for eight months in Singapore with 103 participants. The system collects micro-survey responses, physiological data, and weather data, and delivers threshold-based (temperature >30 °C, noise >70 dBA) or personalized Random-Forest-triggered intervention messages. The paper reports weekly-survey results over the first three weeks of two deployment phases, claiming increases in perceived usefulness and in self-reported noise- and thermal-related adaptive behaviors (location changes, earphone use, thermal adjustments), plus associations with personality, gender, and outdoor preference, and a spatial analysis of message delivery. The authors provide open data and code for reproducibility.","tokens_in":18276,"tokens_out":9117,"duration_ms":77744,"significance":"If the behavioral trends were valid, this would be one of the first field demonstrations of JITAI in the built environment, with practical implications for personalized environmental nudging and occupant-centric building controls. The study's strengths are the large longitudinal deployment (103 participants, >12,000 micro-surveys, >3,600 messages), the open-source platform, and the explicit reproducibility statement. However, the headline behavioral effect sizes are not supported by the reported analysis because response-category definitions and denominators change across weeks, the ranges in the abstract are not derivable from the text, and the design (no control group, self-report only) cannot support causal claims, as the authors themselves acknowledge in Sections 4.7 and 4.8.","major_comments":[{"comment":"The weekly behavioral percentages are computed with different Likert-category combinations across weeks, making the time trends invalid. For example, Phase 2 location changes are 'Sometimes' alone in Week 1 (18%), 'Often and Sometimes' in Week 2 (22%), and 'Always, Often and Sometimes' in Week 3 (26%); Phase 1 earphone use is 'Always and Sometimes' in Week 1 (10%), 'Often and Sometimes' in Week 2 (14%), and three categories in Week 3 (31%). Adding response categories mechanically inflates the combined percentage even when no participant changes behavior. The abstract's and conclusion's ranges (4-11%, 2-17%, 3-13%) therefore do not measure a consistent behavioral outcome. Please recompute all trends with a fixed category definition (e.g., 'Often or more') and report those values, or remove the ranges.","section":"Section 3.1, Figure 9"},{"comment":"The effect-size ranges in the abstract and conclusion are not consistent with the week-by-week numbers in Section 3.1. For instance, Phase 1 location changes go from 10% (Week 1) to 26% (Week 3) and Phase 2 from 18% to 26%; Phase 1 earphone use goes from 10% to 31% and Phase 2 from 17% to 33%; Phase 1 thermal adjustments go from 10% to 26% while Phase 2 goes from 31% to 29%. None of these differences matches the claimed '4-11%', '2-17%', or '3-13%' ranges, and no computation is shown. Please state explicitly how each range was derived, or correct the abstract and conclusion to match the reported data.","section":"Abstract and Section 5"},{"comment":"The percentages are also computed on different denominators across weeks: the weekly N changes (e.g., Phase 1 Week 1 N=39 vs. Week 2 N=48; Phase 2 Week 1 N=45 vs. Week 2 N=55), and the stacked bars in Figure 9 include 'No response' and 'No intervention messages received' categories. Thus a rising percentage of a given response category may reflect changes in response rate or message receipt rather than behavior change. The text should report the denominators and either exclude non-respondents consistently or analyze response rates separately.","section":"Section 3.1, Figure 9"},{"comment":"The causal framing is not supported by the design. The authors acknowledge in Section 4.7 that 'we cannot definitively determine whether the increased behavioral responses over time were driven by exposure to the intervention alone or by changes in ambient discomfort,' and Section 4.8 states that the framework 'lacked measurable observation that would provide objective evidence of a behavior change' and recommends a no-intervention control. Given these limitations, statements in Section 3.1 such as 'the intervention messages were effective in encouraging participants to adopt earphone use' and the title's 'Nudging' overstate what the data can establish. Please revise the abstract, Section 3.1, and Section 5 to present the results as descriptive trends and to consistently qualify any effectiveness language with the acknowledged design limitations.","section":"Sections 4.7, 4.8, and 5"},{"comment":"The Phase 1 versus Phase 2 comparison is confounded by the different message-delivery schedules. According to Figure 13, Phase 1 threshold-based JITAIs were sent only after 50 micro-surveys, whereas Phase 2 threshold messages were sent until 50 micro-surveys and personalized messages thereafter; moreover, only 19 of 55 Phase 2 participants received any personalized message. The weekly behavior percentages in Section 3.1 are phase-level aggregates that do not condition on whether, when, or how many intervention messages each participant received, so differences between phases cannot be attributed to the threshold versus personalized mechanism. Please report the trends separately for message-recipient subgroups or otherwise account for exposure.","section":"Section 4.1, Figure 13"}],"minor_comments":[{"comment":"The caption of Figure 13a states that threshold-based JITAIs were 'only sent after the participant submitted 50 micro-surveys' in Phase 1, which contradicts Section 2.3 and Figure 4, where Phase 1 sends threshold-based JITAI messages from the beginning of data collection. Please correct the caption.","section":"Figure 13 caption"},{"comment":"The ordinal logistic regression and chi-square analyses are mentioned but no coefficients, confidence intervals, or p-values are reported in the text or figures; without these, the reader cannot evaluate the strength of the personality, gender, and preference associations.","section":"Section 3.2"},{"comment":"The word 'significantly' is used in a descriptive context ('the percentage increased significantly to 22%') without any statistical test; please rephrase to report the observed percentage.","section":"Section 3.1"},{"comment":"The caption says participants who enjoy being outdoors 'were less likely to make environmental adjustments,' while the text (Section 3.2) says they 'appeared to be more likely to adjust their location or thermostat for thermal comfort.' One of these statements is inverted; please correct the inconsistency.","section":"Figure 11 caption"},{"comment":"The personalized Random Forest model description lacks hyperparameter details and any validation accuracy; since Phase 2's personalized mechanism is central to the paper, a brief report of model performance (e.g., cross-validated accuracy) would support reproducibility.","section":"Section 2.5"},{"comment":"The sentence 'All participants in this study received the same interventions' is imprecise because Phase 2 participants received personalized messages after 50 micro-surveys while Phase 1 participants did not; clarify that the message content was the same but the trigger mechanism differed.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The open data and reproducibility commitment are commendable and should be preserved. The main risk is that the headline effect sizes are not reproducible from the reported numbers; if the reanalysis with fixed response categories yields materially different trends, the abstract's claims must be revised accordingly. The lack of a control group is a design limitation that the authors already acknowledge; I would not reject on that basis alone, but the causal language needs to be toned down."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a serious feasibility study with a large, open dataset, but the headline effect-size claims are not reliable. The new content is the 103-participant, eight-month deployment of smartwatch JITAI for heat and noise in Singapore, with a threshold-based versus personalized-trigger comparison. The JITAI concept and Cozie platform were published earlier by the same group; this paper adds the field data and weekly-survey analysis.\n\nWhat the paper does well: the platform description is clear, the data and code are openly available, and the limitations sections (4.7 and 4.8) are candid about the lack of a control group and objective behavior measurement. The spatial analysis is a nice addition, and the conscientiousness/gender explorations are framed appropriately as exploratory.\n\nThe soft spots are substantial. Most importantly, the abstract's percentage ranges (4-11%, 2-17%, 3-13%) are not computable from the weekly percentages reported in Section 3.1. The combined response buckets change across weeks — for example, Week 1 might combine \"Always and Sometimes\", Week 2 \"Often and Sometimes\", Week 3 \"Always, Often and Sometimes\" — so the serial percentages do not measure a fixed behavioral outcome. Adding categories in later weeks mechanically inflates the combined percentage even if behavior is unchanged. The clearest red flag: Phase 2 thermal adjustments are reported as 31%, 31%, and 29% across Weeks 1-3 in the text, yet the abstract claims a 3-13% increase in thermal adjustments. Unless the open dataset contains a different, fixed-category calculation, the headline ranges should be recomputed or removed. The paper also has no control group, no objective behavior measurements, and no confidence intervals or p-values for the main percentages, so the causal framing in the abstract overreaches. The authors themselves acknowledge this in Sections 4.7 and 4.8. Finally, the personalized-versus-threshold comparison is underpowered: only 19 of 55 Phase 2 participants received personalized messages, so that comparison is mostly descriptive.\n\nOverall, this is a useful feasibility study and a valuable platform/data resource. The central claim that JITAI nudges changed behavior is not supported as currently presented, but the deployment itself is real and the authors are transparent about many of its limits. A serious referee should engage with it, but the authors need to fix the effect-size calculations, report per-category proportions on a consistent definition, and soften the causal interpretation. I would send it to peer review with that expectation.","headline":"Useful feasibility study with open data, but the abstract's effect-size ranges do not survive a close read of Section 3.1 and one range contradicts the body text.","tokens_in":18868,"tokens_out":3123,"would_cite":true,"duration_ms":29365,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Context-triggered smartwatch messages can nudge people to change location, use earphones, and adjust their environment for heat and noise comfort.","keywords":["Thermal comfort","Noise","Distraction","Wearables","Occupant behavior","Just-in-time adaptive interventions","Smartwatch","Digital twin"],"falsifier":"A controlled trial in which a randomly chosen half of participants either receive no intervention messages or have messages withheld on randomly assigned days, with identical weekly surveys and objective sensors (e.g., GPS-derived location changes and smartwatch sound exposure); if the no-message group shows the same upward trends, the paper's central claim is falsified.","tokens_in":17902,"feed_emoji":"⌚","tokens_out":8492,"duration_ms":80625,"temperature":0.7,"pith_summary":"The paper tries to establish that Just-in-Time Adaptive Interventions (JITAI) — short, context-triggered messages delivered to a smartwatch — can nudge people in real urban settings to take small actions against heat and noise discomfort. In an eight-month Singapore deployment with 103 participants, more than 12,000 micro-surveys and 3,600 intervention messages were collected, and weekly self-reports over the first three weeks showed increases in perceived usefulness (8–19%), noise-related location changes (4–11%), earphone use (2–17%), and thermal adjustments (3–13%). The paper argues this is evidence that giving people the right information at the right place and time can make occupants active partners in their own comfort rather than passive recipients of building services. It also claims that personality traits such as conscientiousness, gender, and environmental preference shape who finds such messages helpful.","feed_headline":"Smartwatch alerts nudge people to dodge heat and noise","feed_subtitle":"In an eight-month Singapore trial, recipients reported more location changes, earphone use, and thermal adjustments.","key_machinery":"The carrying mechanism is a JITAI delivery loop built on the open-source Cozie Apple smartwatch platform. Participants answer micro-surveys on the watch; the platform fuses those self-reports with physiological and activity streams (sound level, heart rate, step count, GPS) and with public weather data; a cloud function then decides whether to push an intervention message. Two trigger logics are tested: threshold-based triggers (outdoor air temperature above 30°C for thermal messages, smartwatch sound meter above 70 dBA for noise messages, capped at four per weekday) and personalized triggers, where a Random Forest classifier — a decision-tree ensemble — trained on each participant's first 50 micro-survey responses predicts the probability of thermal or noise preference by hour of day and sends messages when the predicted non-neutral preference is highest. A weekly in-app survey supplies the outcome measures: perceived helpfulness, annoyance, and self-reported behavioral responses.","core_discovery":"On the paper's own terms, the central finding is that a just-in-time adaptive intervention loop built on smartwatch micro-surveys can shift self-reported occupant behavior in the field. Over the first three weeks of message delivery in two deployment phases, weekly-survey responses moved upward: perceived message helpfulness rose by 8–19 percentage points, reported location changes after noise messages by 4–11 points, earphone use by 2–17 points, and adjustments of location or thermostat for thermal comfort by 3–13 points. The paper also reports that the personalized prediction mechanism, trained on each participant's first 50 micro-surveys, triggered messages for only 19 of 55 phase-2 participants, that annoyance grew in the personalized phase, and that conscientiousness, gender, and outdoor preference were associated with how helpful and actionable participants found the messages. The authors present this as one of the first demonstrations that JITAI is tolerable and potentially behavior-changing in built-environment thermal and aural comfort contexts.","pith_inferences":["Our inference: because the study has no no-message control group and outcomes are self-reported, the observed upward trends could partly reflect growing familiarity with the app or social desirability; a within-subject design that randomly withholds messages on some days would separate the nudge effect from ambient trends.","Our inference: the 70 dBA and 30°C thresholds are tuned to Singapore's climate and street noise; applying the same framework elsewhere would require recalibrating both thresholds and message frequency, so the reported effect sizes are place-specific.","Our inference: the rise in earphone use suggests acoustic JITAI messages may function as a wearable form of personal environmental control, shifting some acoustic comfort management from building design to time- and place-specific behavior; this could be tested by pairing message logs with objective noise-exposure measurements.","Our inference: the personalized model's cold start from the first 50 micro-surveys means participants who respond sparsely get few or no personalized messages; testing models trained on smaller or adaptive sample sizes could broaden who benefits."],"forward_implications":["If the reported trends reflect real behavior change, timely messages could reduce reliance on energy-intensive heating and cooling by helping occupants choose where to sit, what to wear, or whether to use earphones.","Personalized delivery reached only 19 of 55 phase-2 participants and produced more messages, so personalization based on 50 micro-surveys concentrates effectiveness on a subset and risks notification fatigue.","Personality, gender, and outdoor preference correlate with perceived helpfulness, which implies that one-size-fits-all nudging will under-serve some occupants and that tailoring by individual traits is a natural next step.","The spatial concentration of messages around particular campus paths and urban areas suggests the same framework could be extended with geofencing to warn people about heat- or noise-prone zones.","Because the study measured proximal behavior rather than distal outcomes, the authors' framework implies that future work must test whether nudged actions actually improve comfort, productivity, or well-being."],"supporting_citations":[{"why":"Supplies the JITAI components and design principles that structure the study's intervention and discussion.","marker":"Nahum-Shani et al., 2018"},{"why":"Describes the Cozie Apple platform used for micro-survey collection and data transfer.","marker":"Tartarini et al., 2023"},{"why":"Provides the earlier smartwatch micro-survey methodology that grounds the data collection approach.","marker":"Jayathissa et al., 2019"},{"why":"Supplies the longitudinal smartwatch thermal preference data collection approach.","marker":"Quintana et al., 2021"},{"why":"Supplies the random-forest personal comfort prediction approach used for personalized triggers.","marker":"Quintana et al., 2022"},{"why":"Outlines the smartwatch-driven JITAI conceptual framework that this deployment implements.","marker":"Miller et al., 2022"},{"why":"Provides the systematic review of JITAIs for physical activity that motivates applying the method in the built environment.","marker":"Hardeman et al., 2019"},{"why":"Documents that noise and distraction are among the most common workplace complaints, motivating the noise interventions.","marker":"Parkinson et al., 2023"},{"why":"Shows that conventional thermal comfort models are often inaccurate, motivating occupant-empowering alternatives.","marker":"Cheung et al., 2019"}],"fun_headline_variants":["Smartwatch nudges improve heat and noise comfort responses","Adaptive alerts boost thermal and aural comfort actions","Field trial: JITAI messages increase comfort adjustments","Smartwatch prompts shift behavior in heat and noise","Eight-month trial finds smartwatch nudges aid comfort"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the rises in self-reported behavior came from the alert messages themselves, not from normal week-to-week changes in weather, noise, or people's growing familiarity with the app; the study did not include a comparison group of participants who received no alerts.","fun_headline_variants_meta":{"raw":{"variants":["Smartwatch nudges improve heat and noise comfort responses","Adaptive alerts boost thermal and aural comfort actions","Field trial: JITAI messages increase comfort adjustments","Smartwatch prompts shift behavior in heat and noise","Eight-month trial finds smartwatch nudges aid comfort"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000361,"raw_usage":{"total_tokens":1999,"prompt_tokens":1046,"completion_tokens":953,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":878}},"tokens_in":662,"tokens_out":953,"duration_ms":9989,"temperature":1.0,"reasoning_tokens":878,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:55:32.818762+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled trial in which a randomly chosen half of participants either receive no intervention messages or have messages withheld on randomly assigned days, with identical weekly surveys and objective sensors (e.g., GPS-derived location changes and smartwatch sound exposure); if the no-message group shows the same upward trends, the paper's central claim is falsified.","supporting_citations":[{"cited_title":", author Smith, S.N","cited_arxiv_id":null,"evidence_quote":"Supplies the JITAI components and design principles that structure the study's intervention and discussion."},{"cited_title":", author Frei, M","cited_arxiv_id":null,"evidence_quote":"Describes the Cozie Apple platform used for micro-survey collection and data transfer."},{"cited_title":", author Quintana, M","cited_arxiv_id":null,"evidence_quote":"Provides the earlier smartwatch micro-survey methodology that grounds the data collection approach."},{"cited_title":", author Abdelrahman, M","cited_arxiv_id":null,"evidence_quote":"Supplies the longitudinal smartwatch thermal preference data collection approach."},{"cited_title":", author Schiavon, S","cited_arxiv_id":null,"evidence_quote":"Supplies the random-forest personal comfort prediction approach used for personalized triggers."},{"cited_title":", author Chua, Y.X","cited_arxiv_id":null,"evidence_quote":"Outlines the smartwatch-driven JITAI conceptual framework that this deployment implements."},{"cited_title":", author Houghton, J","cited_arxiv_id":null,"evidence_quote":"Provides the systematic review of JITAIs for physical activity that motivates applying the method in the built environment."},{"cited_title":", author Schiavon, S","cited_arxiv_id":null,"evidence_quote":"Documents that noise and distraction are among the most common workplace complaints, motivating the noise interventions."},{"cited_title":", author Schiavon, S","cited_arxiv_id":null,"evidence_quote":"Shows that conventional thermal comfort models are often inaccurate, motivating occupant-empowering alternatives."}],"review_version":1}