{"id":"e73664ab-4663-4013-9764-a2a10a528052","arxiv_id":"2608.08856","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Interviews with 31 older adults show they judge wearable trustworthiness from visible and bodily cues because true sensor quality is hidden, a mismatch the paper calls the observability gap.","lead":"Older adults in China who wore health trackers trusted devices based on price, brand, busy screens, and how readings matched their own bodily feelings, rather than on technical accuracy. The study names this mismatch the 'observability gap' and suggests designs that show signal quality and failure conditions at the moment users decide.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Retrospective self-reports weaken evidence for the calibration-over-time claim, but the paper's exploratory scope and explicit limitations leave the verdict unchanged.","rationale":"The reader's weakest assumption correctly identifies retrospective self-reports as the most load-bearing assumption. My analysis agrees: the paper's central claim about how reliance is calibrated over time rests entirely on participants' recollections, which are subject to memory smoothing and narrative coherence. However, the paper is explicitly exploratory, frames its contribution as an account of subjective judgments, and states this limitation in Section 5.5. The design directions are labeled untested. Given the scope, the limitation does not change the accept verdict. I considered whether the lack of behavioral observation should force a CONDITIONAL verdict, but the paper does not overclaim—it says 'how older adults calibrate reliance' based on what participants reported, not on what was observed. No independent evidence contradicts the account; no internal inconsistency arises. Thus the verdict remains ACCEPT as the reader recommended, with the retrospective self-report issue as a clearly bounded caveat rather than a fatal flaw.","tokens_in":9970,"tokens_out":5386,"duration_ms":57227,"concrete_test":"Run a small longitudinal substudy with 10–15 new participants: log device usage and alert interactions for 8 weeks, collect brief weekly diaries of reliance decisions (e.g., 'Did you act on or ignore an alert?'), and conduct exit interviews. Then compare, per participant, whether the cues named in the exit interview (price, interface density, bodily agreement) correspond to the observed reliance episodes. If named cues do not track logged behavior, the retrospective account is likely reconstructive; if they do track, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is temporal: older adults calibrate reliance through bodily experience and everyday use, accumulating conditional trust over time. The evidence is entirely retrospective interview self-reports (Section 3.3). Participants reconstructed, after at least four months, their initial expectations, memorable breakdowns, and shifts in trust. This design cannot distinguish an actual calibration process from post-hoc narrative reconstruction that makes sense of the current relationship with the device. The paper acknowledges this in Section 5.5: 'The data come from retrospective interviews rather than direct observation... Future work should also combine interviews with diaries, observation, or logs.' Because the claim is framed as an exploratory account of subjective reasoning, the concern is real but bounded: it does not undermine the descriptive contribution, but the stronger reading—that these cues actually drove reliance decisions—is not directly evidenced. The lack of independent device accuracy validation is less relevant, since the paper deliberately does not claim objective accuracy; what is missing is behavioral or longitudinal trace data on reliance episodes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative interview study with 31 older adults in China who had used health-related wearables for at least four months. The authors examine how participants judged whether wearable outputs were reliable enough for everyday use, finding that participants relied on proxy cues such as price and brand, visible interface activity, lived interaction experience, and comparison with bodily sensation. These cues enabled conditional trust, but did not provide access to sensor validity, data continuity, or failure conditions; the authors term this mismatch the 'observability gap.' Based on the analysis, they propose four design directions: surfacing signal quality and sensing gaps, communicating reliability by context, treating human-system fit as part of reliance, and distinguishing automation from medical authority. The paper frames itself as an exploratory account of subjective reasoning rather than a technical assessment of device accuracy.","tokens_in":10084,"tokens_out":3558,"duration_ms":35632,"significance":"If the finding holds, the paper makes a useful contribution to HCI and health-wearable research by shifting attention from adoption and selection to post-adoption inference among older adults. The four inference resources are clearly described and grounded in participant quotes, and the link to work on folk theories and trust in automation is apt. A particular strength is the explicit acknowledgment of methodological limits: Section 5.5 states that data are retrospective interviews, that devices were not independently validated, and that design directions are not yet evaluated. This scoping makes the descriptive contribution appropriately modest and does not overclaim. The observability-gap framing may prove generative for future interface design, even though the design implications are not tested here.","major_comments":[],"minor_comments":[{"comment":"The phrase 'anobservability gap' is missing a space; it should read 'an observability gap.' Similarly, Section 3.1 contains 'usewearable-centered health monitoring systemsto' which appears to be a formatting artifact. Please ensure consistent spacing in the final version.","section":"Section 1, paragraph 2"},{"comment":"The manuscript retains poster-template artifacts, including the header 'Conference’17, July 2017, Washington, DC, USA' and the sentence 'This poster addresses that question.' If the submission is intended as a full paper, these leftovers should be replaced with the target venue's formatting and appropriate wording.","section":"Title page and Section 1"},{"comment":"The statement that 'After about 24 interviews, additional data elaborated existing mechanisms rather than adding new ones' makes a saturation-like claim without supporting detail. Since the paper does not use inter-rater reliability, adding a brief description of how thematic stability was assessed (for example, the range of disconfirming cases examined) would strengthen the reporting of analytic rigor.","section":"Section 3.3"},{"comment":"The age statistic reads 'ages 61 to 76 years, M= 68.4.' Please format the mean consistently (e.g., 'M = 68.4') and ensure the same styling for any other statistical notation.","section":"Section 3.2"}],"recommendation":"accept","confidential_remarks":"The paper relies heavily on citations to the authors' own arXiv preprints (e.g., references 9, 30, 31, 34, 35, 36, 45). In particular, reference 34, 'Living Inside the Black Box: Behavioral Probing and Adaptation in Mandatory Wearable Sensing,' appears topically close to this submission. The manuscript does not discuss its relationship to that preprint. The editor may wish to verify that there is no overlapping or duplicate submission, and to confirm that the cited preprints are in fact publicly available as described."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a clearly written qualitative study with a genuinely useful concept: the observability gap. Second, the stronger claim in the title—older adults calibrate reliance over time through bodily experience—rests on retrospective interviews, so it should be read as an account of how participants narrate their reasoning, not direct evidence of a calibration process.\n\nWhat is new: prior HCI work covers sleep-tracker credibility, app selection, and older-adult adoption. This study moves past adoption and shows how older users continue to judge reliability using visible and embodied cues. The four inference resources (proxy cues, visible signs of activity, lived interaction experience, bodily/context comparison) are well illustrated with quotes, and the paper includes disconfirming cases, which strengthens credibility. The analysis handles the reflexive thematic analysis tradition correctly and the limitations section is unusually honest: it flags the retrospective design, the lack of device validation, the China-specific sample, and the untested design directions. Credit where due.\n\nSoft spots. The main one is exactly what the paper concedes: all evidence comes from self-report. Participants recall after at least four months what they believed and why. That can produce a coherent narrative rather than a true trace. A stronger design would add diaries or logs; the authors say so. This means the descriptive contribution stands, but the subtitle's 'calibrate over time' should be softened. The sample is self-selected and local, so transferability is unknown. Minor point: the reference list leans on the authors' own preprints, some of which are tangential. That is harmless but should be trimmed.\n\nThis paper is for health-wearable HCI researchers and anyone working on trust in automation for older adults. It is a modest but honest contribution. I would send it to review; it deserves a serious referee. The authors should be encouraged to frame the temporal claim as exploratory and to treat the design implications as hypotheses.","headline":"A useful concept and an honest exploratory study, but the calibration-over-time claim rests on retrospective self-reports and should be read as descriptive, not process.","tokens_in":10612,"tokens_out":3360,"would_cite":true,"duration_ms":32306,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Older adults calibrate trust in health wearables from visible cues, not technical evidence.","keywords":["health wearables","older adults","trust calibration","observability gap","conditional trust","qualitative interviews","digital health","self-tracking"],"falsifier":"A field or lab study that logs actual device data and breakdowns alongside older adults' routine use, then compares their self-reported trust rules and reliance decisions with what the device actually did, would settle the claim: if users' trust tracks visible cues but not measured signal quality, the observability gap holds; if reported trust matches actual sensor performance, the gap is an artifact of retrospection.","tokens_in":1288,"feed_emoji":"⌚","tokens_out":2156,"duration_ms":58248,"temperature":0.7,"pith_summary":"This paper argues that older adults decide whether to rely on a health wearable using cues they can see and feel—price, brand, medical-looking design, dense charts, fast updates, smooth interaction, and agreement with bodily sensation—rather than direct evidence of sensor accuracy, data continuity, or failure conditions. The authors call this mismatch the observability gap: the properties that matter for reliance are hidden, so users substitute observable polish and activity for them. Conditional trust can accumulate over months of use without yielding a technically accurate understanding of the system. If this account is right, current wearable interfaces reward visible sophistication and feedback rhythm while leaving users without grounds for knowing when a reading is trustworthy.","feed_headline":"Trust in wearables comes from visible cues, not hidden accuracy","feed_subtitle":"In 31 interviews, older users based trust on price, brand, interface activity, and bodily agreement, not sensor validity.","key_machinery":"The central object is the observability gap, defined as the mismatch between the evidence users can inspect and the hidden properties they need in order to judge reliance. The paper uses six quality domains from prior work as an analytical lens; validity, context-specific reliability, data continuity, and human-system fit define reliance-relevant quality, while safety and alert provenance and privacy and governance are treated as adjacent high-stakes qualities. The machinery works by showing that each of the four inference resources maps to visible or embodied cues, not to these hidden quality domains, so trust accumulation is decoupled from technical understanding.","core_discovery":"The paper's central discovery is that older adults with at least four months of wearable experience form practical, context-bound rules about when to trust a device—for example, trusting sleep scores when they match felt fatigue or heart-rate spikes during exercise—but these rules are calibrated against the body and the interface, not against sensing validity. Four recurrent inference resources assemble trust: proxy cues of investment and reputation, visible signs of capability and autonomy, lived interaction experience, and experiential calibration against body and context. These resources can reinforce or override one another, and they produce a settlement the authors call conditional trust: stable enough for daily use, but fragile because bodily sensation can also justify trusting a wrong number.","pith_inferences":["If the observability gap is general, then adding dense visualization or \"AI\" labels to a wearable could actively increase unwarranted trust; this is a testable design hypothesis the paper does not directly test.","The same cue-substitution logic likely applies to younger users and to other opaque health algorithms, though the paper restricts its claims to older adults in China.","A natural next study would combine interviews with device logs and diaries to compare reported trust changes with actual breakdowns, missing data, and context shifts, as the paper itself names as future work.","Because bodily sensation can reinforce error, the account implies that users with stable but misleading symptoms may overtrust a systematically biased device more than users who check outputs against objective measurements."],"forward_implications":["Designers should place reliability evidence at the moments where users already judge systems—reading a graph, responding to an alert, noticing lag—rather than in manuals or help pages.","Interfaces should surface signal quality and sensing gaps, such as contact quality indicators, wear time summaries, and low-confidence notices, instead of treating continuous updates as proof of correctness.","Reliability should be communicated by context, specifying when and where outputs are most and least trustworthy, rather than through blanket claims such as \"accurate monitoring.\"","Comfort, charging burden, and interaction friction should be treated as part of reliance, because users fold these experiences into global judgments of whether the system is well made.","Systems should distinguish wellness prompts, heuristic suggestions, and clinically grounded alerts by labeling their provenance and status."],"supporting_citations":[{"why":"Supplies the theoretical benchmark of trust calibrated to actual capability versus trust grounded in understanding.","marker":"[25]"},{"why":"Shows that sleep-tracking users struggle to assess how devices produce their claims, the adjacent empirical result this study broadens.","marker":"[28, 29]"},{"why":"Provides the six quality domains, including validity, reliability, robustness, and ecological validity, used as the analytical lens.","marker":"[7, 8, 15]"},{"why":"Establishes that people form folk theories from partial visible cues in opaque systems, a mechanism applied to wearables here.","marker":"[12, 14]"},{"why":"Shows how explanations and interface framing shape acceptance of sensed outputs, supporting the design direction for reliability evidence.","marker":"[1, 37]"},{"why":"Documents adoption and abandonment of activity trackers among older adults, the population context this study extends from adoption to inference.","marker":"[24, 43]"}],"fun_headline_variants":["Older adults trust wearables by body feel, not sensor specs","Trust in wearables: visible cues beat hidden accuracy for older users","Wearable reliance calibrated by experience, not validity checks","The observability gap: why older users trust wearable numbers"],"cache_read_input_tokens":12928,"weakest_assumption_plain":"The load-bearing premise is that participants' retrospective interview accounts accurately reconstruct how their reliance on wearables was calibrated over time; the paper's limitations section states that there were no direct observations, diaries, or device logs, and no independent validation of device accuracy.","fun_headline_variants_meta":{"raw":{"variants":["Older adults trust wearables by body feel, not sensor specs","Trust in wearables: visible cues beat hidden accuracy for older users","Wearable reliance calibrated by experience, not validity checks","The observability gap: why older users trust wearable numbers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000596,"raw_usage":{"total_tokens":2705,"prompt_tokens":776,"completion_tokens":1929,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":392,"completion_tokens_details":{"reasoning_tokens":1858}},"tokens_in":392,"tokens_out":1929,"duration_ms":16222,"temperature":1.0,"reasoning_tokens":1858,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:21:17.833618+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A field or lab study that logs actual device data and breakdowns alongside older adults' routine use, then compares their self-reported trust rules and reliance decisions with what the device actually did, would settle the claim: if users' trust tracks visible cues but not measured signal quality, the observability gap holds; if reported trust matches actual sensor performance, the gap is an artifact of retrospection.","supporting_citations":[],"review_version":1}