REVIEW 4 minor 46 references
Wearing Trust: How Older Adults Calibrate Reliance on Health Wearables Through Bodily Experience and Everyday Use
T0 review · 0 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Older adults calibrate trust in health wearables from visible cues, not technical evidence.
desk verdict A useful concept and an honest exploratory study, but the calibration-over-time claim rests on retrospective self-reports and should be read as descriptive, not process. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the observability gap, defined as the mismatch between the evidence users can inspect and the hidden properties they need in order to judge reliance. The paper uses six quality domains from prior work as an analytical lens; validity, context-specific reliability, data continuity, and human-system fit define reliance-relevant quality, while safety and alert provenance and privacy and governance are treated as adjacent high-stakes qualities. The machinery works by showing that each of the four inference resources maps to visible or embodied cues, not to these hidden quality domains, so trust accumulation is decoupled from technical understanding.
What would settle it
A field or lab study that logs actual device data and breakdowns alongside older adults' routine use, then compares their self-reported trust rules and reliance decisions with what the device actually did, would settle the claim: if users' trust tracks visible cues but not measured signal quality, the observability gap holds; if reported trust matches actual sensor performance, the gap is an artifact of retrospection.
Extended reading notes
Core claim
The paper's central discovery is that older adults with at least four months of wearable experience form practical, context-bound rules about when to trust a device—for example, trusting sleep scores when they match felt fatigue or heart-rate spikes during exercise—but these rules are calibrated against the body and the interface, not against sensing validity. Four recurrent inference resources assemble trust: proxy cues of investment and reputation, visible signs of capability and autonomy, lived interaction experience, and experiential calibration against body and context. These resources can reinforce or override one another, and they produce a settlement the authors call conditional trust: stable enough for daily use, but fragile because bodily sensation can also justify trusting a wrong number.
Load-bearing premise
The load-bearing premise is that participants' retrospective interview accounts accurately reconstruct how their reliance on wearables was calibrated over time; the paper's limitations section states that there were no direct observations, diaries, or device logs, and no independent validation of device accuracy.
Editorial extensions
If this is right
- Designers should place reliability evidence at the moments where users already judge systems—reading a graph, responding to an alert, noticing lag—rather than in manuals or help pages.
- Interfaces should surface signal quality and sensing gaps, such as contact quality indicators, wear time summaries, and low-confidence notices, instead of treating continuous updates as proof of correctness.
- Reliability should be communicated by context, specifying when and where outputs are most and least trustworthy, rather than through blanket claims such as "accurate monitoring."
- Comfort, charging burden, and interaction friction should be treated as part of reliance, because users fold these experiences into global judgments of whether the system is well made.
- Systems should distinguish wellness prompts, heuristic suggestions, and clinically grounded alerts by labeling their provenance and status.
Reading between the lines
- If the observability gap is general, then adding dense visualization or "AI" labels to a wearable could actively increase unwarranted trust; this is a testable design hypothesis the paper does not directly test.
- The same cue-substitution logic likely applies to younger users and to other opaque health algorithms, though the paper restricts its claims to older adults in China.
- A natural next study would combine interviews with device logs and diaries to compare reported trust changes with actual breakdowns, missing data, and context shifts, as the paper itself names as future work.
- Because bodily sensation can reinforce error, the account implies that users with stable but misleading symptoms may overtrust a systematically biased device more than users who check outputs against objective measurements.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a qualitative interview study with 31 older adults in China who had used health-related wearables for at least four months. The authors examine how participants judged whether wearable outputs were reliable enough for everyday use, finding that participants relied on proxy cues such as price and brand, visible interface activity, lived interaction experience, and comparison with bodily sensation. These cues enabled conditional trust, but did not provide access to sensor validity, data continuity, or failure conditions; the authors term this mismatch the 'observability gap.' Based on the analysis, they propose four design directions: surfacing signal quality and sensing gaps, communicating reliability by context, treating human-system fit as part of reliance, and distinguishing automation from medical authority. The paper frames itself as an exploratory account of subjective reasoning rather than a technical assessment of device accuracy.
Significance. If the finding holds, the paper makes a useful contribution to HCI and health-wearable research by shifting attention from adoption and selection to post-adoption inference among older adults. The four inference resources are clearly described and grounded in participant quotes, and the link to work on folk theories and trust in automation is apt. A particular strength is the explicit acknowledgment of methodological limits: Section 5.5 states that data are retrospective interviews, that devices were not independently validated, and that design directions are not yet evaluated. This scoping makes the descriptive contribution appropriately modest and does not overclaim. The observability-gap framing may prove generative for future interface design, even though the design implications are not tested here.
minor comments (4)
- [Section 1, paragraph 2] The phrase 'anobservability gap' is missing a space; it should read 'an observability gap.' Similarly, Section 3.1 contains 'usewearable-centered health monitoring systemsto' which appears to be a formatting artifact. Please ensure consistent spacing in the final version.
- [Title page and Section 1] The manuscript retains poster-template artifacts, including the header 'Conference’17, July 2017, Washington, DC, USA' and the sentence 'This poster addresses that question.' If the submission is intended as a full paper, these leftovers should be replaced with the target venue's formatting and appropriate wording.
- [Section 3.3] The statement that 'After about 24 interviews, additional data elaborated existing mechanisms rather than adding new ones' makes a saturation-like claim without supporting detail. Since the paper does not use inter-rater reliability, adding a brief description of how thematic stability was assessed (for example, the range of disconfirming cases examined) would strengthen the reporting of analytic rigor.
- [Section 3.2] The age statistic reads 'ages 61 to 76 years, M= 68.4.' Please format the mean consistently (e.g., 'M = 68.4') and ensure the same styling for any other statistical notation.
Circularity Check
No circularity: the poster is an inductive qualitative study with no fitted parameters, derived equations, or load-bearing self-referential claims.
full rationale
This paper contains no mathematical derivation, model fitting, or predictive computation, so the circularity patterns involving self-definition, fitted inputs renamed as predictions, or uniqueness theorems do not apply. The central concept, the observability gap, is introduced by the authors as an interpretation of interview data, not derived from or equivalent to an input assumption. The empirical claims are explicitly framed as participants' judgments rather than independently verified device performance (Section 3.1), and the analysis is described as inductive reflexive thematic analysis with themes developed before comparison with prior system-quality dimensions (Sections 3.1 and 3.3). Several references to the authors' own prior work appear in the reference list, but none is load-bearing: they support background claims about ongoing research programs and do not supply the paper's framework or conclusions. The limitations section candidly acknowledges that data come from retrospective interviews rather than direct observation and that design directions are unevaluated (Section 5.5), which is a validity limitation, not circularity. The honest finding is therefore no significant circularity, consistent with the reader's take.
Assumptions & free parameters
assumptions (3)
- domain assumption Retrospective interview self-reports are treated as valid evidence of trust calibration.
- domain assumption The six reliance-relevant quality dimensions borrowed from prior wearable research are an appropriate analytical frame.
- domain assumption The sample of 31 older adults in China adequately captures the phenomenon.
Cite this review
Pith. "Pith review of Wearing Trust: How Older Adults Calibrate Reliance on Health Wearables Through Bodily Experience and Everyday Use." pith.science (2026). https://pith.science/paper/UGHEYZTI
@misc{pith2026260808856,
author = {Pith},
title = {Pith review of: Wearing Trust: How Older Adults Calibrate Reliance on Health Wearables Through Bodily Experience and Everyday Use},
year = {2026},
howpublished = {\url{https://pith.science/paper/UGHEYZTI}},
note = {Machine review of arXiv:2608.08856}
}
read the original abstract
Older adults increasingly use health wearables, yet often cannot inspect the properties that matter for reliance. Through 31 semi-structured interviews in China, we examined how participants judged whether wearable outputs were reliable enough for everyday use. Participants relied on brand and price, visible interface activity, lived interaction experience, and comparison with bodily sensation. These cues supported conditional trust, but did not reveal sensor validity, data continuity, or failure conditions. We describe this mismatch as an observability gap and outline design directions for showing signal quality, reliability by context, human-system fit, and alert provenance.
Reference graph
Works this paper leans on
-
[1]
Y. Abdelaal, M. Aupetit, A. Baggag, and D. Al-Thani. 2024. Exploring the Applica- tions of Explainability in Wearable Data Analytics: Systematic Literature Review. Journal of Medical Internet Research26 (2024), e53863. doi:10.2196/53863
doi:10.2196/53863 2024
-
[2]
Adler, Yuewen Yang, Thalia Viranda, Xuhai Xu, David C
Daniel A. Adler, Yuewen Yang, Thalia Viranda, Xuhai Xu, David C. Mohr, Anna R. Van Meter, Julia C. Tartaglia, Nicholas C. Jacobson, Fei Wang, Deborah Estrin, and Tanzeem Choudhury. 2024. Beyond Detection: Towards Actionable Sensing Research in Clinical Mental Healthcare.Proc. ACM IMWUT8, 4, Article 160 (2024), 33 pages. doi:10.1145/3699755
doi:10.1145/3699755 2024
-
[3]
Riku Arakawa, Hiromu Yakura, Vimal Mollyn, Suzanne Nie, Emma Russell, Dustin P. DeMeo, Haarika A. Reddy, Alexander K. Maytin, Bryan T. Carroll, Jill Fain Lehman, and Mayank Goel. 2023. Prism-tracker: A framework for multimodal procedure tracking using wearable sensors and state transition information with user-driven handling of errors and uncertainty.Pro...
doi:10.1145/3569504 2023
-
[4]
C. I. R. Braem, U. S. Yavuz, H. J. Hermens, and P. H. Veltink. 2024. Missing Data Statistics Provide Causal Insights into Data Loss in Diabetes Health Monitoring by Wearable Sensors.Sensors24, 5 (2024), 1526. doi:10.3390/s24051526
-
[5]
Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology.Qualitative Research in Psychology3, 2 (2006), 77–101. doi:10.1191/1478088706qp063oa
-
[6]
Virginia Braun and Victoria Clarke. 2021. One size fits all? What counts as quality practice in (reflexive) thematic analysis?Qualitative Research in Psychology18, 3 (2021), 328–352. doi:10.1080/14780887.2020.1769238
arXiv 2021
-
[7]
Silvia Canali, Viola Schiaffonati, and Andrea Aliverti. 2022. Challenges and Rec- ommendations for Wearable Devices in Digital Health: Data Quality, Interop- erability, Health Equity, Fairness.PLOS Digital Health1, 10 (2022), e0000104. doi:10.1371/journal.pdig.0000104
-
[8]
B. Carrier, B. Barrios, B. D. Jolley, and J. W. Navalta. 2020. Validity and Reliability of Physiological Data in Applied Settings Measured by Wear- able Technology: A Rapid Systematic Review.Technologies8, 4 (2020), 70. doi:10.3390/technologies8040070
Show all 46 references
-
[9]
Ruiqi Chen, Yibo Meng, Huidi Lu, and Xiaolan Ding. 2026. Between Knowl- edge and Care: A Mixed-Methods Evaluation of Generative AI for T2DM Self-Management from Patient and Physician Perspectives.arXiv preprint arXiv:2607.03720(2026)
2026 arXiv
-
[10]
S. Cho, I. Ensari, C. Weng, M. G. Kahn, and K. Natarajan. 2021. Factors Affecting the Quality of Person-Generated Wearable Device Data and Associated Chal- lenges: Rapid Systematic Review.JMIR mHealth and uHealth9, 3 (2021), e20738. doi:10.2196/20738
2021 doi
-
[11]
S. Cho, C. Weng, M. Kahn, and K. Natarajan. 2021. Identifying Data Quality Dimensions for Person-Generated Wearable Device Data: Multi-Method Study. JMIR mHealth and uHealth9, 12 (2021), e31618. doi:10.2196/31618
2021 doi
-
[12]
Hancock, Megan French, and Sunny Liu
Michael Ann DeVito, Jeremy Birnholtz, Jeffery T. Hancock, Megan French, and Sunny Liu. 2018. How People Form Folk Theories of Social Media Feeds and What It Means for How We Study Self-Presentation. InProc. CHI 2018. ACM, Article 120, 12 pages. doi:10.1145/3173574.3173694
2018
-
[13]
H. Ding, K. Ho, E. Searls, S. Low, Z. Li, S. Rahman, S. Madan, A. Igwe, Z. Popp, A. Burk, H. Wu, Y. Ding, P. Hwang, I. Anda-Duran, V. Kolachalama, K. Gifford, L. Shih, R. Au, and H. Lin. 2024. Assessment of Wearable Device Adherence for Monitoring Physical Activity in Older Ad...
2024 doi
-
[14]
Motahhare Eslami, Karrie Karahalios, Christian Sandvig, Kristen Vaccaro, Aimee Rickman, Kevin Hamilton, and Alex Kirlik. 2016. First I “Like” It, Then I Hide It: Folk Theories of Social Feeds. InProc. CHI 2016. ACM, 2371–2382. doi:10.1145/2858036.2858494
2016
-
[15]
Daniel Fuller, Erica Colwell, Jenna Low, Kassidy Orychock, Matthew Tobin, Bless- ing Simango, Rachel Buote, Daniel Van Heerden, Hao Luan, Karen Cullen, Laura Slade, and Nathan Taylor. 2020. Reliability and Validity of Commercially Available Wearable Devices for Measuring Steps...
2020 doi
-
[16]
Pallavi Rao Gadahad and Anirudha Joshi. 2022. Wearable Activity Track- ers in Managing Routine Health and Fitness of Indian Older Adults: Ex- ploring Barriers to Usage. InProc. NordiCHI 2022. ACM, Article 7, 11 pages. doi:10.1145/3546155.3546645
2022
-
[17]
Gathright, I
R. Gathright, I. Mejia, J. M. Gonzalez, S. I. Hernandez Torres, D. Berard, and E. J. Snider. 2024. Overview of Wearable Healthcare Devices for Clinical Decision Sup- port in the Prehospital Setting.Sensors24, 24 (2024), 8204. doi:10.3390/s24248204
2024 doi
-
[18]
Yao Guo, Xiangyu Liu, Shun Peng, Xinyu Jiang, Ke Xu, Chen Chen, Zeyu Wang, Chenyun Dai, and Wei Chen. 2021. A review of wearable and unobtrusive sensing technologies for chronic disease management.Computers in Biology and Medicine 129 (2021), 104163. doi:10.1016/j.compbiomed.2...
2021
-
[19]
Yunjo Han, Panyu Zhang, Minseo Park, and Uichin Lee. 2024. Systematic Evalua- tion of Personalized Deep Learning Models for Affect Recognition.Proc. ACM IMWUT8, 4, Article 206 (2024), 35 pages. doi:10.1145/3699724
2024 doi
-
[20]
Qingyang He, Weicheng Zheng, Hanxi Bao, Ruiqi Chen, and Xin Tong. 2023. Exploring Designers’ Perceptions and Practices of Collaborating with Generative AI as a Co-Creative Agent in a Multi-Stakeholder Design Process: Take the Do- main of Avatar Design as an Example. InProceedi...
2023
-
[21]
Jennifer Hepburn, Lynn Williams, and Lisa McCann. 2025. Barriers to and Fa- cilitators of Digital Health Technology Adoption Among Older Adults With Chronic Diseases: Updated Systematic Review.JMIR Aging8 (2025), e80000. doi:10.2196/80000
2025 doi
-
[22]
Keogh, J
A. Keogh, J. Dorn, L. Walsh, F. Calvo, and B. Caulfield. 2020. Comparing the Usability and Acceptability of Wearable Sensors Among Older Irish Adults in a Real-World Context: Observational Study.JMIR mHealth and uHealth8, 4 (2020), e15704. doi:10.2196/15704
2020 doi
-
[23]
Mohammad Kianpisheh, Alex Mariakakis, and Khai N. Truong. 2024. exHAR: An interface for helping non-experts develop and debug knowledge-based hu- man activity recognition systems.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies8, 1 (2024), 1–...
2024 doi
-
[24]
Vaughn Rikard, Shelia Cotten, and Wei Peng
Anastasia Kononova, Lin Li, Kendra Kamp, Marie Bowen, R. Vaughn Rikard, Shelia Cotten, and Wei Peng. 2019. The Use of Wearable Activity Trackers Among Older Adults: Focus Group Study of Tracker Perceptions, Motivators, and Barriers in the Maintenance Stage of Behavior Change.J...
2019 doi
-
[25]
Lee and Katrina A
John D. Lee and Katrina A. See. 2004. Trust in Automation: Designing for Appro- priate Reliance.Human Factors46, 1 (2004), 50–80. doi:10.1518/hfes.46.1.50_30392
2004 doi
-
[26]
Jong Ho Lee, Jessica Schroeder, and Daniel A. Epstein. 2021. Understanding and Supporting Self-Tracking App Selection.Proc. ACM IMWUT5, 4, Article 166 (Dec. 2021), 25 pages. doi:10.1145/3494980
2021 doi
-
[27]
Li and Peter Washington
J. Li and Peter Washington. 2024. A Comparison of Personalized and Generalized Approaches to Emotion Recognition Using Consumer Wearable Devices: Machine Learning Study.JMIR AI3 (2024), e52171. doi:10.2196/52171
2024 doi
-
[28]
Zilu Liang and Bernd Ploderer. 2016. Sleep tracking in the real world: a qualitative study into barriers for improving sleep. InProceedings of the 28th Australian Con- ference on Computer-Human Interaction. 537–541. doi:10.1145/3010915.3010988
2016
-
[29]
Zilu Liang and Bernd Ploderer. 2020. How does Fitbit measure brainwaves: a qualitative study into the credibility of sleep-tracking technologies.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies4, 1 (2020), 1–29. doi:10.1145/3380994
2020 doi
-
[30]
Shu Liu, Soujanya Ponnapalli, Shreya Shankar, Sepanta Zeighami, Alan Zhu, Shubham Agarwal, Ruiqi Chen, Samion Suwito, Shuo Yuan, Ion Stoica, et al. 2025. Supporting Our AI Overlords: Redesigning Data Systems to Be Agent-First.arXiv preprint arXiv:2509.00997(2025)
2025
-
[31]
What’s Happening
Xuewen Luo, Fan Ding, Rishikesh Panda, Ruiqi Chen, Junnyong Loo, and Shuyun Zhang. 2025. “What’s Happening”—A Human-Centered Multimodal Interpreter Explaining the Actions of Autonomous Vehicles. InProceedings of the Winter Conference on Applications of Computer Vision. 1163–1170
2025
-
[32]
Lakmal Meegahapola, William Droz, Peter Kun, Amalia De Götzen, Chaitanya Nutakki, Shyam Diwakar, Salvador Ruiz Correa, Donglei Song, Hao Xu, Miriam Bidoglia, George Gaskell, Altangerel Chagnaa, Amarsanaa Ganbold, Tsolmon Zundui, Carlo Caprini, Daniele Miorandi, Alethia Hume, J...
-
[33]
Yibo Meng, Bingyi Liu, Ruiqi Chen, Xin Chen, and Yan Guan. 2026. 52-Hz Whale Song: An Embodied VR Experience for Exploring Misunderstanding and Empathy. InProceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems. 1–5. doi:10.1145/3...
2026
-
[34]
Yibo Meng, Bingyi Liu, Ruiqi Chen, Xiaolan Ding, and Shuai Ma. 2026. Living Inside the Black Box: Behavioral Probing and Adaptation in Mandatory Wearable Sensing.arXiv preprint arXiv:2607.09009(2026)
2026 arXiv
-
[35]
Yibo Meng, Ruiqi Chen, Zhiming Liu, and Xiaolan Ding. 2026. TibetCPR: A Multimodal Tactile Feedback System to Enhance Cardiopulmonary Resuscitation Training in High-Altitude Regions of Tibet.arXiv preprint arXiv:2606.07765(2026)
2026 arXiv
-
[36]
Yibo Meng, Ruiqi Chen, Zhuoran Lu, Shuai Ma, and Chengxi Zang. 2025. Tracing Generative AI in Digital Art: A Longitudinal Study of Chinese Painters’ Attitudes, Practices, and Identity Negotiation.arXiv preprint arXiv:2511.03117(2025)
2025
-
[37]
Kelly, Simon D’Alfonso, and Reeva Lederman
Joshua Newn, Ryan M. Kelly, Simon D’Alfonso, and Reeva Lederman. 2022. Examining and Promoting Explainable Recommendations for Personal Sensing Technology Acceptance.Proc. ACM IMWUT6, 3, Article 133 (2022), 27 pages. doi:10.1145/3550297
2022 doi
-
[38]
Clarke, and Brad Holschuh
Robert Pettys-Baker, Megan E. Clarke, and Brad Holschuh. 2024. Functional Now, Wearable Later: Examining the Design Practices of Wearable Technologists. In Proc. ISWC 2024. ACM, 71–81. doi:10.1145/3675095.3676615 5 Conference’17, July 2017, Washington, DC, USA Yibo Meng, Bingy...
2024
-
[39]
Roos and George M
Lydia G. Roos and George M. Slavich. 2023. Wearable Technologies for Health Re- search: Opportunities, Limitations, and Practical and Conceptual Considerations. Brain, Behavior, and Immunity113 (2023), 444–452. doi:10.1016/j.bbi.2023.08.008
2023 doi
-
[40]
Kavous Salehzadeh Niksirat, Lev Velykoivanenko, Noé Zufferey, Mauro Cherubini, Kévin Huguenin, and Mathias Humbert. 2024. Wearable Activity Trackers: A Survey on Utility, Privacy, and Security.ACM Computing Surveys56, 7, Article 183 (2024), 40 pages. doi:10.1145/3645091
2024 doi
-
[41]
Stuart, M
S. Stuart, M. de Kok, B. O’Searcoid, H. Morrisroe, I. B. Serban, F. Jagers, R. Dulos, S. Houben, L. van de Peppel, and J. van den Brand. 2024. Critical Design Con- siderations for Longer-Term Wear and Comfort of On-Body Medical Devices. Bioengineering11, 11 (2024), 1058. doi:1...
2024 doi
-
[42]
Jonas Van Der Donckt, Niels Vandenbussche, Shuo Chen, Maria Stojchevska, Mathias De Brouwer, Brecht Steenwinckel, Koen Paemeleire, Femke Ongenae, and Sofie Van Hoecke. 2024. Mitigating Data Quality Challenges in Ambulatory Wrist-Worn Wearable Monitoring Through Analytical and ...
2024 doi
-
[43]
Dimitri Vargemidis, Kathrin Gerling, Vero Vanden Abeele, Luc Geurts, and Katta Spiel. 2021. Irrelevant Gadgets or a Source of Worry: Exploring Wearable Activity Trackers with Older Adults.ACM Transactions on Accessible Computing14, 3, Article 16 (2021), 28 pages. doi:10.1145/3473463
2021 doi
-
[44]
Ruijing Wang, Onur Asan, and Ting Liao. 2025. Investigating the Role of Wear- able Devices in Facilitating Telehealth Adoption Among the Aging Population: Mediation Analysis of US National Data.JMIR Medical Informatics13 (2025), e68559. doi:10.2196/68559
2025 doi
-
[45]
C., Pan Hui, and Xin Tong
Keyi Zeng, Jingyang Lin, Ruiqi Chen, Ray L. C., Pan Hui, and Xin Tong. 2025. Parental Perceptions of Children’s d/Deaf Identity Shaping Technology Use: A Qualitative Study on Communication Technologies in Mixed-Hearing Families. InProceedings of the Extended Abstracts of the C...
2025
-
[2023]
doi:10.1145/3569483
Generalization and personalization of mobile sensing-based mood inference models: an analysis of college students in eight countries.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies6, 4 (2023), 1–32. doi:10.1145/3569483
2023 doi
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.