Pith. sign in

REVIEW 2 major objections 1 minor 31 references

Longitudinal Multimodal Sensing of Physical Activity and Well-Being in Older Adults

T0 review · 2 major / 1 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read Sensed signals predict observable behaviors like activity levels far better than abstract clinical outcomes like sleep apnea severity.

desk verdict Small real-world older-adult sensing study shows an expected predictability gradient but the N=66 sample undercuts claims about its generality. read the letter →

arxiv 2606.00345 v1 pith:HX75H5XY submitted 2026-05-29 cs.LG

classification cs.LG
keywords multimodalsensinglongitudinaldataolderadultswearablesensorspredictivemodelingactivitylevelssleepapneaexplainabilityanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that predictive performance in longitudinal multimodal sensing forms a clear gradient tied to how directly a health target aligns with the collected signals. In a real-world study of 66 older adults, tasks like activity level prediction reach macro-F1 scores around 65 percent while sleep apnea severity classification stays challenging even after beating baselines. Historical features from past data prove the strongest predictors across tasks, showing the value of repeated measurements over time. This setup highlights limits when moving from directly sensed behaviors to more removed clinical outcomes in an older population rarely studied this way.

What carries the argument

The unified evaluation framework spanning tasks with increasing levels of observability from sensed signals, which isolates the effect of signal-target alignment on model performance.

What would settle it

A replication with a larger cohort showing no performance difference across the activity, sleep duration, and sleep apnea tasks would falsify the claimed gradient.

Watch

Extended reading notes

Core claim

A unified evaluation framework applied to tasks with increasing levels of observability demonstrates a predictability gradient: highly observable behavioral targets achieve robust performance while more abstract outcomes remain challenging, with historical features consistently emerging as the most informative predictors and underscoring the central role of longitudinal information.

Load-bearing premise

The chosen tasks genuinely represent increasing levels of observability from the sensed signals and the 66-participant dataset supports general claims about predictability gradients.

Editorial extensions

If this is right

  • Models achieve highest accuracy on directly observable targets such as activity levels.
  • Incorporating historical features improves predictions for every task examined.
  • Multimodal sensing yields consistent gains over baselines even on harder targets.
  • Longitudinal data collection is required to capture the most informative predictors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Sensing systems for older adults may achieve more reliable results by focusing first on directly measurable behaviors rather than complex clinical scores.
  • Extending the framework to include additional sensor modalities could test whether the observability gradient persists or narrows.
  • Deployment in clinical decision support would likely prioritize tasks where the gradient favors high predictability.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript reports results from a longitudinal multimodal sensing study of 66 older adults in real-world conditions, combining wearable sensors, behavioral monitoring, and clinical assessments. It evaluates predictive performance on three tasks chosen to span increasing levels of observability (Activity Levels prediction with macro-F1 of 65%, Sleep Duration estimation, and Sleep Apnea Severity classification), claims a clear predictability gradient aligned with observability, and uses explainability analysis to show that historical features are the most informative predictors.

Significance. If the reported gradient holds after proper validation, the work would be useful for informing sensor-based health monitoring systems targeted at older adults, an underrepresented population in longitudinal studies. The real-world, into-the-wild data collection and unified evaluation framework across tasks are strengths; the emphasis on longitudinal information via historical features is also a constructive finding.

major comments (2)
  1. [Abstract] Abstract: the claim of a 'clear gradient of predictability' across the three tasks is load-bearing for the central contribution, yet the abstract (and available description) provides no statistical tests comparing task performances, no power analysis, and no external validation; with N=66 and high inter-individual variability typical in older-adult cohorts, observed differences could arise from sampling artifacts rather than systematic observability alignment.
  2. [Abstract] Abstract: the assumption that Activity Levels, Sleep Duration, and Sleep Apnea Severity genuinely represent increasing levels of observability from the multimodal signals is not justified or tested; without explicit mapping from sensor features to each target or ablation showing signal-target alignment, the gradient interpretation remains an unverified modeling choice.
minor comments (1)
  1. The abstract states 'consistent improvements over baseline models' without naming the baselines, reporting their scores, or indicating the magnitude of gains, which would help readers assess practical significance.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive comments. We address each major comment below and indicate where revisions will be made to the manuscript.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the claim of a 'clear gradient of predictability' across the three tasks is load-bearing for the central contribution, yet the abstract (and available description) provides no statistical tests comparing task performances, no power analysis, and no external validation; with N=66 and high inter-individual variability typical in older-adult cohorts, observed differences could arise from sampling artifacts rather than systematic observability alignment.

    Authors: We agree that the abstract would be strengthened by supporting statistical evidence for the reported performance differences. The full manuscript presents the macro-F1 and other metrics for each task but does not include formal pairwise comparisons. We will add bootstrap confidence intervals or appropriate statistical tests for differences between tasks, update the abstract language to reflect only those supported by the data, and explicitly note the absence of a priori power analysis as a limitation. External validation is not possible with this single-cohort dataset; we will add a limitations paragraph on generalizability while retaining the internal unified evaluation framework as a contribution. revision: yes

  2. Referee: [Abstract] Abstract: the assumption that Activity Levels, Sleep Duration, and Sleep Apnea Severity genuinely represent increasing levels of observability from the multimodal signals is not justified or tested; without explicit mapping from sensor features to each target or ablation showing signal-target alignment, the gradient interpretation remains an unverified modeling choice.

    Authors: The task selection was motivated by domain considerations of how directly each outcome aligns with the available sensor modalities (accelerometry for activity, wearable-derived estimates for sleep duration, and clinical diagnosis for apnea severity). We acknowledge that the manuscript does not provide an explicit feature-to-target mapping or ablation study to validate this ordering. We will add a methods subsection with justification based on sensor characteristics and include supporting ablation or feature-importance results to ground the observability gradient interpretation. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical observational study with no derivations

full rationale

The paper is a standard empirical ML study on a longitudinal dataset of 66 older adults. It reports model performance (macro-F1 scores) across three tasks and notes that historical features rank highest in explainability analysis. No equations, first-principles derivations, fitted parameters renamed as predictions, uniqueness theorems, or self-citation chains appear in the abstract or described content. The claimed predictability gradient is an observed empirical ordering, not a quantity forced by construction from the inputs. This matches the default expectation for non-circular empirical work.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract describes an empirical sensing study and contains no mathematical model, fitted parameters, axioms, or postulated entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Longitudinal Multimodal Sensing of Physical Activity and Well-Being in Older Adults." pith.science (2026). https://pith.science/paper/HX75H5XY

@misc{pith2026260600345,
  author       = {Pith},
  title        = {Pith review of: Longitudinal Multimodal Sensing of Physical Activity and Well-Being in Older Adults},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HX75H5XY}},
  note         = {Machine review of arXiv:2606.00345}
}
read the original abstract

Wearable and mobile sensing technologies enable continuous monitoring of human behavior and health in real-world settings. However, predictive modeling in longitudinal multimodal data remains challenging, particularly when targeting complex or clinically derived outcomes. In this work, we present a longitudinal multimodal study of 66 older adults conducted in real-world conditions and combining wearable sensing, behavioral monitoring, and clinical assessments. This setting provides a rare opportunity to study an underrepresented population in long-term, into-the-wild conditions. Building on this dataset, we investigate how the alignment between sensed signals and target variables affects predictive performance across health-related tasks. We design a unified evaluation framework spanning tasks with increasing levels of observability, including Activity Levels prediction, Sleep Duration estimation, and Sleep Apnea Severity classification. Our results reveal a clear gradient of predictability: highly observable behavioral targets achieve robust performance (macro-F1 65%), while more abstract outcomes remain challenging despite consistent improvements over baseline models. Moreover, through explainability analysis, we show that historical features consistently emerge as the most informative predictors, highlighting the central role of longitudinal information.

Figures

Figures reproduced from arXiv: 2606.00345 by the authors.

Figure 1
Figure 1. Experimental protocol. such as the World Health Organization (WHO) guidelines on physical activity for older adults [28], which recommends 150–300 minutes of moderate physical activity per week. This roughly corresponds to an additional 2,000–4,000 steps per day that, when combined with habitual daily activity, translates to approximately 6,000–8,000 total steps per day in older adults. Although step count is a conv… view at source ↗
Figure 2
Figure 2. Data collection architecture recommended measurement frequency of at least 3 times per week. In addition to body weight, the scale performs bioelectrical impedance analysis (BIA) to estimate body composition param￾eters, including fat mass, muscle mass, bone mass, and total body water. It is worth noting that the clinical measurements of body composition relies on a certified medical device, reporting thus a more ac… view at source ↗
Figure 2
Figure 2. A daily scheduled task programmatically authenticates to the Garmin and Withings [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figures from the paper (9 more)
Figure 3
Figure 3. Figure 3: Dataset overview throughout the study duration. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png]
Figure 4
Figure 4. Figure 4: Comparison of data validity distributions between full (All) and APA periods across [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: CCDF visualizations of cross-modal data availability. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Changes in users’ functional capabilities from baseline to follow-up assessment (a). [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Average weekly trends during Phase 1 stratified by users’ baseline BMI. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Global class distributions across predictive tasks. [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Cross-subject performance distribution across predictive tasks. [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: Comparative heatmap of the top 15 most important features across LR, RF, and [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: Comparative heatmap of the top 15 most important features across LR, RF, and [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

31 extracted references · 2 canonical work pages

  1. [1]

    Transformer-based recognition of activ- ities of daily living from wearable sensor data

    Gabriela Augustinov, Muhammad Adeel Nisar, Fr´ ed´ eric Li, Amir Tabatabaei, Marcin Grze- gorzek, Keywan Sohrabi, and Sebastian Fudickar. Transformer-based recognition of activ- ities of daily living from wearable sensor data. InProceedings of the 7th International Workshop on Sensor-Based Activity Recognition and Artificial Intelligence, iWOAR ’22, New Y...

  2. [2]

    Timed up and go test and risk of falls in older adults: a systematic review.The journal of nutrition, health & aging, 15(10):933–938, 2011

    Olivier Beauchet, Bruno Fantino, Gilles Allali, SW Muir, Manuel Montero-Odasso, and C´ edric Annweiler. Timed up and go test and risk of falls in older adults: a systematic review.The journal of nutrition, health & aging, 15(10):933–938, 2011

  3. [3]

    Lifetrace: A longitudinal multimodal dataset on daily 22 physical activity, well-being, and habits.Proc

    Francesco Bombassei De Bona, Ioana Andreea Cˆ ampanu, Marc Langheinrich, Martin Gjoreski, and Georgiana Juravle. Lifetrace: A longitudinal multimodal dataset on daily 22 physical activity, well-being, and habits.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 9(4), December 2025

  4. [4]

    Cornet and Richard J

    Victor P. Cornet and Richard J. Holden. Systematic review of smartphone-based passive sensing for health and wellbeing.Journal of Biomedical Informatics, 77:120–132, 2018

  5. [5]

    The six-minute walk test.Respiratory Care, 48(8):783–785, 2003

    Paul L Enright. The six-minute walk test.Respiratory Care, 48(8):783–785, 2003

  6. [6]

    Hypertension in older adults: What is the target blood pressure?Cleve

    Anu Garg and Barbara J Messinger-Rapport. Hypertension in older adults: What is the target blood pressure?Cleve. Clin. J. Med., 85(3):193–195, March 2018

  7. [7]

    Sex differences in the association between sleep duration and frailty in older adults: evidence from the knhanes study.BMC geriatrics, 24(1):434, 2024

    Beomman Ha, Mijin Han, Wi-Young So, and Seonho Kim. Sex differences in the association between sleep duration and frailty in older adults: evidence from the knhanes study.BMC geriatrics, 24(1):434, 2024

  8. [8]

    Gabriella M Harari, Nicholas D Lane, Rui Wang, Benjamin S Crosier, Andrew T Campbell, and Samuel D Gosling. Using smartphones to collect behavioral data in psychological sci- ence: Opportunities, practical considerations, and challenges.Perspectives on Psychological Science, 11(6):838–854, 2016

Show all 31 references
  1. [9]

    The short-form version of the depression anxiety stress scales (dass-21): Construct validity and normative data in a large non-clinical sample

    Julie D Henry and John R Crawford. The short-form version of the depression anxiety stress scales (dass-21): Construct validity and normative data in a large non-clinical sample. British journal of clinical psychology, 44(2):227–239, 2005

  2. [10]

    What is adapted physical activity?, 2020

    International Federation of Adapted Physical Activity. What is adapted physical activity?, 2020

  3. [11]

    Validity of the international physical activity questionnaire short form (ipaq-sf): A systematic review

    Paul H Lee, Duncan J Macfarlane, Tai Hing Lam, and Sunita M Stewart. Validity of the international physical activity questionnaire short form (ipaq-sf): A systematic review. International journal of behavioral nutrition and physical activity, 8(1):115, 2011

  4. [12]

    Mobile health applications for older adults: a systematic review of interface and persuasive feature design.J Am Med Inform Assoc, 28(11):2483–2501, October 2021

    Na Liu, Jiamin Yin, Sharon Swee-Lin Tan, Kee Yuan Ngiam, and Hock Hai Teo. Mobile health applications for older adults: a systematic review of interface and persuasive feature design.J Am Med Inform Assoc, 28(11):2483–2501, October 2021

  5. [13]

    Lundberg, Gabriel G

    Scott M. Lundberg, Gabriel G. Erion, and Su-In Lee. Consistent individualized feature attribution for tree ensembles.CoRR, abs/1802.03888, 2018

  6. [14]

    A unified approach to interpreting model predictions

    Scott M Lundberg and Su-In Lee. A unified approach to interpreting model predictions. Advances in neural information processing systems, 30, 2017

  7. [15]

    Mattingly et al

    S. Mattingly et al. The tesserae project: Large-scale, longitudinal, in situ, multimodal sensing of information workers. InProceedings of the CHI Conference on Human Factors in Computing Systems, 2019

  8. [16]

    Tatyana Mollayeva, Pravheen Thurairajah, Kirsteen Burton, Shirin Mollayeva, Colin M Shapiro, and Angela Colantonio. The pittsburgh sleep quality index as a screening tool for sleep dysfunction in clinical and non-clinical samples: A systematic review and meta- analysis.Sleep m...

  9. [17]

    Mundnich et al

    K. Mundnich et al. Tiles-2018: A longitudinal physiological and behavioral dataset of hospital workers.arXiv preprint arXiv:2003.08474, 2020

  10. [18]

    Jukka-Pekka Onnela and Scott L. Rauch. Harnessing smartphone-based digital phenotyp- ing to enhance behavioral and mental health.Neuropsychopharmacology, 41(7):1691–1696, Jun 2016. 23

  11. [19]

    Ellis, Sally Andrews, and Adam Joinson

    Lukasz Piwek, David A. Ellis, Sally Andrews, and Adam Joinson. The rise of consumer health wearables: Promises and barriers.PLOS Medicine, 13(2):1–9, 02 2016

  12. [20]

    Towards customizable foun- dation models for human activity recognition with wearable devices.Proc

    Minghui Qiu, Cekai Weng, Mingming Fan, and Kaishun Wu. Towards customizable foun- dation models for human activity recognition with wearable devices.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 9(3), September 2025

  13. [21]

    Unobtrusive perceived sleep quality monitoring in the wild.Proc

    Alvise Dei Rossi, Davide Marzorati, Tiziano Gerosa, Radoslava ˇSvihrov´ a, Silvia Santini, and Francesca Faraci. Unobtrusive perceived sleep quality monitoring in the wild.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 9(3), September 2025

  14. [22]

    Multi-task self-supervised learning for human activity detection.Proc

    Aaqib Saeed, Tanir Ozcelebi, and Johan Lukkien. Multi-task self-supervised learning for human activity detection.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 3(2), June 2019

  15. [23]

    The mini-mental state examination: a compre- hensive review.Journal of the American Geriatrics Society, 40(9):922–935, 1992

    Tom N Tombaugh and Nancy J McIntyre. The mini-mental state examination: a compre- hensive review.Journal of the American Geriatrics Society, 40(9):922–935, 1992

  16. [24]

    Vaizman, K

    Y. Vaizman, K. Ellis, and G. Lanckriet. Extrasensory: A multimodal dataset for context recognition in the wild. InProceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2017

  17. [25]

    The mini nutritional assessment (mna) and its use in grading the nutritional state of elderly patients.Nutrition, 15(2):116–122, 1999

    Bruno Vellas, Yves Guigoz, Philip J Garry, Fati Nourhashemi, David Bennahum, Sylvie Lauque, and Jean-Louis Albarede. The mini nutritional assessment (mna) and its use in grading the nutritional state of elderly patients.Nutrition, 15(2):116–122, 1999

  18. [26]

    Campbell

    Rui Wang, Fanglin Chen, Zhenyu Chen, Tianxing Li, Gabriella Harari, Stefanie Tignor, Xia Zhou, Dror Ben-Zeev, and Andrew T. Campbell. Studentlife: assessing mental health, academic performance and behavioral trends of college students using smartphones. In Proceedings of the 2...

  19. [27]

    Winter, Robyn J

    Jennifer E. Winter, Robyn J. MacInnis, Natalie Wattanapenpaiboon, and Caryl A. Nowson. Bmi and all-cause mortality in older adults: a meta-analysis.Obesity, 22(1):–, 2014

  20. [28]

    Who guidelines on physical activity and sedentary behaviour, 2020

    World Health Organization. Who guidelines on physical activity and sedentary behaviour, 2020

  21. [29]

    Wearable foundation models should go beyond static encoders, 2026

    Yu Yvonne Wu, Yuwei Zhang, Hyungjun Yoon, Ting Dang, Dimitris Spathis, Tong Xia, Qiang Yang, Jing Han, Dong Ma, Sung-Ju Lee, and Cecilia Mascolo. Wearable foundation models should go beyond static encoders, 2026

  22. [30]

    Kuehn, Jeremy F

    Xuhai Xu, Xin Liu, Han Zhang, Weichen Wang, Subigya Nepal, Yasaman Sefidgar, Woosuk Seo, Kevin S. Kuehn, Jeremy F. Huckins, Margaret E. Morris, Paula S. Nurius, Eve A. Riskin, Shwetak Patel, Tim Althoff, Andrew Campbell, Anind K. Dey, and Jennifer Mankoff. Globem: Cross-datase...

  23. [31]

    D. Y. Zhang, D. W. An, Y. L. Yu, J. D. Melgarejo, J. Boggia, D. S. Martens, T. W. Hansen, K. Asayama, T. Ohkubo, K. Stolarz-Skrzypek, S. Malyutina, E. Casiglia, L. Lind, G. E. Maestre, J. G. Wang, Y. Imai, K. Kawecka-Jaszcz, E. Sandoya, M. Rajzer, T. S. Nawrot, E. O’Brien, W. ...

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.