REVIEW 3 cited by
UniMTS: Unified Pre-training for Motion Time Series
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Motion time series collected from mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR due to their low-power, always-on nature. However, given security and privacy concerns, building large-scale motion time series datasets remains difficult, preventing the development of pre-trained models for human activity analysis. Typically, existing models are trained and tested on the same dataset, leading to poor generalizability across variations in device location, device mounting orientation and human activity type. In this paper, we introduce UniMTS, the first unified pre-training procedure for motion time series that generalizes across diverse device latent factors and activities. Specifically, we employ a contrastive learning framework that aligns motion time series with text descriptions enriched by large language models. This helps the model learn the semantics of time series to generalize across activities. Given the absence of large-scale motion time series data, we derive and synthesize time series from existing motion skeleton data with all-joint coverage. Spatio-temporal graph networks are utilized to capture the relationships across joints for generalization across different device locations. We further design rotation-invariant augmentation to make the model agnostic to changes in device mounting orientations. Our model shows exceptional generalizability across 18 motion time series classification benchmark datasets, outperforming the best baselines by 340% in the zero-shot setting, 16.3% in the few-shot setting, and 9.2% in the full-shot setting.
Forward citations
Cited by 3 Pith papers
-
Foundation models for movement data: Are they ready for prime-time?
Foundation models for accelerometer data do not consistently beat supervised baselines on standard activity recognition, but some lead on fall/stress detection and as frozen feature extractors.
-
Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision
Frozen egocentric video features can discriminate freezing of gait above chance, but underperform IMU-based models, and the evidence for complementary visual information is only qualitative.
-
Foundation Models for CPS-IoT: Opportunities and Challenges
Current foundation models fall short on CPS-IoT needs in resource efficiency, spatial generalization, long-term context, and knowledge integration; the paper proposes desiderata and a community roadmap.
Discussion (0). Continue with ORCID to comment.