Pith. sign in

REVIEW 17 cited by

SensorLM: Learning the Language of Wearable Sensors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.09108 v1 pith:KUPSGVJO submitted 2025-06-10 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords datasensorlmsensorlanguagewearablelearningreal-worldsensor-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BCG-FM: A Foundation Model for Ambient Cardiac Health Sensing

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    BCG-FM, the first foundation model for ambient BCG, achieves 3.26-year MAE on biological age estimation and discriminates 15 health conditions using frozen embeddings from participant-level contrastive pretraining on ...

  2. OpenWatch: A Multimodal Benchmark for Hand Gesture Recognition on Smartwatches

    cs.HC 2026-05 conditional novelty 7.0 of 10

    OpenWatch provides the first open multimodal smartwatch gesture dataset and benchmark, with MixToken and NormWear-Lora methods reaching 90% F1-score using 223k parameters versus 66% for 136M-parameter foundation models.

  3. Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors

    cs.AI 2026-04 unverdicted novelty 7.0 of 10

    LLMs exhibit authority inversion by prioritizing natural-language user claims over numerical sensor data in conflicts, diagnosed with new geometric metrics and mitigated via layer-level calibration.

  4. SleepLM: Natural-Language Intelligence for Human Sleep

    cs.AI 2026-02 conditional novelty 7.0 of 10

    A sleep-language foundation model trained with contrastive, captioning, and reconstruction objectives outperforms general LLMs and fine-tuned VLMs on zero-shot sleep staging, event localization, and cross-modal retrieval.

  5. HEARTS: Benchmarking LLM Reasoning on Health Time Series

    cs.LG 2026-02 conditional novelty 7.0 of 10

    A 110-task benchmark across 20 health signal modalities shows current LLMs underperform specialized models and depend on simple heuristics rather than robust time-series reasoning.

  6. Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Across 14 sensor generation settings, flow-matching models are the strongest overall baseline, while demographic conditioning, time-frequency modeling, and moderate synthetic augmentation improve hard regimes and down...

  7. A robust PPG foundation model using multimodal physiological supervision

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    A PPG foundation model pretrained via multimodal ECG/respiratory contrastive sample selection on ICU data improves performance on 14 of 15 downstream tasks including field-like data while using 3x fewer subjects.

  8. MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    MyoSem is a multimodal alignment framework that maps EMG signals to text-based action semantics for bidirectional retrieval and improved generalization in hand action understanding.

  9. TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    TimeSRL uses semantic abstractions from time-series data optimized via reinforcement learning to achieve better cross-dataset generalization than standard ML or LLM baselines in mental health prediction.

  10. Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    IMU-to-4D uses wearable IMU data and repurposed LLMs to predict coherent 4D human motion plus coarse scene structure, outperforming cascaded state-of-the-art pipelines in temporal stability.

  11. Sonata: A Hybrid World Model for Inertial Kinematics under Clinical Data Scarcity

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Sonata is a small hybrid world model pre-trained to predict future IMU states that outperforms autoregressive baselines on clinical discrimination, fall-risk prediction, and cross-cohort transfer while fitting on-devi...

  12. RF-LEGO: Modularized Signal Processing-Deep Learning Co-Design for RF Sensing via Deep Unrolling

    cs.DC 2026-04 unverdicted novelty 6.0 of 10

    RF-LEGO turns signal processing algorithms into trainable modular DL modules via deep unrolling, outperforming pure SP and DL baselines in RF sensing while preserving interpretability.

  13. OSF: On Pre-training and Scaling of Sleep Foundation Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Channel-masked self-supervised pretraining on a 166,500-hour multi-source sleep corpus yields OSF, which generalizes better to missing channels and scales with data and model size.

  14. Gravity-Aware Hierarchical Routing for Lightweight SensorLLM on Human Activity Recognition

    eess.SP 2026-06 unverdicted novelty 5.0 of 10

    Introduces a lightweight gravity-aware routing head that improves macro-F1 on static classes in compressed SensorLLM for human activity recognition on the MHealth dataset.

  15. Wearable AI in the Era of Large Sensor Models

    eess.SP 2026-04 unverdicted novelty 5.0 of 10

    Large Sensor Models trained on large-scale multimodal wearable data can provide a scalable, general framework for wearable AI by learning transferable representations across modalities and tasks.

  16. Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook

    eess.SP 2026-04 accept novelty 5.0 of 10

    The survey organizes foundation models for sensor-based HAR into a lifecycle taxonomy and identifies three trajectories: HAR-specific models from scratch, adaptation of general time-series models, and integration with...

  17. PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning

    cs.CL 2026-02 unverdicted novelty 5.0 of 10

    PulseLM aggregates PPG data from 16 sources into 1M segments and 2.5M QA pairs for 12 tasks, providing a standardized benchmark for PPG-text multimodal learning.

Pith tools