REVIEW 17 cited by
SensorLM: Learning the Language of Wearable Sensors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks.
Forward citations
Cited by 17 Pith papers
-
BCG-FM: A Foundation Model for Ambient Cardiac Health Sensing
BCG-FM, the first foundation model for ambient BCG, achieves 3.26-year MAE on biological age estimation and discriminates 15 health conditions using frozen embeddings from participant-level contrastive pretraining on ...
-
OpenWatch: A Multimodal Benchmark for Hand Gesture Recognition on Smartwatches
OpenWatch provides the first open multimodal smartwatch gesture dataset and benchmark, with MixToken and NormWear-Lora methods reaching 90% F1-score using 223k parameters versus 66% for 136M-parameter foundation models.
-
Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors
LLMs exhibit authority inversion by prioritizing natural-language user claims over numerical sensor data in conflicts, diagnosed with new geometric metrics and mitigated via layer-level calibration.
-
SleepLM: Natural-Language Intelligence for Human Sleep
A sleep-language foundation model trained with contrastive, captioning, and reconstruction objectives outperforms general LLMs and fine-tuned VLMs on zero-shot sleep staging, event localization, and cross-modal retrieval.
-
HEARTS: Benchmarking LLM Reasoning on Health Time Series
A 110-task benchmark across 20 health signal modalities shows current LLMs underperform specialized models and depend on simple heuristics rather than robust time-series reasoning.
-
Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series
Across 14 sensor generation settings, flow-matching models are the strongest overall baseline, while demographic conditioning, time-frequency modeling, and moderate synthetic augmentation improve hard regimes and down...
-
A robust PPG foundation model using multimodal physiological supervision
A PPG foundation model pretrained via multimodal ECG/respiratory contrastive sample selection on ICU data improves performance on 14 of 15 downstream tasks including field-like data while using 3x fewer subjects.
-
MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding
MyoSem is a multimodal alignment framework that maps EMG signals to text-based action semantics for bidirectional retrieval and improved generalization in hand action understanding.
-
TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health
TimeSRL uses semantic abstractions from time-series data optimized via reinforcement learning to achieve better cross-dataset generalization than standard ML or LLM baselines in mental health prediction.
-
Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs
IMU-to-4D uses wearable IMU data and repurposed LLMs to predict coherent 4D human motion plus coarse scene structure, outperforming cascaded state-of-the-art pipelines in temporal stability.
-
Sonata: A Hybrid World Model for Inertial Kinematics under Clinical Data Scarcity
Sonata is a small hybrid world model pre-trained to predict future IMU states that outperforms autoregressive baselines on clinical discrimination, fall-risk prediction, and cross-cohort transfer while fitting on-devi...
-
RF-LEGO: Modularized Signal Processing-Deep Learning Co-Design for RF Sensing via Deep Unrolling
RF-LEGO turns signal processing algorithms into trainable modular DL modules via deep unrolling, outperforming pure SP and DL baselines in RF sensing while preserving interpretability.
-
OSF: On Pre-training and Scaling of Sleep Foundation Models
Channel-masked self-supervised pretraining on a 166,500-hour multi-source sleep corpus yields OSF, which generalizes better to missing channels and scales with data and model size.
-
Gravity-Aware Hierarchical Routing for Lightweight SensorLLM on Human Activity Recognition
Introduces a lightweight gravity-aware routing head that improves macro-F1 on static classes in compressed SensorLLM for human activity recognition on the MHealth dataset.
-
Wearable AI in the Era of Large Sensor Models
Large Sensor Models trained on large-scale multimodal wearable data can provide a scalable, general framework for wearable AI by learning transferable representations across modalities and tasks.
-
Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook
The survey organizes foundation models for sensor-based HAR into a lifecycle taxonomy and identifies three trajectories: HAR-specific models from scratch, adaptation of general time-series models, and integration with...
-
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
PulseLM aggregates PPG data from 16 sources into 1M segments and 2.5M QA pairs for 12 tasks, providing a standardized benchmark for PPG-text multimodal learning.
Discussion (0). Sign in to comment.