Pith. sign in

REVIEW 10 cited by

SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17766 v1 pith:2EHGQJT6 submitted 2024-05-28 cs.LG cs.AIeess.SP

SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals

classification cs.LG cs.AIeess.SP
keywords sleepmulti-modalsleepfmlearningauprcaurocbraincontrastive
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Sleep is a complex physiological process evaluated through various modalities recording electrical brain, cardiac, and respiratory activities. We curate a large polysomnography dataset from over 14,000 participants comprising over 100,000 hours of multi-modal sleep recordings. Leveraging this extensive dataset, we developed SleepFM, the first multi-modal foundation model for sleep analysis. We show that a novel leave-one-out approach for contrastive learning significantly improves downstream task performance compared to representations from standard pairwise contrastive learning. A logistic regression model trained on SleepFM's learned embeddings outperforms an end-to-end trained convolutional neural network (CNN) on sleep stage classification (macro AUROC 0.88 vs 0.72 and macro AUPRC 0.72 vs 0.48) and sleep disordered breathing detection (AUROC 0.85 vs 0.69 and AUPRC 0.77 vs 0.61). Notably, the learned embeddings achieve 48% top-1 average accuracy in retrieving the corresponding recording clips of other modalities from 90,000 candidates. This work demonstrates the value of holistic multi-modal sleep modeling to fully capture the richness of sleep recordings. SleepFM is open source and available at https://github.com/rthapa84/sleepfm-codebase.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces

    cs.LG 2026-05 unverdicted novelty 7.0

    NeuroAtlas benchmarks foundation models on 42 EEG datasets and reports that EEG-specific models do not consistently outperform generic time-series models, standard metrics miss clinical utility, and rankings vary by domain.

  2. Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning

    cs.LG 2026-05 unverdicted novelty 7.0

    xMAE pretrains biosignal representations via masked cross-modal reconstruction of temporally ordered signals like ECG and PPG, outperforming baselines on 15 of 19 downstream tasks including cardiovascular prediction a...

  3. Agentic AI-enabled discovery across large-scale sleep physiology

    cs.MA 2026-07 conditional novelty 6.0

    A human-guided agentic AI pipeline analyzed ~124,000 polysomnograms and reported five sleep findings, including an association between reduced N2 brain coupling and incident Parkinson's (HR 1.48) and Alzheimer's (HR 1...

  4. OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

    q-bio.NC 2026-04 unverdicted novelty 6.0

    OmniMouse demonstrates data-driven scaling in multi-task brain models on a 150B-token neural dataset, achieving SOTA across prediction, decoding, and forecasting while model size gains saturate.

  5. PRISM-CTG: A Foundation Model for Cardiotocography Analysis with Multi-View SSL

    cs.LG 2026-04 unverdicted novelty 6.0

    PRISM-CTG is the first large-scale foundation model for cardiotocography that uses multi-view self-supervised learning on unlabeled data to learn transferable representations, outperforming baselines on seven downstre...

  6. OSF: On Pre-training and Scaling of Sleep Foundation Models

    cs.LG 2026-02 conditional novelty 6.0

    Channel-masked self-supervised pretraining on a 166,500-hour multi-source sleep corpus yields OSF, which generalizes better to missing channels and scales with data and model size.

  7. SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures

    cs.AI 2026-02 conditional novelty 6.0

    SleepMaMi, a dual-encoder sleep foundation model pretrained on 158K hours of PSG, matches or beats existing sleep foundation models on staging, apnea segmentation, and disease prediction.

  8. Wearable AI in the Era of Large Sensor Models

    eess.SP 2026-04 unverdicted novelty 5.0

    Large Sensor Models trained on large-scale multimodal wearable data can provide a scalable, general framework for wearable AI by learning transferable representations across modalities and tasks.

  9. Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook

    eess.SP 2026-04 accept novelty 5.0

    The survey organizes foundation models for sensor-based HAR into a lifecycle taxonomy and identifies three trajectories: HAR-specific models from scratch, adaptation of general time-series models, and integration with...

  10. From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning

    eess.AS 2026-07 unverdicted novelty 3.0

    A survey that organizes audio SSL into five objective paradigms, relates their demands to architectural biases, and interprets downstream applications as tests of generalization.