Pith. sign in

REVIEW 3 major objections 5 minor 36 references

Non-EEG sleep staging is limited by missing cortical information, not by model capacity.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 14:06 UTC pith:A3JP3MWF

load-bearing objection Solid decomposition of non-EEG sleep staging, but the headline EEG-gap number rests on a cross-dataset comparison that should be fixed before it is cited as a measurement of cortical information. the 3 major comments →

arxiv 2607.19441 v1 pith:A3JP3MWF submitted 2026-07-21 cs.HC

How Far Can Wearable-Compatible Signals Go? A Controlled Decomposition of Non-EEG Sleep Staging

classification cs.HC
keywords sleep stagingwearable sensingnon-EEG signalscontrolled decompositioncardiorespiratory physiologyconfidence-based abstentionCohen's kappaMamba2
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks how much sleep-stage information can be recovered from non-EEG signals and where the performance loss actually originates. Using a signal-quality ladder—consumer wrist HR/ACC, laboratory ECG/respiration/SpO2, and EEG+EOG—the authors hold a single compact model fixed and decompose the pipeline into signal source, physiological representation, temporal prior, and decision layers. They find that adding richer physiology helps most (+0.078 kappa), temporal smoothing helps modestly (+0.040), and the remaining gap to EEG is much larger (+0.304). They conclude that non-EEG staging is information-limited, not model-limited, and that confidence-based abstention is a practical way to make wearable staging trustworthy.

Core claim

Using the same Mamba2 model across all tiers, the paper reports that laboratory cardiorespiratory signals (ECG, respiration, SpO2) reach kappa=0.492 with Viterbi decoding, while EEG+EOG on Sleep-EDF-20 reaches kappa=0.796. The gap of 0.304 is interpreted as the information lost when cortical signals are absent, because the summed contribution of all non-EEG layers is only about +0.089. Consumer HR/ACC reaches only kappa=0.255, so the wearable penalty is split between degraded signal quality and the fundamental absence of EEG. Finally, the model's per-epoch confidence is well-calibrated: dropping the 20% lowest-confidence epochs raises kappa from 0.452 to 0.512, and at 50% coverage kappa=0.62

What carries the argument

The four-layer controlled decomposition framework, which separates sleep staging into signal source, physiological representation, temporal prior, and decision layers, and quantifies each layer's marginal contribution as the change in Cohen's kappa when that layer is added. The same compact Mamba2 state-space model with multi-scale temporal evidence aggregation is applied across all signal tiers, making signal modality the only variable. The coverage-kappa abstention curve, which ranks epochs by maximum softmax probability and computes agreement at decreasing coverage levels, converts model confidence into an operational tool for deciding when to report a stage and when to abstain.

Load-bearing premise

The 0.304 gap is treated as pure signal-modality loss, but it is measured across different datasets and feature pipelines, so cross-cohort and cross-feature differences are assumed negligible.

What would settle it

Simultaneously record EEG and non-EEG signals in the same subjects, run both through identical feature engineering and the same compact model, and compare the within-subject gap; a gap much smaller than 0.304 would show the headline number was inflated by dataset differences.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Investing in larger or deeper sequence models for non-EEG staging is unlikely to close the gap; richer physiological features and confidence calibration are the productive levers.
  • Wearable sleep reports should separate high-confidence from low-confidence epochs and consider merging N1 into a Light Sleep category, since N1 is intrinsically hard to stage without EEG.
  • The same controlled-decomposition protocol could be applied to raw PPG waveforms to quantify how much of the consumer penalty is due to derived signals rather than sensing hardware.
  • The modality ceiling implies that non-EEG staging should be positioned as sleep-structure and trend monitoring, not as a replacement for EEG-based clinical staging.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the EEG reference comes from a different cohort with different sensors and feature pipelines, a same-subject paired recording might produce a different magnitude of the modality gap; the within-SHHS ablations still support the representation-over-model conclusion.
  • Extending the decomposition to a transformer or a deeper CNN would test whether the layer contributions are architecture-dependent, as the paper's single-architecture limitation suggests.
  • If the coverage-kappa curve holds in a clinical cohort, abstention could be used to triage which nights deserve full PSG referral, increasing the utility of at-home monitoring.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a four-layer controlled decomposition framework for non-EEG sleep staging, separating signal source, physiological representation, temporal prior, and decision layers. Using the same compact Mamba2 model across a signal-quality ladder (Apple Watch HR/ACC, SHHS ECG/respiratory/SpO2, and Sleep-EDF-20 EEG+EOG), it reports that physiological representation gives the largest within-pipeline gain (Δκ=+0.078/0.089), temporal decoding adds only Δκ=+0.040, and the residual gap to EEG+EOG (Δκ=+0.304) is attributed to signal modality rather than model capacity. Confidence-based abstention is shown to improve κ from 0.452 to 0.512 at 80% coverage, and a label-shuffled control collapses to κ≈0. The paper argues that non-EEG sleep staging is limited by per-epoch physiological information content, not by temporal modeling or model capacity.

Significance. If the central claims hold, the paper contributes a useful diagnostic methodology for wearable sleep staging and provides evidence that non-EEG staging may be inherently bounded by autonomic/respiratory surrogates. The subject-disjoint splits, label-shuffled negative control, and channel ablations are well-designed internal controls that support the within-SHHS decomposition. The abstention analysis is a practical strength with clear translational relevance. However, the headline EEG-gap conclusion depends on a cross-dataset comparison (SHHS non-EEG vs. Sleep-EDF-20 EEG+EOG), which is not a controlled measurement of signal modality. The paper's central claim is therefore not yet established, though the internal ablations are valuable and the gap can be directly tested using SHHS EEG/EOG channels already available under the same protocol.

major comments (3)
  1. [§IV-A, Table II] Table II is internally inconsistent with the text and Table III. Row 1 is labeled 'ECG physiology alone (Viterbi)' with κ=0.403, row 2 '+ Resp/SpO2 physiology' gives κ=0.452 (Δ=+0.049), and row 3 '+ Viterbi temporal prior' gives κ=0.492 (Δ=+0.040). But the text and Table III report ECG argmax=0.373, combined argmax=0.452, and combined Viterbi=0.492. The table appears to mix argmax and Viterbi stages, so the marginal gains are not additive as claimed (0.403→0.452→0.492 is not the same as 0.373→0.452→0.492). The representation gain should be computed from argmax (0.452−0.373=+0.079), not 0.403→0.452. Please correct the table and ensure the four-layer decomposition's marginal Δκ values are consistent with the reported numbers in Tables III and the text.
  2. [§III-C, Layer 3] The Viterbi transition weight λ is selected as the 'best Viterbi result per channel configuration' from λ∈{0.1,0.3,0.5,1.0}. No validation procedure is described; if λ is chosen based on test-fold κ, the reported temporal-prior gain (Δκ=+0.040) is optimistically biased. This is load-bearing for the conclusion that temporal modeling contributes only modestly and that the per-epoch representation is the binding constraint. Report results for all λ values or a λ chosen on a held-out validation split. Table III states 'Viterbi at λ=1.0', which suggests a fixed choice, but Section III-C says best per configuration; please disambiguate.
  3. [§IV-A, Table II row 4; §V; §VI] The headline claim that Δκ=+0.304 is attributable to 'signal modality rather than model capacity' rests on comparing SHHS non-EEG results with Sleep-EDF-20 EEG+EOG results. This comparison confounds signal modality with dataset identity, cohort characteristics (20 healthy young subjects vs. 195 older community subjects with suspected sleep-disordered breathing), feature engineering, and scoring environment. The Limitations section acknowledges this but asserts the EEG ceiling is 'conservative'; that is only one direction of the confound. Cleaner subjects and labels in Sleep-EDF-20 could inflate the EEG κ relative to what the same architecture would achieve on SHHS EEG/EOG. Since SHHS PSG includes EEG/EOG channels under the same protocol and labels used for the non-EEG arms, the paper should run the identical Mamba2 pipeline on SHHS EEG/EOG to produce a matched-cohort reference. Without t
minor comments (5)
  1. [Abstract and §IV-A] 'Reflects missing cortical information rather than temporal modeling alone' is too strong given the cross-dataset reference; the abstract should say 'is consistent with missing cortical information' or similar until a matched-cohort EEG comparison is provided.
  2. [Table II] Row 4 uses the word 'Irrecoverable' for the EEG/EOG ceiling. This is an overstatement even under the authors' interpretation; the gap is 'not recovered by the tested non-EEG features and temporal model,' not proven irrecoverable.
  3. [§IV-A] The text states the representation gain is Δκ=+0.089 in Viterbi, but Table II reports +0.049 for the same step. Please align these values.
  4. [§III-B] The SHHS cohort is described as 'rpoint200' but the analysis uses 195 subjects; the dataset description would benefit from explaining why 5 of the 200 subjects were excluded and whether any sensitivity analysis was performed.
  5. [§IV-C, Fig. 3] The coverage-κ curve would benefit from reporting the number/percentage of epochs abstained at each coverage level per stage, especially for N1, to support the 'physiologically structured uncertainty' claim beyond mean confidence.

Circularity Check

0 steps flagged

No significant circularity: held-out ablations, negative control, and external reference; the cross-dataset EEG gap is a validity confound, not a circular derivation.

full rationale

The paper's central claim (the residual gap Δκ=+0.304 reflects missing cortical information) is an interpretation of a measured cross-dataset difference, not a quantity derived from the model's own fitted parameters. The within-SHHS four-layer decomposition is produced by held-out five-fold subject-disjoint evaluation; the representation and Viterbi contributions (Δκ=+0.078 and +0.040) are empirical ablations, not fitted to the output. The label-shuffled negative control (κ=−0.003) is an independently specified check, and the subject-overlap audit verifies no leakage. No parameter is fit to the target claim and then renamed as a prediction. The 'EEG+EOG ceiling' is an external dataset (Sleep-EDF-20) used as an upper anchor; the paper's Limitations explicitly acknowledge this is cross-dataset and uses simpler spectral features. That is a confound for the claim's validity, not a circular derivation: the result would be the same number regardless of the interpretation, and it could be falsified by running the same pipeline on SHHS EEG/EOG. There are no substantive self-citations referenced as authority: Mamba2 [1],[2] are architecture citations, not uniqueness claims or prior 'predictions' by this author. Per-stage confidence and abstention analyses are empirical summaries. Therefore no circularity is present; score 0.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

No new physical or architectural entities are invented. The central claims rest on dataset-label fidelity, cross-dataset comparability of the EEG reference, softmax-confidence validity, and the standard Markov transition assumption—the second of which is the most fragile.

free parameters (3)
  • Viterbi transition weight λ = 1.0 (reported as best from {0.1, 0.3, 0.5, 1.0})
    Controls the strength of the temporal prior; selecting the best value per channel configuration without a validation-based protocol can optimistically bias the reported temporal gain Δκ=+0.040.
  • Laplace smoothing α for transition probabilities = 1.0
    Hand-set smoothing constant for transition log-probabilities; conventional and low-impact but still a choice.
  • Mamba2 architecture and training hyperparameters = d_model=64, d_state=16, d_head=8, kernels {1,3,5,7}, lr=3e-3, wd=1e-4, epochs=10, seed 42
    Fixed across tiers so not a confound for layer comparisons, but no sensitivity analysis is shown; absolute κ values could shift with different hyperparameters.
axioms (4)
  • domain assumption AASM five-class sleep-stage labels in Apple Watch Sleep-Accel, SHHS, and Sleep-EDF-20 are accurate ground truth.
    All κ and F1 computations treat the provided PSG annotations as ground truth without inter-rater correction.
  • domain assumption Sleep-EDF-20 EEG+EOG performance can serve as a cross-dataset reference ceiling for SHHS non-EEG performance.
    The Δκ=+0.304 modality-gap interpretation assumes dataset, cohort, sensor, and feature-pipeline differences are negligible, which is questionable.
  • domain assumption Maximum softmax probability is a valid confidence measure for abstention.
    The coverage-κ curve relies on this; no calibration curve or comparison with ensemble/ MC-dropout uncertainty is provided in the main results.
  • domain assumption Viterbi transition probabilities estimated from training-fold labels approximate the true sleep-stage transition structure.
    The temporal-prior layer uses this standard HMM-style assumption; violations could bias the temporal contribution estimate.

pith-pipeline@v1.3.0-alltime-deepseek · 11713 in / 10777 out tokens · 92708 ms · 2026-08-01T14:06:16.516403+00:00 · methodology

0 comments
read the original abstract

Consumer wearables increasingly infer sleep stages from signals including heart rate, accelerometry, and photoplethysmography. However, existing studies often report end-to-end performance under a fixed signal setting, making it difficult to determine whether the observed performance comes from genuine physiological decoding, temporal priors, or dataset-specific confounds. To address this limitation, we introduce a four-layer controlled decomposition framework for non-EEG sleep staging, covering signal source, physiological representation, temporal prior, and decision layers. The framework is evaluated across a signal-quality ladder spanning Apple Watch Sleep-Accel ($N=31$), the Sleep Heart Health Study ($N=195$, laboratory ECG, respiratory, and SpO$_2$ signals), and Sleep-EDF-20 as an EEG+EOG reference, using the same compact Mamba2 model throughout. Laboratory cardiorespiratory signals reach $\kappa=0.492$, while EEG+EOG reaches $\kappa=0.796$, leaving a residual gap of $\Delta\kappa=+0.304$ that reflects missing cortical information rather than temporal modeling alone. Consumer HR/ACC reaches only $\kappa=0.255$, quantifying the additional penalty of derived wearable signals and real-world sensing constraints. Confidence-based abstention provides a calibrated operating mode: removing the 20% lowest-confidence epochs increases $\kappa$ from $0.452$ to $0.512$, while a label-shuffled control collapses to $\kappa=-0.003$. These results support non-EEG sleep staging as coarse, confidence-aware sleep-structure monitoring rather than EEG-equivalent five-class clinical staging.

Figures

Figures reproduced from arXiv: 2607.19441 by Yi Wang.

Figure 1
Figure 1. Figure 1: Overview of the controlled decomposition framework for non-EEG sleep staging. The framework separates the staging pipeline into four independently [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Experimental workflow for the controlled decomposition study. Signals from consumer wearable, laboratory non-EEG, and EEG+EOG reference [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Coverage-κ curve for the SHHS combined model (argmax). Epochs are sorted by decreasing softmax confidence. Shaded region: 95% bootstrap CI. Dashed horizontal lines: EEG+EOG reference (κ = 0.796) and technologist inter-rater agreement (κ ≈ 0.80). TABLE IV PER-STAGE MEAN MODEL CONFIDENCE (MAX SOFTMAX PROBABILITY). Stage Mean confidence N epochs Wake 0.771 50,626 N1 0.630 7,349 N2 0.643 82,853 N3 0.706 25,948… view at source ↗
Figure 4
Figure 4. Figure 4: Normalized confusion matrix for SHHS combined model (pooled [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Signal-availability ladder. Horizontal bars show per-subject Cohen’s [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

36 extracted references · 4 linked inside Pith

  1. [1]

    Mamba: Linear-time sequence modeling with selective state spaces,

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,”arXiv preprint arXiv:2312.00752, 2023, accepted at Conference on Language Modeling (COLM) 2024

  2. [2]

    Transformers are SSMs: Generalized models and ef- ficient algorithms through structured state space duality,

    T. Dao and A. Gu, “Transformers are SSMs: Generalized models and ef- ficient algorithms through structured state space duality,” inInternational Conference on Machine Learning (ICML), 2024, arXiv:2405.21060

  3. [3]

    Harnessing electroencephalography con- nectomes for cognitive and clinical neuroscience,

    Y . Zhang and Z. S. Chen, “Harnessing electroencephalography con- nectomes for cognitive and clinical neuroscience,”Nature Biomedical Engineering, vol. 9, no. 8, pp. 1186–1201, 2025

  4. [4]

    Machine learning prediction on spatial and environmental perception and work efficiency using electroencephalography including cross-subject scenarios,

    T. Yu, J. Li, Y . Jin, W. Wu, X. Ma, W. Xu, and S. Lu, “Machine learning prediction on spatial and environmental perception and work efficiency using electroencephalography including cross-subject scenarios,”Jour- nal of Building Engineering, vol. 99, p. 111644, 2025

  5. [5]

    Electrooculography dataset for objective spatial naviga- tion assessment in healthy participants,

    M. Zibandehpoor, F. Alizadehziri, A. A. Larki, S. Teymouri, and M. Delrobaei, “Electrooculography dataset for objective spatial naviga- tion assessment in healthy participants,”Scientific Data, vol. 12, no. 1, p. 553, 2025

  6. [6]

    Comprehensive human locomotion and electromyography dataset: Gait120,

    J. Boo, D. Seo, M. Kim, and S. Koo, “Comprehensive human locomotion and electromyography dataset: Gait120,”Scientific Data, vol. 12, no. 1, p. 1023, 2025

  7. [7]

    From pose to muscle: Multimodal learning for piano hand muscle electromyography,

    R. Liu, Y . Peng, T. Oku, C.-C. Liao, E. Wu, S. Furuya, and H. Koike, “From pose to muscle: Multimodal learning for piano hand muscle electromyography,”Advances in Neural Information Processing Systems, vol. 38, pp. 115 558–115 587, 2026

  8. [8]

    Polysomnography in transition: Reassessing its role in the future of sleep medicine,

    D. Leger, C. Mutti, A. Rouen, and L. Parrino, “Polysomnography in transition: Reassessing its role in the future of sleep medicine,”Journal of Sleep Research, vol. 34, no. 6, p. e70217, 2025

  9. [9]

    Severity classification of obstructive sleep apnea using aasm and separ criteria: A cross-sectional reclassification analysis,

    J. Nieto-Pino, E. Retamal-Riquelme, M. Henriquez-Beltr ´an, M. Otto- Ya˜nez, R. Torres-Castro, and G. Labarca, “Severity classification of obstructive sleep apnea using aasm and separ criteria: A cross-sectional reclassification analysis,”European Archives of Oto-Rhino-Laryngology, vol. 283, no. 2, pp. 1279–1287, 2026

  10. [10]

    R. B. Berry, R. Brooks, C. Gamaldo, S. M. Harding, R. M. Lloyd, S. F. Quan, M. T. Troester, and B. V . Vaughn,The AASM Manual for the Scoring of Sleep and Associated Events: Rules, Terminology and Technical Specifications, Version 2.4. American Academy of Sleep Medicine, 2017

  11. [11]

    Deepsleepnet: A model for automatic sleep stage scoring based on raw single-channel eeg,

    A. Supratak, H. Dong, C. Wu, and Y . Guo, “Deepsleepnet: A model for automatic sleep stage scoring based on raw single-channel eeg,”IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 25, no. 11, pp. 1998–2008, 2017

  12. [12]

    U-sleep: resilient high-frequency sleep staging,

    M. Perslev, S. Darkner, L. Kempfner, M. Nikolic, P. J. Jennum, and C. Igel, “U-sleep: resilient high-frequency sleep staging,”npj Digital Medicine, vol. 4, no. 1, p. 72, 2021

  13. [13]

    XSleepNet: Multi-view sequential model for automatic sleep staging,

    H. Phan, O. Y . Chen, P. Koch, Y . Liu, K. Mikkelsen, M. De V os, P. Maass, and A. Miklody, “XSleepNet: Multi-view sequential model for automatic sleep staging,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 9, pp. 5903–5915, 2022

  14. [14]

    Privacy in consumer wearable technologies: a living systematic analysis of data policies across leading manufacturers,

    C. Doherty, M. Baldwin, R. Lambe, M. Altini, and B. Caulfield, “Privacy in consumer wearable technologies: a living systematic analysis of data policies across leading manufacturers,”npj Digital Medicine, vol. 8, no. 1, p. 363, 2025

  15. [15]

    Heart rate variability in normal and pathological sleep,

    E. Tobaldini, L. Nobili, S. Strada, K. R. Casali, A. Braghiroli, and N. Montano, “Heart rate variability in normal and pathological sleep,” Frontiers in Physiology, vol. 4, p. 294, 2013

  16. [16]

    Autonomic activity during human sleep as a function of time and sleep stage,

    J. Trinder, J. Kleiman, M. Carrington, S. Smith, S. Breen, N. Tan, and Y . Kim, “Autonomic activity during human sleep as a function of time and sleep stage,”Journal of Sleep Research, vol. 10, no. 4, pp. 253–264, 2001

  17. [17]

    Deep learning for automated sleep staging using instantaneous heart rate,

    N. Sridhar, A. Shoeb, P. Stephens, A. Kini, J. Barber, J. Getchius, S. Perkins, and J. Waugh, “Deep learning for automated sleep staging using instantaneous heart rate,”npj Digital Medicine, vol. 3, no. 1, p. 106, 2020

  18. [18]

    Sleep stage classification from heart-rate variability using long short-term memory neural networks,

    M. Radha, P. Fonseca, A. Moreau, M. Ross, A. Cerny, P. Anderer, X. Long, and R. M. Aarts, “Sleep stage classification from heart-rate variability using long short-term memory neural networks,”Scientific Reports, vol. 9, p. 14149, 2019

  19. [19]

    Sleep staging from electrocardiography and respiration with deep learning,

    H. Sun, W. Ganglberger, E. Panneerselvam, M. J. Leone, S. A. Quadri, B. Goparaju, R. A. Tesh, O. Akeju, R. J. Thomas, and M. B. Westover, “Sleep staging from electrocardiography and respiration with deep learning,”arXiv preprint arXiv:1908.11463, 2019

  20. [20]

    Tinysleepnet: An efficient deep learning model for sleep stage scoring based on raw single-channel eeg,

    A. Supratak and Y . Guo, “Tinysleepnet: An efficient deep learning model for sleep stage scoring based on raw single-channel eeg,”Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 641–644, 2020

  21. [21]

    SleepTransformer: automatic sleep staging with interpretability and un- certainty quantification,

    H. Phan, K. Mikkelsen, O. Y . Chen, P. Koch, A. Mertins, and M. De V os, “SleepTransformer: automatic sleep staging with interpretability and un- certainty quantification,”IEEE Transactions on Biomedical Engineering, vol. 69, no. 8, pp. 2456–2467, 2022

  22. [22]

    A systematic review of sensing technologies for wearable sleep staging,

    S. A. Imtiaz, “A systematic review of sensing technologies for wearable sleep staging,”Sensors, vol. 21, no. 5, p. 1562, 2021

  23. [23]

    Sleep stage prediction with raw acceleration and photoplethysmography heart rate data derived from a consumer wearable device,

    O. Walch, Y . Huang, D. Forger, and C. Goldstein, “Sleep stage prediction with raw acceleration and photoplethysmography heart rate data derived from a consumer wearable device,”Sleep, vol. 42, no. 12, 2019

  24. [24]

    Performance of seven consumer sleep-tracking devices compared with polysomnography,

    E. D. Chinoy, J. A. Cuellar, K. E. Huwa, J. T. Jameson, C. H. Watson, S. C. Bessman, D. A. Hirsch, A. D. Cooper, S. P. A. Drummond, and R. R. Markwald, “Performance of seven consumer sleep-tracking devices compared with polysomnography,”Sleep, vol. 44, no. 5, p. zsaa291, 2021

  25. [25]

    The promise of sleep: A multi-sensor approach for accurate sleep stage detection using the Oura ring,

    M. Altini and H. Kinnunen, “The promise of sleep: A multi-sensor approach for accurate sleep stage detection using the Oura ring,” Sensors, vol. 21, no. 13, p. 4302, 2021

  26. [26]

    An evaluation of cardiorespiratory and movement features with respect to sleep-stage classification,

    T. Willemen, D. Van Deun, V . Verhaert, M. Vandekerckhove, V . Ex- adaktylos, J. Verbraecken, S. Van Huffel, B. Haex, and J. Vander Sloten, “An evaluation of cardiorespiratory and movement features with respect to sleep-stage classification,”IEEE Journal of Biomedical and Health Informatics, vol. 18, no. 2, pp. 661–669, 2014

  27. [27]

    SleepPPG-Net: a deep learning algorithm for robust sleep staging from continuous photoplethysmography,

    K. Kotzen, P. H. Charlton, S. Salabi, L. Amar, A. Landesberg, and J. A. Behar, “SleepPPG-Net: a deep learning algorithm for robust sleep staging from continuous photoplethysmography,”arXiv preprint arXiv:2202.05735, 2022

  28. [28]

    The sleep heart health study: design, rationale, and methods,

    S. F. Quan, B. V . Howard, C. Iber, J. P. Kiley, F. J. Nieto, G. T. O’Connor, D. M. Rapoport, S. Redline, J. Robbins, J. M. Sametet al., “The sleep heart health study: design, rationale, and methods,”Sleep, vol. 20, no. 12, pp. 1077–1085, 1997

  29. [29]

    The National Sleep Research Resource: towards a sleep data commons,

    G.-Q. Zhang, L. Cui, R. Mueller, S. Tao, M. Kim, M. Rueschman, S. Mariani, D. Mobley, and S. Redline, “The National Sleep Research Resource: towards a sleep data commons,”Journal of the American Medical Informatics Association, vol. 25, no. 10, pp. 1351–1358, 2018

  30. [30]

    Analysis of a sleep-dependent neuronal feedback loop: the slow-wave microcontinuity of the EEG,

    B. Kemp, A. H. Zwinderman, B. Tuk, H. A. Kamphuisen, and J. J. Obery´e, “Analysis of a sleep-dependent neuronal feedback loop: the slow-wave microcontinuity of the EEG,”IEEE Transactions on Biomed- ical Engineering, vol. 47, no. 9, pp. 1185–1194, 2000

  31. [31]

    PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals,

    A. L. Goldberger, L. A. Amaral, L. Glass, J. M. Hausdorff, P. C. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley, “PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals,”Circulation, vol. 101, no. 23, pp. e215–e220, 2000

  32. [32]

    A real-time QRS detection algorithm,

    J. Pan and W. J. Tompkins, “A real-time QRS detection algorithm,”IEEE Transactions on Biomedical Engineering, vol. 32, no. 3, pp. 230–236, 1985

  33. [33]

    Error bounds for convolutional codes and an asymptoti- cally optimum decoding algorithm,

    A. J. Viterbi, “Error bounds for convolutional codes and an asymptoti- cally optimum decoding algorithm,”IEEE Transactions on Information Theory, vol. 13, no. 2, pp. 260–269, 1967

  34. [34]

    A coefficient of agreement for nominal scales,

    J. Cohen, “A coefficient of agreement for nominal scales,”Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960

  35. [35]

    Consumer sleep technology: An american academy of sleep medicine position statement,

    S. Khosla, M. C. Deak, D. Gault, C. A. Goldstein, D. Hwang, Y . Kwon, D. O’Hearn, S. Schutte-Rodin, M. Yurcheshen, I. M. Rosenet al., “Consumer sleep technology: An american academy of sleep medicine position statement,”Journal of Clinical Sleep Medicine, vol. 14, no. 5, pp. 877–880, 2018

  36. [36]

    Wearable technologies for developing sleep and circadian biomarkers: a summary of workshop discussions,

    C. M. Depner, P. C. Cheng, J. K. Devine, S. Khosla, M. de Zambotti, R. Robillard, A. Vakulin, and S. P. Drummond, “Wearable technologies for developing sleep and circadian biomarkers: a summary of workshop discussions,”Sleep, vol. 43, no. 2, p. zsz254, 2020