Pith. sign in

REVIEW 3 major objections 6 minor 87 references

Exploring the Alignment of Perceived and Measured Sleep Quality with Working Memory using Consumer Wearables

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Day-to-day changes in REM sleep, overnight heart rate, bedtimes, and N-back scores are the strongest predictors of how people rate their sleep, while REM sleep and the self-rating together best predict working-memory performance.

desk verdict A useful public dataset and a mostly sound weak-correlation story, but the 'three subgroups' novelty is built on in-sample model selection and needs out-of-sample validation before it can carry the paper. read the letter →

arxiv 2507.19491 v1 pith:CVTO3CVY submitted 2025-05-31 cs.HC cs.CY

classification cs.HCcs.CY
keywords wearablesleeptrackingself-assessmentREMworkingmemoryN-backtaskecologicalmomentaryassessmentOuraringin-the-wildstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a consumer sleep tracker tells users something they do not already feel about their sleep. Across 4–8 weeks, 29 participants rated each night's sleep against the previous night, wore an Oura ring, and took a daily 3-back working-memory test. The paper argues that day-to-day changes in REM (rapid eye movement) sleep duration, overnight heart rate, bedtimes, and N-back scores are the strongest predictors of how people rate their sleep, while REM sleep from the current and prior night plus the self-rating best predict working-memory performance. The associations are real but modest: the model explains $R^2 = 0.15$ of the variance in self-ratings, and the direct correlation between self-rating and N-back score is $\rho = 0.08$. The study also finds that users split into three groups by which sleep signals their self-reports track, so the informational value of a sleep tracker is not the same for everyone.

What carries the argument

The central mechanism is the day-pair difference design: every sleep feature is converted into a change from the previous night, matched to a five-point comparative self-rating, so the outcome is a first derivative of sleep quality rather than a level. This is paired with reverse regression—modeling the subjective rating as the outcome rather than the sensor data—with participant as a random effect, and cross-checked by an ordinal cumulative link model and a Bayesian regression. The N-back score itself is a composite of accuracy and response time, $S = C + \frac{T_{\max}-T}{T_{\max}-T_{\min}}$, with correctness and time each scaled to $[0,1]$, analyzed only within participants because N-back is not treated as comparable between individuals. Gaussian mixture clustering on the sleep features then splits participants into three groups to test whether the same predictors hold for everyone.

What would settle it

Run overnight polysomnography in parallel with the ring for a subset of these participants and check whether night-to-night changes in ring-derived REM duration match polysomnography-derived changes. If the ring's REM differences do not track the reference measurements, the paper's REM-based conclusions lose their foundation; alternatively, a pre-registered replication that swaps the order so the N-back test comes before the sleep-rating question would test whether the self-rating-to-N-back link survives.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that subjective sleep quality is not a single internal feeling but a mixture of signals, and a wearable's sensor stream captures part of that mixture while missing the rest. Using a daily comparative question ('How was your sleep compared to yesterday?') avoids the floor and ceiling effects that plague absolute scales. Mixed-effects, ordinal, and Bayesian models converge on the same feature set: a night rated as better tends to be one with more REM sleep than the previous night, a lower average overnight heart rate, a later bedtime start and end, and a higher N-back score; together these account for $R^2 = 0.15$ of the variance in self-ratings (0.18 with participant random effects). For working memory, the strongest predictors are absolute REM duration, the change in REM duration from the prior night (negative coefficient), and the participant's self-rating, with marginal $R^2 = 0.46$. The authors conclude that sensor readings and self-reports carry partly non-overlapping information, and that grouping participants by their sleep-feature profiles reveals three subgroups with different sensitivities to these markers.

Load-bearing premise

The load-bearing premise is that the Oura ring's nightly sleep-stage estimates, especially REM, are accurate enough to track real within-person day-to-day changes; if the ring's REM numbers are noisy or biased for some participants, the reported links between REM differences and both self-rated sleep quality and working-memory scores would be distorted.

Editorial extensions

If this is right

  • If correct, sleep trackers offer genuine but partial information gain: self-ratings do not fully capture what sensors measure, and sensors do not fully capture what self-ratings capture.
  • REM sleep acts over at least two nights: both absolute REM duration and the change from the prior night enter the working-memory model, so single-night summaries miss part of the effect.
  • Because self-assessment itself predicts N-back performance beyond REM measures, asking users how they slept adds predictive information that the ring alone does not provide.
  • The three user groups imply that a one-size-fits-all wearable interface will over-serve some users and under-serve others; designs that identify which signals a user tracks could tailor feedback.
  • For HCI, translating sleep data into predicted cognitive readiness may be the actionable format that keeps users engaged with trackers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test the paper leaves implicit: whether the self-rating-to-N-back link is causal or an artifact of question order and expectation, since participants rated sleep before the test and the correlation is only $\rho = 0.08$; a version with randomized feedback or reversed order would separate these.
  • The three subgroups' different sensitivities to bedtime start versus bedtime end hint at chronotype differences; a larger study with morning-eveningness questionnaires could make the grouping interpretable rather than purely data-driven.
  • If the ring's REM stage estimates are validated per-participant against polysomnography, the same modeling pipeline could be rerun on those validated days, which would tell whether the REM findings survive measurement error.
  • The public dataset invites direct replication with other wearables or clinical-grade devices, and the day-pair difference format could be reused as a standard template for comparative sleep self-reports.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper reports an in-the-wild study of 29 university students who wore an Oura ring for 4–8 weeks, answered a daily relative sleep-quality self-assessment, and completed a 3-back working-memory task. The authors compute Spearman correlations between sleep features and self-assessment/N-back scores, then fit mixed-effects, ordinal, and Bayesian models to identify predictors of self-assessment and N-back performance. They further divide participants into subgroups via a Gaussian mixture model and claim that self-report sensitivity toward sleep markers differs across three identified groups. The main reported findings are that differences in REM sleep duration, nocturnal heart rate, bedtimes, and N-back scores predict sleep self-assessment, and that REM sleep and self-assessment predict N-back scores. The dataset is promised to be made publicly available.

Significance. If the findings hold, the paper would provide a useful longitudinal dataset and a modest confirmation that consumer wearables partially align with subjective sleep experience, with the additional observation that this alignment may be heterogeneous across users. The data-sharing commitment and the within-subject, in-the-wild design are strengths. However, the central novelty—the three-group differential-information-gain claim—rests on an in-sample model-selection procedure without out-of-sample validation, and the abstract's wording ('highly predict') substantially overstates the reported effect sizes (R²_marginal = 0.15 for self-assessment; rho = 0.08 for the N-back–self-report association). The subgroup analysis is exploratory, and the authors themselves concede in Section 5.1 that the number of three groups may be an artifact of data limitations. The paper's contribution is therefore best viewed as a descriptive, exploratory study rather than a confirmatory demonstration of distinct user archetypes; with appropriate reanalysis or tempering of claims, it could be a worthwhile addition to the HCI/sleep-tracking literature.

major comments (3)
  1. [Section 4.2, Table 7] The number of GMM subgroups is selected by comparing in-sample R² values (0.157/0.177, 0.185/0.208, 0.163/0.190 for n=2, 3, 4) and choosing n=3 because it gives the highest R². These values come from models fit and evaluated on the same participants, with no cross-validation, bootstrap, or correction for model selection. Since the three-group claim is the paper's central novelty ('self-report sensitivity towards sleep markers differs among participants'), this conclusion is not supported as it stands. The group sizes are small (5, 11, 13), and several coefficients in Table 7 have large standard errors (e.g., Group 1 N-Back coefficient 0.348, SE 0.409). The authors' own admission in Section 5.1 that 'there are indications that the currently identified number of 3 groups comes from data limitations' underscores that this result is fragile. I recommend adding out-of-sample validation such as repeated K-fold cross-validation or bootstrap stability of group assignments, or explicitly reframing the subgroup analysis as exploratory and removing the three-group claim from the abstract.
  2. [Abstract and Section 4.2, Table 3] The abstract claims that the identified features 'highly predict sleep self-assessment in significance and effect size,' but Table 3 reports R²_marginal = 0.15 and R²_conditional = 0.18, and Table 2 shows a Spearman correlation of only 0.08 between N-back score and self-assessment. These are weak-to-modest in-sample associations, which is also acknowledged in the body text ('weak correlations'). The language in the abstract and conclusion should be calibrated to the observed effect sizes, or the authors should provide out-of-sample predictive metrics (e.g., cross-validated R² or RMSE) to justify a stronger claim.
  3. [Section 4.2, Tables 3 and 4] The mixed-effects models are built by iteratively removing features to minimize VIF and AIC on the same dataset used to evaluate them, and the reported p-values and R² are not adjusted for this selection. This inflates the apparent significance and effect sizes. A notable discrepancy: Table 2 reports no significant raw correlation between REM sleep duration and N-back score (rho = 0.036, p = 0.162), yet Table 4 reports REM sleep duration and its difference as significant predictors with R²_marginal = 0.46. This pattern is consistent with selection overfitting. The authors should either prespecify their predictors, use a proper post-selection inference procedure, or clearly label these models as exploratory and report the number of candidate models considered.
minor comments (6)
  1. [Section 3.3, Eq. (1)] Please clarify whether T_max and T_min in Eq. (1) are computed per participant, per day, or globally across the whole dataset; the interpretation of the composite N-back score depends on this choice.
  2. [Section 3.2 and Section 4.2] The Likert responses are described as '1) Much better ... 5) Much worse' but are later mapped symmetrically to +2x, +x, 0, −x, −2x. Please state explicitly where this mapping is applied and whether the arbitrary scale factor x affects any reported coefficients or R² values.
  3. [Table 2] The right-hand column header reads 'Pearson correlation of N-Back score,' but the text states that Spearman's rank correlation is used because of skewness. Please correct this inconsistency.
  4. [Section 3.4, Table 1] The term 'Consecutive Day-Pairs' is used in Table 1 but is not formally defined until Section 3.5; moving the definition earlier would improve readability.
  5. [References] References [28] and [29] appear to be the same paper (Kainec et al., 2024) listed twice; please deduplicate.
  6. [Figure 6] The group color coding (blue, beige, red) may be difficult to distinguish for color-blind readers; consider adding hatching or direct labels.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: regressions are empirical and self-contained; self-citations are background only; in-sample 'prediction' wording and group selection are validation concerns, not by-construction reductions.

full rationale

The paper's derivation chain is empirical rather than definitional. Daily sleep self-assessment is a subjective comparative rating collected before participants viewed Oura data (Section 3.2); the sensor features come from the ring's independent measurements; and N-back scores are defined by the task composite in Eq. 1. No quantity is defined in terms of the outcome it is claimed to predict, and no fitted parameter is redeployed as a prediction of the same data: the mixed-effects (Table 3), ordinal, and Bayesian models estimate associations from the full sample, and the abstract's 'highly predict' wording merely summarizes in-sample coefficients and R^2 (marginal 0.15), which is a reporting/validation weakness, not a circular reduction. The mutual significance of self-assessment and N-back in Tables 3 and 4 is a bidirectional empirical correlation; the paper explicitly flags self-bias and behavioral confirmation as alternative explanations (Section 5.3), so the association is not presented as a derivational circle. The three-group claim (Section 4.2) selects n by in-sample R^2 comparison without cross-validation, and the authors concede 'there are indications that the currently identified number of 3 groups comes from data limitations'; this is an overfitting/model-selection concern, since the GMM groups and group-specific coefficients are estimated, not assumed. Self-citations ([6], [44], [45], [53]) support only background claims (in-the-wild Oura procedures, user perception vs. data) and are not load-bearing for the regressions. The Oura validity assumption rests on external validation studies [12,13,38,65], and the central derivation is self-contained, so circularity is essentially absent.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The study contributes an empirical dataset and standard statistical analyses. It does not introduce new physical entities or fitted physical constants. The load-bearing assumptions are measurement-level: Oura's sleep estimates, N-back as a working memory proxy, and the single-item relative self-rating. The main data-driven choice is the number of subgroups in the Gaussian mixture model, selected on the same data used for the group mixed models.

free parameters (2)
  • Number of GMM subgroups = 3
    Chosen as the k giving the highest pseudo-R^2 for the group models on the same data (marginal 0.185 for k=3 vs 0.157 for k=2, 0.163 for k=4); no validation or stability check.
  • Likert-to-numeric scale factor = x (arbitrary)
    The five relative response options are mapped to -2x, -x, 0, x, 2x. The authors note x is arbitrary; coefficient tests are invariant to x, so this is not load-bearing.
assumptions (3)
  • domain assumption Oura ring sleep-stage and physiological estimates are accurate enough for within-person day-to-day comparisons.
    Section 3.1 relies on the Oura ring's proprietary algorithms for sleep stages, HR, HRV, and the Oura sleep score; prior validation studies are cited but no per-participant PSG or ECG ground truth is collected in this study.
  • domain assumption The 3-back task score reflects working memory performance within subjects.
    Section 2.2 and 3.3 justify N-back as a working memory measure via the literature, but no independent cognitive test is administered in this study.
  • domain assumption Relative single-item self-assessment ('compared to yesterday') is a meaningful measure of perceived sleep quality without floor/ceiling effects.
    Section 3.2 argues this design removes floor/ceiling effects; the paper offers no validation that participants interpret the anchors consistently within or across languages.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring the Alignment of Perceived and Measured Sleep Quality with Working Memory using Consumer Wearables." pith.science (2026). https://pith.science/paper/CVTO3CVY

@misc{pith2026250719491,
  author       = {Pith},
  title        = {Pith review of: Exploring the Alignment of Perceived and Measured Sleep Quality with Working Memory using Consumer Wearables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CVTO3CVY}},
  note         = {Machine review of arXiv:2507.19491}
}
read the original abstract

Wearable devices offer detailed sleep-tracking data. However, whether this information enhances our understanding of sleep or simply quantifies already-known patterns remains unclear. This work explores the relationship between subjective sleep self-assessments and sensor data from an Oura ring over 4--8 weeks in-the-wild. 29 participants rated their sleep quality daily compared to the previous night and completed a working memory task. Our findings reveal that differences in REM sleep, nocturnal heart rate, N-Back scores, and bedtimes highly predict sleep self-assessment in significance and effect size. For N-Back performance, REM sleep duration, prior night's REM sleep, and sleep self-assessment are the strongest predictors. We demonstrate that self-report sensitivity towards sleep markers differs among participants. We identify three groups, highlighting that sleep trackers provide more information gain for some users than others. Additionally, we make all experiment data publicly available.

Figures

Figures reproduced from arXiv: 2507.19491 by the authors.

Figure 1
Figure 1. Examples of Data Presented in the Oura Ring Application. On the left is summary data that gives an overview of both [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Experiment app. Left: Sleep assessment question. Center: [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Data availability per participant against days from experiment start. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Change in REM sleep duration for participants and their self-assessment. Green and red candlesticks depict the [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Histograms for all five self-report answer options displaying their relative frequency against the days of the week, [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Cumulative self-reported sleep quality (first column) and the features included in the mixed effects models (columns [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

87 extracted references · 65 canonical work pages

  1. [1]

    Nouran Abdalazim, Leonardo Alchieri, Lidia Alecci, and Silvia Santini. 2023. BiHeartS: Bilateral Heart Rate from Multiple Devices and Body Positions for Sleep Measurement Dataset. https://doi.org/10.48550/arXiv.2308.06811 arXiv:2308.06811 [cs]

  2. [2]

    Glenn Affleck, Susan Urrows, Howard Tennen, Pamela Higgins, and Micha Abeles. 1996. Sequential daily relations of sleep, pain intensity, and attention to pain among women with fibromyalgia. Pain 68, 2-3 (1996), 363–368

  3. [3]

    Hirotugu Akaike. 2011. Akaike’s Information Criterion. In International Ency- clopedia of Statistical Science , Miodrag Lovric (Ed.). Springer, Berlin, Heidelberg, 25–25. https://doi.org/10.1007/978-3-642-04898-2_110

  4. [4]

    Lidia Alecci, Nouran Abdalazim, Leonardo Alchieri, Shkurta Gashi, and Silvia Santini. 2023. On the Mismatch between Measured and Perceived Sleep Quality. In Adjunct Proceedings of the 2022 ACM International Joint Conference on Pervasive and Ubiquitous Computing and the 2022 ACM International Symposium on Wear- able Computers (UbiComp/ISWC ’22 Adjunct) . A...

  5. [5]

    Eneko Antón, Manuel Carreiras, and Jon Andoni Duñabeitia. 2019. The Impact of Bilingualism on Executive Functions and Working Memory in Young Adults.PLoS ONE 14, 2 (Feb. 2019), e0206770. https://doi.org/10.1371/journal.pone.0206770

  6. [6]

    Shota Arai, Andrew Vargo, Benjamin Tag, and Koichi Kise. 2024. In-the-Wild Exploration of the Impact of the Lunar Cycle on Sleep in a University Cohort with Oura Rings. In Proceedings of the Augmented Humans International Conference 2024 (AHs ’24) . Association for Computing Machinery, New York, NY, USA, 259–262. https://doi.org/10.1145/3652920.3653044 15...

  7. [7]

    Zawadzki

    Amber Carmen Arroyo and Matthew J. Zawadzki. 2022. The Implementation of Behavior Change Techniques in mHealth Apps for Sleep: Systematic Review. JMIR mHealth and uHealth 10, 4 (April 2022), e33527. https://doi.org/10.2196/ 33527

  8. [8]

    Alan Baddeley. 2003. Working memory: looking back and looking forward. Nature Reviews Neuroscience 4, 10 (2003), 829–839. https://doi.org/10.1038/ nrn1201

Show all 87 references
  1. [9]

    Christine Blome and Matthias Augustin. 2016. Measuring change in subjective wellbeing: Methods to quantify recall bias and recalibration response shift. Working Paper 2016/12. HCHE Research Paper. https://www.econstor.eu/handle/10419/ 145973

  2. [10]

    Michael H Bonnet. 1989. The effect of sleep fragmentation on sleep and per- formance in younger and older subjects. Neurobiology of aging 10, 1 (1989), 21–25

  3. [11]

    Noelia Calvo, Agustín Ibáñez, and Adolfo M. García. 2016. The Impact of Bilin- gualism on Working Memory: A Null Effect on the Whole May Not Be So on the Parts. Frontiers in Psychology 7 (Feb. 2016). https://doi.org/10.3389/fpsyg.2016. 00265

  4. [12]

    Rui Cao, Iman Azimi, Fatemeh Sarhaddi, Hannakaisa Niela-Vilen, Anna Axelin, Pasi Liljeberg, and Amir M. Rahmani. 2022. Accuracy Assessment of Oura Ring Nocturnal Heart Rate and Heart Rate Variability in Comparison With Electrocardiography in Time and Frequency Domains: Compreh...

  5. [13]

    Nicholas I. Y. N. Chee, Shohreh Ghorbani, Hosein Aghayan Golkashani, Ruth L. F. Leong, Ju Lynn Ong, and Michael W. L. Chee. 2021. Multi-Night Validation of a Sleep Tracking Ring in Adolescents Compared with a Research Actigraph and Polysomnography. Nature and Science of Sleep ...

  6. [14]

    Lee, Bongshin Lee, Wanda Pratt, and Julie A

    Eun Kyoung Choe, Nicole B. Lee, Bongshin Lee, Wanda Pratt, and Julie A. Kientz. 2014. Understanding quantified-selfers’ practices in collecting and exploring personal data. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada)...

  7. [15]

    Michelle A Cretikos, Rinaldo Bellomo, Ken Hillman, Jack Chen, Simon Finfer, and Arthas Flabouris. 2008. Respiratory Rate: The Neglected Vital Sign. Medical Journal of Australia 188, 11 (2008), 657–659. https://doi.org/10.5694/j.1326- 5377.2008.tb01825.x

  8. [16]

    Cudney, Benicio N

    Lauren E. Cudney, Benicio N. Frey, Randi E. McCabe, and Sheryl M. Green

  9. [17]

    Dixon, Anna L

    William G. Dixon, Anna L. Beukenhorst, Belay B. Yimer, Louise Cook, Antonio Gasparrini, Tal El-Hay, Bruce Hellman, Ben James, Ana M. Vicedo-Cabrera, Malcolm Maclure, Ricardo Silva, John Ainsworth, Huai Leng Pisaniello, Thomas House, Mark Lunt, Carolyn Gamble, Caroline Sanders,...

  10. [18]

    Druce, John McBeth, Sabine N

    Katie L. Druce, John McBeth, Sabine N. van der Veer, David A. Selby, Bertie Vidgen, Konstantinos Georgatzis, Bruce Hellman, Rashmi Lakshminarayana, Afiqul Chowdhury, David M. Schultz, Caroline Sanders, Jamie C. Sergeant, and William G. Dixon. 2017. Recruitment and Ongoing Enga...

  11. [19]

    https://doi.org/10.1038/s41746-019-0180-3 Number: 1 Publisher: Nature Publishing Group

  12. [20]

    Victoria A Felix, Nathalie A Campsen, Abbey White, and Walter C Buboltz. 2017. College students’ prevalence of sleep hygiene awareness and practices.Advances in Social Sciences Research Journal 4, 4 (2017)

  13. [22]

    David Fischer. 1999. Capturing the Patient’s View of Change as a Clinical Outcome Measure. JAMA 282, 12 (Sept. 1999), 1157. https://doi.org/10.1001/ jama.282.12.1157

  14. [23]

    Mehta, Akhila Reddy, and Mellar Davis

    Shagufta Firdous, Andrea Berger, Waqas Jehangir, Carlos Fernandez, Bertrand Behm, Zankhana Y. Mehta, Akhila Reddy, and Mellar Davis. 2021. How Should We Assess Pain: Do Patients Prefer a Quantitative or Qualitative Scale? A Study of Patient Preferences. American Journal of Hos...

  15. [24]

    Gajewski, Eva Hanisch, Michael Falkenstein, Sven Thönes, and Edmund Wascher

    Patrick D. Gajewski, Eva Hanisch, Michael Falkenstein, Sven Thönes, and Edmund Wascher. 2018. What Does the N-Back Task Measure as We Get Older? Relations Between Working-Memory Measures and Other Cognitive Functions Across the Lifespan. Frontiers in Psychology 9 (Nov. 2018), ...

  16. [25]

    Frenda and Kimberly M

    Steven J. Frenda and Kimberly M. Fenn. 2016. Sleep Less, Think Worse: The Effect of Sleep Deprivation on Working Memory.Journal of Applied Research in Memory and Cognition 5, 4 (2016), 463–469. https://doi.org/10.1016/j.jarmac.2016.10.001 Working Memory in the Wild: Applied Re...

  17. [26]

    Rúben Gouveia, Evangelos Karapanos, and Marc Hassenzahl. 2015. How do we engage with activity trackers? A longitudinal study of Habito. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing (Osaka, Japan) (UbiComp ’15). Association...

  18. [27]

    Debus, Francesca Gasparini, and Silvia Santini

    Shkurta Gashi, Lidia Alecci, Elena Di Lascio, Maike E. Debus, Francesca Gasparini, and Silvia Santini. 2022. The Role of Model Personalization for Sleep Stage and Sleep Quality Recognition Using Wearables. IEEE Pervasive Computing 21, 2 (April 2022), 69–77. https://doi.org/10....

  19. [29]

    Jaeggi, Martin Buschkuehl, Walter J

    Susanne M. Jaeggi, Martin Buschkuehl, Walter J. Perrig, and Beat Meier. 2010. The Concurrent Validity of the N-back Task as a Working Memory Measure. Memory 18, 4 (May 2010), 394–412. https://doi.org/10.1080/09658211003702171

  20. [30]

    Wayne K Kirchner. 1958. Age differences in short-term retention of rapidly changing information. Journal of Experimental Psychology 55, 4 (1958), 352. https://doi.org/10.1037/h0043688

  21. [31]

    Kainec, Jamie Caccavaro, Morgan Barnes, Chloe Hoff, Annika Berlin, and Rebecca M

    Kyle A. Kainec, Jamie Caccavaro, Morgan Barnes, Chloe Hoff, Annika Berlin, and Rebecca M. C. Spencer. 2024. Evaluating Accuracy in Five Commercial Sleep- Tracking Devices Compared to Research-Grade Actigraphy and Polysomnogra- phy. Sensors 24, 2 (Jan. 2024), 635. https://doi.o...

  22. [32]

    Kenichi Kuriyama, Kazuo Mishima, Hiroyuki Suzuki, Sayaka Aritake, and Makoto Uchiyama. 2008. Sleep Accelerates the Improvement in Working Memory Performance. The Journal of Neuroscience 28, 40 (Oct. 2008), 10145–10150. https://doi.org/10.1523/JNEUROSCI.2039-08.2008

  23. [33]

    Elina Kuosmanen, Aku Visuri, Saba Kheirinejad, Niels van Berkel, Heli Koskimäki, Denzil Ferreira, and Simo Hosio. 2022. How Does Sleep Tracking Influence Your Life? Experiences from a Longitudinal Field Study with a Wearable Ring. Proc. ACM Hum.-Comput. Interact. 6, MHCI, Arti...

  24. [34]

    Ian Li, Anind Dey, and Jodi Forlizzi. 2010. A stage-based model of personal infor- matics systems. InProceedings of the SIGCHI Conference on Human Factors in Com- puting Systems (Atlanta, Georgia, USA) (CHI ’10). Association for Computing Ma- chinery, New York, NY, USA, 557–56...

  25. [35]

    Kenichi Kuriyama, Kazuo Mishima, Hiroyuki Suzuki, Sayaka Aritake, and Makoto Uchiyama. 2008. Sleep Accelerates the Improvement in Working Memory Perfor- mance. The Journal of Neuroscience: The Official Journal of the Society for Neuro- science 28, 40 (Oct. 2008), 10145–10150. ...

  26. [36]

    Liddell and John K

    Torrin M. Liddell and John K. Kruschke. 2018. Analyzing ordinal data with metric models: What could possibly go wrong? Journal of Experimental Social Psychology 79 (Nov. 2018), 328–348. https://doi.org/10.1016/j.jesp.2018.08.009

  27. [37]

    Zilu Liang, Bernd Ploderer, Wanyu Liu, Yukiko Nagata, James Bailey, Lars Kulik, and Yuxuan Li. 2016. SleepExplorer: a visualization tool to make sense of correla- tions between personal sleep data and contextual factors.Personal and Ubiquitous Computing 20, 6 (01 Nov 2016), 98...

  28. [38]

    Milad Asgari Mehrabadi, Iman Azimi, Fatemeh Sarhaddi, Anna Axelin, Han- nakaisa Niela-Vilén, Saana Myllyntausta, Sari Stenholm, Nikil Dutt, Pasi Liljeberg, and Amir M. Rahmani. 2020. Sleep Tracking of a Commercially Available Smart Ring and Smartwatch Against Medical-Grade Act...

  29. [39]

    Matthews, Sanjay R

    Karen A. Matthews, Sanjay R. Patel, Elizabeth J. Pantesco, Daniel J. Buysse, Thomas W. Kamarck, Laisze Lee, and Martica H. Hall. 2018. Similarities and differences in estimates of sleep duration by polysomnography, actigraphy, diary, and self-reported habitual sleep in a commu...

  30. [40]

    Seiko Miyata, Akiko Noda, Kunihiro Iwamoto, Naoko Kawano, Masato Okuda, and Norio Ozaki. 2013. Poor Sleep Quality Impairs Cognitive Performance in Older Adults. Journal of Sleep Research 22, 5 (Oct. 2013), 535–541. https: //doi.org/10.1111/jsr.12054

  31. [41]

    Michael, Maryanne Garry, and Irving Kirsch

    Robert B. Michael, Maryanne Garry, and Irving Kirsch. 2012. Suggestion, Cogni- tion, and Behavior. Current Directions in Psychological Science 21, 3 (June 2012), 151–156. https://doi.org/10.1177/0963721412446369

  32. [42]

    Davis, Paul Karoly, Patrick Finan, Howard Tennen, and Mark P

    Chung Jung Mun, Hye Won Suk, Mary C. Davis, Paul Karoly, Patrick Finan, Howard Tennen, and Mark P. Jensen. 2019. Investigating intraindividual pain variability. PAIN 160, 11 (Nov. 2019), 2415–2429. https://doi.org/10.1097/j.pain. 0000000000001626 Publisher: Ovid Technologies (...

  33. [43]

    Harvey Moldofsky. 2001. Sleep and pain. Sleep medicine reviews 5, 5 (2001), 385–396. 16 Exploring the Alignment of Perceived and Measured Sleep Quality with Working Memory

  34. [44]

    Peter Neigel, Andrew Vargo, Yusuke Komatsu, Chris Blakely, and Koichi Kise

  35. [45]

    Shinichi Nakagawa and Holger Schielzeth. 2013. A General and Simple Method for Obtaining R2 from Generalized Linear Mixed-Effects Models. Methods in Ecology and Evolution 4, 2 (2013), 133–142. https://doi.org/10.1111/j.2041-210x. 2012.00261.x

  36. [46]

    Romain Nith, Yun Ho, and Pedro Lopes. 2024. SplitBody: Reducing Mental Workload while Multitasking via Muscle Stimulation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, N...

  37. [47]

    Oreel, Philippe Delespaul, Iris D

    Tom H. Oreel, Philippe Delespaul, Iris D. Hartog, José P. S. Henriques, Justine E. Netjes, Alexander B. A. Vonk, Jorrit Lemkes, Michael Scherer-Rath, Hanneke W. M. Van Laarhoven, Mirjam A. G. Sprangers, and Pythia T. Nieuwkerk. 2020. Ecological momentary assessment versus retr...

  38. [48]

    Vargo, Benjamin Tag, and Koichi Kise

    Peter Neigel, Andrew W. Vargo, Benjamin Tag, and Koichi Kise. 2024. Using Wearables to Unobtrusively Identify Periods of Stress in a Real University En- vironment. In ACM International Symposium on Wearable Computing (ISWC) . ACM. https://doi.org/10.1145/3675095.3676608

  39. [49]

    Oura Health Oy. 2021. Oura Blog: Body Temperature. https://support.ouraring. com/hc/en-us/articles/360025587493-Body-Temperature Accessed 2024-02-01 21:49:52

  40. [50]

    Owen, Kathryn M

    Adrian M. Owen, Kathryn M. McMillan, Angela R. Laird, and Ed Bullmore. 2005. N-back working memory paradigm: A meta-analysis of normative functional neuroimaging studies. Human Brain Mapping 25, 1 (2005), 46–59. https://doi. org/10.1002/hbm.20131

  41. [51]

    Antti Oulasvirta and Pertti Saariluoma. 2004. Long-term working memory and interrupting messages in human-computer interaction. Behaviour & Information Technology 23, 1 (2004), 53–64. https://doi.org/10.1080/01449290310001644859

  42. [52]

    June J Pilcher and Amy S Walters. 1997. How sleep deprivation affects psycho- logical variables related to college students’ cognitive performance. Journal of American College Health 46, 3 (1997), 121–126

  43. [53]

    Nolasco, Andrew Vargo, Yusuke Komatsu, Motoi Iwata, and Koichi Kise

    Hannah R. Nolasco, Andrew Vargo, Yusuke Komatsu, Motoi Iwata, and Koichi Kise. 2023. Perception Versus Reality: How User Self-reflections Compare to Actual Data. In Human-Computer Interaction – INTERACT 2023 , José Ab- delnour Nocera, Marta Kristín Lárusdóttir, Helen Petrie, A...

  44. [54]

    Ziyi Peng, Cimin Dai, Yi Ba, Liwei Zhang, Yongcong Shao, and Jianquan Tian

  45. [55]

    Patel, Julie A

    Ruth Ravichandran, Sang-Wha Sien, Shwetak N. Patel, Julie A. Kientz, and Laura R. Pina. 2017. Making Sense of Sleep Sensors: How Sleep Sensing Tech- nologies Support and Undermine Sleep Health. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (De...

  46. [56]

    Amirreza Sajjadieh, Ali Shahsavari, Ali Safaei, Thomas Penzel, Christoph Schoebel, Ingo Fietze, Nafiseh Mozafarian, Babak Amra, and Roya Kelishadi

  47. [57]

    Whitehouse, Andrew Judge, Alex J

    Adrian Sayers, Michael R. Whitehouse, Andrew Judge, Alex J. MacGregor, Ash- ley W. Blom, and Yoav Ben-Shlomo. 2020. Analysis of change in patient-reported outcome measures with floor and ceiling effects using the multilevel Tobit model: a simulation study and an example from a...

  48. [58]

    Amon Rapp and Federica Cena. 2016. Personal informatics for everyday life: How users without prior self-tracking experience engage with personal data. International Journal of Human-Computer Studies 94 (10 2016), 1–17. https: //doi.org/10.1016/J.IJHCS.2016.05.006

  49. [59]

    Florian Schmiedek, Martin Lövdén, and Ulman Lindenberger. 2014. A Task Is a Task Is a Task: Putting Complex Span, n-Back, and Other Working Memory Indicators in Psychometric Context. Frontiers in Psychology 5 (Dec. 2014). https: //doi.org/10.3389/fpsyg.2014.01475

  50. [60]

    Maureen Schmitter-Edgecombe, Catherine Luna, Brooke Beech, Shenghai Dai, and Diane J. Cook. 2024. Capturing Cognitive Capacity in the Everyday Environ- ment across a Continuum of Cognitive Decline Using a Smartwatch N-Back Task and Ecological Momentary Assessment. Neuropsychol...

  51. [61]

    Tanaffos 19, 2 (Nov

    The Association of Sleep Duration and Quality with Heart Rate Vari- ability and Blood Pressure. Tanaffos 19, 2 (Nov. 2020), 135–143. https: //www.ncbi.nlm.nih.gov/pmc/articles/PMC7680518/

  52. [62]

    Scott, Thomas L

    Alexander J. Scott, Thomas L. Webb, Marrissa Martyn-St James, Georgina Rowse, and Scott Weich. 2021. Improving sleep quality leads to better mental health: A meta-analysis of randomised controlled trials. Sleep Medicine Reviews 60 (2021), 101556. https://doi.org/10.1016/j.smrv...

  53. [63]

    Demin, Katharina Lederer, Thomas Penzel, and Ingo Fietze

    Julia Schlagintweit, Naima Laharnar, Martin Glos, Maria Zemann, Artem V. Demin, Katharina Lederer, Thomas Penzel, and Ingo Fietze. 2023. Effects of Sleep Fragmentation and Partial Sleep Restriction on Heart Rate Variability during Night. Scientific Reports 13, 1 (April 2023), ...

  54. [64]

    Oltmanns, and Aaron J

    Jiyoung Song, Esther Howe, Joshua R. Oltmanns, and Aaron J. Fisher. 2023. Examining the Concurrent and Predictive Validity of Single Items in Ecological Momentary Assessments. Assessment 30, 5 (July 2023), 1662–1671. https://doi. org/10.1177/10731911221113563 Publisher: SAGE P...

  55. [65]

    Thomas Svensson, Kaushalya Madhawa, Hoang Nt, Ung-il Chung, and Akiko Kishi Svensson. 2024. Validity and Reliability of the Oura Ring Gen- eration 3 (Gen3) with Oura Sleep Staging Algorithm 2.0 (OSSA 2.0) When Com- pared to Multi-Night Ambulatory Polysomnography: A Validation ...

  56. [66]

    Schwartz, Rita Bode, Nicholas Repucci, Janine Becker, Mirjam A

    Carolyn E. Schwartz, Rita Bode, Nicholas Repucci, Janine Becker, Mirjam A. G. Sprangers, and Peter M. Fayers. 2006. The clinical significance of adaptation to changing health: A meta-analysis of response shift. Quality of Life Research 15, 9 (Nov. 2006), 1533–1550. https://doi...

  57. [67]

    Nicole K. Y. Tang, Ptolemy D. W. Banks, and Adam N. Sanborn. 2023. Judgement of Sleep Quality of the Previous Night Changes as the Day Unfolds: A Prospective Experience Sampling Study. Journal of Sleep Research 32, 3 (2023), e13764. https: //doi.org/10.1111/jsr.13764

  58. [68]

    Mark Snyder and Julie A. Haugen. 1995. Why Does Behavioral Confirmation Occur? A Functional Perspective on the Role of the Target. Personality and Social Psychology Bulletin 21, 9 (Sept. 1995), 963–974. https://doi.org/10.1177/ 0146167295219010

  59. [69]

    Teirlinck, Dieke S

    Carolien H. Teirlinck, Dieke S. Sonneveld, Sita M. A. Bierma-Zeinstra, and Pim A. J. Luijsterburg. 2019. Daily Pain Measurements and Retrospective Pain Measurements in Hip Osteoarthritis Patients With Intermittent Pain. Arthri- tis Care & Research 71, 6 (2019), 768–776. https:...

  60. [70]

    Aloe, and Betsy Jane Becker

    Christopher Glen Thompson, Rae Seon Kim, Ariel M. Aloe, and Betsy Jane Becker. 2017. Extracting the Variance Inflation Factor and Other Multicollinearity Diagnostics from Typical Regression Results. Basic and Applied Social Psychology 39, 2 (March 2017), 81–90. https://doi.org...

  61. [71]

    Melanie Swan. 2009. Emerging patient-driven health care models: an examination of health social networks, consumer personalized medicine and quantified self- tracking. International journal of environmental research and public health 6 (2009), 492–525. Issue 2. https://doi.org...

  62. [72]

    Lynn Marie Trotti. 2017. Waking up Is the Hardest Thing I Do All Day: Sleep Inertia and Sleep Drunkenness. Sleep medicine reviews 35 (Oct. 2017), 76–84. https://doi.org/10.1016/j.smrv.2016.08.005

  63. [73]

    Oura Team. 2024. Technology in the Oura Ring. https://ouraring.com/blog/ring- technology/ Accessed 2024-02-01, 21:18:39

  64. [74]

    Niels van Berkel, Jorge Goncalves, Peter Koval, Simo Hosio, Tilman Dingler, Denzil Ferreira, and Vassilis Kostakos. 2019. Context-Informed Scheduling and Analysis: Improving Accuracy of Mobile Self-Reports. In Proceedings of the 2019 CHI Conference on Human Factors in Computin...

  65. [75]

    Omer Van den Bergh and Marta Walentynowicz. 2016. Accuracy and bias in retrospective symptom reporting. Current Opinion in Psychiatry 29, 5 (Sept. 2016), 302–308. https://doi.org/10.1097/YCO.0000000000000267

  66. [76]

    Tin, Mia Austria, Gabriel Ogbennaya, Susan Chimonas, Paulin Andréll, Thomas M

    Amy L. Tin, Mia Austria, Gabriel Ogbennaya, Susan Chimonas, Paulin Andréll, Thomas M. Atkinson, Andrew J. Vickers, and Sigrid V. Carlsson. 2023. Pain as bad as you can imagine or extremely severe pain? A randomized controlled trial comparing two pain scale anchors. Journal of ...

  67. [77]

    Matthew Walker. 2017. Why we sleep: Unlocking the power of sleep and dreams . Simon and Schuster

  68. [78]

    Niels van Berkel, Denzil Ferreira, and Vassilis Kostakos. 2017. The Experience Sampling Method on Mobile Devices. ACM Comput. Surv. 50, 6 (Dec. 2017), 93:1–93:40. https://doi.org/10.1145/3123988

  69. [79]

    Eunju Yang. 2017. Bilinguals’ Working Memory (WM) Advantage and Their Dual Language Practices. Brain Sciences 7, 7 (July 2017), 86. https://doi.org/10. 3390/brainsci7070086

  70. [80]

    Hyeryeon Yi. 2013. Sleep quality and its associated factors in adults. Journal of Korean Public Health Nursing 27, 1 (2013), 76–88

  71. [81]

    van Dijk, Willem van Rhenen, Jaap M

    Dela M. van Dijk, Willem van Rhenen, Jaap M. J. Murre, and Esmée Verwijk

  72. [82]

    PloS One 15, 4 (2020), e0231906

    Cognitive Functioning, Sleep Quality, and Work Performance in Non- Clinical Burnout: The Role of Working Memory. PloS One 15, 4 (2020), e0231906. https://doi.org/10.1371/journal.pone.0231906 17 Neigel et al

  73. [83]

    Matúš Šimkovic and Birgit Träuble. 2019. Robustness of statistical methods when measure is affected by ceiling and/or floor effect. PLOS ONE 14, 8 (Aug. 2019), e0220889. https://doi.org/10.1371/journal.pone.0220889 Publisher: Public Library of Science. 18

  74. [84]

    McArdle, and Timothy A

    Lijuan Wang, Zhiyong Zhang, John J. McArdle, and Timothy A. Salt- house. 2008. Investigating Ceiling Effects in Longitudinal Data Anal- ysis. Multivariate Behavioral Research 43, 3 (Sept. 2008), 476–496. https://doi.org/10.1080/00273170802285941 Publisher: Routledge _eprint: h...

  75. [87]

    Zhiwei Zhang. 2014. Reverse Regression: A Method for Joint Analysis of Multiple Endpoints in Randomized Clinical Trials. Statistica Sinica 24, 4 (2014), 1753–1769. https://www.jstor.org/stable/24310968 Publisher: Institute of Statistical Science, Academia Sinica

  76. [88]

    Aleksandra Żmijewska, Wojciech Kopacz, Mikołaj Wojtas, Monika Maleszewska, Mateusz Sztybór, Marcin Kapica, Maria Krzyżanowska, Karen Głogowska, Julia Piątkiewicz, and Gabriela Nowak. 2024. The Association Between Heart Rate Variability and Sleep Quality -a Narrative Review. Jo...

  77. [2020]

    Frontiers in Neuroscience 14 (May 2020), 469

    Effect of Sleep Deprivation on the Working Memory-Related N2-P3 Com- ponents of the Event-Related Potential Waveform. Frontiers in Neuroscience 14 (May 2020), 469. https://doi.org/10.3389/fnins.2020.00469

  78. [2022]

    Journal of clinical sleep medicine: JCSM: official publication of the American Academy of Sleep Medicine 18, 3 (March 2022), 927–936

    Investigating the Relationship between Objective Measures of Sleep and Self-Report Sleep Quality in Healthy Adults: A Review. Journal of clinical sleep medicine: JCSM: official publication of the American Academy of Sleep Medicine 18, 3 (March 2022), 927–936. https://doi.org/1...

  79. [2023]

    In the Wild

    Exploring Users’ Ability to Choose a Proper Fit in Smart-Rings: A Year- Long “In the Wild” Study. In Human-Computer Interaction – INTERACT 2023 (Lecture Notes in Computer Science, Vol. 141415) , José Abdelnour Nocera, Marta Kristín Lárusdóttir, Helen Petrie, Antonio Piccinno, ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.