{"id":"f6cf161b-a5f9-4523-8d22-8a16f1e959f9","arxiv_id":"2506.08574","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A broadly benchmarked ensemble of deep sleep staging models, trained on a large harmonized data pool, matches or exceeds published state-of-the-art and human-expert consensus alignment on out-of-domain datasets.","lead":"SLEEPYLAND is an open benchmark that trains three deep learning sleep staging models on about 220,000 hours of harmonized recordings and combines them into an ensemble, SOMNUS, that generally beats individual models and several published state-of-the-art results on 24 datasets. The paper also reports demographic and clinical bias in model predictions and shows that ensemble disagreement flags epochs where human scorers disagree.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's SOTA-superiority claim is contradicted by Table 3: on PHYS and SEDF-SC, SOMNUS (out-of-domain) scores below U-Sleep trained in-domain, so the 'even when ID vs OOD' claim needs qualification.","rationale":"I read the central claim as the ensemble's superiority over previous state-of-the-art methods, especially the striking sub-claim that this can hold when SOTA models had in-domain training while SOMNUS did not. The evidence for that sub-claim is Table 3. Inspecting Table 3, the pattern is not uniform: for PHYS and SEDF-SC, SOMNUS is below U-Sleep trained in-domain (PHYS MF1 0.73 vs 0.79; SEDF-SC 0.75 vs 0.79). These are not edge cases; they are exactly the 'trained ID vs OOD' scenario the abstract singles out. The sentence 'Notably, SOMNUS surpasses previous state-of-the-art methods, even including cases...' therefore overstates what the table shows. The same table is built from published results with different evaluation protocols; the caption acknowledges SOMNUS values are re-aggregated to match each original study. This makes 'surpasses' a comparison of numbers, not a comparison of models under fair evaluation, which is ironic given the paper's stated goal. I would not reject the paper: the release of Docker images, pretrained weights, and the large standardized benchmark are real contributions, and the internal benchmarking of SOMNUS against its own constituent models (94.9% better, never significantly worse) is credible and well supported. The reader's consensus and label concern is legitimate for the human-expert section but is secondary to the central SOTA claim: even if consensus were the wrong clinical target, the resource claim and ensemble benefit over individual models would stand. My proposed check directly tests whether the SOTA claim survives head-to-head; if it does not, the conditional verdict should remain, with the abstract rewritten. Because the reader's CONDITIONAL verdict already accommodates such a fix, I leave the verdict unchanged.","tokens_in":76777,"tokens_out":6943,"duration_ms":81435,"concrete_test":"Run publicly released SOTA checkpoints (YASA, U-Sleep, SleepTransformer, DeepResNet where available) through the SLEEPYLAND inference pipeline on the exact SLEEPYLAND test splits for all Table 3 datasets, computing recording-level macro-F1 with identical preprocessing and channel handling. Then count wins and losses, specifically the subset of rows where the SOTA model was trained in-domain and SOMNUS is out-of-domain. If PHYS and SEDF-SC still show losses under this controlled protocol, the abstract claim is false as written and must be limited to 'in most, but not all, cases'; if the controlled comparison reverses those rows, the current superiority claim is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that SOMNUS 'surpasses previous state-of-the-art methods, even including cases where compared models were trained in-domain while SOMNUS treated the same data as out-of-domain' is the load-bearing assertion, and Table 3 is its sole evidence. That table contains direct counterexamples. In the U-Sleep rows for PHYS and SEDF-SC, U-Sleep is marked as trained in-domain (✓) while SOMNUS is marked as out-of-domain (✗), yet SOMNUS reports macro-F1 0.73 vs 0.79 on PHYS and 0.75 vs 0.79 on SEDF-SC. The SleepTransformer row for SEDF-SC similarly favors the in-domain model on MF1 (0.79 vs 0.75). The abstract's wording implies SOMNUS wins such comparisons; the paper's own numbers show it loses some of them. The comparison is also not controlled: Table 3 mixes dataset-level and recording-level metrics and adapts SOMNUS values to match each original study's aggregation, as its caption states. Published SOTA numbers come from different test splits, channel sets, and preprocessing, so 'surpasses previous SoA' is asserted without a head-to-head protocol. The resource itself (open Docker images, weights, 24-dataset benchmark) is valuable and independently checkable, but this specific superlative claim should be softened to 'matches or exceeds in most compared settings.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces SLEEPYLAND, an open-source Docker-based framework for benchmarking automatic sleep staging models. The authors train U-Sleep, DeepResNet, and SleepTransformer on harmonized NSRR data (about 27,494 recordings) under a unified protocol, evaluate them across 24 in-domain and out-of-domain datasets in single- and multi-channel EEG/EOG configurations, and introduce SOMNUS, a soft-voting ensemble across architectures and channel setups. They report SOMNUS macro-F1 scores between 68.7% and 87.2%, superiority over individual models in 94.9% of comparisons, and state-of-the-art performance, including cases where compared models were trained in-domain while SOMNUS was out-of-domain. Using GAMLSS on the Bern Sleep-Wake Registry (N=6,633), they quantify biases in performance and clinical markers by age, gender, AHI, and PLMI. On the multi-scored DOD-H and DOD-O datasets, they compare SOMNUS against five human scorers and SimpleSleepNet, report higher agreement with the scorer consensus than any individual expert, and train logistic-regression predictors of scorer disagreement from ensemble entropy and pairwise model divergence (ROC AUC up to 0.828). The manuscript honestly concludes that clinical use is not yet possible, but several headline claims, especially the state-of-the-art superiority wording, go beyond what the evidence supports.","tokens_in":77077,"tokens_out":7834,"duration_ms":93664,"significance":"SLEEPYLAND has the potential to become a valuable community resource: it publicly releases Docker images, pre-trained weights, and a GUI, and it benchmarks three architectures on an unprecedentedly large harmonized data pool with a common training/evaluation protocol. The within-framework comparisons are carefully done, with paired Wilcoxon tests adjusted for multiple comparisons and both recording-level and dataset-level metrics reported, and the GAMLSS bias analysis over a large clinical registry is a substantive contribution. The disagreement-prediction result is falsifiable and evaluated with leave-one-recording-out cross-validation. The main caveat is that the 'surpasses previous SOTA' superlative in the Abstract and Discussion is not supported by Table 3, which contains direct counterexamples and is not a controlled head-to-head comparison. The resource itself and the ensemble-over-individual-models result remain valuable once this claim is appropriately qualified.","major_comments":[{"comment":"The Abstract's claim that 'SOMNUS surpasses previous SoA methods, even including cases where compared models were trained ID while SOMNUS treated the same data as OOD' is contradicted by Table 3. On PHYS, U-Sleep (in-domain) MF1 is 0.79 versus SOMNUS (out-of-domain) 0.73; on SEDF-SC, U-Sleep (in-domain) MF1 is 0.79 versus SOMNUS 0.75, and SleepTransformer (in-domain) MF1 is 0.79 versus SOMNUS 0.75. The Discussion's statement that SOMNUS 'consistently exceeds' previously reported SOTA performance is likewise too strong. The Results text ('SOMNUS often matches or closely approaches their performance') is accurate. Please revise the Abstract and Discussion to say that SOMNUS matches or exceeds prior methods in most compared settings, and remove the unqualified in-domain-versus-out-of-domain superiority formulation.","section":"Abstract and Table 3"},{"comment":"The SOTA comparison in Table 3 is not a controlled head-to-head benchmark. As the caption states, SOMNUS metrics are computed using the same aggregation as each original study, and the table mixes recording-level and dataset-level metrics, different test splits, channel sets, and preprocessing pipelines (e.g., PHYS, SEDF-SC). A published number from another study cannot establish 'surpasses previous state-of-the-art' without a common evaluation protocol. The within-framework comparisons (SOMNUS versus individual models in Table 2 and Supplementary Tables 4-19) are properly controlled and support the ensemble claim; the SOTA-superiority claim should either be backed by a protocol-matched comparison on identical test recordings or explicitly presented as a non-head-to-head literature reference.","section":"Table 3 and 'Benchmarking, ensembling and generalization'"},{"comment":"The comparison against human experts rests on the assumption that the four-scorer consensus, or soft consensus, is the appropriate reference for 'better' staging. This is a legitimate operationalization for consensus alignment, but it is not evidence of objective clinical correctness, and the Abstract's 'SOMNUS exceeds the best human scorer' should be qualified as 'alignment with the consensus of the other scorers.' Note also the asymmetry in the reference set: for each human scorer, the consensus excludes that scorer, whereas for SOMNUS the consensus is derived from the four most reliable scorers. With at most five scorers and only 25 and 55 recordings in DOD-H and DOD-O, the comparison is statistically informative but clinically limited; this limitation should be stated explicitly where Table 6 is discussed.","section":"Model-ensemble versus human-ensemble; Methods Eqs. (2)-(4), Table 6"}],"minor_comments":[{"comment":"The footer contains a typo: 'dateset used during training' should be 'dataset used during training.'","section":"Table 3 caption"},{"comment":"The phrase 'outperforming individual models in 94.9% of cases' is not precisely defined in the main text; please specify the comparison unit, such as architecture-by-channel-configuration-by-dataset combinations, and the total number of comparisons.","section":"Abstract and Results"},{"comment":"The logistic-regression disagreement predictor is described only as using features from 'both sources of variability'; please define the exact feature vector (entropy plus mean, standard deviation, and maximum of the pairwise cosine distances) and the binary target threshold for scorer 'disagreement' in one place in the Methods.","section":"Ensemble disagreement as an indicator of scoring ambiguity"},{"comment":"The DOD-O SimpleSleepNet row contains irregular spacing in the reported values (e.g., '55.4 ± 16.8' and '89.7 ± 10.5'); please normalize the formatting across all rows.","section":"Table 6"},{"comment":"For U-Sleep, the text states that the initial learning rate was increased from the original 10^-7 to 10^-5; since this is a 100-fold change, please also report the learning-rate schedule and state whether the change was motivated by the larger training pool or by stability.","section":"Methods, Sleep staging algorithms"},{"comment":"The GAMLSS result tables use blank cells to denote excluded predictors; a brief footnote directly under each table, rather than only in the body text, would improve readability.","section":"Tables 4 and 5"}],"recommendation":"major_revision","confidential_remarks":"To the editor: this is a strong resource paper, and the core benchmarking, ensemble evaluation, and bias analysis are largely sound. The central problem is the overclaiming of SOTA superiority, which is concentrated in the Abstract and Discussion and can be fixed by softening the wording; I see no grounds for rejection. The authors should be asked either to provide a protocol-matched SOTA comparison or to explicitly label Table 3 as a literature reference rather than a head-to-head benchmark."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, SLEEPYLAND is a real resource: public Docker images, pretrained weights for U-Sleep, DeepResNet, and SleepTransformer trained on roughly 220k hours of harmonized NSRR data, a standardized 24-dataset evaluation, and a GAMLSS bias audit on 6,633 BSWR recordings. That part is solid, reproducible, and independently checkable. Second, the abstract's headline SOTA claim overstates the paper's own Table 3. On PHYS and SEDF-SC, SOMNUS as out-of-domain scores below U-Sleep trained in-domain (0.73 vs 0.79 MF1 on both), and SleepTransformer on SEDF-SC also beats SOMNUS on MF1 (0.79 vs 0.75). So \"SOMNUS surpasses previous state-of-the-art methods, even including cases where compared models were trained in-domain while SOMNUS treated the same data as out-of-domain\" is contradicted by their own numbers. The Results text says \"often matches or closely approaches,\" which is accurate; the abstract needs to be softened to that.\n\nWhat is genuinely new is the resource, not the learning method: public weights, the Dockerized framework, the 24-dataset benchmark, the bias audit, and the ensemble-disagreement metrics for flagging ambiguous epochs. The claim that SOMNUS beats individual models in 94.9% of cases is well supported by paired Wilcoxon tests (significant in 72.7%, never significantly worse), and the disagreement metrics with leave-one-recording-out CV are a fair, useful addition. The citation pattern is fine; the self-citations are to their own prior bias-framework work and are relevant.\n\nThe soft spots: Table 3 compares against published numbers without re-running the same protocol, and the metric aggregation is adapted per study. They disclose this in the caption, but it means \"surpasses SOTA\" is not a controlled claim. And the human-expert comparison defines the target as the majority/soft consensus of at most five scorers. Showing SOMNUS aligns with that consensus better than any single expert is a reasonable and clinically defensible goal, but \"exceeds the best human scorer\" in the abstract is the wrong phrasing — an ensemble is structurally likely to approximate a consensus. The body text actually frames it correctly in places.\n\nWho is this for: anyone building or evaluating sleep staging models, and anyone doing bias auditing of clinical AI. It deserves a serious referee. I'd recommend accepting after the SOTA claim is calibrated, the human-expert framing is tightened to \"agreement with consensus,\" and ideally one head-to-head re-evaluation under a single protocol.","headline":"A genuinely useful open benchmarking resource whose abstract overclaims SOTA superiority; the paper's own Table 3 contains counterexamples, and the human-expert comparison needs a consensus-framing caveat.","tokens_in":77630,"tokens_out":3063,"would_cite":true,"duration_ms":37012,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SOMNUS, a soft-voting ensemble trained on about 220,000 hours of heterogeneous sleep recordings, outperforms individual models in 94.9% of comparisons across 24 datasets and matches scorer consensus better than any individual expert on…","keywords":["automatic sleep staging","deep learning","model benchmarking","ensemble learning","soft-voting","model bias quantification","inter-scorer variability","sleep staging evaluation"],"falsifier":"Take a set of PSG recordings scored by a large independent panel (say 20 or more physicians) with an adjudicated gold-standard hypnogram, run the released SOMNUS on them out-of-domain, and check whether its agreement with that adjudicated reference still exceeds the best individual expert's; if the margin shrinks or reverses, the reported human-beating results depend on the five-scorer majority reference rather than on superior staging.","tokens_in":76596,"feed_emoji":"😴","tokens_out":17325,"duration_ms":158235,"temperature":0.7,"pith_summary":"The paper argues that the bottleneck for clinical adoption of automatic sleep staging is not model architecture but fair evaluation: researchers need a common benchmark with harmonized data, standardized splits, and open pre-trained weights. To that end it builds SLEEPYLAND, an open-source framework trained on roughly 220,000 hours of in-domain and 84,000 hours of out-of-domain recordings from the NSRR repository, and it introduces SOMNUS, a soft-voting ensemble that combines three architectures (U-Sleep, DeepResNet, SleepTransformer) across single- and multi-channel EEG/EOG setups. The central claim is that SOMNUS achieves macro-F1 scores between 68.7% and 87.2% across twenty-four datasets, outperforms individual models in 94.9% of cases, and even beats previous state-of-the-art models that were trained on the evaluation data while SOMNUS treated that data as out-of-domain. On two multi-scored out-of-domain datasets it exceeds the best human scorer's agreement with the scorer consensus, and its internal disagreement metrics predict where human scorers disagree with ROC AUC up to 0.828. A sympathetic reader would care because, if right, a simple averaging ensemble over diverse training data becomes a strong default for sleep staging and a template for fair model evaluation across clinical signal-annotation tasks.","feed_headline":"Ensemble sleep stager out-scores experts on data it never trained on","feed_subtitle":"Soft voting across three architectures gives a dependable default for sleep staging and flags ambiguous epochs.","key_machinery":"The load-bearing object is SOMNUS, a soft-voting ensemble: the average of per-epoch probability distributions over the five sleep stages (Wake, N1, N2, N3, REM) across all trained models, where the models vary in architecture (convolutional U-Sleep, residual-recurrent DeepResNet, attention-based SleepTransformer) and in channel configuration (EEG-only, EOG-only, EEG+EOG), with a majority vote across channel derivations at test time. Soft voting is what allows disagreement to be measured: the ensemble's Shannon entropy and the pairwise cosine distances between constituent predictions quantify model uncertainty, and the paper shows these two quantities track human scorer disagreement. The other supporting machinery is methodological: a unified preprocessing, training, and split protocol borrowed from U-Sleep and applied to harmonized NSRR data; the GAMLSS distributional-regression framework used to model bias of performance metrics and predicted clinical markers (like total sleep time and wake after sleep onset) as functions of age, gender, AHI, and PLMI; and the multi-scorer evaluation protocol of Equations (2)–(4), which builds soft- and hard-consensus labels from at most five physicians and computes agreement via Cohen's $\\kappa$ and hypnodensity cosine similarity (ACS).","core_discovery":"SLEEPYLAND trains three state-of-the-art architectures — U-Sleep, DeepResNet, and SleepTransformer — under a single harmonized protocol on the largest collection of NSRR polysomnograms assembled to date, and it releases the weights openly. SOMNUS (Soft-voting Over Multiple Networks for Unified Sleep-staging) averages the per-epoch probability vectors of all architecture and channel-configuration variants (Equation (1)). The paper's central claim is that this untrained, simple averaging ensemble generalizes better than any constituent: it reaches recording-level macro-F1 between 68.7% and 87.2% across 17 in-domain and 7 out-of-domain datasets, is the best model in 94.9% of pairwise comparisons and is never significantly worse, and it matches or exceeds previously published state-of-the-art scores even in the hardest comparison, where the competitor was trained in-domain and SOMNUS saw the same data only at test time. On the multi-scored Dreem Open Datasets (each recording staged by five physicians), SOMNUS — never trained on those datasets — agrees with the expert majority better than any individual scorer ($\\kappa = 0.89$ on the healthy cohort, 0.85 on the OSA cohort; macro-F1 85.2% vs 80.8% and 80.2% vs 75.9%), and its per-epoch entropy together with inter-model divergence flags epochs where the human scorers themselves disagree (ROC AUC up to 0.828). The paper further claims, from a GAMLSS analysis on the 6,633-recording Bern Sleep-Wake Registry subset, that all models and the ensemble carry comparable age-, gender-, AHI-, and PLMI-related biases, so ensembling improves robustness but does not resolve systemic bias.","pith_inferences":["With only five sources in the consensus equations, SOMNUS can beat each individual expert by approximating a majority vote even if it never captures a statistically better signal; a fairer clinical benchmark would compare against an adjudicated or much larger panel of scorers, which would likely shrink the reported gap.","The same disagreement metric could be inverted at deployment: instead of returning one hypnogram, the ensemble could route only high-entropy epochs to a human reader, turning the human-uncertainty proxy into a workflow tool that cuts manual review cost.","If data diversity is the dominant factor here, the principle transfers to other clinical annotation tasks with similar inter-rater disagreement, such as EEG event scoring or medical image segmentation: a soft vote of heterogeneous models on a large harmonized pool is a natural baseline before designing new architectures.","Because the models are released openly, they invite fine-tuning on local scorer styles; the paper's bias results imply such fine-tuning risks entrenching local bias unless the reference labels are themselves scrutinized."],"forward_implications":["If SOMNUS's 94.9% win rate over individual models holds, an untrained soft vote across architectures and channel setups is the most dependable default for plug-and-play sleep staging across heterogeneous clinics and hardware.","Size and diversity of training data, not architecture novelty, drive out-of-domain generalization: models trained on a single dataset can fail catastrophically on other data, while SOMNUS stays stable across all 24 datasets.","Releasing pre-trained weights lets clinics test models on their own recordings with data staying on-site, making fair evaluation and local validation feasible without transferring sensitive polysomnograms.","Ensemble entropy and inter-model divergence can be used in practice to flag ambiguous epochs for expert review, providing a data-driven proxy for human uncertainty.","Bias quantification shows no single architecture consistently minimizes demographic and clinical bias, so ensemble accuracy gains do not translate into bias reduction; bias-aware training and standardized reporting of clinical markers are needed before routine clinical use."],"supporting_citations":[{"why":"Supplies the U-Sleep convolutional architecture and its training and data-sampling protocol, which SLEEPYLAND adopts as the unified training procedure.","marker":"[9]"},{"why":"Provides the DeepResNet architecture (residual blocks with recurrent temporal processing), one of the three ensemble constituents whose in-domain-trained results SOMNUS must beat.","marker":"[10]"},{"why":"Provides the attention-based SleepTransformer architecture and its spectrogram-based setup, the third ensemble constituent.","marker":"[16]"},{"why":"The NSRR repository is the source of the roughly 220,000 hours of harmonized in-domain training data and of most out-of-domain test cohorts.","marker":"[13]"},{"why":"The Bern Sleep-Wake Registry (BSWR) is the large private out-of-domain dataset on which the GAMLSS model-bias analysis is run.","marker":"[17]"},{"why":"The Dreem Open Datasets (DOD-H and DOD-O), each recording scored by five physicians, provide the multi-scored data and the SimpleSleepNet baseline against which SOMNUS's human-consensus alignment is measured.","marker":"[18]"},{"why":"Introduces the GAMLSS-based algorithmic-bias quantification framework used to estimate age, gender, AHI, and PLMI effects on performance and clinical markers.","marker":"[41, 42]"},{"why":"Defines the multi-scored soft-consensus and the averaged cosine similarity (ACS) metric used to measure how closely SOMNUS tracks collective human annotations.","marker":"[12]"}],"fun_headline_variants":["Ensemble sleep stager beats top human scorers on unseen data","Soft voting across architectures yields more reliable sleep staging","Open-source sleep staging framework generalizes across 24 datasets","Ensemble model flags ambiguous sleep epochs, matching expert disagreement","Sleep staging ensemble outperforms individual models in 95% of tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison of models against human experts rests on treating the majority or soft consensus of a handful of physician scorers (at most five, Equations (2)–(4)) as the ground truth for correct sleep staging, so a model that approximates that majority can beat every individual expert even when its output is not objectively more accurate.","fun_headline_variants_meta":{"raw":{"variants":["Ensemble sleep stager beats top human scorers on unseen data","Soft voting across architectures yields more reliable sleep staging","Open-source sleep staging framework generalizes across 24 datasets","Ensemble model flags ambiguous sleep epochs, matching expert disagreement","Sleep staging ensemble outperforms individual models in 95% of tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000675,"raw_usage":{"total_tokens":3255,"prompt_tokens":1314,"completion_tokens":1941,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":930,"completion_tokens_details":{"reasoning_tokens":1858}},"tokens_in":930,"tokens_out":1941,"duration_ms":15634,"temperature":1.0,"reasoning_tokens":1858,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:07:42.965592+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of PSG recordings scored by a large independent panel (say 20 or more physicians) with an adjudicated gold-standard hypnogram, run the released SOMNUS on them out-of-domain, and check whether its agreement with that adjudicated reference still exceeds the best individual expert's; if the margin shrinks or reverses, the reported human-beating results depend on the five-scorer majority reference rather than on superior staging.","supporting_citations":[{"cited_title":"Sleep 35(6), 757–767 (2012)","cited_arxiv_id":null,"evidence_quote":"Supplies the U-Sleep convolutional architecture and its training and data-sampling protocol, which SLEEPYLAND adopts as the unified training procedure."},{"cited_title":"Sleep 38(6), 877–888 (2015)","cited_arxiv_id":null,"evidence_quote":"Provides the DeepResNet architecture (residual blocks with recurrent temporal processing), one of the three ensemble constituents whose in-domain-trained results SOMNUS must beat."},{"cited_title":"Scientiﬁc Data 9(1), 421 (2022)","cited_arxiv_id":null,"evidence_quote":"Provides the attention-based SleepTransformer architecture and its spectrogram-based setup, the third ensemble constituent."},{"cited_title":"Sleep 38(3), 411–421 (2015)","cited_arxiv_id":null,"evidence_quote":"The NSRR repository is the source of the roughly 220,000 hours of harmonized in-domain training data and of most out-of-domain test cohorts."},{"cited_title":": The sleep heart health study: design, rationale, and methods","cited_arxiv_id":null,"evidence_quote":"The Bern Sleep-Wake Registry (BSWR) is the large private out-of-domain dataset on which the GAMLSS model-bias analysis is run."},{"cited_title":": Appendicular bone density and age predict hip fracture in women","cited_arxiv_id":null,"evidence_quote":"The Dreem Open Datasets (DOD-H and DOD-O), each recording scored by five physicians, provide the multi-scored data and the SimpleSleepNet baseline against which SOMNUS's human-consensus alignment is measured."},{"cited_title":"Journal of the American Geriatrics Society 59(12), 2217–2225 (2011)","cited_arxiv_id":null,"evidence_quote":"Defines the multi-scored soft-consensus and the averaged cosine similarity (ACS) metric used to measure how closely SOMNUS tracks collective human annotations."}],"review_version":1}