{"id":"8935162e-4f49-49f3-b071-f9071c2a8a17","arxiv_id":"2606.18506","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Causal discovery on two large cohorts identifies five recurrent domains and yields an SRS with up to 2.5× better alignment to perceived recovery than AHI.","lead":"The paper proposes a causal-discovery framework that derives a hierarchical Sleep Recovery Score from five PSG physiological domains and reports up to 2.5 times stronger alignment with patient-reported recovery than the standard AHI. A smart generalist might read it to understand how clinical sleep data could translate into better wearable-based recovery tracking in connected health.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Validity of two-stage screening (physiology constraints + constrained LLM auditing) in removing all structural confounders and construct overlaps is unverified.","rationale":"Reader's weakest assumption directly identifies the critical untested link in the causal pipeline. Full-text description does not add independent verification (e.g., no machine-checked causal criteria or ablation), so the concern stands and moves the verdict from UNVERDICTED to CONDITIONAL pending the proposed check.","tokens_in":1784,"tokens_out":294,"duration_ms":21068,"concrete_test":"Recompute the DAGs, domain selection, and SRS correlation using only the physiology-based constraints (omitting the LLM auditing step entirely); if the set of five domains changes or the alignment improvement falls below 1.5×, the screening step is load-bearing for the central claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline claim requires that the five recurrent domains and resulting SRS are free of bias so that the 2.5× alignment improvement over AHI can be attributed to the causal structure rather than residual confounding or variable overlap. The method relies on a two-stage process whose second stage uses LLM-assisted auditing; no sensitivity analysis, inter-rater comparison with human experts, or explicit back-door criterion checks are described to confirm completeness of removal. In observational PSG data, even modest residual confounding would invalidate the DAG-derived domains and the downstream correlation claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a causal-discovery-guided framework for constructing an interpretable Sleep Recovery Score (SRS) from polysomnography (PSG) data. Using DAG learning on two cohorts (MESA: n=1540; MrOS: n=825), it identifies five recurrent domains: respiratory burden, hypoxic burden, sleep fragmentation, sleep architecture, and autonomic regulation. A two-stage screening process (physiology-based constraints combined with constrained LLM-assisted auditing) is used to remove structural confounders and construct-overlapping variables. The resulting SRS is reported to show up to 2.5× stronger alignment with patient-reported outcomes (PROs) such as perceived recovery compared to the Apnea-Hypopnea Index (AHI), with potential applications in connected health technologies.","tokens_in":1939,"tokens_out":698,"duration_ms":29161,"significance":"If the results hold after addressing verification concerns, this could represent a meaningful advance in sleep medicine by providing a more comprehensive, interpretable, and bias-aware metric than AHI that links multimodal physiology to functional recovery. The mapping to wearable sensing streams is a practical strength, and the emphasis on causal discovery and domain structure offers a template for similar efforts in other health domains. The work credits the use of large population cohorts and recurrent domain identification across them.","major_comments":[{"comment":"The two-stage screening process is presented as removing all structural confounders and construct-overlapping variables, but no sensitivity analysis, inter-rater reliability with human experts, or explicit checks against the back-door criterion are described. This is load-bearing for the claim that the 2.5× alignment improvement is due to the causal structure rather than residual confounding, as even modest bias in observational PSG data could affect the DAG-derived domains and downstream correlations.","section":"Methods (two-stage screening process)"},{"comment":"The abstract and results claim up to 2.5× stronger alignment with perceived recovery than AHI, but specific details on the metric used (e.g., correlation coefficient, regression R²), error bars, cohort-specific values, and whether evaluation was on held-out data are needed to assess if post-hoc choices influenced the result. Without these, it is difficult to evaluate the robustness of the central quantitative claim.","section":"Results (alignment with PROs)"},{"comment":"The application of DAG learning to identify candidate drivers raises the possibility of data-driven tuning if the domain selection or SRS weights were optimized on the same data used for evaluation. Clarification on the separation between discovery and validation steps is required to rule out circularity in the reported improvement.","section":"Methods (DAG learning and domain selection)"}],"minor_comments":[{"comment":"The abstract mentions 'constrained LLM-assisted auditing' without specifying the constraints or the LLM model used; adding these details would improve reproducibility.","section":"Abstract"},{"comment":"Consider adding a limitations section explicitly addressing the observational nature of the data and potential unmeasured confounding.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The paper's fit to a machine learning journal is reasonable given the causal discovery component, but the primary contribution appears more applied to sleep medicine; ensure the novelty in the ML method is clearly distinguished from prior causal discovery applications in health data."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and detailed comments on our manuscript. We address each major comment below and will revise the manuscript to improve transparency and robustness where the concerns are valid.","responses":[{"response":"We agree that the manuscript would benefit from additional verification of the two-stage screening process. While the process integrates physiology-based constraints and constrained LLM-assisted auditing to enhance plausibility and remove confounders, we did not report sensitivity analyses or explicit back-door criterion applications. In revision, we will add sensitivity analyses varying the constraint thresholds and auditing parameters, report inter-rater reliability for the auditing step, and include a limitations discussion on observational data constraints.","revision_made":"yes","referee_comment":"[Methods (two-stage screening process)] The two-stage screening process is presented as removing all structural confounders and construct-overlapping variables, but no sensitivity analysis, inter-rater reliability with human experts, or explicit checks against the back-door criterion are described. This is load-bearing for the claim that the 2.5× alignment improvement is due to the causal structure rather than residual confounding, as even modest bias in observational PSG data could affect the DAG-derived domains and downstream correlations."},{"response":"We will revise the results section to explicitly detail the alignment metric (including whether correlation or R²), provide error bars or confidence intervals, report cohort-specific values for both MESA and MrOS, and clarify the evaluation procedure including any separation from discovery or use of held-out data. This addresses the need for transparency on the 2.5× claim.","revision_made":"yes","referee_comment":"[Results (alignment with PROs)] The abstract and results claim up to 2.5× stronger alignment with perceived recovery than AHI, but specific details on the metric used (e.g., correlation coefficient, regression R²), error bars, cohort-specific values, and whether evaluation was on held-out data are needed to assess if post-hoc choices influenced the result. Without these, it is difficult to evaluate the robustness of the central quantitative claim."},{"response":"DAG learning was applied independently to each cohort to identify recurrent domains across MESA and MrOS, with domain selection driven by recurrence rather than direct optimization against PRO alignment. SRS weights followed from the causal structure. We will add explicit text clarifying this separation of discovery and evaluation steps. To further mitigate concerns, we will also report a sensitivity analysis using a held-out subset.","revision_made":"partial","referee_comment":"[Methods (DAG learning and domain selection)] The application of DAG learning to identify candidate drivers raises the possibility of data-driven tuning if the domain selection or SRS weights were optimized on the same data used for evaluation. Clarification on the separation between discovery and validation steps is required to rule out circularity in the reported improvement."}],"tokens_in":1604,"tokens_out":609,"duration_ms":46182,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this work derives a hierarchical Sleep Recovery Score from causal discovery on MESA and MrOS data, reporting stronger correlation with patient-reported recovery than the usual AHI. The five domains (respiratory burden, hypoxic burden, fragmentation, architecture, autonomic regulation) come up in both cohorts.\n\nThe paper does a few things cleanly. Using two independent large cohorts to check recurrence of the domains is a solid move. The intent to produce an interpretable score that could transfer to wearables is practical. The two-stage screening that mixes physiology rules with constrained LLM auditing is a reasonable attempt to keep the DAGs from picking up obvious junk variables.\n\nThe soft spots sit mostly in the screening and the headline number. The abstract gives no sensitivity checks, no inter-rater numbers against human experts, and no back-door diagnostics to show the LLM step actually removed the structural confounders it claims to target. The 2.5x alignment figure appears without error bars, cohort breakdowns, or confirmation that the domain weights were not fitted on the same data used for evaluation. If those details are missing from the full text as well, the improvement could partly reflect post-hoc selection rather than cleaner causal structure.\n\nThis is for people working on sleep metrics, PRO alignment, or connected-health sensing who want alternatives to AHI. A reader who needs concrete new scoring ideas or domain lists could pull something useful, but anyone wanting to adopt the SRS would need the missing validation steps.\n\nIt deserves peer review so the screening procedure and the alignment statistics can be examined directly.","headline":"The paper builds a new SRS via DAG learning on two PSG cohorts and claims 2.5x better PRO alignment than AHI, but the LLM-assisted screening step lacks visible validation.","tokens_in":2438,"tokens_out":399,"would_cite":false,"duration_ms":28579,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A causal-discovery framework derives a Sleep Recovery Score from PSG data that aligns up to 2.5 times better with patient-perceived recovery than the Apnea-Hypopnea Index.","keywords":["sleep recovery score","causal discovery","polysomnography","apnea-hypopnea index","patient-reported outcomes","directed acyclic graphs","connected health","wearable sensing"],"falsifier":"In an independent cohort the Sleep Recovery Score fails to show stronger correlation with patient-reported recovery outcomes than AHI once the same two-stage screening is applied.","tokens_in":2681,"feed_emoji":"🛌","tokens_out":499,"duration_ms":18567,"temperature":0.7,"pith_summary":"The paper establishes that five recurrent physiological domains—respiratory burden, hypoxic burden, sleep fragmentation, sleep architecture, and autonomic regulation—can be extracted from standard polysomnography recordings via directed acyclic graph learning. These domains are screened for bias through a two-stage process of physiology constraints plus constrained LLM auditing, then combined into a hierarchical Sleep Recovery Score. Across two large cohorts the resulting score tracks patient-reported recovery more closely than the conventional AHI summary. The approach is designed to translate directly to emerging wearable streams such as ECG, oximetry, and sleep staging, offering a structured alternative to event-counting indices that often fail to capture functional recovery.","feed_headline":"Sleep Recovery Score aligns 2.5 times better with patient reports than AHI","feed_subtitle":"Causal discovery on two large PSG cohorts yields five domains whose composite score tracks perceived recovery more closely than event counts","key_machinery":"The two-stage screening process that applies physiology-based constraints followed by constrained LLM-assisted auditing to produce bias-free DAG-derived domains for the Sleep Recovery Score.","core_discovery":"Directed acyclic graph learning applied to multimodal PSG recordings from the MESA and MrOS cohorts identifies five recurrent physiological domains associated with recovery; after removal of structural confounders and construct-overlapping variables via a two-stage physiology-plus-LLM screening process, these domains combine into a hierarchical Sleep Recovery Score whose alignment with perceived recovery reaches up to 2.5 times that of the Apnea-Hypopnea Index.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["DAG learning on PSG identifies five domains for Sleep Recovery Score","Five domains drive 2.5x aligned Sleep Recovery Score in two cohorts","Causal discovery yields five domains with 2.5x recovery score alignment","Sleep Recovery Score from five PSG domains reaches 2.5x AHI alignment"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The two-stage screening process correctly identifies and removes every structural confounder and construct-overlapping variable so that the resulting domains are free of bias.","fun_headline_variants_meta":{"raw":{"variants":["DAG learning on PSG identifies five domains for Sleep Recovery Score","Five domains drive 2.5x aligned Sleep Recovery Score in two cohorts","Causal discovery yields five domains with 2.5x recovery score alignment","Sleep Recovery Score from five PSG domains reaches 2.5x AHI alignment"]},"model":"grok-4.3","cost_usd":0.007641,"raw_usage":{"total_tokens":3530,"prompt_tokens":732,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":76412000,"prompt_tokens_details":{"text_tokens":732,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2720,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":732,"tokens_out":78,"duration_ms":22167,"temperature":1.0,"reasoning_tokens":2720,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T01:04:56.489134+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"In an independent cohort the Sleep Recovery Score fails to show stronger correlation with patient-reported recovery outcomes than AHI once the same two-stage screening is applied.","supporting_citations":[],"review_version":1}