{"id":"44fd1899-5909-473e-9380-9f0162e3189d","arxiv_id":"1908.00960","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AHI agreement between sleep monitors and reference polysomnography should be assessed with qualitative and threshold-weighted metrics rather than correlation alone, and the paper provides a ranking function and Shiny app for that.","lead":"This paper argues that comparing sleep-study devices by correlation alone is misleading and demonstrates a menu of quantitative and qualitative agreement metrics. It adds a hand-tuned ranking function and a free Shiny web app that computes these metrics for clinicians.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The eMAE's default weights are arbitrary and the Table 2 amplification is a product of that choice; until a sensitivity analysis shows robustness, the recommendation to report eMAE with 1.5/0.5 defaults is not yet supported.","rationale":"The paper is an honest methods-and-software contribution: it correctly argues that correlation coefficients alone are insufficient, assembles a useful suite of agreement metrics, and makes the tool available as a Shiny app. The most load-bearing part of the new proposal is the eMAE, because the abstract and discussion invite adoption of this metric for device validation. The clinical meaning of eMAE is carried entirely by arbitrary default weights; the authors themselves flag this in Section 2.2 ('we decided to set them arbitrarily'), so the concern is not manufactured. What would make it land is showing whether eMAE-based conclusions are stable under plausible weight choices. If they are stable, the default values are a presentation issue; if not, the claim that eMAE highlights clinically important errors is not empirically supported. This does not affect the qualitative-analysis or multi-metric recommendations, so the verdict remains CONDITIONAL, matching the reader. I partially agree with the reader: the same weakest assumption is identified, but the sharper version here points to a concrete instability test rather than a general need for clinical validation.","tokens_in":26231,"tokens_out":5107,"duration_ms":51922,"concrete_test":"Re-run the Table 2 analysis on both datasets under alternative weight regimes: (i) hotspots/midpoints set to 1.2/0.8; (ii) 2.0/0.2; (iii) weights based on distance to the nearest threshold from either measurement; and (iv) A(e) identically 1. If the eMAE-to-MAE relative difference (3.31x vs 2.42x) or the ordering of cubic/sinusoidal/linear approximations changes materially under any plausible alternative, then the default weighting is not a stable summary and should be reported only with sensitivity analysis or as a clearly user-configurable option.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central methodological novelty is the eMAE ranking function (Eq. 1, Section 2.2). Its clinical meaning rests entirely on hand-set weights: 1.5 at hotspots, 0.5 at midpoints, and 0.5 for within-subrange errors, values the authors state were chosen arbitrarily. Because eMAE is a weighted average of absolute errors, these weights do not measure a pre-existing clinical property; they define one. The observed 3.31x vs 2.42x relative difference between datasets in Table 2 is therefore an artifact of the weight set, not evidence that eMAE captures clinically important errors. Moreover, B(ref) weights errors only by the reference AHI, so a 1-unit error contributes about 1.5 to eMAE when ref=5 but only about 0.25 when ref=10 (because A=0.5 also applies within the same subrange). This six-fold difference is not clinically motivated by any outcome data. Without a sensitivity analysis or external anchor, the recommendation that validation studies report eMAE with the default weights is not yet supported. This concern is addressable and does not undermine the broader multi-metric proposal.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the Pearson correlation coefficient, which is frequently used to validate AHI-measuring devices, is insufficient because it does not reflect clinical significance. The authors review a range of quantitative and qualitative alternatives, including Bland-Altman plots, linear-model parameters, Spearman's rho, Lin's concordance correlation coefficient, MAE, accuracy, sensitivity/specificity, Cohen's Kappa, and multi-class AUC. As a new contribution, they propose an 'extended mean absolute error' (eMAE) that weights AHI errors by a ranking function emphasizing values near clinical thresholds (5, 15, 30 for adults), and they implement all methods in a Shiny web application. The paper demonstrates the approach on two example datasets and reports that eMAE stresses errors more than plain MAE, especially for the less concordant dataset.","tokens_in":26494,"tokens_out":4101,"duration_ms":41264,"significance":"The paper makes a useful and largely correct critical point: correlation alone is an inadequate validation metric for AHI devices, and multi-metric agreement assessment, including qualitative classification metrics, is more clinically relevant. The Shiny application is a practical, accessible tool that could benefit clinicians and statisticians. However, the central novelty, the eMAE ranking function, is currently supported only by hand-set weights that the authors explicitly state were chosen arbitrarily, and the reported advantages of eMAE over MAE are consequences of that choice rather than demonstrated properties. The comparisons in Tables 2 and 3 also lack uncertainty quantification. If the authors add sensitivity analysis, uncertainty measures, and either external clinical anchoring or a clearly exploratory framing, the paper could become a solid methods-and-software contribution.","major_comments":[{"comment":"The clinical meaning of eMAE rests entirely on the hand-set weights B(ref) and A(e): 1.5 at hotspots, 0.5 at midpoints, and 0.5 for within-subrange paired values. The paper states in Section 2.2 that 'the chosen values may be different, but we decided to set them arbitrarily.' Because eMAE is a weighted mean absolute error, the observed amplification (Table 2: eMAE cubic 3.31x versus MAE 2.42x) is a direct consequence of this arbitrary weight set, not evidence that eMAE captures clinically important errors. In particular, for a unit error at ref=5 the weight is about 1.5, while at ref=10 it is about 0.5 because the A(e)=0.5 factor applies within the same subrange, producing a six-fold difference that is not clinically motivated. The manuscript should include a sensitivity analysis over plausible weight choices and either anchor the default weights to external clinical criteria or explicitly present eMAE as an exploratory, user-weighted metric rather than a recommended diagnostic standard.","section":"Section 2.2, Eq. (1)"},{"comment":"Tables 2 and 3 report point estimates for MAE/eMAE, accuracy, Cohen's Kappa, and multi-class AUC without any measure of uncertainty. The central comparison in Table 2 (e.g., 2.42x versus 3.31x relative differences) is used to conclude that eMAE stresses errors more than MAE, but with sample sizes of 71 and 304 these differences could plausibly fall within sampling variability. Bootstrap confidence intervals, standard errors, or other uncertainty measures should be provided for all reported metrics, and the comparisons between datasets should account for that uncertainty before drawing conclusions about eMAE's behavior.","section":"Section 3, Tables 2 and 3"},{"comment":"The claim that eMAE assesses clinical significance is not validated against any clinical outcome, treatment decision, or expert-based criterion. The paper demonstrates only that eMAE values differ from MAE values under the chosen weights; it does not show that eMAE better predicts clinical management or patient outcomes than ordinary MAE or simple accuracy. Without such an external anchor, the ranking function is a descriptive re-weighting of absolute error, and the stated objective of assessing 'clinical significance' alongside 'statistical significance' is not yet supported. The authors should either provide criterion validation linking eMAE to clinical endpoints or substantially soften the interpretation, presenting eMAE as one possible exploratory weighting scheme rather than a clinically validated metric.","section":"Sections 2.2 and 4"}],"minor_comments":[{"comment":"The notation A(e) is used without defining e; the text refers to 'points (e) in the same subrange,' which is unclear. Define e_i = res_i - ref_i, or state explicitly that A depends on the error being within a single AHI subrange.","section":"Section 2.2, Eq. (1)"},{"comment":"References [8] and [31] are the same paper (Yalamanchali et al., 2013) and should be merged or cross-referenced to avoid duplication.","section":"References"},{"comment":"The phrase 'lo-factors' should be written as 'LOF factors' or 'local outlier factors' to match the cited density-based local outlier method.","section":"Section 4, Discussion"},{"comment":"The statement that the cubic ranking function stresses errors 'little more than sinusoidal and linear, as the mean value of ranking function throughout the range is the lowest' is not self-evident from the figures; provide the mean weighting factors or a direct numerical justification.","section":"Section 3, after Table 2"},{"comment":"In the submitted PDF, Figures 4-6 have overlapping or repeated axis annotations and equations, making them hard to read; please regenerate the figures with clearer layout and smaller annotation fonts.","section":"Figures 4-6"}],"recommendation":"major_revision","confidential_remarks":"The paper's critique of correlation and its multi-metric framework are solid, and the Shiny application is a useful contribution. My main reservation is that the eMAE novelty, which is the paper's claimed new methodological contribution, relies on arbitrary weights and lacks uncertainty quantification or clinical validation. These issues are addressable with sensitivity analysis and careful reframing, so I recommend major revision rather than rejection. The authors should also be aware that the duplicate references and figure readability issues, while minor, should be cleaned up."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful methods-and-software paper with an honest weakness at its core. The critique of correlation as the primary validation tool for AHI devices is sound and well anchored, and the Shiny app is genuinely useful as a way to get clinicians and statisticians looking at multiple agreement metrics instead of one r value. That part deserves credit. The paper will likely change how some sleep labs report device validation, and that would be an improvement.\n\nThe genuinely new piece is the ranking-function-weighted eMAE. I agree with the reader and the stress-test note: the weights (1.5 at hotspots, 0.5 at midpoints) are arbitrary, and the authors say so explicitly in Section 2.2. That means the Table 2 result — eMAE amplifying the difference between datasets from 2.42x to 3.31x — is a construction artifact, not evidence that eMAE captures clinically important errors. The paper presents no sensitivity analysis, no anchor to clinical outcomes, and no justification for why a 1-unit error near AHI 5 should contribute six times more than the same error near AHI 10. The claim in the abstract that the approach is reliable for pediatric studies is also unsupported: no pediatric data are shown.\n\nOther soft spots are minor but real: Tables 2 and 3 have no uncertainty measures, so the reader cannot tell whether the reported differences are meaningful. The Wilcoxon p-value of 0.04 is brushed off as \"probably due to higher N,\" which is hand-wavy. And the app is hosted but the source code is not shipped, so the reproducibility is limited to the deployed version.\n\nNone of this sinks the paper. The multi-metric proposal is sensible, the qualitative analysis (accuracy, Kappa, multiclass AUC) is a clear improvement over correlation-only reporting, and the two worked examples are helpful. The problem is that the paper oversells eMAE as a clinically grounded innovation when it is a transparently hand-tuned heuristic. A revision that adds a sensitivity analysis over weights, reports confidence intervals or bootstrap results, drops the unsupported pediatric claim, and releases the source code would make this a solid contribution. As it stands, the recommendation to report eMAE with the 1.5/0.5 defaults is not yet supported.\n\nThis is a paper for sleep researchers and device developers, not for statisticians looking for a new estimator. I would send it to reviewers who understand both clinical sleep medicine and measurement agreement. With a moderate revision it could be worth publishing; without it, the eMAE part should not be presented as established.","headline":"A useful methods-and-software paper for sleep-device validation, but the new eMAE metric's weights are arbitrary and the headline comparison is an artifact of that choice.","tokens_in":26977,"tokens_out":1860,"would_cite":false,"duration_ms":21506,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Correlation alone is insufficient to validate AHI-measuring devices; the paper's eMAE metric weights errors near clinical thresholds.","keywords":["Apnea-Hypopnea Index","correlation coefficient","Bland-Altman analysis","clinical significance","ranking function","Shiny web application","mean absolute error","device validation"],"falsifier":"A decisive test would use paired device-and-reference AHI data together with the actual clinical decisions those values produced (e.g., CPAP prescribed or not, treatment escalation). Compute both MAE and eMAE for each candidate device, then compare which metric better ranks devices by the observed rate of clinically harmful disagreement (a decision that would change if the reference value were used). If eMAE does not outperform plain MAE in predicting or ranking mismanagement, or if the arbitrary 1.5/0.5 weights can be varied without changing device rankings, the claim that eMAE captures clinical significance would fail.","tokens_in":26062,"feed_emoji":"😴","tokens_out":9097,"duration_ms":80755,"temperature":0.7,"pith_summary":"This paper contends that validation studies for sleep apnea devices lean too heavily on the correlation coefficient between a device's Apnea-Hypopnea Index (AHI) and the reference polysomnography value, because high correlation can hide systematic offsets and clinically harmful misclassification. To address that, it assembles a toolbox of quantitative agreement measures (regression intercept and slope, Spearman's rho, Lin's concordance correlation coefficient, Bland-Altman analysis, mean absolute error) and qualitative classification measures (accuracy, Cohen's kappa, multi-class ROC AUC), and it argues that both kinds should be reported. Its new contribution is a ranking function that multiplies each absolute AHI error by a weight depending on where the reference value sits relative to clinical severity thresholds, and the resulting extended mean absolute error (eMAE) is proposed as a single clinically weighted error score. The paper also provides a Shiny web application that computes all of these statistics from two columns of AHI values, which is why a clinician or device developer could adopt the approach without writing code.","feed_headline":"Correlation alone is not enough to validate sleep monitors","feed_subtitle":"eMAE up-weights AHI errors near severity thresholds, where small differences change treatment.","key_machinery":"The load-bearing object is the ranking function $B$ used inside the extended mean absolute error, $\\mathrm{eMAE} = \\frac{1}{n}\\sum_{i=1}^{n} A(e_i)\\, B(\\mathrm{ref}_i)\\, |\\mathrm{res}_i - \\mathrm{ref}_i|$. $B$ is built from the clinical thresholds of the AHI severity scale: it assigns weight 1.5 at each threshold ('hotspot', e.g., 5, 15, 30 for adults), weight 0.5 at the midpoint of each subrange, and weight 0.5 at twice the highest threshold, then interpolates between these points with a cubic, sinusoidal, or linear curve, with cubic as the default and the most restrictive. The companion factor $A$ is 0.5 when both measurements fall in the same severity subrange and 1.0 otherwise. The mechanism does the work of the paper's argument: it makes the error metric penalize boundary-crossing errors more heavily and within-subrange differences less heavily, so the eMAE's ranking of devices is intended to track clinical mismanagement rather than raw numerical distance.","core_discovery":"The paper's central claim is that correlation alone is not sufficient to establish that a device measures AHI reliably; agreement should be judged both statistically and clinically. It defines clinical significance through the established severity subranges (normal/mild/moderate/severe; thresholds 5, 15, 30 for adults, and 1, 5, 10 for children), so two values that fall in the same subrange are clinically concordant even if numerically far apart. To give raw-error metrics clinical meaning, it introduces the eMAE, in which each absolute difference is weighted by $B(\\mathrm{ref}_i)$, a ranking function equal to 1.5 at each threshold, 0.5 at subrange midpoints, and 0.5 beyond twice the highest threshold, with cubic (default), sinusoidal, or linear interpolation between anchors; an extra factor $A(e_i)$ halves the weight when both measurements are in the same subrange. Applied to two published datasets, eMAE separates a weak portable monitor (MAE 20.21, eMAE 11.90) from a stronger tracheal-sound method (MAE 5.91, eMAE 2.76) more sharply than MAE alone, and the qualitative metrics (accuracy 57.7% versus 84.2%, kappa 0.32 versus 0.76, multi-class AUC 0.733 versus 0.939) reveal clinical consequences that correlation coefficients alone would not show.","pith_inferences":["The eMAE's weights are arbitrary by the authors' own statement, so the obvious next step is empirical calibration: regress actual treatment decisions on AHI errors at various locations to estimate a data-driven weighting function, then compare it with the hand-set 1.5/0.5 weights.","The same ranking-function construction could be applied to any ordinal clinical scale with treatment thresholds (e.g., hypertension stages, hemoglobin A1c categories, tumor grading), giving a general family of clinically weighted error metrics.","The ratio eMAE/MAE could be read as a boundary-concentration index: a high ratio indicates errors cluster near thresholds, which is exactly the situation where average error understates clinical risk.","A sensitivity analysis over the three interpolation shapes, as the app already permits, would show whether device rankings are stable or an artifact of the chosen interpolation."],"forward_implications":["Sleep-device validation studies that report only Pearson correlation should be considered incomplete; the paper's framework implies that accuracy, Cohen's kappa, and multi-class AUC should be reported alongside quantitative agreement metrics.","A device can show a statistically non-significant median difference (Wilcoxon p = 0.95) while still misclassifying roughly 42% of patients, so statistical and clinical significance must be assessed together rather than interchangeably.","The eMAE, especially with the default cubic ranking function, provides a single number that penalizes boundary-crossing errors more than plain MAE, making the relative gap between a weak and a strong device appear larger (2.42x for MAE versus 3.31x for cubic eMAE in the paper's two example datasets).","Because thresholds are adjustable in the Shiny application, the same framework transfers to pediatric AHI thresholds (1, 5, 10) and to any device comparison where clinically defined subranges exist.","When reporting eMAE, the chosen interpolation shape (cubic, sinusoidal, or linear) should be stated, since it changes the value and the shape's restrictiveness."],"supporting_citations":[{"why":"Supplies the clinical AHI severity subranges (5, 15, 30 for adults; 1, 5, 10 for children) that define the thresholds used by the ranking function.","marker":"[7]"},{"why":"Provides Lin's concordance correlation coefficient, one of the alternative quantitative agreement metrics the paper recommends alongside correlation.","marker":"[10]"},{"why":"Introduces the Bland-Altman plot and limits of agreement, a central alternative technique implemented and interpreted in the paper.","marker":"[13]"},{"why":"First demonstration dataset (71 hospitalized patients, portable monitor versus PSG); serves as the poorly concordant example and the app's default data.","marker":"[28]"},{"why":"Second demonstration dataset (304 testing points, deep-neural-network tracheal sound analysis versus PSG); serves as the better concordant comparison.","marker":"[29]"},{"why":"Supports the paper's central methodological stance by recommending that correlational analyses be accompanied by qualitative (categorical) analysis during validity testing.","marker":"[34]"}],"fun_headline_variants":["Correlation isn't enough: new metric weights sleep apnea severity","Beyond correlation: clinical significance for sleep apnea devices","eMAE: a smarter way to compare sleep monitors","Sleep monitor validation: shift from correlation to clinical impact"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hand-set weights—1.5 at the severity thresholds, 0.5 at subrange midpoints, and 0.5 for points inside the same subrange—genuinely reflect how much an AHI error matters for patient management, a connection the paper asserts rather than demonstrates, since it states that the chosen values may be different but were set arbitrarily.","fun_headline_variants_meta":{"raw":{"variants":["Correlation isn't enough: new metric weights sleep apnea severity","Beyond correlation: clinical significance for sleep apnea devices","eMAE: a smarter way to compare sleep monitors","Sleep monitor validation: shift from correlation to clinical impact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1374,"prompt_tokens":997,"completion_tokens":377,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":613,"tokens_out":377,"duration_ms":4518,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:26:22.566404+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would use paired device-and-reference AHI data together with the actual clinical decisions those values produced (e.g., CPAP prescribed or not, treatment escalation). Compute both MAE and eMAE for each candidate device, then compare which metric better ranks devices by the observed rate of clinically harmful disagreement (a decision that would change if the reference value were used). If eMAE does not outperform plain MAE in predicting or ranking mismanagement, or if the arbitrary 1.5/0.5 weights can be varied without changing device rankings, the claim that eMAE captures clinical significance would fail.","supporting_citations":[{"cited_title":"The new AASM criteria for scoring hypopneas: impact on the apnea hypopnea index,","cited_arxiv_id":null,"evidence_quote":"Supplies the clinical AHI severity subranges (5, 15, 30 for adults; 1, 5, 10 for children) that define the thresholds used by the ranking function."},{"cited_title":"A concordance correlation coeﬃcient to evaluate reproducibility","cited_arxiv_id":null,"evidence_quote":"Provides Lin's concordance correlation coefficient, one of the alternative quantitative agreement metrics the paper recommends alongside correlation."},{"cited_title":"Statistical methods for assessing agreement between two methods of clinical measurement","cited_arxiv_id":null,"evidence_quote":"Introduces the Bland-Altman plot and limits of agreement, a central alternative technique implemented and interpreted in the paper."},{"cited_title":"The Accuracy of Portable Monitoring in Diagnosing Signiﬁcant Sleep Disor- dered Breathing in Hospitalized Patients","cited_arxiv_id":null,"evidence_quote":"First demonstration dataset (71 hospitalized patients, portable monitor versus PSG); serves as the poorly concordant example and the app's default data."},{"cited_title":"Tracheal Sound Analysis Using a Deep Neural Network to Detect Sleep Apnea","cited_arxiv_id":null,"evidence_quote":"Second demonstration dataset (304 testing points, deep-neural-network tracheal sound analysis versus PSG); serves as the better concordant comparison."},{"cited_title":"Methodolog- ical strategies in using home sleep apnea testing in research and practice","cited_arxiv_id":null,"evidence_quote":"Supports the paper's central methodological stance by recommending that correlational analyses be accompanied by qualitative (categorical) analysis during validity testing."}],"review_version":1}