{"id":"1869aafb-4e38-4f3a-b189-cefc7c790da4","arxiv_id":"2607.19798","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A radar-only system reconstructs spirometry volume curves and detects bronchodilator response in children without per-subject calibration, achieving FVC correlation 0.902 and 89.5% BDR classification accuracy on 58 trials.","lead":"SpiRadar uses millimeter-wave radar to track chest movement during forced breathing and estimates the spirometry curves that normally require a mouthpiece, reporting 0.23 L average error and correlations above 0.84 for FVC and FEV1 in 39 children. It is proposed as a calibration-free path to pediatric asthma monitoring and home-based lung-function tests.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported LOOCV results are likely optimistic: polynomial order, LASSO regularization, and feature set are selected using all 58 trials, so held-out subjects are not unseen for model selection.","rationale":"The reader's weakest_assumption was the log-interpolation correction. That is a legitimate modeling concern, but it is secondary: the correction operates on radar data alone, and its failure would produce bias attributable to a stated physiological assumption. The more load-bearing issue is that the evaluation protocol itself appears to leak information through model selection. The paper selects I, gamma, and the feature vector based on aggregate performance over all trials, including the held-out subjects, before reporting LOOCV metrics. This is a standard but serious methodological flaw: the 'unseen subject' in each fold is not unseen for hyperparameter and feature selection, so the reported generalization numbers are optimistically biased in an unknown amount. The central claim depends on subject-level LOOCV, and this flaw directly undermines it. I recommend keeping the reader's CONDITIONAL verdict because the method could still be viable, but the condition should explicitly require nested validation and external replication before the headline performance is accepted as evidence of calibration-free generalization.","tokens_in":18661,"tokens_out":5773,"duration_ms":68540,"concrete_test":"Re-run the evaluation with fully nested model selection. For each outer LOOCV fold, use only the 38 training subjects to select I in {1,...,7}, gamma in {0, 0.01, 0.1, 1, 10, 100, 1000}, and feature set (with/without FVC_R) via an inner LOOCV, then train on those 38 subjects and evaluate on the held-out subject. Report the outer-fold mean curve RMSE, FVC/FEV1 correlations, and BDR accuracy. If these differ materially from the reported 0.23 L, 0.902, 0.846, and 89.5%—for example, RMSE >0.30 L or FVC r <0.85—the reported numbers are not valid evidence of calibration-free generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section V-D.1 selects I=3 from Fig. 11, V-D.2 selects gamma=10 from Fig. 12, and V-D.3 removes FVC_R after inspecting its correlations with target parameters. These analyses appear to use the full 58-trial dataset, i.e., the same trials that are later held out in the subject-level LOOCV of Sections IV-D/V. No nested selection is described. As a result, the per-fold test subjects are not unseen for model selection: the polynomial degree, sparsity penalty, and feature vector were chosen with knowledge of their aggregate effect on the test folds. This is selection leakage that biases every headline number (mean RMSE 0.23±0.12 L, FVC r=0.902, FEV1 r=0.846, BDR accuracy 89.5%) in an unknown direction and magnitude. The paper's central claim—calibration-free generalization to unseen subjects—rests on the very LOOCV whose integrity is compromised. I focus on this rather than the log-interpolation step (Section III-A.1 step 4): that correction uses only radar-derived endpoints and slope, so at worst it is a physiological modeling assumption; the selection leakage, by contrast, invalidates the inference from the reported numbers before any such assumption is considered.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SpiRadar, a mmWave FMCW radar framework for non-contact spirometry. The method preprocesses radar thoracic displacement, enforces monotonicity via a logarithmic interpolation correction, and models the volume curve as a feature-dependent polynomial in displacement with coefficients learned through a LASSO-regularized linear model. The authors validate on a pediatric cohort of 39 subjects (58 trials) using subject-level leave-one-out cross-validation, reporting curve reconstruction RMSE of 0.23±0.12 L, FVC correlation r=0.902, FEV1 correlation r=0.846, and BDR classification accuracy of 89.5%. They compare against linear/polynomial curve-fitting baselines and a direct parameter-regression baseline.","tokens_in":19028,"tokens_out":3743,"duration_ms":40796,"significance":"If the reported results are unbiased, this would be a meaningful step toward non-contact spirometry, with particular value for pediatric asthma monitoring. The study has notable strengths: simultaneous external spirometer ground truth, subject-level partitioning that avoids trial-level leakage for patients with paired pre/post-BD measurements, and baseline methods that receive the same preprocessed inputs. The sparse polynomial framework is a sensible and well-motivated way to combine radar-derived displacement with anthropometric features. However, the central generalization claim is compromised by non-nested model selection, and an important preprocessing assumption is not directly validated.","major_comments":[{"comment":"The reported LOOCV is not nested. The polynomial order I=3 is selected from Fig. 11, the LASSO weight gamma=10 from Fig. 12, and the feature FVC_R is removed after inspecting its correlations in Section V-D.3. These analyses appear to use the full 58-trial LOOCV folds. The held-out subjects are therefore not unseen with respect to these design choices; the headline numbers (mean RMSE 0.23±0.12 L, FVC r=0.902, FEV1 r=0.846, BDR accuracy 89.5%) are the result of selecting a configuration that minimizes error on the same folds used for evaluation. The bias is in an unknown direction and magnitude. This is a load-bearing issue because the paper's central claim is calibration-free generalization to unseen subjects. Please fix hyperparameters a priori, or perform an inner subject-level LOOCV for model selection and an outer LOOCV for evaluation, or otherwise re-estimate the final performance a","section":"Sections V-D.1, V-D.2, V-D.3 and IV-D"},{"comment":"The logarithmic interpolation correction replaces non-monotonic radar displacement segments with a logarithmic curve whose parameters are fixed by endpoints and prior slope. This enforces Assumption B-1 (monotonicity) rather than testing it. The ablation in V-D.4 (Table VI) shows that the correction improves downstream metrics, but it compares two versions of the same preprocessing; it does not establish that the non-monotonic segments contain no volume-related information. If those segments reflect real thoracic motion (e.g., chest-wall recoil, upper-body movement, or multi-organ displacement), every reconstructed curve and derived parameter is biased toward the assumed shape. This is a physiological modeling assumption central to the pipeline. Please report the proportion and magnitude of corrected segments, and analyze error as a function of correction magnitude, or discuss this as a","section":"Section III-A.1, step 4 (and Assumption B-1)"}],"minor_comments":[{"comment":"The author names contain spacing artifacts: 'Y onathan Eder', 'Y onathan Kvich', 'Y onina C. Eldar' should be 'Yonathan', 'Yonina'.","section":"Author list and metadata"},{"comment":"Typo: 'ANOV A F-statistics' should be 'ANOVA F-statistics'.","section":"Section V-D.3"},{"comment":"Inconsistent notation: 'FEV1' vs 'FEV 1' in text and tables; unify to a single style, e.g., 'FEV₁'.","section":"Throughout"},{"comment":"The caption says 'box height' indicates the mean, which is unconventional for box plots; clarify the actual box plot elements (e.g., IQR, whiskers) or use a different visualization.","section":"Section V-A / Fig. 5"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern lands: the reported numbers are likely optimistic because hyperparameters and the feature set are selected using the same LOOCV folds that produce the final evaluation. This is fixable within the manuscript's scope by re-running with pre-fixed hyperparameters or nested CV. If the corrected numbers remain strong, the paper could be acceptable. The log-interpolation assumption should also be addressed, but the more urgent issue is the selection leakage. I do not see evidence of intentional misconduct; the missing nested validation is a methodological gap that the authors should be asked to close."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: SpiRadar is a solid proof-of-concept for contact-free spirometry, but the reported performance is probably inflated. The core idea — a feature-dependent sparse polynomial that maps radar chest displacement to spirometry curves, trained across subjects and tested with subject-level LOOCV — is genuinely new for this application, and the BDR classification adds a nice clinical angle. The paper is also well written and transparent about its pipeline.\n\nThe good parts: unlike earlier radar/webcam spirometry that needed per-subject calibration or only estimated parameters directly, SpiRadar reconstructs the whole volume-time curve and then derives FVC, FEV1, etc. The 39-subject pediatric cohort with paired pre/post-bronchodilator trials is a real effort, and the baselines all see the same preprocessed data. That part is fair. The self-cited radar localization work is used as a component, not as the claimed result, so no circularity issue there.\n\nThe soft spot is the model selection. In Sections V-D.1 through V-D.3, polynomial order, LASSO gamma, and the feature set are chosen using the full 58-trial dataset, including the trials that are later held out in the subject-level LOOCV. That is selection leakage. The LOOCV is then reported as if each test subject were unseen, but the model saw aggregate information about all trials when picking I=3, gamma=10, and dropping FVC_R. The direction of bias is unknown, but the headline numbers — RMSE 0.23 L, FVC r=0.902, FEV1 r=0.846, BDR 89.5% — almost certainly overstate generalization. This is not a minor quibble; it undermines the central claim of calibration-free generalization.\n\nOther concerns are smaller. The log-interpolation preprocessing (Section III-A.1 step 4) forces monotonicity on the radar signal, which is a physiological assumption that could bias curves if the non-monotonic segments carry real information. The cohort is a single center with only 39 subjects, and no code or data is released. Those are standard limitations for a study like this, not fatal flaws.\n\nOverall, this is a credible and creative application that deserves a serious referee, but the paper needs major revision before the numbers can be trusted. The authors should run nested cross-validation or hold out a separate validation cohort for hyperparameter selection, and ideally add external data. I'd send it to peer review with that expectation, and I'd cite the framework in my own work even while noting the validation issue. A reading group could usefully discuss it as a case study in how easy it is to leak selection into cross-validation.","headline":"A promising radar-spirometry framework with a real validation flaw: the headline numbers are likely optimistic because model selection used the same data that the LOOCV later tests.","tokens_in":720,"tokens_out":1031,"would_cite":true,"duration_ms":30234,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SpiRadar claims that mmWave radar alone can reconstruct forced spirometry curves in children with 0.23 L root-mean-square error and 89.5% bronchodilator-response classification accuracy, without per-subject calibration.","keywords":["non-contact spirometry","mmWave FMCW radar","pediatric asthma","sparse optimization","polynomial transformation","bronchodilator response","thoracic displacement","pulmonary function testing"],"falsifier":"Record simultaneous spirometer and radar data while subjects perform forced exhalations with deliberately inserted brief pauses or partial inspiratory efforts (making the true volume-time curve non-monotonic). If SpiRadar's log-interpolation correction replaces those real segments with smooth monotonic curves, the reconstructed FEV1 and curve RMSE will degrade exactly in the corrected window compared to a version with the correction disabled; if the correction is only removing true artifacts, performance will improve. Comparing corrected vs uncorrected reconstructions on such data would settle","tokens_in":18603,"feed_emoji":"📡","tokens_out":6326,"duration_ms":59639,"temperature":0.7,"pith_summary":"The paper sets out to show that a millimeter-wave radar pointed at a child's chest can replace the mouthpiece-and-nose-clip spirometer for measuring forced expiration, producing full volume-time and flow-volume curves rather than just a few parameters. The core idea is that a feature-dependent polynomial—radar displacement raised to powers, with coefficients that depend on age, height, weight, and radar-derived features—can be learned once on a training cohort and applied to new subjects with no individual calibration. On 39 children (58 trials, including asthma patients before and after bronchodilator), subject-level leave-one-out validation yields a curve reconstruction error of 0.23 L, FVC correlation of 0.90, FEV1 correlation of 0.85, and 89.5% accuracy in classifying bronchodilator response. If true, this would make spirometry feasible for young children, home monitoring, and telehealth scenarios where cooperation or contact is a barrier.","feed_headline":"Radar-only spirometry hits 0.23 L curve error in kids","feed_subtitle":"Chest-motion radar maps to volume-time curves with FVC r=0.90 and 89.5% bronchodilator-response accuracy.","key_machinery":"The central object is the feature-dependent polynomial model s = V C α (equivalently s = D c with D = αᵀ ⊗ V): a Vandermonde matrix V built from powers of the radar displacement vector, a learned sparse coefficient matrix C, and a feature vector α of anthropometric and radar-derived values. This makes the radar-to-volume relation subject-adaptable without per-subject calibration, because the coefficients are shared across the population while the features personalize the curve for each subject. Sparsity (LASSO, solved with FISTA) prunes the expanded IQ-dimensional coefficient space, and the logarithmic interpolation correction enforces the assumed monotonicity of the radar trace.","core_discovery":"The paper's central claim is that the mapping from radar-measured thoracic displacement v to spirometric volume s during forced expiration can be captured by a cubic polynomial whose coefficients are linear functions of subject features, and that the combined coefficient vector can be learned sparsely from a multi-subject cohort and applied to unseen subjects. Concretely, s = V C α, where V is the Vandermonde matrix of displacement powers, C is the learned coefficient matrix, and α includes anthropometric and radar-derived features; via the Kronecker product the model linearizes to s = D c, and c is recovered with ℓ1-regularized least squares. The framework also relies on a preprocessing ste","pith_inferences":["If the calibration-free mapping holds beyond the clinic, it opens the door to longitudinal home monitoring where day-to-day FEV1 trend, rather than a single clinic reading, becomes the diagnostic signal.","The feature-dependent polynomial might transfer to other non-contact displacement sources (e.g., camera-based chest tracking), since the model depends only on a displacement trace and features, not on radar-specific quantities.","A direct test with simultaneous contact and radar measurement during deliberately interrupted exhalations would clarify whether the monotonicity assumption is a helpful physiological prior or a source of bias in abnormal breathing patterns.","Pediatric asthma monitoring at home would need robustness to posture changes and device placement; the paper does not test these, so the next most valuable experiment is a longitudinal study with natural movement."],"forward_implications":["Spirometry curves and key parameters (FVC, FEV1, ratio, PEF) can be obtained from radar alone in a calibration-free manner, removing the mouthpiece and nose-clip requirement that limits pediatric cooperation.","Clinicians would gain access to the flow-volume loop morphology non-invasively, preserving diagnostic information (obstructive vs restrictive patterns) that direct parameter regression discards.","BDR assessment at 89.5% accuracy could support asthma diagnosis and medication-response decisions without repeated contact measurements.","Subject-level LOOCV indicates the learned polynomial coefficients generalize to children not seen in training, a prerequisite for home monitoring after a one-time population-level training.","The framework's performance is fairly insensitive to polynomial order ≥3 and regularization over three orders of magnitude, easing deployment tuning."],"fun_headline_variants":["Radar-only spirometry: 0.23 L curve error, no mouthpiece","Radar-based spirometry achieves 89.5% BDR accuracy in asthma kids","Sparse polynomial maps radar chest motion to spirometry curves","No-contact radar spirometry: 0.23 L RMSE in pediatric cohort"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The preprocessing step assumes that any non-monotonic segment of the radar chest-displacement trace during forced exhalation is an artifact rather than real volume change, and replaces those segments with a logarithmic curve fixed by the trace's endpoints and prior slope; if real expiratory pauses or reversals exist, every reconstructed curve and parameter inherits the assumed shape.","fun_headline_variants_meta":{"raw":{"variants":["Radar-only spirometry: 0.23 L curve error, no mouthpiece","Radar-based spirometry achieves 89.5% BDR accuracy in asthma kids","Sparse polynomial maps radar chest motion to spirometry curves","No-contact radar spirometry: 0.23 L RMSE in pediatric cohort"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001293,"raw_usage":{"total_tokens":5159,"prompt_tokens":829,"completion_tokens":4330,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":4246}},"tokens_in":573,"tokens_out":4330,"duration_ms":31341,"temperature":1.0,"reasoning_tokens":4246,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T11:40:36.556472+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record simultaneous spirometer and radar data while subjects perform forced exhalations with deliberately inserted brief pauses or partial inspiratory efforts (making the true volume-time curve non-monotonic). If SpiRadar's log-interpolation correction replaces those real segments with smooth monotonic curves, the reconstructed FEV1 and curve RMSE will degrade exactly in the corrected window compared to a version with the correction disabled; if the correction is only removing true artifacts, performance will improve. Comparing corrected vs uncorrected reconstructions on such data would settle","supporting_citations":[],"review_version":1}