{"id":"05bbc9e5-5796-4467-91de-fd91856c96a4","arxiv_id":"2411.14656","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"mmWave radar can reliably measure sit-to-stand duration and trunk motion, but not knee motion, when compared with Kinect and wearable sensors in 45 healthy adults.","lead":"This paper tests whether a 60 GHz millimeter-wave radar can track sit-to-stand movements well enough for fall-risk screening without cameras or body sensors. It compares radar-derived trunk and knee angles against Kinect and wearable sensors in 45 healthy adults, finding good agreement for trunk motion but poor agreement for knee motion.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Radar trunk reliability claim overstates independent agreement: R-W ICCs are moderate (TrunkROM 0.72) to poor (TrunkExtensionPeakVelocity 0.38), and K-R ICCs are inflated by Kinect-labeled training.","rationale":"The reader identified the Kinect-as-ground-truth circularity as the weakest assumption; that is a real flaw because K-R agreement is partly self-consistency. However, the most decisive problem for the central claim is that the one independent comparison, radar vs. wearables, yields only moderate-to-poor ICCs for the trunk features that the conclusion calls 'high agreement.' This is an internal inconsistency in the paper's interpretation, not just a missing external reference. A re-analysis of the existing data with and without outlier removal, plus confidence intervals, would settle whether the R-W agreements are actually high. The verdict remains CONDITIONAL because the study provides useful comparative data and honest limitations, but the central wording overreaches beyond what the evidence supports.","tokens_in":13614,"tokens_out":4824,"duration_ms":45651,"concrete_test":"Reanalyze the matched STS feature dataset used for Table II: compute R-W ICCs for TrunkROM, TrunkFlexionPeakVelocity, and TrunkExtensionPeakVelocity (a) with the current Z-score outlier removal, (b) without any outlier removal, and (c) with 95% confidence intervals, reporting the number of excluded points per feature. If TrunkExtensionPeakVelocity R-W remains below 0.5 or the ICCs drop substantially without outlier removal, the phrase 'high agreement with wearable sensors' in the conclusion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section V Conclusions and Section IV-D Discussion) that mmWave radar 'remains a reliable tool for assessing trunk movements, showing high agreement with the Kinect and wearable sensors' is not supported by the independent evidence. In Section III-D, the radar pose model is trained on Kinect skeletons, so the K-R ICCs (TrunkROM 0.9295, TrunkFlexionPeakVelocity 0.8723, TrunkExtensionPeakVelocity 0.7862) partly reflect self-consistency and cannot by themselves establish true measurement accuracy. The only independent comparison, radar versus wearables (R-W), yields TrunkROM ICC 0.7218, TrunkFlexionPeakVelocity 0.6209, and TrunkExtensionPeakVelocity 0.3787 (Table II). These values are not uniformly 'high'; the extension-velocity value is poor. Additionally, Section III-G removes outliers with |Z|>3 without reporting how many, and no confidence intervals are given, so the reported ICCs may be optimistic. Thus, the load-bearing assumption that the radar reliably measures trunk movements against an independent reference is insecure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a comparative study of mmWave radar, Kinect, and wearable inertial sensors for Sit-to-Stand (STS) motion analysis. A 60 GHz FMCW radar collects point clouds; a deep learning pose estimation model (mmPose-FK) predicts a 17-joint skeleton using Kinect body tracking as the training target. Inverse kinematics converts skeletons to joint angles; STS repetitions are segmented; features (duration, trunk ROM, trunk peak velocities, waist-thigh ROM, knee ROM) are extracted and compared across sensor pairs using ICC and Bland-Altman. The paper reports high agreement for duration and trunk-related features, low agreement for knee ROM, and concludes that radar is reliable for trunk-level STS assessment.","tokens_in":13809,"tokens_out":4320,"duration_ms":42314,"significance":"The work addresses a real need: non-contact, privacy-preserving motion analysis for fall-risk assessment. It builds on the authors' prior mmPose-FK work and provides an end-to-end pipeline from radar point clouds to STS features, with a 45-participant dataset. The honest reporting of the low knee ICC and the explicit acknowledgment of Kinect lower-leg jitter are commendable. The independent radar-vs-wearable agreement for trunk ROM (0.72) suggests that radar captures some trunk-level information, but the evidence is not as strong as the paper claims. The central limitation—Kinect serving as both training target and comparison reference—undermines the K-R agreement as a validation, and the absence of a gold-standard reference system means the accuracy of radar-derived angles remains unestablished. The paper is a useful feasibility study, but its conclusions require tempering.","major_comments":[{"comment":"The radar pose model is trained on Kinect skeletons (Section III-B explicitly states that \"the Kinect skeleton serves as the learning target for the radar\"), and the same Kinect is used as the reference for the K-R ICCs. The K-R trunk ICCs (TrunkROM 0.9295, TrunkFlexionPeakVelocity 0.8723) therefore partly quantify the model's ability to reproduce its training target, not its measurement accuracy against an independent standard. The only independent comparison, radar versus wearables, gives TrunkROM ICC 0.7218, TrunkFlexionPeakVelocity 0.6209, and TrunkExtensionPeakVelocity 0.3787 (Table II). The Section IV-D conclusion that radar \"remains a reliable tool for assessing trunk movements, showing high agreement\" is not supported by the independent evidence, particularly for extension peak velocity. Please either soften the claim to reflect moderate trunk ROM agreement and limited extension-velocity agreement, or add an external validation against a motion-capture reference.","section":"Sections III-B, IV-B and Table II"},{"comment":"Outliers are removed using |Z|>3 before computing ICCs, but the number of removed data points per feature is not reported. Post-hoc outlier removal can inflate ICC estimates, especially with a limited number of participants. Please report the number of removed repetitions for each feature and provide a sensitivity analysis by also reporting ICCs computed without outlier removal. In addition, Table II gives only point estimates; please add confidence intervals and the effective sample size (number of participants and repetitions) for each ICC so the precision of these agreement measures can be assessed.","section":"Section III-G"},{"comment":"The synchronization of wearable signals is performed by cross-correlating the knee-angle signals, which is the very feature later used to evaluate KneeROM. This introduces a form of circularity: the R-W knee ICC can be inflated by the alignment procedure, and the time shift derived from knee-angle correlation also affects the temporal alignment of all other features. Although the knee ICCs are low across all pairs, the procedure should be justified as independent of the measured outcomes. Please report the distribution of the estimated lags and run a robustness check using an alternative alignment target (e.g., trunk/waist angle) or a synchronization method that does not rely on the evaluated signal.","section":"Section III-D"},{"comment":"No gold-standard motion capture system is used to validate the Kinect skeleton that serves both as the training target and as the comparison reference for radar. The paper itself acknowledges \"the Kinect skeleton's abnormal jittering in the lower leg\" and suggests using a VICON system in future work. Without an external reference, the paper should not interpret the K-R agreement as evidence of measurement accuracy; the claims should be limited to cross-modal agreement. Please add an explicit statement that the current study validates agreement but not absolute accuracy, and clearly separate \"learning fidelity\" (K-R) from \"cross-modal agreement\" (R-W) throughout the discussion and conclusion.","section":"Sections IV-A and IV-D"}],"minor_comments":[{"comment":"The phrase \"they are sensitivity to lighting conditions\" should read \"they are sensitive to lighting conditions\" (also appears as \"sensitivity\" in the same context in the Introduction).","section":"Sections I and II-B"},{"comment":"The Kinect skeleton is repeatedly called the \"ground truth\" for the radar, but Section IV-B later notes the Kinect skeleton's lower-leg jitter. Using the term \"reference\" instead of \"ground truth\" would be more precise and avoid overstating the reference's accuracy.","section":"Section III-D"},{"comment":"The repetition-matching step accepts start-time differences up to 0.5 seconds; please report how many repetitions were excluded by this criterion and whether the excluded repetitions were distributed evenly across sensors and participants.","section":"Section III-F"},{"comment":"The table would benefit from sample sizes and confidence intervals for each ICC value, or at least a footnote stating the number of STS repetitions and participants used for each comparison.","section":"Table II"},{"comment":"The statement \"the current azimuth and elevation angle resolution is 29 degrees\" is useful but it is not supported by any measurement or derivation in the text; please provide a reference to the sensor datasheet or a calculation.","section":"Section IV-D"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful feasibility study, but the strength of the central claim exceeds what the validation design can support. The authors may wish to reframe the contribution as an exploration of radar-based STS feature extraction rather than as a validation of radar accuracy. I would also encourage the editor to check that the novelty relative to the authors' own prior mmPose-FK work is clearly articulated, as several methodological components (pose estimation and filtering) are reused from earlier publications."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The takeaway: this is a competent, honest engineering comparison, but the paper's central 'reliable for trunk' claim is a step too far. The only independent radar evidence is moderate at best, and the radar-Kinect agreement is partly self-consistency because Kinect labels trained the model.\n\nWhat's genuinely new: it's the first point-cloud-level radar STS analysis that extracts joint-level features (as opposed to Doppler/range-Doppler), and it compares against both Kinect and wearables. The authors reuse their mmPose-FK skeleton model and apply IK to get angles. The methods are clearly described, and they report ICCs and Bland-Altman. Good credit: they are upfront that knee agreement is poor, that Kinect lower-leg jitter is a likely cause, and that healthy adults only were enrolled. They also explicitly suggest VICON for future work.\n\nThe soft spots are structural. Kinect is both the training target for the radar pose model and the reference for the K-R comparison, so those high ICCs (TrunkROM 0.93) are partially circular. The independent R-W numbers tell a more modest story: TrunkROM 0.72, TrunkFlexionPeakVelocity 0.62, TrunkExtensionPeakVelocity 0.38. That last value is poor, and calling all trunk agreement 'high' is not justified. Also, outliers are removed by |Z|>3 without reporting how many, no confidence intervals are given for the ICCs, and the wearable synchronization uses cross-correlation on knee angle—the very signal they later say fails. These are not fatal to the paper's contribution, but they should be addressed.\n\nWho it's for: people working on RF-based human motion analysis or sensor comparison for fall-risk assessment will find a useful reference. It's not a clinical validation. I'd send it to peer review, with a clear request to tone down the conclusions and add the missing analysis.","headline":"Useful first point-cloud radar STS comparison, but the 'reliable trunk' claim outruns the independent evidence.","tokens_in":14414,"tokens_out":2396,"would_cite":true,"duration_ms":23724,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 60 GHz mmWave radar can count sit-to-stand repetitions and track trunk motion at the same level of agreement as Kinect and wearable sensors, but it does not capture knee motion reliably.","keywords":["sit-to-stand","mmWave radar","pose estimation","inverse kinematics","fall risk assessment","Kinect","wearable sensors","intraclass correlation coefficient"],"falsifier":"Collect simultaneous radar, Kinect, and marker-based optical motion capture data for the same sit-to-stand protocol, and compare trunk and knee features. If radar-versus-optical trunk range-of-motion ICC falls far below the reported radar-Kinect value of 0.9295, or if the radar-versus-optical knee ICC stays near zero while the optical reference is stable, the paper's claim that radar is reliable for trunk assessment is not supported.","tokens_in":13374,"feed_emoji":"📡","tokens_out":5736,"duration_ms":52266,"temperature":0.7,"pith_summary":"This paper tries to establish that a 60 GHz millimeter-wave radar, which works from a contact-free point cloud, can serve as a useful alternative for sit-to-stand analysis, at least for the coarse features most relevant to fall risk assessment. It trains a deep-learning pose estimator on Kinect skeletons, converts the recovered 17-joint skeleton into sagittal-plane trunk and knee angles via inverse kinematics, and compares these with Kinect and five wearable gyroscope sensors on 45 healthy adults. The result, if correct, is that radar reliably counts STS repetitions and measures trunk range of motion and trunk peak velocities, with intraclass correlations roughly 0.72 to 0.96, while knee range of motion is not captured reliably. The study therefore argues that radar is a viable non-contact, privacy-preserving option for monitoring the gross movement component of fall risk, with full-body precision still awaiting better ground truth or sensor fusion.","feed_headline":"Radar matches Kinect on sit-to-stand trunk tracking","feed_subtitle":"In 45 healthy adults, a 60 GHz radar agreed with Kinect and wearables on STS duration and trunk range; knee-range readings diverged.","key_machinery":"The load-bearing mechanism is the mmPose-FK pipeline: a 60-64 GHz MIMO radar produces range-Doppler-angle point clouds, a deep neural network learned from Kinect body tracking regresses a 17-joint skeleton, forward-kinematic constraints and a smoothing filter stabilize it, and inverse kinematics with Rodrigues rotations convert the skeleton into a time series of relative joint angles for 12 joints around three axes. From those angles the pipeline segments each STS repetition, extracts duration, trunk range of motion, trunk flexion and extension peak velocities, waist-thigh range, and knee range, and scores sensor agreement with intraclass correlation and Bland-Altman analysis. The Kinect skeleton is simultaneously the training label and the main comparison reference, which is why the knee-joint jitter weakens both the learned radar model and the Kinect-radar agreement.","core_discovery":"The central discovery claimed is that mmWave radar point clouds, after deep-learning pose estimation and inverse kinematics, produce STS features that agree strongly with Kinect and wearable sensors for movement duration and trunk-level motion, but not for knee-level motion. Reported ICCs are 0.9556 (Kinect-radar), 0.9634 (Kinect-wearable), and 0.9526 (radar-wearable) for STS duration; trunk range-of-motion ICCs are 0.9295, 0.7865, and 0.7218 for the three sensor pairs; knee range-of-motion ICCs are only 0.3064, 0.0697, and 0.0456. The authors attribute the knee failure largely to jitter in the Kinect skeleton's lower leg, which also serves as the training target for the radar model, and conclude that radar remains a reliable tool for trunk-focused fall risk assessment rather than for fine joint detail.","pith_inferences":["Beyond the paper, a natural extension is to test whether radar trunk features can separate older adults at high fall risk from younger controls using the same protocol; the current healthy sample cannot answer that.","Because the radar and Kinect share a coordinate calibration and the radar learns from Kinect, the high Kinect-radar trunk ICC may partly reflect training fidelity rather than independent measurement accuracy; an independent reference would disentangle these.","The radar's limited angular resolution suggests the knee failure is plausibly a hardware limit rather than a fundamental one, so denser antenna arrays may recover lower-leg tracking.","A testable extension of the multi-sensor fusion argument is to combine radar trunk features with wearable shank gyroscope data, potentially recovering knee-level detail that radar alone misses."],"forward_implications":["Automated STS repetition counting in a 30-second chair-stand test can be done by radar placed a few meters away, without wearable devices, with agreement effectively equal to Kinect and wearables.","Trunk range-of-motion and trunk peak-velocity features, the ones most tied to fall risk, can be extracted from radar with useful agreement for cohort-level assessment.","Knee range-of-motion from radar is not yet usable for clinical decisions, and sensor comparisons reporting knee-level agreement must treat Kinect's lower-leg jitter as a confound.","Because the radar learns from Kinect, radar accuracy is bounded by Kinect accuracy, making a marker-based optical motion capture reference the natural next validation step."],"supporting_citations":[{"why":"Supplies the mmPose-FK pose-estimation model that turns radar point clouds into a 17-joint skeleton, the core of the measurement pipeline.","marker":"[15]"},{"why":"Earlier CNN-based work that established real-time skeletal pose estimation from mmWave radar point clouds and is extended here.","marker":"[41]"},{"why":"Natural-language-processing variant for precise skeletal pose estimation, part of the skeleton-recovery line the paper builds on.","marker":"[42]"},{"why":"Filtering approach used to stabilize the recovered skeleton and smooth angle and velocity signals before feature extraction.","marker":"[43]"},{"why":"Provides the intraclass correlation and Bland-Altman statistical protocol recommended and used for the sensor agreement comparisons.","marker":"[44]"}],"fun_headline_variants":["Radar matches Kinect on sit-to-stand trunk, not knees","mmWave radar: trunk tracking matches Kinect, knee fails","Non-contact radar rivals Kinect for sit-to-stand trunk","60 GHz radar: good on trunk, bad on knees in STS","Radar sees sit-to-stand trunk, blinds on knee range"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Kinect's built-in body tracking is treated as ground truth for training the radar pose model and as the reference for comparison, despite the paper reporting that the Kinect skeleton jitters in the lower leg; if that reference is biased, the radar-Kinect agreement does not establish true measurement accuracy.","fun_headline_variants_meta":{"raw":{"variants":["Radar matches Kinect on sit-to-stand trunk, not knees","mmWave radar: trunk tracking matches Kinect, knee fails","Non-contact radar rivals Kinect for sit-to-stand trunk","60 GHz radar: good on trunk, bad on knees in STS","Radar sees sit-to-stand trunk, blinds on knee range"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1646,"prompt_tokens":933,"completion_tokens":713,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":623}},"tokens_in":549,"tokens_out":713,"duration_ms":6785,"temperature":1.0,"reasoning_tokens":623,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:02:51.954437+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect simultaneous radar, Kinect, and marker-based optical motion capture data for the same sit-to-stand protocol, and compare trunk and knee features. If radar-versus-optical trunk range-of-motion ICC falls far below the reported radar-Kinect value of 0.9295, or if the radar-versus-optical knee ICC stays near zero while the optical reference is stable, the paper's claim that radar is reliable for trunk assessment is not supported.","supporting_citations":[{"cited_title":"mmpose-fk: A forward kinematics approach to dynamic skeletal pose estimation using mmwave radars,","cited_arxiv_id":null,"evidence_quote":"Supplies the mmPose-FK pose-estimation model that turns radar point clouds into a 17-joint skeleton, the core of the measurement pipeline."},{"cited_title":"mm-pose: Real-time human skeletal posture esti- mation using mmwave radars and cnns,","cited_arxiv_id":null,"evidence_quote":"Earlier CNN-based work that established real-time skeletal pose estimation from mmWave radar point clouds and is extended here."},{"cited_title":"mmpose-nlp: A natural language processing approach to precise skeletal pose estimation using mmwave radars,","cited_arxiv_id":null,"evidence_quote":"Natural-language-processing variant for precise skeletal pose estimation, part of the skeleton-recovery line the paper builds on."},{"cited_title":"Stabilizing skeletal pose estimation using mmwave radar via dynamic model and filtering,","cited_arxiv_id":null,"evidence_quote":"Filtering approach used to stabilize the recovered skeleton and smooth angle and velocity signals before feature extraction."},{"cited_title":"The effect of vibratory stimulation on the timed- up-and-go mobility test: A pilot study for sensory-related fall risk assessment,","cited_arxiv_id":null,"evidence_quote":"Provides the intraclass correlation and Bland-Altman statistical protocol recommended and used for the sensor agreement comparisons."}],"review_version":1}