{"id":"4b6e7773-3bfc-46f4-830d-171ed8e629d6","arxiv_id":"2608.07064","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new multi-modality dataset with over 22,000 walking samples from 27 participants in three indoor scenes, plus a benchmark pipeline, shows Wi-Fi and acoustic sensing work better together.","lead":"This paper introduces XGait, a dataset of synchronized Wi-Fi, acoustic, and camera recordings of 27 people walking in three indoor spaces. It is meant to help researchers build and compare systems that track people and recognize who they are from wireless signals alone.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The tracking benchmark's load-bearing premise is the accuracy of the vision-based ground truth, and no independent error analysis is given for the YOLOv8/back-projection pipeline in Sections 3.1 and 5.1.","rationale":"Good-faith reading: this is a dataset paper whose central claim has two components, that a new public multi-modality dataset exists with the claimed properties, and that benchmarks on it demonstrate complementary strengths of Wi-Fi and acoustics. The second component is what the validation section supports, and it is where the reader's weakest assumption applies. I agree with the reader that the camera-based ground truth is the least secure premise; it enters at Sections 3.1 and 5.1 and underlies every tracking metric. The paper honestly restricts ground truth to usable clusters in the home and meeting-room scenarios, but that selection mechanism itself needs validation. Secondary concerns, the small tracking subset (four subjects), subject-closed identity protocol, and absence of error bars, are real but are scope limitations rather than point failures. If the proposed ground-truth accuracy check passes, the dataset contribution and the direction of the complementarity findings likely survive; if it fails, the tracking numbers in Section 5.2 must be recomputed. Therefore the reader's conditional verdict is appropriate, and I would not move it.","tokens_in":929,"tokens_out":817,"duration_ms":63248,"concrete_test":"Select roughly 100 trials stratified across the three scenarios and path clusters, then independently annotate the foot-ground contact point in each video frame (or use a heel-mounted AprilTag of known size as a fiducial) and recompute the world-coordinate trajectory with the same calibration. Report per-scenario RMSE, bias, and error autocorrelation of the YOLOv8-based ground truth relative to this reference. If the median ground-truth error is not below about 0.2 m, one fifth of the 1 m metric used in Figs. 6-8, or errors correlate with turns or LoS/NLoS conditions, recompute the CDFs and fusion-win ratios in Section 5.2 with the corrected reference and test whether the Wi-Fi-domination and fusion-gain conclusions persist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"XGait's tracking benchmark rests on the claim that camera-derived trajectories are accurate enough to serve as reference. Section 3.1 and Section 5.1 describe ground truth as YOLOv8 foot keypoints back-projected to the floor plane via a single Canon EOS 80D camera, with AprilTag-derived extrinsics, but no independent accuracy assessment is reported. This is load-bearing because the quantitative support for the paper's central complementarity finding consists of CDFs at a 1 m threshold (Section 5.2, Figs. 6-8) and per-trajectory fusion-win ratios. Systematic errors in the reference trajectory, from foot-keypoint phase or heel-strike bias, floor-plane assumption errors, or pose-estimate degradation under occlusion, would change both the absolute errors and the ranking of Wi-Fi versus acoustic versus fusion. The manuscript itself notes that usable vision data is available only for selected path clusters in the home and meeting-room scenarios (Section 3.2), so the evaluation subset is selected on the basis of visual reliability yet never validated against a second reference. Without an error characterization of the ground-truth pipeline, the reported complementarity could reflect the spatial error structure of the vision system rather than wireless sensing physics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces XGait, a multi-modality wireless sensing dataset that synchronously records Wi-Fi CSI, active acoustic signals, and camera video across three indoor scenarios, with 27 participants and more than 22K walking samples. The authors propose a unified Doppler spectrogram representation and a benchmark pipeline for two tasks: indoor tracking and identity recognition. They report that Wi-Fi is more reliable for trajectory tracking, acoustics provides stronger and more trajectory-robust biometric signatures, and fusion is most beneficial in complex trajectories and challenging environments. The dataset and code are publicly released.","tokens_in":71908,"tokens_out":4882,"duration_ms":49664,"significance":"If the claims hold, XGait is a valuable community resource: it is, to my knowledge, the first public dataset combining Wi-Fi CSI, active acoustic echoes, and camera ground truth for joint tracking and identification. The release of raw data, the unified spectrogram representation, the standardized benchmark pipeline, and the extensive cross-scenario evaluation are all strengths that could enable reproducible comparison and new research on modality complementarity. The physical grounding of the Doppler representation and the explicit treatment of temporal alignment are also useful contributions. The significance is conditional, however, on the accuracy of the vision-based ground truth and on the representativeness of the evaluation subset.","major_comments":[{"comment":"The tracking benchmark uses camera-derived trajectories from YOLOv8 pose estimation and AprilTag back-projection as ground truth, but no independent accuracy assessment of this pipeline is reported. The quantitative evidence for the central complementarity claim consists of CDFs at a 1 m threshold (Figs. 6-8) and per-trajectory fusion win ratios; systematic errors in the reference trajectories, due to foot-keypoint bias, floor-plane assumptions, or occlusion, would change both absolute errors and the relative ranking of Wi-Fi, acoustic, and fusion. Please report validation of the vision ground truth against a second reference (e.g., a person-worn marker, a second camera, or a laser/IMU tracker) per scenario, including error bounds, failure rates, and the spatial regions where the reference is reliable.","section":"Sections 3.1 and 5.1"},{"comment":"Tracking performance is evaluated on only 4 out of 27 participants (3 male, 1 female), and only on path clusters with manually selected reliable vision annotations in the home and meeting-room scenarios. Because the headline finding is that fusion benefits increase with environmental complexity, and Fig. 8(d) claims consistent performance across individuals, this small and non-random subset is load-bearing evidence. Please either extend the tracking evaluation to more participants, or explicitly rephrase the cross-scenario claims as exploratory and report per-participant variability so readers can judge how much the 4-user subset supports the conclusions.","section":"Section 5.1"},{"comment":"The claim that acoustic sensing offers stronger biometric discrimination is not uniformly supported by the reported experiments. In the laboratory random-split PPVP results (Fig. 11(c)) Wi-Fi is better, and in the meeting-room cross-trajectory setting (Section 5.3(3)) the acoustic modality degrades more than Wi-Fi; the acoustic advantage appears mainly in cross-trajectory Doppler-based experiments and in clothing-variation fine-tuning. The conclusion should be conditioned on the feature paradigm and scenario, with the conflicting results explicitly reconciled, since the paper's abstract presents complementary strengths as a general finding.","section":"Section 5.3 and abstract"}],"minor_comments":[{"comment":"The text states 22,288 valid Wi-Fi CSI recordings, while Table 2 reports 22,284; please reconcile the count.","section":"Section 3.3 and Table 2"},{"comment":"The identity recognition protocol is described as subject-closed, which is reasonable for the household-occupant scenario, but the paper should state more prominently that the reported identity accuracies are closed-set re-identification rates rather than open-set identification performance.","section":"Section 5.1"},{"comment":"The weight matrix W is said to be 'based on link reliability,' but no procedure for setting it is given; please specify how W is computed and whether the same W is used for Wi-Fi-only, acoustic-only, and fusion configurations.","section":"Section 4.2, Eq. (5)"},{"comment":"The text says the benchmark pipeline 'will be released alongside the dataset,' while the abstract states the dataset and code are already available; please align these statements and provide the exact release status at the time of publication.","section":"Section 4.4"},{"comment":"Several figure labels and legends render poorly in the PDF, making it hard to read the path-cluster names, subject IDs, and confusion-matrix axes; please ensure all subfigures are legible in the camera-ready version.","section":"Figures 3, 9, and 14"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is strong and timely, and I expect the paper to be citable and useful after revision. My main concern is the absent validation of the vision-based ground truth, which is the load-bearing reference for the tracking benchmark; the 4-user tracking subset and the partially conflicting identity results also need to be addressed. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"XGait is worth knowing about: it is the first public dataset I know of that pairs Wi-Fi CSI, active acoustic, and vision ground truth for both tracking and identification, with 27 subjects and 22K samples across three scenes. The dataset design is careful—multi-node bistatic acoustic setup, loose trajectory constraints, and an honest description of where vision data is unusable. The unified Doppler spectrogram is a sensible standardization, not a deep technical leap, and the benchmark pipeline is a reasonable starting point. Releasing raw data and code is real evidence and should count in the paper's favor.\n\nThe soft spots are mostly where the evaluation outruns the data. The tracking benchmark uses four subjects—3 male, 1 female—and only paths with reliable vision annotations. The paper says this in Sections 3.2 and 5.1, so it is not hidden, but it does mean the cross-scene complementarity trend is suggestive rather than established. The larger issue is that the vision-based ground truth itself gets no independent error analysis. YOLOv8 foot keypoints back-projected to a floor plane via a single camera with AprilTag extrinsics; if that reference has systematic bias—heel-strike phase, floor-plane error, occlusion-induced pose failure—it changes both the absolute errors and the ranking of Wi-Fi versus acoustic versus fusion. The paper selects for visual reliability and then never validates the reference against anything else. That is load-bearing for the central claim, and it should be addressed with a calibration check or a second reference (even a few manual annotations or a laser rangefinder sweep).\n\nLesser issues: the identity recognition uses a subject-closed protocol, which is fine for a household-occupant scenario but should be labeled as such; there are no error bars or variance estimates on the headline percentages; and the exact TFRSP parameters and alignment details are not fully specified, so exact reproduction will depend on the code release. None of these are fatal. The summary claim that Wi-Fi is better for tracking and acoustics for identity is a reasonable reading of the data they collected, and the fusion-win pattern across environments is plausible.\n\nWho is this for: anyone in ubiquitous computing or wireless sensing who needs a public benchmark with two complementary modalities. I would send it to a serious referee. The dataset contribution stands on its own; the evaluation claims need tightening, not redoing.","headline":"A genuinely useful multi-modal dataset, but the headline complementarity finding rests on a vision ground-truth pipeline that gets no error analysis.","tokens_in":72473,"tokens_out":2045,"would_cite":true,"duration_ms":20935,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"XGait is introduced as the first public dataset that records the same indoor walks with Wi-Fi CSI, active acoustics, and camera ground truth, from 27 participants and more than 22,000 samples across three indoor environments, and it uses…","keywords":["wireless sensing","indoor tracking","human identification","Wi-Fi CSI","acoustic sensing","Doppler spectrogram","multi-modality dataset","gait recognition"],"falsifier":"Re-run the ground-truth extraction on the released videos for a subset of trials and compare the back-projected foot positions against manually annotated or motion-capture positions, especially in occluded non-line-of-sight segments; if the vision reference error is comparable to the roughly one-meter tracking differences reported, the modality comparisons would not be resolvable.","tokens_in":71475,"feed_emoji":"📡","tokens_out":3715,"duration_ms":34546,"temperature":0.7,"pith_summary":"XGait is presented as the first public dataset that captures the same indoor walks with Wi-Fi channel state information, active acoustic echoes, and camera-based ground truth, from 27 participants and more than 22,000 samples across a laboratory, a home, and a meeting room. The paper's central claim is that these two wireless modalities are complementary rather than redundant: Wi-Fi gives the more reliable basis for trajectory tracking, while acoustic Doppler signatures are more stable and discriminative for identity recognition. To make the comparison fair, the authors map both signal types into a shared Doppler spectrogram and provide a benchmark pipeline for alignment, feature construction, tracking, and recognition. The importance, if the claims hold, is that wireless sensing researchers get a common ground on which to test generalization across environments and trajectories, instead of relying on small single-modality collections.","feed_headline":"Wi-Fi plus sound tracks people and IDs them indoors","feed_subtitle":"22,000 synchronized walks with camera ground truth show Wi-Fi is best for tracking, acoustics for identity.","key_machinery":"The unifying device is the Doppler spectrogram: Wi-Fi CSI and acoustic echoes are both converted into shared time-frequency spectra using the time-frequency reassignment spectrum (TFRSP), after modality-specific pre-processing such as CSI-ratio cleaning for Wi-Fi and quadrature demodulation plus resampling for acoustics. Torso motion is captured as a path length change rate (PLCR) sequence, which serves as a modality-invariant temporal anchor for aligning unsynchronized links. Tracking solves a weighted least-squares velocity projection from multi-link signed PLCRs, while identity recognition uses either spectrum-driven deep features or the model-based polar-coordinate velocity profile (PPVP).","core_discovery":"The paper claims that a multi-modality dataset with vision ground truth can reveal and quantify modality complementarity in wireless sensing. The empirical discovery is a task-oriented division of labor: Wi-Fi provides the stronger baseline for trajectory tracking, acoustic signals provide finer spectral granularity that better survives cross-trajectory and clothing-induced shifts in identity recognition, and fusion of the two yields selective gains that grow as environmental complexity and non-line-of-sight conditions increase. The paper also reports that the proposed Doppler-spectrogram representation and benchmark pipeline make these comparisons reproducible, and that a model-based descriptor such as the polar-coordinate velocity profile inherits tracking errors, which explains why spectrum-driven recognition features are generally more robust.","pith_inferences":["Inference: the unified Doppler-spectrogram representation likely extends to other Doppler-based sensing modalities such as millimeter-wave radar, making XGait a template for future multi-modality benchmarks.","Inference: the observed 'selective gain' of fusion implies that a confidence-weighted or dynamically switching fusion rule could outperform static fusion, a testable extension on the released data.","Inference: the acoustic modality's robustness to clothing variation suggests that commodity-speaker identity systems remain viable even when Wi-Fi features degrade.","Inference: the sharp cross-scene zero-shot collapse points to multipath geometry as the dominant covariate, so scene-agnostic representations should be evaluated explicitly on XGait."],"forward_implications":["Researchers can benchmark both indoor tracking and identity recognition on the same recordings, with vision-derived trajectories as reference, enabling direct cross-modal comparison.","Fusion gains become more pronounced as environments grow more complex and occluded, reaching about 48% of trajectories in the meeting-room scenario, which implies adaptive fusion is worth pursuing.","Acoustic sensing appears better suited for identity recognition, while Wi-Fi is the more reliable tracking baseline under the tested conditions.","Model-based features like PPVP inherit tracking errors, so recognition evaluations should report both spectrum-driven and model-driven results.","Cross-scene zero-shot transfer collapses to near-chance levels, but a small amount of fine-tuning recovers quickly, suggesting that scene geometry rather than identity information dominates the domain shift."],"supporting_citations":[{"why":"Supplies the Linux CSI Tool used to acquire the 30-subcarrier Wi-Fi CSI measurements that form the Wi-Fi modality.","marker":"[8]"},{"why":"Provides the AprilTag fiducial markers used to estimate camera pose for vision-based ground-truth trajectory reconstruction.","marker":"[24]"},{"why":"Contributes the PLCR-based tracking formalism and the Widar-style dataset naming convention that XGait builds upon.","marker":"[25]"},{"why":"Defines the polar-coordinate velocity profile used as the primary model-based identity descriptor in the benchmark.","marker":"[28]"},{"why":"Serves as a prior Wi-Fi-only identification dataset that XGait compares against in the dataset overview.","marker":"[31]"},{"why":"Introduces the AcousticID system whose cycle-level gait attributes motivate the model-based feature paradigm.","marker":"[44]"},{"why":"Gives the CSI-ratio pre-processing that stabilizes Wi-Fi phase before Doppler spectrogram generation.","marker":"[49]"},{"why":"Provides the body-coordinate velocity profile that the benchmark references and contrasts with PPVP.","marker":"[54]"}],"fun_headline_variants":["22K walks with Wi-Fi and sound: best of both for indoor ID","Wi-Fi tracks, sound IDs: new dataset proves it","Vision-labeled Wi-Fi+audio dataset shows track vs ID split","Multimodal sensing: Wi-Fi leads tracking, audio leads ID","New dataset: Wi-Fi for tracking, acoustic for identification"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The vision-based ground-truth trajectories, produced by YOLOv8 pose estimation and AprilTag back-projection, are assumed accurate enough to serve as reference for tracking errors, but the paper gives no independent error analysis of this ground-truth pipeline.","fun_headline_variants_meta":{"raw":{"variants":["22K walks with Wi-Fi and sound: best of both for indoor ID","Wi-Fi tracks, sound IDs: new dataset proves it","Vision-labeled Wi-Fi+audio dataset shows track vs ID split","Multimodal sensing: Wi-Fi leads tracking, audio leads ID","New dataset: Wi-Fi for tracking, acoustic for identification"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000828,"raw_usage":{"total_tokens":3610,"prompt_tokens":928,"completion_tokens":2682,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":2594}},"tokens_in":544,"tokens_out":2682,"duration_ms":18310,"temperature":1.0,"reasoning_tokens":2594,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:28:27.213832+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the ground-truth extraction on the released videos for a subset of trials and compare the back-projected foot positions against manually annotated or motion-capture positions, especially in occluded non-line-of-sight segments; if the vision reference error is comparable to the roughly one-meter tracking differences reported, the modality comparisons would not be resolvable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Linux CSI Tool used to acquire the 30-subcarrier Wi-Fi CSI measurements that form the Wi-Fi modality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the PLCR-based tracking formalism and the Widar-style dataset naming convention that XGait builds upon."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the polar-coordinate velocity profile used as the primary model-based identity descriptor in the benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Serves as a prior Wi-Fi-only identification dataset that XGait compares against in the dataset overview."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the AcousticID system whose cycle-level gait attributes motivate the model-based feature paradigm."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the body-coordinate velocity profile that the benchmark references and contrasts with PPVP."}],"review_version":1}