{"id":"74fe1917-e126-4c78-b59f-965d639f46df","arxiv_id":"2506.12997","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A per-delay-bin Doppler velocity representation plus an order- and repetition-invariant classifier improves cross-user Wi-Fi hand gesture recognition.","lead":"This paper introduces a Wi-Fi sensing method that turns channel measurements into per-path Doppler velocity traces and classifies them with a model that ignores the random ordering of those traces. It reports better cross-user hand gesture recognition than prior methods on a new public dataset, especially when a few calibration samples from the target user are available.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Delay-resolution premise fails: UTHAMO's 2.4-GHz channel has c/BW≈7–15 m, so the 52 delay bins are not resolved path directions; Eq. (38)/(43) are unsupported unless delay-spread data are provided.","rationale":"The paper has two intertwined claims: an empirical improvement in cross-user hand-gesture classification, and a theoretical model in which each delay bin is a distinct Doppler velocity projection. The empirical claim is supported by a released dataset, a LOSO protocol, ablations, and a plausible classifier; those are real strengths. The load-bearing weakness is the delay-resolution premise. The reader's identified assumption, κ≫1 in Lemma 3, is relevant but sits after a more fundamental step: Eq. (30) shows each delay bin is a weighted mixture of all physical paths, and with c/BW≈7–15 m in a 6 m×5.6 m room those paths are not separated. The high-index bins in Fig. 6 correspond to impossible excess path lengths, suggesting the method is feeding normalized sidelobe/noise bins into the classifier. A single delay-spread/energy-concentration audit on the dataset would settle whether the decomposition is physically real or merely a learned transformation. The verdict should remain conditional: the manuscript needs either delay-resolution measurements supporting the decomposition or a substantially softened theoretical claim.","tokens_in":21225,"tokens_out":9135,"duration_ms":112379,"concrete_test":"Use the released UTHAMO data to compute the IDFT delay profile (Eq. 26) for each AP and a static segment; measure the number of bins containing 90% of the energy and the RMS delay spread, and compare the PSD peak SNR of bins 17/35/47 to the noise floor. If 90% energy lies in ≤2 bins (delay spread well below 1/BW) or the high-index bins are noise-dominated, the claimed per-bin Doppler projection decomposition is falsified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central theoretical claim is Eq. (38): each delay-bin phase gives the projection v(s)·m_i of the user's velocity onto a single scatterer direction. This requires that h(s;τ_i) in Eq. (30) is dominated by one resolved multipath component. But Eq. (30) is a superposition over all paths weighted by sinc_N(Δf(τ_i−τ_j)); paths are separable only if |τ_i−τ_j|>1/BW (Section III-C). The UTHAMO link is 2.4-GHz 802.11ac with 20-MHz (or at most 40-MHz) bandwidth, giving c/BW≈15 m (or 7.5 m), comparable to the 6 m×5.6 m room. Most physical paths therefore fall in the same delay bin; the high-index bins shown in Fig. 6 (e.g., 17, 35, 47) correspond to delays ≥850 ns, i.e., excess path lengths ≥255 m, which cannot exist in this room and must be sinc sidelobes or noise. After normalization, even noise-only bins become unit-variance inputs to MORIC. Consequently, V_r(s) is not a set of distinct 'virtual-camera' velocity observations; the PSD peak of a mixture is not v(s)·m_i (Eq. 43), and the κ≫1 assumption in Lemma 3 is secondary. The paper reports no delay-spread or per-bin SNR audit, so the central representation is unsupported. This does not invalidate the empirical accuracy, but it removes the physical interpretation that motivates the method.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MORIC, a Wi-Fi CSI-based human activity recognition pipeline. It sanitizes CSI phase, applies an IDFT across subcarriers to obtain a delay profile, and estimates a Doppler velocity per delay bin from the peak of the bin's power spectral density. These per-bin velocities are treated as unordered, possibly repeated projections of the user's 3-D velocity onto unknown multipath directions ('virtual cameras'). A classifier based on random convolutional kernels, per-bin MLPs, and max-pooling is designed to be invariant to ordering and repetition. The theoretical part derives Eq. (38)/(43), connecting per-bin Doppler velocity to v(s)·m_i under a von Mises-Fisher concentration assumption. Experiments on a new six-user hand-gesture dataset (UTHAMO) report leave-one-subject-out accuracies of 56.3% for four classes with all APs versus 36.1% for the best baseline, and 98.8% with ten calibration samples per class.","tokens_in":21571,"tokens_out":9823,"duration_ms":105494,"significance":"If the physical interpretation holds, the delay-Doppler decomposition is an elegant way to turn multipath from a nuisance into multiple observation angles, and MORIC's permutation-invariant architecture is a sensible response to the unordered nature of the representation. The paper contributes a public dataset, explicit derivations, and an extensive comparison against several baselines. However, the central interpretation rests on delay resolvability and angular concentration assumptions that are not validated in the measured environment; the empirical gains may survive a weakened interpretation, but the motivating physical claim does not yet have adequate support.","major_comments":[{"comment":"The decomposition into 'multipath components' indexed by delay bin requires that physical paths be resolvable, i.e., |τ_i - τ_j| > 1/BW, as the manuscript itself states near Eq. (31). The paper never specifies the UTHAMO bandwidth; the 52 data subcarriers on a 2.4-GHz 802.11ac link imply the 20-MHz mode, for which Δτ_min = 50 ns and Δd_min = c/BW ≈ 15 m. In the 6 m × 5.6 m room, most physical paths differ by far less than 15 m and therefore fall into the same delay bin, so h(s; τ_i) in Eq. (30) is a mixture of unresolved paths with different directions and Doppler shifts, not a single path with a well-defined mean direction m_i. The high-index bins shown in Fig. 6 (17, 35, 47) correspond to delays of 850 ns, 1.75 μs, and 2.35 μs, i.e., excess path lengths of 255 m, 525 m, and 705 m, which cannot exist in this room and must be sinc sidelobes or noise. Because Section III-G normalizes every retained bin to unit variance, these non-physical bins enter the classifier as pseudo-velocity features. The central representation V_r(s) is therefore unsupported unless the authors provide a measured power-delay profile and a per-bin SNR audit showing which bins contain real paths, or they reduce the physical claim to an empirical feature construction.","section":"III-C, Eqs. (27)-(31), Fig. 6"},{"comment":"The derivation of the projected velocity identity relies on Lemma 3, which requires κ ≫ 1 for the von Mises-Fisher distribution p_i_Ω. No estimate of κ is provided for the measured environment, and indoor sub-7 GHz scattering in a small furnished office is plausibly diffuse or multimodal rather than sharply concentrated. If κ is moderate or the directional distribution has multiple modes, the ratio in Eq. (34) is a weighted average over different directions, not v(s)·m_i, and the PSD peak in Eq. (43) does not correspond to a single mean direction. This is not an external disagreement but an unvalidated assumption inside the derivation. A concrete test would be to estimate κ (or the angular spread) per delay bin from the UTHAMO measurements, or to evaluate Eq. (34) and the PSD peak location for κ in the range 1-20 and show that the velocity-projection interpretation remains accurate.","section":"III-D, Lemma 3, Eqs. (36) and (43)"},{"comment":"Equation (41) gives S(f; τ_i) = E_{r∼p_i_Ω}[|β_i(r)|^2 δ(f − v(s)·r/λ)]. The text then claims that the PSD peak occurs at f = v(s)·m_i/λ. This inference is not justified by Lemma 3, which is a statement about the mean of a smooth function on the sphere, whereas here the integrand contains a Dirac delta and the quantity of interest is the mode of an expectation. If |β_i(r)|^2 varies over the support of p_i_Ω, the argmax of S(f; τ_i) is biased relative to v(s)·m_i/λ; Lemma 3 applied naively would give the mean Doppler, not the peak Doppler. The paper should either state and justify the additional condition that |β_i(r)|^2 is slowly varying over the concentration support, or quantify the resulting velocity bias.","section":"III-E, Eqs. (41)-(43)"}],"minor_comments":[{"comment":"State the Wi-Fi bandwidth (20 vs 40 MHz) and subcarrier spacing explicitly; all delay-bin interpretations depend on this value.","section":"IV-A"},{"comment":"The symbol v_r(s; τ_i) is used both for the estimated phase derivative and for the true projection v(s)·m_i; use separate notation (e.g., a hat) until the approximation is invoked.","section":"III-D, Eqs. (37)-(38)"},{"comment":"The statement that the PSD 'behaves as a probability density function on the Doppler frequency f' is imprecise, because S(f; τ_i) is an unnormalized weighted expectation that depends on |β_i(r)|^2.","section":"After Eq. (41)"},{"comment":"The Fig. 7 caption contains a typo ('peojections'), and the captions of Figs. 10-13 contain OCR artifacts ('/uni...') that must be cleaned in the published version.","section":"Fig. 7 caption; Figs. 10-13 captions"},{"comment":"The '–' entry for APNSS + APSC under All APs is unexplained; state that this baseline does not support multi-AP fusion.","section":"Table I"},{"comment":"Clarify whether the SNR threshold is 2 dB or a linear SNR of 2; the notation SNRdB with 'less than or equal to 2' is ambiguous.","section":"Eq. (45)"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline case. The empirical comparison is promising, but the paper's motivating physical interpretation is not currently supported by the data: the delay-resolution premise fails for a 20-MHz link in a 6 m × 5.6 m room, and the κ and |β| assumptions behind Eq. (43) are unvalidated. If the authors can add a measured power-delay-profile and per-bin SNR audit, estimate or bound κ, and quantify the PSD-peak bias, the paper could be acceptable; otherwise they should reframe the method as an empirical representation and soften the virtual-camera claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my read. The paper's substance: it builds a delay-profile representation of Wi-Fi CSI, estimates per-delay-bin Doppler velocity via PSD peaks, and feeds the unordered set of projections through a ROCKET-based order/repetition-invariant classifier (MORIC). That combination is new, the dataset is public, and the cross-user four-class result (56.3% vs 36.1% for the best baseline) is a real improvement, even if absolute accuracy is modest. The ablation study is honest and useful: phase compensation, max pooling, PSD estimation, and the classifier all contribute. I believe the empirical claim, as far as it goes.\n\nThe math is mostly standard (IDFT, Dirichlet kernel, von Mises–Fisher asymptotics) and not circular. But the stress-test concern is on target: with 20 MHz bandwidth, delay resolution is c/BW ≈ 15 m, comparable to the room size. So most physical paths fall in the same few delay bins, and the high-index bins in Fig. 6 (17, 35, 47) cannot correspond to physically plausible excess path lengths in a 6 m × 5.6 m room; they are leakage or noise. Eq. (38) and (43) present each bin's PSD peak as v(s)·m_i, which requires the bin to be dominated by a single direction with sharply concentrated angular spread. The paper itself acknowledges bins may contain unresolved paths and then proceeds as if each bin had a well-defined mean direction. It never reports delay spread or a per-bin SNR audit, and κ is never estimated. So the “virtual cameras” interpretation is not supported as stated. The method might still work as a learned feature over a fixed filter bank; the physical narrative needs to be backed by data or softened.\n\nThat said, the flaw is in the interpretation, not in the evaluation. The experiments are reasonably careful: LOSO cross-validation, multiple APs and orientations, confusion matrices, ablations, and a public dataset. The main empirical soft spots are one environment, six subjects, and no comparison to a strong cross-domain baseline such as Widar3.0; the reader flagged that, and I agree. The calibration results are striking but amount to few-shot adaptation rather than pure zero-shot generalization; the paper should be clearer about that.\n\nWho this is for: people working in Wi-Fi sensing and HAR will want to read it. It deserves a serious referee. I recommend sending it to peer review with a request for delay-spread evidence and κ justification, plus an additional baseline. A motivated revision can fix the gap between the math and the measured channel.\n\nRecommendation: accept for review, do not desk reject.","headline":"A useful method with a shaky physical story: the delay–Doppler decomposition is empirically strong, but the delay-resolution premise needs validation before the “per-path virtual camera” interpretation is taken at face value.","tokens_in":22102,"tokens_out":2630,"would_cite":true,"duration_ms":29558,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multipath echoes as motion cameras: Wi-Fi gesture recognition hits 56.3% cross-user accuracy","keywords":["Wi-Fi sensing","human activity recognition","channel state information","Doppler velocity","delay-Doppler decomposition","multipath","order-invariant classification","cross-user generalization"],"falsifier":"Place a Wi-Fi transmitter and receiver with one reflector moving at a known constant velocity, and measure the per-delay Doppler PSD in two rooms: one with a single dominant reflector cluster and one with deliberately diffuse scattering. If the PSD peak is not at $v(s)\\cdot m_i/\\lambda$, or is multi-modal, in the diffuse room, the projection identity on which the representation rests fails under realistic conditions.","tokens_in":21023,"feed_emoji":"📡","tokens_out":9443,"duration_ms":88465,"temperature":0.7,"pith_summary":"Wi-Fi-based human activity recognition works in controlled settings but degrades sharply when the user, room, or hardware changes, because methods lean on environment-specific magnitude and phase patterns. This paper argues that the right representation is motion itself: it decomposes CSI into delay bins and, in each bin, reads a Doppler velocity that is the projection of the user's 3-D velocity onto an unknown direction set by that multipath path. Because those paths arrive in random order and repeat unpredictably, the paper introduces MORIC, a classifier built from random convolutional kernels and max-pooling across delay bins, which is invariant to order and repetition. On a new four-gesture hand-motion dataset, MORIC reaches 56.3% average cross-user accuracy with all access points, versus 36.1% for the best baseline, and a few calibration samples per class push accuracy above 98% with a single access point. A sympathetic reader would take the paper to be establishing that per-path Doppler projections, rather than aggregated Doppler or raw CSI, are the route to generalizable Wi-Fi sensing.","feed_headline":"Multipath echoes as cameras: Wi-Fi activity accuracy hits 56.3%","feed_subtitle":"By splitting CSI into per-delay Doppler velocities and ignoring their order, MORIC doubles the prior cross-user score.","key_machinery":"The central mechanism is the delay-Doppler decomposition: an N-point IDFT over CSI subcarriers separates the channel into N resolvable delay bins, each carrying a sinc-weighted multipath component $h(s;\\tau_i)$. Doppler velocity per bin is estimated not by differentiating phase but by locating the peak of the power spectral density, which Lemma 3 (von Mises–Fisher concentration with $\\kappa\\gg1$) identifies as $f^* = v(s)\\cdot m_i/\\lambda$. The classification engine is MORIC: random 1-D convolutional kernels turn each velocity time series into a feature vector, K shared MLP heads reduce each vector, and max-pooling across delay bins yields a single representation invariant to the order and repetition of the projections.","core_discovery":"The paper's central claim is that each resolvable multipath delay bin carries a distinct, physically meaningful Doppler velocity projection of the same 3-D motion. After an IDFT over subcarriers, the channel at delay $\\tau_i$ is modeled as a superposition of scatterers whose directions follow a von Mises–Fisher distribution, and for a sharply concentrated distribution ($\\kappa\\gg1$) both the phase derivative and the PSD peak reduce to $v_r(s;\\tau_i)=v(s)\\cdot m_i$, the projection of the velocity vector onto the mean arrival direction of that path; equations (38) and (43) state this identity. Since the directions $m_i$ are random and unordered, the set of projections is treated as a bag, and MORIC's shared MLP heads and max-pooling over delay bins make the classifier invariant to which view comes first or repeats. The claim is supported by leave-one-subject-out experiments on the UTHAMO dataset, where this representation outperforms magnitude-, phase-, and aggregated-Doppler baselines.","pith_inferences":["Beyond the paper, one can test the geometric story directly: move a single reflector at a known velocity in rooms with one dominant reflector cluster versus deliberately diffuse scattering, and check whether the per-delay PSD peak stays single-peaked and lands at $v(s)\\cdot m_i/\\lambda$.","Beyond the paper, since max-pooling discards which delay bin produced which projection, an experiment that sorts or clusters the projections before pooling would reveal whether the bag assumption is correct or whether bin identity carries useful information.","Beyond the paper, the virtual-camera analogy implies a geometry-dependent ceiling: access points that view the motion from similar directions should add little information, so adding elevation diversity to AP placement should sharpen the confusable push–pull versus left–right pair."],"forward_implications":["With all five access points, MORIC reaches $56.3\\%\\pm 9.1\\%$ four-class cross-user accuracy against $36.1\\%\\pm 5.1\\%$ for the strongest baseline (the CSI ratio model) at orientation $180^\\circ$ of the UTHAMO dataset.","A single access point where the user blocks the line of sight (SAP 5) already gives $51.5\\%$ cross-user accuracy, showing that multi-AP fusion is helpful but not required.","Calibration with 4, 6, or 10 samples per gesture class raises single-AP accuracy to $88.5\\%$, $92.5\\%$, and $98.8\\%$, respectively, making few-shot on-device adaptation plausible.","Removing the max-pooling stage drops accuracy from $56.3\\%$ to $42.5\\%$, confirming that order and repetition invariance is load-bearing for the method.","Binary experiments show the representation separates up–down from circle at $91.7\\%$ and push–pull from up–down at $89.2\\%$, while push–pull versus left–right remains the hardest pair."],"supporting_citations":[{"why":"Supplies the linear-regression phase compensation that removes STO and SFO before the delay-Doppler transformation.","marker":"[24]"},{"why":"Supplies the von Mises–Fisher scatterer model on the sphere used in Lemma 3.","marker":"[25]"},{"why":"Supplies the instantaneous-frequency identity used in Lemma 2 to relate phase derivative to velocity.","marker":"[26]"},{"why":"Supports PSD-based Doppler velocity estimation as more robust than phase differentiation.","marker":"[27]"},{"why":"Supports the Gaussian-profile approximation of the Doppler PSD under concentrated scatterer distributions.","marker":"[29]"},{"why":"Supplies the random convolutional kernel transform used for feature extraction in MORIC.","marker":"[32]"},{"why":"Provides the aggregated-Doppler baseline (SHARP) that MORIC outperforms by preserving per-path projections.","marker":"[18]"},{"why":"Provides the CSI ratio model, the strongest baseline in the comparison, which MORIC surpasses by 20 points.","marker":"[15]"},{"why":"Supplies the CSI extraction toolkit used to collect the UTHAMO dataset.","marker":"[35]"}],"fun_headline_variants":["Multipath delays as velocity views for Wi-Fi activity","Bag of Doppler projections: robust Wi-Fi activity","From echo delays to velocity: Wi-Fi HAR without order","MORIC: order-blind Doppler projections for Wi-Fi HAR","Delay-Doppler decomposition generalizes Wi-Fi HAR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the scatterers contributing to each delay bin arrive from a tightly clustered direction, so the Doppler power spectrum has a single peak at $v(s)\\cdot m_i/\\lambda$; if reflections arrive from many directions at once, that peak is not a clean velocity projection.","fun_headline_variants_meta":{"raw":{"variants":["Multipath delays as velocity views for Wi-Fi activity","Bag of Doppler projections: robust Wi-Fi activity","From echo delays to velocity: Wi-Fi HAR without order","MORIC: order-blind Doppler projections for Wi-Fi HAR","Delay-Doppler decomposition generalizes Wi-Fi HAR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1472,"prompt_tokens":1024,"completion_tokens":448,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":369}},"tokens_in":640,"tokens_out":448,"duration_ms":4921,"temperature":1.0,"reasoning_tokens":369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:36:17.329857+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Place a Wi-Fi transmitter and receiver with one reflector moving at a known constant velocity, and measure the per-delay Doppler PSD in two rooms: one with a single dominant reflector cluster and one with deliberately diffuse scattering. If the PSD peak is not at $v(s)\\cdot m_i/\\lambda$, or is multi-modal, in the diffuse room, the projection identity on which the representation rests fails under realistic conditions.","supporting_citations":[{"cited_title":"Decimeter ranging with channel state information,","cited_arxiv_id":null,"evidence_quote":"Supplies the linear-regression phase compensation that removes STO and SFO before the delay-Doppler transformation."},{"cited_title":"Correlation properties in channels with von mises-fisher distribution of scatterers,","cited_arxiv_id":null,"evidence_quote":"Supplies the von Mises–Fisher scatterer model on the sphere used in Lemma 3."},{"cited_title":"Estimating and interpreting the instantaneous frequency of a signal-part 1: Fundamentals,","cited_arxiv_id":null,"evidence_quote":"Supplies the instantaneous-frequency identity used in Lemma 2 to relate phase derivative to velocity."},{"cited_title":"Adaptive spectral doppler estimation,","cited_arxiv_id":null,"evidence_quote":"Supports PSD-based Doppler velocity estimation as more robust than phase differentiation."},{"cited_title":"Doppler power spectrum in channels with von mises-fisher distribution of scatterers,","cited_arxiv_id":null,"evidence_quote":"Supports the Gaussian-profile approximation of the Doppler PSD under concentrated scatterer distributions."},{"cited_title":"Rocket: exceptionally fast and accurate time series classification using random convolutional kernels,","cited_arxiv_id":null,"evidence_quote":"Supplies the random convolutional kernel transform used for feature extraction in MORIC."},{"cited_title":"Sharp: Environment and person independent activity recognition with commodity ieee 802.11 access points,","cited_arxiv_id":null,"evidence_quote":"Provides the aggregated-Doppler baseline (SHARP) that MORIC outperforms by preserving per-path projections."}],"review_version":1}