{"id":"d2580a43-b9f3-49c1-845f-a6c2d56fb83a","arxiv_id":"2608.08439","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"FSTC-Encoder factorizes RF sensing into feature, spatial, and temporal correlation stages and reports the best mean cross-domain accuracy on Widar3.0 (92.15%) and strong results on CSI-Bench and XRF55 with a single shared backbone.","lead":"FSTC-Encoder is a new neural architecture for radio-frequency sensing that separates feature, spatial, and temporal modeling, so one backbone can serve Wi-Fi, radar, and RFID tasks. It reports strong cross-domain and cross-modality results on Widar3.0, CSI-Bench, and XRF55, including cutting the RFID performance gap through cross-RF training.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Widar3.0 headline result is not input-matched: FSTC uses RSSI+dCSI extra features versus baselines, so the 92.15% claim may reflect input advantage rather than the factorized architecture.","rationale":"The reader's formal weakest assumption concerns the set-based spatial encoder's exchangeability and whether permutation-invariant aggregation preserves behavior-relevant geometry. That is a reasonable theoretical concern, but the empirical results on a geometry-sensitive task (Room-Level Localization, 99.81% Accuracy) weaken it substantially. I find the input mismatch in the headline Widar3.0 comparison more load-bearing: it directly affects the abstract's central quantitative claim and the causal attribution to architecture. The reader did flag the input-mismatched Widar3.0 comparison in the rationale, so there is partial agreement, but it was not identified as the weakest assumption. My recommendation is to keep the conditional posture: the paper's core idea is plausible, the ablations are informative, and the cross-dataset breadth is real, but the Widar3.0 evidence needs a matched-input baseline before the domain-robustness claim can be accepted as architecture-driven. I am not proposing rejection because the concern is testable and the CSI-Bench matched-input ablation provides partial support for the architecture's ability to use richer features effectively.","tokens_in":12275,"tokens_out":4069,"duration_ms":46544,"concrete_test":"Run Wi-CBR (hard) and PatchTST with exactly the same three input feature families (CSI→RSSI, DFS, dCSI) and identical preprocessing under the same multi-fold CL, CO, and CE splits, reporting mean Accuracy and per-fold standard deviation. If Wi-CBR (hard) rises from 89.20 to within noise of 92.15, the Widar3.0 superiority claim is not established as an architectural effect. Also report fold-wise variance for FSTC to assess whether the 2.95-point gap is meaningful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim attributes high domain robustness to the factorized architecture, and the abstract leads with the 92.15% mean Accuracy on Widar3.0. The primary comparison in Table 1 is not input-matched. FSTC-Encoder processes three feature families (CSI→RSSI, DFS, dCSI), while the strongest baseline Wi-CBR (hard) uses two (CSI→Phase, DFS) and PatchTST uses one (CSI→DFS). The reported improvement over Wi-CBR (hard) is only 2.95 points (92.15 vs. 89.20). Because RSSI and dCSI are additional signal quantities, this gap could be explained by strictly more information entering the model rather than by the feature–spatial–temporal factorization. The matched-input control in Table 5 is run on CSI-Bench HAR, not on the Widar3.0 multi-factor protocols, and it shows that PatchTST degrades when given all three families in that setting; it does not show that a baseline cannot exploit the extra families under the particular Widar3.0 splits. The conclusion's statement that matched-input analyses attribute gains to the architecture rather than richer inputs therefore overreaches for the headline dataset. This is the most load-bearing weakness because the 92.15% number is the first quantitative claim in the abstract and the main evidence for domain robustness. The argument is not internally inconsistent, but the causal attribution is underdetermined by the presented evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"FSTC-Encoder factorizes RF sensing representation learning into feature, spatial, and temporal correlation encoding, with a structure-aware feature encoder, a set-based spatial encoder, and a hierarchical temporal encoder. The architecture is evaluated on Widar3.0, CSI-Bench, and XRF55, reporting 92.15% mean accuracy on Widar3.0 multi-factor protocols, a 57.94% CD average Weighted-F1 on CSI-Bench HAR, first place on three of four CSI-Bench tasks, and a reduction of the XRF55 cross-modality gap from 18.85% to 12.93%. The paper claims these results demonstrate domain robustness, task generality, and modality extensibility, and it uses ablations to attribute the gains to the factorized architecture rather than to richer inputs or larger models.","tokens_in":12550,"tokens_out":7423,"duration_ms":76371,"significance":"The factorization idea is well motivated and the architecture is clearly specified, with a clean interface between stages and a common spatial–temporal backbone across modalities. The evaluation spans multiple datasets, tasks, and RF modalities, and the matched-input ablation in Table 5 is a genuine control that goes beyond the common practice of simply adding input families. The structural ablation shows that each stage contributes. If the central claims survive the experimental concerns below, the paper would be a useful contribution to generalizable RF sensing. However, the current evidence does not yet establish the attribution claim for the headline Widar3.0 result, and the lack of uncertainty quantification makes the small reported margins difficult to interpret.","major_comments":[{"comment":"The headline Widar3.0 comparison is not input-matched. FSTC-Encoder is fed CSI→RSSI, DFS, and dCSI, while Wi-CBR (hard) uses CSI→Phase and DFS and PatchTST uses CSI→DFS. The reported advantage over Wi-CBR (hard) is 2.95 points (92.15 vs. 89.20). Since RSSI and dCSI are additional signal quantities, the margin could be explained by additional input information rather than by the factorized architecture. The matched-input ablation in Table 5 is conducted on CSI-Bench HAR, not on Widar3.0, and it shows only that PatchTST is harmed by adding all three families in that dataset. The conclusion's sentence that matched-input analyses attribute gains to the architecture rather than richer inputs therefore overreaches for the main dataset. Please add a matched-input Widar3.0 comparison (for example, a strong baseline trained with the same three feature families) or explicitly restrict the attribution claim to the CSI-Bench setting.","section":"Table 1 and §Cross-Domain Generalization on Widar3.0"},{"comment":"All empirical tables report single point estimates. For example, Table 1 reports only a mean of 92.15 for FSTC-Encoder against 89.20 for Wi-CBR (hard), and Table 2 reports a CD Average of 57.94 against 55.66 for the next-best method. No standard deviations, fold-level results, number of runs, or significance tests are given. In the absence of these, claimed margins of a few points cannot be distinguished from noise. Please report variance across folds and/or multiple seeds, and where appropriate a paired test across the cross-domain protocols.","section":"All experimental tables"},{"comment":"The spatial encoder is implemented as a permutation-invariant aggregation over the set of receivers, links, channels, or tags, using shared observation encoding, order-invariant statistics, and set attention with only an optional scalar quality bias β_r. This assumes that receiver identity, antenna geometry, and link topology carry no behavior-relevant information beyond what can be recovered from order-invariant statistics and the bias. The paper provides no ablation or comparison against an order-aware or geometry-aware spatial encoder (for example, one that consumes link coordinates or receiver indices) on any of the three datasets. Because the factorization argument relies on the spatial stage discarding only configuration-specific spatial layout, this assumption should be tested or at least explicitly discussed with evidence.","section":"Eqs. (9)–(13), §Set-Based Spatial Correlation Encoding"}],"minor_comments":[{"comment":"The method section leaves the values of L_w, S_w, δ, the dimensions D_F, D_S, D_T, D_Z, the number of layers and heads, and the training procedure (optimizer, learning rate, batch size, epochs, seeds) unspecified. The free parameters δ, L_w, S_w, and β_r are named but not assigned values. Please provide these details in an appendix or supplementary material, as the empirical claims are otherwise difficult to reproduce.","section":"§Method, hyperparameters"},{"comment":"The description 'holding out each condition of the target factor in turn while allowing the remaining factors to vary' does not state how many folds are used or how conditions are partitioned. Please add this detail so the reported 92.15% mean Accuracy is interpretable.","section":"§Experimental Setup, Widar3.0"},{"comment":"The phrase 'multi-factor Widar3.0 protocols' in the abstract could be read as a single protocol in which all factors vary simultaneously, whereas the experiments report a mean over cross-location, cross-orientation, and cross-environment protocols. Consider rewording to avoid ambiguity.","section":"Abstract and Table 1"},{"comment":"The sentence 'Its weaker CDev result may partly reflect information loss when reconstructing the multi-link structure from the flattened official release' is a limitation that is not further analyzed. Please provide supporting evidence or temper the claim of a balanced cross-domain profile.","section":"§In-the-Wild Domain Shifts on CSI-Bench"},{"comment":"The cross-RF DML training for FSTC-Multi is not described. A brief explanation of the mutual-learning objective and how the three modalities are used in training is needed to interpret the gap reduction from 18.85% to 12.93%.","section":"Table 4 and §Modality Extensibility"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the architecture proposal is sensible. The input-matched concern on Widar3.0 is the most important issue and is addressable with additional experiments. I would also ask the authors to report uncertainties across folds and seeds, as the current single-point tables make several comparisons unassessable. I do not see evidence of self-citation or novelty problems."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The FSTC-Encoder paper is a genuinely useful piece of engineering. The feature–spatial–temporal decomposition is clean, and the evaluation breadth—Widar3.0, CSI-Bench, XRF55, spanning WiFi, mmWave, and RFID—is unusual and welcome. The structural ablation showing temporal encoding contributes 23 points on Widar3.0, and the scale control showing FSTC has 2M parameters, both help. The matched-input comparison on CSI-Bench (Table 5) is the right kind of control, and it does show that PatchTST chokes when given all three feature families while FSTC benefits.\n\nThe soft spot is the headline number. Table 1 gives FSTC three feature families (RSSI, DFS, dCSI) while Wi-CBR (hard) gets two (phase, DFS) and PatchTST gets one. The 2.95-point edge over Wi-CBR on Widar3.0 could therefore come from strictly more input information rather than from the architecture. The matched-input control is run on CSI-Bench HAR, not on the Widar3.0 multi-factor protocols, so the conclusion that gains are architectural overreaches for the abstract's first quantitative claim. The authors should rerun the Widar3.0 comparison with matched features, or temper the causal language.\n\nEverything else is proportionate. All tables report single point estimates with no error bars, and several margins are small, so the statistical significance is unclear. No code or hyperparameters are released, which is a reproducibility gap. The set-based spatial encoder's exchangeability assumption—treating receivers, links, and tags as an unordered set—is a reasonable modeling choice, but the paper does not test whether receiver identity or geometry carries behavior-relevant information that the permutation-invariant aggregation discards.\n\nThis paper is for RF sensing researchers working on cross-domain and cross-modal generalization. It deserves peer review, but I would not accept it in current form. The input-matching issue on Widar3.0 and the missing error bars need addressing first. Conditional acceptance is the right posture.","headline":"Solid factorized architecture with broad evaluation, but the headline Widar3.0 gain is confounded by richer input features; the authors need a matched-input rerun.","tokens_in":13088,"tokens_out":2406,"would_cite":true,"duration_ms":24535,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes that one factorized encoder can serve as a reusable RF sensing backbone, and reports cross-domain, cross-task, and cross-modality results supporting that claim.","keywords":["RF sensing","domain generalization","cross-modal sensing","WiFi CSI","mmWave radar","RFID","set-based spatial encoding","human activity recognition"],"falsifier":"In a sensing setup where the same physical motion produces identical per-sensor signal statistics at two sensor locations but only one location's identity indicates the correct class, remove sensor identity from the input and compare accuracy with a model that receives sensor position; if the position-aware model wins, the set-based encoder's exchangeability assumption is the bottleneck. Also, permuting sensor order at test time should leave FSTC-Encoder's predictions unchanged, and any deviation would falsify the claimed permutation invariance.","tokens_in":12061,"feed_emoji":"📡","tokens_out":6015,"duration_ms":58469,"temperature":0.7,"pith_summary":"The paper tries to establish that a single factorized encoder, FSTC-Encoder, can serve as a reusable backbone for radio-frequency sensing across WiFi, millimeter-wave radar, and RFID, and across tasks from gesture recognition to breathing detection, even when users, environments, devices, locations, and orientations change. It argues that the obstacle to such generality is not model capacity but the coupling of behavior-relevant correlations with acquisition-specific feature geometry and spatial layout. The proposed solution standardizes the interfaces between feature, spatial, and temporal correlation stages rather than standardizing the raw inputs themselves, so only the feature configuration and task head change while the spatial-temporal backbone stays fixed. If true, this would let one sensing model be deployed in new settings without per-dataset redesign.","feed_headline":"One encoder handles WiFi, radar, and RFID sensing alike","feed_subtitle":"Separating signal structure from behavior lets one backbone survive new rooms, users, and radio types.","key_machinery":"The load-bearing mechanism is the factorization itself, enforced by three interfaces. First, structure-aware feature encoding separates branch-specific operators for each internal dependency structure (projection plus depthwise temporal convolution for low-dimensional descriptors; axis-wise convolution and masked pooling for ordered axes; paired embedding plus aggregation for coupled components) and then fuses them with sample-level gating and a joint projection. Second, set-based spatial encoding treats all spatial observations as exchangeable: each unit goes through the same residual encoder, and aggregation combines three order-invariant statistics (mean, max, standard deviation) with a set attention whose query is the spatial mean and whose values are the per-unit encodings, plus an optional receiver-quality bias. Third, hierarchical temporal encoding computes per-frame differences to expose motion transitions, applies multi-kernel depthwise convolutions, splits the sequence into overlapping windows modeled by a shared Local RoPE Transformer, and then models the window sequence with a Global RoPE Transformer using physical timestamps as scalar rotary coordinates. The output is a sequence of information-rich tokens that any lightweight task head can pool.","core_discovery":"FSTC-Encoder claims that human behavior induces recurring correlations in RF signals—feature, spatial, and temporal—that can be learned separately from the measuring configuration. It decomposes representation learning into three stages: structure-aware feature encoding maps each signal family (low-dimensional descriptors, ordered axes like subcarriers or Doppler bins, pair-coupled components like I/Q or sine/cosine) into a common temporal-spatial tensor; set-based spatial encoding treats receivers, links, channels, or tags as an unordered set and aggregates them with shared per-observation encoding, order-invariant statistics, and set attention; hierarchical temporal encoding separates motion-transition enhancement, within-window local attention with rotary position encodings, and cross-window global modeling using physical timestamps. With this factorization, the paper reports 92.15% mean accuracy under its harder multi-factor cross-domain protocols on Widar3.0, the best average cross-domain Weighted-F1 on CSI-Bench, first place on three of four CSI-Bench tasks, and a reduction of the cross-modality performance gap from 18.85% to 12.93% on XRF55 through cross-RF training.","pith_inferences":["Testable extension: freeze the spatial-temporal backbone and attach a feature branch for a completely new RF modality, such as UWB or acoustic sensing, to see whether the demonstrated transfer across WiFi, mmWave, and RFID extends further.","The set-based spatial encoder implies a stronger invariance claim: predictions should be identical under arbitrary permutation of receiver order, so a direct check on existing benchmarks would reveal whether residual performance differences come from non-exchangeable information like antenna geometry.","The cross-window physical-time modeling suggests the architecture should handle irregularly sampled temporal inputs without resampling, because timestamps enter as rotary coordinates rather than fixed grid positions.","The reduction of the modality gap with cross-RF training hints that the feature-spatial-temporal interfaces may also serve as a common space for unsupervised alignment across radio types, which would matter for zero-shot deployment."],"forward_implications":["One trained FSTC backbone can be reused for a new RF sensing task by swapping only the feature configuration and task head, without redesigning spatial or temporal modules.","Cross-domain deployment becomes more predictable: the worst cross-domain performance improves along with the average, so behavior under unseen environments and users is more reliable.","Cross-RF learning transfers knowledge from stronger modalities like WiFi and mmWave to weaker ones like RFID, reducing the largest per-modality gap rather than only the average.","Input quantity alone is not the driver: with matched feature families, a general sequence model degrades when given more feature families, whereas FSTC-Encoder improves, showing that structured modeling, not scale, is doing the work.","Adding a new RF modality reduces to designing one structure-aware feature branch, a smaller engineering surface than rebuilding the entire model."],"supporting_citations":[{"why":"Supplies the Widar3.0 dataset and its physics-based BVP baseline, whose official splits are near saturation and motivate the paper's harder multi-factor cross-domain protocols.","marker":"(Zhang et al. 2022)"},{"why":"Supplies CSI-Bench, the in-the-wild multi-task WiFi benchmark used for cross-device, cross-environment, cross-user, and task-generality evaluation.","marker":"(Zhu et al. 2025)"},{"why":"Supplies XRF55, the synchronized WiFi, millimeter-wave radar, and RFID dataset, and defines the cross-modality gap metric and the DML training baseline.","marker":"(Wang et al. 2024)"},{"why":"Provides Wi-CBR, the strongest RF-specific baseline and the central comparison for both Widar3.0 and CSI-Bench experiments.","marker":"(Zhang et al. 2026)"},{"why":"SCL is the closest prior work, explicitly modeling dimensional, spatial, and temporal correlation graphs; the paper contrasts its fixed graph construction with its own modular factorization.","marker":"(Liu et al. 2024b)"},{"why":"Provides PatchTST, the general sequence baseline used in the matched-input and model-scale controls that attribute gains to structure rather than input quantity.","marker":"(Nie et al. 2023)"},{"why":"Provides the standard Transformer baseline used in the general sequence model comparison on CSI-Bench.","marker":"(Vaswani et al. 2017)"}],"fun_headline_variants":["One backbone unifies WiFi, radar, RFID sensing","FSTC-Encoder slashes cross-modality gap by a third","RF sensing that adapts across devices and radio types","Same encoder, any radio: 92% accuracy across domains","Decomposing RF signals to generalize across sensing tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The spatial encoding stage assumes receivers, links, channels, and tags are interchangeable: it keeps only order-invariant statistics and set attention, so any behavior-relevant information tied to a sensor's identity or geometric position must be recoverable without that identity.","fun_headline_variants_meta":{"raw":{"variants":["One backbone unifies WiFi, radar, RFID sensing","FSTC-Encoder slashes cross-modality gap by a third","RF sensing that adapts across devices and radio types","Same encoder, any radio: 92% accuracy across domains","Decomposing RF signals to generalize across sensing tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000143,"raw_usage":{"total_tokens":1189,"prompt_tokens":979,"completion_tokens":210,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":128}},"tokens_in":595,"tokens_out":210,"duration_ms":3040,"temperature":1.0,"reasoning_tokens":128,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:35:20.004772+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In a sensing setup where the same physical motion produces identical per-sensor signal statistics at two sensor locations but only one location's identity indicates the correct class, remove sensor identity from the input and compare accuracy with a model that receives sensor position; if the position-aware model wins, the set-based encoder's exchangeability assumption is the bottleneck. Also, permuting sensor order at test time should leave FSTC-Encoder's predictions unchanged, and any deviation would falsify the claimed permutation invariance.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies CSI-Bench, the in-the-wild multi-task WiFi benchmark used for cross-device, cross-environment, cross-user, and task-generality evaluation."},{"cited_title":"H.; Sinthong, P.; and Kalagnanam, J","cited_arxiv_id":null,"evidence_quote":"Provides PatchTST, the general sequence baseline used in the matched-input and model-scale controls that attribute gains to structure rather than input quantity."}],"review_version":1}