{"id":"c1e592af-6f83-4051-989e-8f3f4c6e2330","arxiv_id":"2604.02056","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Restoring a fixed N-slot fusion interface with aggregated source-to-target proxy tokens improves missing-modality HAR robustness over skip, imputation, and distillation baselines on XRF55, MM-Fi, and OctoNet.","lead":"COMPASS fills missing sensor slots with proxy tokens so a fixed fusion head always sees the same complete input layout. This improves activity recognition when WiFi, radar, cameras, or other sensors drop out in real deployments.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the boundary the authors and reader already flag.","rationale":"The paper’s strongest claim is not that sum fusion is optimal for all multimodal tasks, but that interface completeness is an effective principle for missing-modality HAR when slots are completed before a fixed fusion operator. Evidence for that scoped claim is consistent across three datasets, all subsets, controlled missing-modality baselines, and ablations; the space-stabilization term and proxy completion each move accuracy in the expected direction. The reader correctly identifies the main soft spot—global tokens and sum fusion may be insufficient when structure is sequence-level or a dominant modality needs expressive cross-attention—but the manuscript already treats this as a boundary condition rather than a universal guarantee. That does not undermine the reported HAR gains or the framing contribution. Absent code is a practical gap for reproducibility, not a correctness failure of the argument. Therefore the CONDITIONAL verdict with high confidence stands; no adjustment is warranted from this stress pass.","tokens_in":17192,"tokens_out":491,"duration_ms":6882,"concrete_test":"Re-run the XRF55 7-scenario protocol with the same encoders/projections but replace sum fusion by a single shared multi-head cross-attention fusion head over the completed N slots (real or proxy); if COMPASS’s average advantage over X-Fi and the controlled baselines collapses by more than ~3–4 pp while single-modality gains remain, the interface-completion principle is doing less work than claimed and sum fusion is a material confounder.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is empirical and scoped to HAR classification with global tokens: restoring a fixed N-slot interface via directed proxy completion plus sum fusion improves robustness under missing modalities. The reader’s weakest assumption (global mean-pooled tokens + uniform average + parameter-free sum may not restore sequence-level or learned cross-slot interactions) is real but already stated by the authors in §4.3 and the conclusion, and is not required for the HAR results as reported. Tables 1–5, ablations (Tables 6–7), multi-seed stats, and controlled baselines support the claim inside that scope; dominant-modality cases where X-Fi wins are disclosed. No hidden derivation failure, circularity, or unacknowledged confound appears load-bearing for the stated claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"COMPASS frames missing-modality multimodal sensing as a fusion-interface mismatch problem and proposes restoring a fixed N-slot input before prediction. Observed modalities fill their slots with real projected tokens; missing slots are filled by directed source-to-target proxy generators in a shared latent space, with multi-source estimates mean-aggregated into one filler. The fusion head is a parameter-free sum over the completed slots. Training uses synthetic modality masking, slot-compatibility L2 supervision against held-out real slots, VICReg-style space stabilization, and an auxiliary per-proxy task loss. On XRF55, MM-Fi, and OctoNet HAR, a single model is evaluated across all non-empty modality subsets, with multi-seed means, ablations of proxies/masking/losses, and controlled re-implementations of SMIL-style imputation, PTA distillation, and CMPT-style translation under matched fine-tuning. Gains are largest in low-modality and strong-modality-absent regimes; dominant-modality cases where cross-attention (X-Fi) remains competitive are disclosed.","tokens_in":17452,"tokens_out":1275,"duration_ms":17249,"significance":"If the results hold, the paper offers a clean and practical design principle for ubiquitous multimodal sensing: keep the fusion interface fixed and push missingness handling into pre-fusion slot completion, rather than subset-dependent fusion or full reconstruction. The contribution is well scoped to HAR classification with global tokens, and the empirical package is stronger than typical missing-modality sensing papers: full-subset tables on three public benchmarks, multi-seed statistics, matched controlled baselines (Table 5), and ablations that separate zero-fill, proxy-only, and full objectives (Tables 6–7). The shared-generator scalability note and corrected OctoNet ToF comparison are also scientifically careful. Within HAR, this is a solid systems/methods contribution rather than a foundational theoretical advance, but the interface-completeness framing is useful and falsifiable.","major_comments":[{"comment":"§3.3–3.4 and Table 6: the central claim attributes robustness primarily to restoring a canonical N-slot interface, yet the largest single ablation drop on XRF55 is removing space stabilization (−2.3pp), while zero-fill already beats X-Fi’s 7-scenario average (76.9 vs 72.1). Please add a short attribution analysis that separates (i) fixed-layout fusion, (ii) proxy generation, and (iii) shared-space/training objectives—e.g., fixed-layout with zeros vs fixed-layout with proxies under identical losses—so readers can judge how much of the gain is truly “interface completeness” versus representation geometry and supervision.","section":null},{"comment":"§4.2 / Table 2 (MM-Fi): X-Fi remains better on R-only, I+R, and D+R. The paper notes this as a dominant-modality boundary, but the main narrative still presents COMPASS as broadly superior. Please quantify and discuss this trade-off more systematically (e.g., when the strongest observed modality exceeds a unimodal accuracy threshold, does sum fusion + proxies underperform expressive cross-attention?), and state the operating regime of the method more precisely in the abstract/conclusion rather than only in §4.3.","section":null},{"comment":"§3.4 training protocol: all supervision for slot completion assumes fully observed training samples with synthetic masking. This is standard, but the paper’s deployment motivation includes sensors that may also be missing in training data. Either evaluate a realistic incomplete-training setting (random missingness in the training set without real-slot targets for some samples) or explicitly limit the claim to “complete training, incomplete inference,” which is currently understated relative to the introduction’s broader framing.","section":null}],"minor_comments":[{"comment":"Throughout the manuscript PDF/source, many figure captions and section headings appear as garbled replacement characters (e.g., “�������”). Ensure the camera-ready text renders COMPASS and all figure labels correctly.","section":null},{"comment":"§3.2, Eq. (13): uniform averaging is default; the text says learned weighting “does not consistently improve,” but no table is shown. A one-row ablation (mean vs attention/confidence weighting) would make this claim checkable.","section":null},{"comment":"Table 8 and the shared-generator paragraph in §4.3: the shared O(1) generator result is important for the scalability claim but is only briefly reported. Consider promoting numbers (params, OctoNet/XRF55 averages) into a small table.","section":null},{"comment":"Figure 3: the compactness–separation scatter is informative; please define the exact formulas for “cross-modal compactness” and “inter-class separation” in the caption or appendix so the plot is reproducible.","section":null},{"comment":"Notation: V, O, M, and slot index conventions are clear, but ¯u_i vs ˜u_i vs u_{i→j} could be summarized in a short symbol table for readers skimming §3.","section":null},{"comment":"Related work §2.1–2.2 is generally fair; a brief explicit contrast with prompt-based missing-modality methods (beyond CMPT) on whether they preserve a fixed fusion layout would sharpen positioning.","section":null}],"recommendation":"minor_revision","confidential_remarks":"I agree with the reader’s CONDITIONAL/HIGH assessment: the central empirical claim is supported inside the stated HAR/global-token scope, and the authors already flag the main boundary conditions. No load-bearing circularity or derivation failure. The three major comments are fixable with additional analysis and clearer claim scoping rather than new theory. Fit for a solid multimodal sensing / ubiquitous computing venue; not a top-theory CV breakthrough, but methodologically careful enough for acceptance after minor revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: missing modalities break the fusion interface, not just the information budget, and restoring a fixed N-slot layout with directed source-to-target proxies plus mean aggregation beats branch-skipping and several standard recovery baselines on three HAR benchmarks.\n\nWhat is actually new is the package, not any single gadget. They fix one slot per modality, fill observed slots with real mean-pooled tokens, complete missing slots with directed single-layer Transformer maps from each observed source, average those estimates, and always fuse the same N tokens with parameter-free sum. Training is synthetic masking plus slot L2 compatibility, VICReg-style space stabilization, and a light per-proxy task head. Related work (CMPT, SMIL, distillation, X-Fi) is cited and then controlled against under matched fine-tuning; that is the right way to do it.\n\nThe evidence is the strong part. Full-subset tables on XRF55, MM-Fi, and OctoNet (31 combinations), multi-seed means, ablations on proxy vs zero-fill, p_drop, and each loss term, plus a corrected ToF expander for the X-Fi collapse. Gains are largest when a strong modality is absent (e.g., WiFi+RFID on XRF55: 86.3 vs 58.1). Dominant-modality cases where X-Fi’s cross-attention wins are reported, not buried. Geometry plots support the compatibility story without overclaiming shared/private decomposition.\n\nSoft spots are real but proportional. Sum fusion and one global token per modality deliberately isolate the interface claim; they will not carry sequence-level pose or dense regression, and the authors say so. Pairwise generators are O(N^{2}) in storage (active cost is |O|·|M|); they show a shared-generator variant holds accuracy, so this is manageable for 3–5 modalities, not a hidden bomb. No public code in the manuscript is the practical gap for a systems paper. Hyperparameters (p_drop, λs, VICReg coeffs) are free but ablated enough that they do not look load-bearing.\n\nThis is for people who ship multimodal sensing under sensor dropouts. It is not a foundational theory paper; significance is subfield-level and engineering-clear. Math and citation pattern look clean; no circularity. I would send it to peer review and would cite the interface-completion framing and the controlled tables if I work on missing-modality HAR.","headline":"Solid empirical methods paper: interface-complete fusion via directed proxies is a useful framing and the HAR results are broad and carefully controlled, with the main limits already disclosed.","tokens_in":18014,"tokens_out":598,"would_cite":true,"duration_ms":7065,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Missing sensors break the fusion interface, not just the data; filling every modality slot with a proxy restores it.","keywords":["multimodal sensing","human activity recognition","missing modalities","fusion-interface completion","proxy tokens","shared latent space","ubiquitous sensing"],"falsifier":"On a HAR benchmark where a dominant modality already carries almost all the signal, or on a dense spatial task such as pose estimation, a cross-attention or sequence-level fusion baseline would outperform the completed-slot sum fusion under the same missingness patterns.","tokens_in":18091,"feed_emoji":"📡","tokens_out":583,"duration_ms":5717,"temperature":0.7,"pith_summary":"Multimodal sensing systems for human activity recognition train a fusion head on a fixed set of modality slots, but at deployment sensors drop out and that interface changes. Compass treats this as a structural problem: each modality keeps a permanent slot; observed sensors put real tokens in their slots, and missing ones get proxy tokens estimated from the sensors that remain. Source-specific proxies for the same missing slot are averaged into one filler so the same simple fusion operator always sees a complete N-slot input. Training simulates dropouts, forces proxies to stay compatible with real slots, and stabilizes the shared representation space so the fillers stay useful for recognition. On three HAR benchmarks the approach improves accuracy under single- and multi-missing patterns, especially when strong modalities are absent, and beats controlled imputation, distillation, and translation baselines. The claim is that restoring the fusion interface is a simple, portable principle for robust multimodal sensing.","feed_headline":"Fill missing sensor slots with proxies, keep fusion fixed","feed_subtitle":"One N-slot interface beats skip, impute, and distill under real sensor dropouts","key_machinery":"Interface-complete fusion via directed source-to-target proxy generators: each missing slot is filled by averaging single-layer Transformer estimates from every observed source into one proxy token, so fusion always receives a fixed N-slot layout.","core_discovery":"Missing modalities create a fusion-interface mismatch as well as information loss. Completing every canonical modality slot—with a real token when the sensor is present and with an aggregated source-to-target proxy token when it is not—lets one lightweight fusion head operate unchanged under arbitrary missingness and yields stronger HAR robustness than branch-skipping, feature imputation, distillation, or translation-style proxies.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Proxy tokens fill missing slots so fusion stays fixed","Complete every sensor slot for one unchanging fusion head","Aggregated proxies restore canonical slots under dropouts","Slot completion beats skip impute and distill for sensing","Target-slot proxies keep multimodal fusion interface fixed"],"cache_read_input_tokens":128,"weakest_assumption_plain":"A single global token per modality, completed by simple averaging of directed maps and fused by plain summation, is enough to restore the interactions the fusion head needs.","fun_headline_variants_meta":{"raw":{"variants":["Proxy tokens fill missing slots so fusion stays fixed","Complete every sensor slot for one unchanging fusion head","Aggregated proxies restore canonical slots under dropouts","Slot completion beats skip impute and distill for sensing","Target-slot proxies keep multimodal fusion interface fixed"]},"model":"grok-4.5","effort":"low","cost_usd":0.004204,"raw_usage":{"total_tokens":1262,"prompt_tokens":746,"num_sources_used":0,"completion_tokens":56,"cost_in_usd_ticks":42040000,"prompt_tokens_details":{"text_tokens":746,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":460,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":746,"tokens_out":56,"duration_ms":3766,"temperature":1.0,"reasoning_tokens":460,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T14:01:16.541832+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a HAR benchmark where a dominant modality already carries almost all the signal, or on a dense spatial task such as pose estimation, a cross-attention or sequence-level fusion baseline would outperform the completed-slot sum fusion under the same missingness patterns.","supporting_citations":[],"review_version":1}