{"id":"778432a9-6a52-4af7-8eca-31473b84b0f6","arxiv_id":"2607.03907","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"MCR²-guided multi-modal feature extraction plus BCD resource allocation yields higher human-activity recognition accuracy than single-modal and heuristic baselines under the same delay/energy limits.","lead":"The paper builds a multi-modal ISCC system that extracts compact features on devices with the MCR² criterion and jointly allocates sensing, communication and computation resources at the edge to maximize activity-recognition accuracy under tight delay and energy budgets. It matters because single-modality edge sensing is brittle; the work shows how to keep multi-modal robustness when bandwidth and energy are scarce.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"The claim that BCD-maximized MCR^{2} (Eq. 12) yields higher accuracy rests on an empirically observed monotonic proxy that is validated only for the XRF55 HAR subset and two classifiers (Fig. 6).","rationale":"The reader correctly isolates the single softest link in the central claim: the unproven generality of the MCR^{2}–accuracy relationship that justifies every resource decision. The rest of the paper is internally consistent—Lemma 1 is the standard variational representation of –log det, the constraints of (21) are convex (Appendix A), Theorem 1 holds, and the reported accuracy curves on XRF55 are higher than the three stated baselines under identical budgets. No derivation error, circular definition, or experimental contradiction appears. Because the concern is already flagged and the empirical support on the tested task is solid, the CONDITIONAL verdict (pending broader validation and code) needs no adjustment.","tokens_in":23697,"tokens_out":583,"duration_ms":40645,"concrete_test":"Extract the final N matrices produced by Algorithm 3 for the six delay budgets of Fig. 9; evaluate both ΔR(N) and true test-set accuracy (SVM and MLP) on those six points plus 20 random feasible N of comparable communication cost. If Spearman rank correlation between ΔR and accuracy falls below 0.85, or if any baseline N yields higher accuracy than the BCD N of equal or lower cost, the proxy is not faithful at the claimed operating points.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The resource-allocation engine (Algorithms 2–3) never sees classifier accuracy; it maximizes the closed-form coding-rate reduction ΔR(N) constructed from clean-training covariances Σ, Σ_l (Eq. 12). The only evidence that larger ΔR implies higher SVM/MLP accuracy after quantization is the scatter in Fig. 6, generated by sweeping arbitrary distortion matrices on the same eight-class XRF55 split. Because the BCD solution produces a highly structured family of N (feature-wise, modality-coupled, delay-tight), it is possible that the operating points actually used in Figs. 9–13 lie in a region where the proxy ranking diverges from true accuracy ranking. If that occurs, the reported gains over equal-time / device-level / single-modality baselines would be an artifact of the surrogate rather than a genuine improvement in sensing performance. No analytic proof of monotonicity under the Gaussian-mixture model after diagonal quantization noise is supplied, nor is the correlation re-checked on the precise N returned by the optimizer.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a task-oriented multi-modal ISCC framework in which IoT devices extract compact features under the maximal coding rate reduction (MCR²) criterion and an edge server performs joint multi-modal inference. MCR² is also used as a differentiable sensing metric (Eq. 12) built from estimated multi-modal covariances, leading to a sensing-accuracy maximization problem under delay, energy, and successful-transmission constraints. After an equivalent reformulation via an auxiliary-matrix identity (Lemma 1), the problem is solved by a BCD algorithm with an alternating inner loop for quantization-bit and communication-time allocation (Algorithms 2–3). On an eight-class subset of the public XRF55 human-activity dataset, with SVM and MLP edge classifiers, the scheme is reported to outperform device-level quantization, equal-time allocation, single-modality, and a semantic JSCC baseline under tight resource budgets.","tokens_in":23996,"tokens_out":1380,"duration_ms":23496,"significance":"Multi-modal ISCC is a timely and under-explored direction for 6G edge intelligence; treating heterogeneous sensing modalities under shared delay/energy budgets is practically relevant. The dual use of MCR² for device-side feature learning and as a closed-form edge-side metric is a clear methodological contribution relative to cross-entropy extractors and black-box accuracy. The optimization path is standard but carefully executed: convexity of the reformulated constraints is argued (Appendix A), monotonic improvement of the AO loop is proven (Theorem 1), and complexity is stated. Empirical evaluation is comparatively thorough—public data, two classifiers, multiple resource sweeps (delay, bandwidth, energy, feature dimension, number of classes), a semantic-communication baseline, and a cross-modal correlation check (Table III). If the reported accuracy gains hold under broader conditions, the work is a solid systems-level contribution to task-oriented multi-modal ISCC.","major_comments":[{"comment":"§VI-B and Fig. 6 establish that sensing accuracy rises monotonically with ΔR when N is generated by sweeping arbitrary distortion levels. The resource allocator (Algorithms 2–3), however, never sees classifier accuracy; it maximizes the closed-form ΔR(N) of Eq. (12) and produces a structured family of feature-wise, modality-coupled N under tight delay. The main claims in Figs. 9–13 therefore rest on the assumption that the same ranking holds at these optimized operating points. Please re-evaluate the ΔR–accuracy scatter (SVM and MLP) specifically on the N matrices returned by Algorithm 3 under the resource settings of the main experiments, and report whether the ranking versus the three system-level baselines is preserved. Without this check, the gap between the surrogate used for optimization and the accuracy used for evaluation remains incompletely closed.","section":null},{"comment":"§III-D and Eq. (12): Σ and Σ_l are estimated once from clean, offline multi-modal training features and then held fixed during online allocation. The text asserts that off-diagonal blocks encode cross-modal correlation and that the optimizer therefore avoids redundant modalities, with supporting evidence only in the all-WiFi vs multi-modal comparison of Table III. Please clarify the sensitivity of the allocated (T_tran_k, N_blk_k) and of final accuracy to mismatch between these offline covariances and the online feature statistics (e.g., different environments, partial modality dropout, or distribution shift). A short sensitivity study or explicit limitation statement is needed, because the central resource-allocation claim depends on the fidelity of these fixed second-order statistics.","section":null},{"comment":"§VI-C.2, single-modality baseline: multi-modal devices use d_k = 20 (total D = 60), while the single-modality (WiFi-only) baseline uses feature dimension 24, described as “empirically determined as the optimal value.” Under the same total communication budget this is not an apples-to-apples comparison of information content versus resource use. Please either (i) report single-modality accuracy also at d = 20 and at d = 60 (matching total multi-modal dimension), or (ii) justify why 24 is the appropriate comparator and show that the multi-modal gain is not an artifact of unequal total feature dimension.","section":null}],"minor_comments":[{"comment":"Table I and several places in the text use “Sening power” / “sening”; correct to “Sensing”.","section":null},{"comment":"Fig. 3 rendering is corrupted (“Vo l (…”) and the geometric packing illustration is hard to read; please regenerate with clear labels for W, W′, Z1, Z2 and the white-ball interpretation of ΔR.","section":null},{"comment":"§III-A: edge computation delay is ignored “due to abundant resources.” A one-sentence bound or reference to the MLP/SVM inference cost on the edge server would make the delay model more complete.","section":null},{"comment":"Notation: N is used both for the full quantization-distortion matrix and (in places) in a way that can be confused with the Gaussian N(·); consider a distinct symbol for the distortion matrix.","section":null},{"comment":"§VI-C.1 JSCC baseline: “we adjust the model size … while keeping its original framework unchanged” is underspecified. State the resulting feature/bit budget and how it was matched to the 0.03 s and 0.09 s delay points.","section":null},{"comment":"Related work on multi-modal semantic / task-oriented communication is appropriate; a brief pointer to other MCR² / rate-reduction uses in communications (if any) would help position the metric choice.","section":null},{"comment":"Algorithm 3 complexity O(I1(D³ + I0 Σ d_k³ + …)) is given; stating typical (I0, I1) used in the experiments would aid reproducibility.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The central empirical claim is supported by true accuracy (not only the surrogate), so I do not view the proxy gap as grounds for rejection or major_revision; requesting the optimized-N re-check and the single-modality dimension control should be sufficient. Scope fits TWC / similar venues well. No integrity concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean engineering extension of single-modal ISCC to the multi-modal case. The actual novelty is not MCR^{2} itself (Yu et al. 2020) but the decision to use it both as the device-side extractor loss and as the closed-form edge-side objective (Eq. 12) that already folds in cross-modal covariance blocks, then solving the resulting joint quantization/time/power problem with a standard BCD + AO loop under delay and energy constraints.\n\nWhat they do well: the reformulation (Lemma 1 + auxiliary matrices) is textbook and correctly applied; Theorem 1 and the appendices check out; the XRF55 experiments (8 classes, SVM + MLP, sweeps over delay, bandwidth, energy, feature dim, number of classes) are thorough and consistently favor the proposal over equal-time, device-level quantization, and single-modality baselines, especially when resources are tight. The cross-modal correlation experiment in Table III is a nice sanity check that the optimizer is not just chasing channel quality. Fig. 6 shows the MCR^{2}–accuracy correlation is monotonic on the tested points, so the circularity burden is low.\n\nSoft spots, in proportion: the stress-test concern is real but not fatal. The optimizer never sees classifier accuracy; it maximizes ΔR(N) built from clean-training covariances. Fig. 6 only sweeps arbitrary N, not the structured family returned by Algorithms 2–3. No analytic monotonicity proof under the Gaussian-mixture + diagonal quantization model is given, and everything is still one HAR task on one public subset. Free parameters (d_k, ε, δ_k^{2}, sensing powers) are fixed by hand. No code. These are ordinary limitations for a systems paper, not load-bearing cracks.\n\nWho it is for: people already working on ISCC / task-oriented edge AI who need a concrete multi-modal scheduler that respects cross-modal redundancy. Not a foundational result, but a usable advance. I would send it to peer review; the math is sound, the experiments are honest, and the contribution is clear enough for referees to evaluate. Worth a look if you are in the subfield; I would cite the formulation if I am writing a multi-modal ISCC paper this year.","headline":"Solid multi-modal ISCC systems paper that turns MCR^{2} into a tractable resource-allocation objective; gains look real under tight budgets, with the usual surrogate-metric caveat.","tokens_in":24608,"tokens_out":566,"would_cite":true,"duration_ms":5515,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Multi-modal ISCC that uses maximal coding rate reduction for both feature extraction and sensing evaluation beats single-modality and equal-resource baselines under tight delay and energy limits.","keywords":["integrated sensing-communication-computation","multi-modal sensing","maximal coding rate reduction","task-oriented communications","edge intelligence","human activity recognition","resource allocation","block coordinate descent"],"falsifier":"On the same XRF55 activity subset, generate a family of quantization-noise matrices that cover the operating range used by the optimizer, plot true SVM/MLP accuracy against the MCR^{2} value of each matrix, and check whether the monotonic relationship claimed in Figure 6 fails for any realistic distortion level.","tokens_in":24585,"feed_emoji":"📡","tokens_out":700,"duration_ms":5955,"temperature":0.7,"pith_summary":"Existing edge-intelligence designs that fuse sensing, communication, and computation usually rely on one sensing modality. That choice leaves the system brittle to occlusions, noise, and sensor failures. This paper argues that multi-modal sensing can restore robustness, but only if the extra data volume and the correlations among modalities are handled explicitly. The authors therefore place a feature extractor at each device that is trained with the maximal coding rate reduction (MCR^{2}) criterion: features of different activity classes are driven into large, well-separated subspaces while same-class features stay compact. The same MCR^{2} quantity, evaluated on the noisy concatenated features that arrive at the edge, is then used as a differentiable proxy for recognition accuracy. Under delay and energy budgets the resulting non-convex resource-allocation problem is rewritten into an equivalent form that a block-coordinate algorithm can solve efficiently. On a public multi-modal human-activity data set the scheme consistently records higher recognition accuracy than equal-time allocation, uniform quantization, and single-modality baselines.","feed_headline":"Multi-modal edge sensing beats single-sensor baselines under tight budgets","feed_subtitle":"MCR^{2}-driven features plus joint bit-and-time allocation raise activity-recognition accuracy when delay and energy are scarce.","key_machinery":"Maximal coding rate reduction (MCR^{2}): the difference between the coding rate of the whole feature matrix and the weighted sum of the coding rates of its class-conditional sub-matrices. It both trains the device-side extractors and, after substitution of the estimated multi-modal covariances plus quantization noise, becomes the objective that the BCD resource allocator maximizes.","core_discovery":"When each device extracts features with the MCR^{2} objective and the edge server treats the same MCR^{2} value computed on the recovered multi-modal features as the sensing metric, jointly optimizing quantization bits, transmit power, and TDMA slots under common delay and energy constraints yields higher human-activity recognition accuracy than device-level quantization, equal-time allocation, or any single-modality scheme that uses the same total resources.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["MCR² multi-modal ISCC lifts HAR accuracy under tight delay-energy budgets","Task-oriented multi-modal edge ISCC beats single-sensor baselines","MCR² features plus joint bit-time allocation raise multi-modal sensing accuracy","Multi-modal ISCC with MCR² metric outperforms equal-time and single-modality schemes","Device MCR² extractors enable robust edge activity recognition under resource limits"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The coding-rate-reduction number computed from estimated multi-modal covariances stays a faithful, monotonic stand-in for actual classifier accuracy once quantization and channel noise are present.","fun_headline_variants_meta":{"raw":{"variants":["MCR² multi-modal ISCC lifts HAR accuracy under tight delay-energy budgets","Task-oriented multi-modal edge ISCC beats single-sensor baselines","MCR² features plus joint bit-time allocation raise multi-modal sensing accuracy","Multi-modal ISCC with MCR² metric outperforms equal-time and single-modality schemes","Device MCR² extractors enable robust edge activity recognition under resource limits"]},"model":"grok-4.5","effort":"low","cost_usd":0.00365,"raw_usage":{"total_tokens":1239,"prompt_tokens":850,"num_sources_used":0,"completion_tokens":89,"cost_in_usd_ticks":36500000,"prompt_tokens_details":{"text_tokens":850,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":300,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":850,"tokens_out":89,"duration_ms":18051,"temperature":1.0,"reasoning_tokens":300,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T23:06:14.740121+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same XRF55 activity subset, generate a family of quantization-noise matrices that cover the operating range used by the optimizer, plot true SVM/MLP accuracy against the MCR^{2} value of each matrix, and check whether the monotonic relationship claimed in Figure 6 fails for any realistic distortion level.","supporting_citations":[],"review_version":1}