{"id":"9efc5b9f-46ea-4b84-9523-8adccd7f60fd","arxiv_id":"2506.22374","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A sheaf-Laplacian regularized decentralized multimodal federated learning algorithm with local attention is proposed and claimed to converge to a stationary point, with reported gains on two wireless tasks.","lead":"This paper proposes Sheaf-DMFL and Sheaf-DMFL-Att, decentralized learning frameworks that let wireless devices with different sensor types (LiDAR, camera, RF, GPS) train shared feature extractors while modeling client relationships with sheaf theory. The authors provide a convergence proof and test on blockage prediction and mmWave beamforming, reporting accuracy gains over federated baselines.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's O(1/R) guarantee does not follow from the proof: the P-gradient term in Eq. (42) is evaluated at (θ̃^{r+1}, P^r), while the claimed bound in Eq. (44) is for (θ̃^r, P^r), and no bound on the mismatch is given.","rationale":"The paper's main advertised contribution is Theorem 1, so the soundness of its proof is the load-bearing issue. Reading the appendix in full, the proof contains index/argument mismatches that prevent the conclusion from following. The reader's rationale already flags the Eq. (42) issue and the partial omission in Lemma 2; however, the reader's listed weakest assumption is Assumption 2 (connectivity of modality subgraphs). In my reading, connectivity is a standard condition and is satisfied in the experiments; the more decisive problem is internal to the proof, so I only partially agree with the reader's framing. I am not raising a consensus disagreement: even granting all four assumptions, the theorem is not established by the provided argument. The concrete test is an independent re-derivation of the telescoping sum with consistent indices, which would settle whether the missing Lipschitz/consensus estimates can be supplied. If they can, the paper rises to 'conditional'; if not, rejection stands. Since the reader's rejection remains correct, the verdict is unchanged.","tokens_in":20667,"tokens_out":7491,"duration_ms":75495,"concrete_test":"Re-derive the proof of Theorem 1 from Eq. (28)–(44) with exact iterates, without replacing ∇_P Ψ(θ̃^{r+1},P^r) by ∇_P Ψ(θ̃^r,P^r). The decisive check is analytic: from the stated assumptions alone, attempt to bound ∑_r [∥∇_P Ψ(θ̃^{r+1},P^r)∥² − ∥∇_P Ψ(θ̃^r,P^r)∥²] by the negative descent terms in (35). If no such bound exists (e.g., because ∇_P Ψ is not Lipschitz in θ under Assumptions 1–4), then Theorem 1 is unproven. For corroboration, run the exact Algorithm 1 on a two-client, one-edge quadratic instance and test whether (1/R)∑∥∇Ψ(θ^r,P^r)∥² ≤ (Ψ⁰−Ψ*)/(ρR) holds; a violation would confirm the gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's advertised formal guarantee is Theorem 1, the O(1/R) convergence of ∥∇Ψ(θ̃^r,P^r)∥². The proof in Appendix B does not establish this. In Eq. (42), the P-update contributes a negative term proportional to ∥∇_P Ψ(θ̃^{r+1},P^r)∥²_F, with the task-specific parameter at r+1. The telescoped sum (43) keeps this index, but the final statement (44) defines the averaged gradient norm using ∥∇_P Ψ(θ̃^r,P^r)∥²_F. No Lipschitz bound on ∇_P Ψ with respect to θ is stated, so the mismatch cannot be absorbed. A second, equally serious index mismatch occurs in Eq. (33): the ω-descent is written as −α∇_ω Ψ(θ̃^r,P^r), but the actual update (11)/(29) uses local gradients ∇_{ω_i} f_i(θ_i^r) computed before encoder averaging; identifying these two requires a consensus-error bound that the proof never provides. Lemma 2's proof is explicitly omitted ('detailed derivation is omitted due to space limitations') and its claimed descent inequality requires gradients of f at θ̃^r while the encoder update (15)–(16) uses local gradients at θ^r followed by gossip averaging. These are not mere presentation issues: they are missing steps in the central theorem.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes Sheaf-DMFL and Sheaf-DMFL-Att, decentralized multimodal federated learning algorithms in which clients with heterogeneous modality sets share per-modality feature encoders via gossip aggregation and couple their task-specific heads through a learnable cellular-sheaf Laplacian regularizer. The central theoretical claim (Theorem 1) is an O(1/R) bound on the average squared gradient norm of the global objective Ψ under smoothness and connectedness assumptions. The experimental section evaluates the methods on link blockage prediction and mmWave beamforming benchmarks, reporting accuracy improvements over DSGD, DMML-KD, and local training baselines.","tokens_in":21040,"tokens_out":9688,"duration_ms":85931,"significance":"If the convergence guarantee were established, the paper would make a useful contribution by extending sheaf-based multi-task learning to decentralized multimodal settings with partial modality overlap. The problem is well motivated and the empirical comparisons on real-world datasets are a strength. However, the advertised theoretical guarantee, which is the paper's main novelty beyond the earlier conference version [22], is not proven: the proof of Theorem 1 contains an invalid identification of gradient terms, an unproved and generally false matrix inequality, an index mismatch in the final telescoping argument, and the central Lemma 2 is asserted with an omitted proof. These are load-bearing defects, not presentation issues. The empirical evaluation also lacks multi-seed statistics for the main accuracy curves and the code is not released, which limits the confidence one can place in the reported gains.","major_comments":[{"comment":"The equality ⟨∇ωΨ(θ̃^r, P^r), ω^{r+1}−ω^r⟩ = −α‖∇ωΨ(θ̃^r, P^r)‖² is not justified by the update rule (29), which uses ∇f(ω^r) evaluated at the un-averaged local parameters. Identifying ∇f(ω^r) with ∇ωΨ(θ̃^r, P^r) requires a consensus-error bound between θ^r and θ̃^r that is never stated or proved. Consequently the descent inequality (35) does not follow.","section":"Appendix B, Eq. (33)"},{"comment":"The proof of Lemma 2 is explicitly truncated ('The detailed derivation is omitted due to space limitations'). This lemma is the foundation of the theorem, since it gives the descent inequality (20) for the modified parameter vector. The sketch in (27)–(28) does not control the difference between the local gradients ∇_{ϕ_i,k} f_i(θ_i^r) used in the updates (15)–(16) and the gradients ∇_{ϕ̄_k} f(θ̃^r) that appear in the bound, so the lemma is unsubstantiated.","section":"Appendix A, Lemma 2"},{"comment":"Inequality (37) asserts that zeroing out entries of the matrix M = P^r − ηλP^rω^{r+1}(ω^{r+1})^T through the Hadamard product with H yields a Gram matrix dominated in the PSD order by M^T M. This is false in general: entrywise masking can increase the largest eigenvalue of the Gram matrix. Because this inequality is used to obtain (41), the claimed descent for the restriction-map update is not established.","section":"Appendix B, Eq. (37)"},{"comment":"The telescoping argument exhibits an index mismatch: the negative P-gradient term in (42) and (43) is evaluated at (θ̃^{r+1}, P^r), while the final averaged gradient norm in (44) is defined with ‖∇PΨ(θ̃^r, P^r)‖²_F. No Lipschitz bound on ∇PΨ with respect to θ is stated, so the two terms cannot be equated or bounded by one another. Thus Theorem 1's O(1/R) claim does not follow from the proof.","section":"Appendix B, Eqs. (42)–(44)"}],"minor_comments":[{"comment":"The baseline is introduced as 'DMML-KL [27]' but Figures 5, 7, 9 and Table I label it 'DMML-KD'; the naming should be unified.","section":"Section V"},{"comment":"Figures 5 and 7 show single-run accuracy curves without error bars or a statement about the number of seeds; Table II reports means and standard deviations, so the same reporting should be used for the main results.","section":"Section V-A and V-B"},{"comment":"The hyperparameters α, ηφ, ηβ, η, λ, and the attention parameters are not listed in the experimental settings; without these values the experiments are difficult to reproduce.","section":"Section V-A and V-B"},{"comment":"In Eq. (42) the coefficient of the P-gradient descent term contains ηλN D_ω²/2, whereas the definition of ρ in Eq. (44) uses ηλD_ω²/2; this inconsistency should be resolved.","section":"Appendix B, Eqs. (42) and (44)"},{"comment":"The statement after Eq. (8) defines restriction maps P_{ij} with dimensions R^{d_{ij}×d_i}, but the compression factor γ in Section III-B gives dimensions ⌊γ·(d_i+d_j)/2⌋×d_i; the relation between the two definitions should be clarified.","section":"Section III-B"},{"comment":"Assumption 2 requires each modality subgraph to be connected; the graphs in Figures 4 and 6 satisfy this by construction, so the experiments do not probe the behavior of the algorithm when this assumption fails. The paper should at least comment on this limitation.","section":"Assumption 2 and Section V"}],"recommendation":"reject","confidential_remarks":"The theoretical proof follows the same template as [20] and [25], both with overlapping authorship, and the sheaf machinery is imported from [20]. This is not itself a reason to reject, but it raises a novelty question: the contribution beyond [22] is largely the convergence analysis, and that analysis is currently incomplete. If the authors can supply a correct proof (including the missing Lemma 2 argument and a valid treatment of the P-update), the paper could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, the core theoretical claim, Theorem 1, is not established by the proof as written. Second, the empirical evaluation on two real 6G datasets is genuine but lacks the reporting standard that would make the performance claims fully credible.\n\nWhat is new: the paper applies the sheaf-Laplacian regularizer from Ben Issaid et al. to task-specific heads in decentralized multimodal federated learning, adds a partially shared encoder architecture and a local attention fusion variant. That combination is not in the cited prior work. The attention variant is a sensible extension, and the experiments on DeepSense blockage prediction and drone beam prediction show consistent gains over DSGD, a distillation baseline, and local training, with an ablation of projection compression factor and initialization. The problem is well motivated and the paper is clearly written.\n\nThe soft spots are real and load-bearing. In Eq. (42) the P-gradient term is evaluated at (θ̃^{r+1}, P^r), but the final bound (44) claims a norm at (θ̃^r, P^r). There is no Lipschitz bound on ∇_P Ψ with respect to θ, so that index shift cannot be absorbed. Equation (33) identifies the global gradient ∇_ω Ψ(θ̃^r, P^r) with the actual local update (11), which uses gradients at θ^r before encoder averaging; that requires a consensus-error bound that the proof never provides. Lemma 2 says 'the detailed derivation is omitted due to space limitations,' which is exactly the step that would need to be checked. These are not presentation issues; they are missing steps in the central advertised contribution.\n\nThe experimental reporting also falls short: accuracy curves are single-seed with no error bars (aside from Table II), no code is provided, and the baseline name switches between DMML-KL and DMML-KD. Assumption 2, connectivity of each modality subgraph, is untested because the experimental graphs connect modality groups by construction.\n\nThat said, this is not a paper to ignore. The problem setup is legitimate, the combination is new, and the real-data evaluation is a serious attempt. My recommendation: send it to a serious referee, but instruct them to focus on the convergence proof. As written, Theorem 1 does not go through. If the authors can repair the proof or honestly weaken the claim, the paper is salvageable.","headline":"The paper's advertised O(1/R) convergence guarantee does not follow from its own proof, and the empirical work, while genuine, is under-reported.","tokens_in":21554,"tokens_out":4234,"would_cite":false,"duration_ms":41448,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that decentralized multimodal federated learning converges to a stationary point at O(1/R) when task relationships are modeled by a learnable sheaf structure, and shows the resulting algorithms outperform decentralized…","keywords":["decentralized learning","multimodal federated learning","sheaf theory","sheaf Laplacian","attention-based fusion","convergence analysis","link blockage prediction","mmWave beamforming"],"falsifier":"Run Sheaf-DMFL-Att on a network where one modality is split across two disconnected clusters, with all other assumptions satisfied, and measure the average squared gradient norm of $\\Psi$ over $R$ rounds: if the bound still decays at $O(1/R)$, the connectivity assumption is not necessary; if the encoders drift apart or the bound fails, the assumption is doing the work.","tokens_in":20440,"feed_emoji":"📡","tokens_out":9091,"duration_ms":86694,"temperature":0.7,"pith_summary":"This paper tries to establish that devices with different sensor modalities can train a shared model together over a peer-to-peer network, without a central server, and that the training is provably convergent. It places a sheaf structure—vector spaces on communication links with learnable projections between clients—on top of each client's task-specific layer, so clients with different modality combinations learn how much to align with each other. The main formal claim is Theorem 1: under smoothness, boundedness, and connectivity assumptions, the attention-based variant Sheaf-DMFL-Att reaches a stationary point of a global objective at rate $O(1/R)$. If true, this gives decentralized multimodal federated learning the same style of convergence guarantee already available for centralized and single-modality decentralized methods, which matters in wireless systems where a parameter server is a single point of failure.","feed_headline":"Decentralized multimodal learning gets a provable convergence rate","feed_subtitle":"A sheaf structure lets devices with different sensors train together, and the proof shows O(1/R) convergence.","key_machinery":"The load-bearing object is a cellular sheaf placed on the communication graph: each client's task-specific head $\\omega_i$ is a stalk over a node, each edge carries a lower-dimensional comparison space, and learnable restriction maps $P_{ij}$ project $\\omega_i$ and $P_{ji}\\omega_j$ into that space. The sheaf Laplacian regularizer $\\frac{\\lambda}{2}\\sum_{(i,j)\\in E}\\|P_{ij}\\omega_i - P_{ji}\\omega_j\\|^2$ penalizes disagreement between neighboring tasks after projection, so the system learns not only the models but also how tasks should be compared. Around this sits the partially shared architecture: modality encoders are averaged across clients using Metropolis-Hastings mixing matrices $W_k$, and attention weights $\\alpha_{i,k}$ fuse modalities locally. The convergence proof's key device is the modified parameter vector $\\tilde{\\theta}_i^r$ in which shared encoders are replaced by their network averages, which decouples the encoder consensus dynamics (Lemma 1) from the head and attention updates (Lemma 2) and lets the whole system telescope into the $O(1/R)$ bound.","core_discovery":"On its own terms, the paper's central discovery is that multimodal heterogeneity in a decentralized network can be modeled as multi-task learning with a learnable task-relationship structure, and that the resulting algorithm carries a worst-case convergence guarantee. Sheaf-DMFL-Att trains shared modality encoders by gossip averaging over modality-specific subgraphs, fuses the encoders' outputs through per-client attention weights, and then aligns the task-specific heads through a sheaf Laplacian regularizer. The proof tracks a modified global parameter vector in which each client uses the average encoder for each modality rather than its local encoder; Lemma 1 shows these averages move like gradient descent because the mixing matrices are doubly stochastic, and Lemma 2 turns L-smoothness into a one-step descent inequality. Theorem 1 then yields $\\frac{1}{R}\\sum_{r=0}^{R-1}\\|\\nabla\\Psi(\\tilde{\\theta}^r,P^r)\\|^2 \\le \\frac{\\Psi(\\tilde{\\theta}^0,P^0)-\\Psi^*}{\\rho R}$, i.e. the average squared gradient norm of the global objective shrinks at the standard $O(1/R)$ non-convex rate.","pith_inferences":["The convergence proof is modular: the sheaf regularizer only touches the task-specific heads, so the $O(1/R)$ bound likely survives if the attention layer is replaced by any local fusion mechanism that keeps the loss L-smooth and the attention parameters bounded.","Because Assumption 2 requires each modality's client subgraph to be connected, deployments where a sensor type appears in isolated clusters should either split that modality into separate tasks or add communication links before applying this algorithm.","The theorem's bound is stated for the averaged-encoder objective $\\Psi(\\tilde{\\theta},P)$, not for the actual local encoders each client deploys; measuring the gap between the averaged model's stationary point and per-client personalized performance would be a natural next step.","The learned attention weights $\\alpha_{i,k}$ could be read as online estimates of modality reliability: if one sensor is occluded or noisy, the weights should shift toward the reliable modality, and tracking that shift could turn into a sensor-quality monitoring tool."],"forward_implications":["Under the stated assumptions, Sheaf-DMFL-Att converges to a stationary point of the global objective at rate $O(1/R)$, matching the standard non-convex decentralized SGD rate.","No central parameter server is required; training runs over peer-to-peer links, with each client exchanging only projected head parameters and encoder aggregates rather than raw data.","The attention mechanism gives a concrete fix for negative transfer: in the beam prediction task, GPS-only clients improve when multimodal neighbors use attention-weighted fusion, because the gradients they receive are no longer dominated by a single strong modality.","The experiments report faster convergence and higher accuracy than DSGD, DMML-KD, and local-only training on blockage prediction and mmWave beamforming, approaching the centralized upper bound.","The design is tunable: larger projection dimension $\\gamma$ and structured initialization of the restriction maps improve accuracy at the cost of memory, so the framework admits a memory-accuracy trade-off."],"supporting_citations":[{"why":"Supplies the sheaf-theoretic decentralized multi-task learning machinery: the sheaf Laplacian, learnable restriction maps, and the update rules this paper adapts to multimodal clients.","marker":"[20]"},{"why":"Supplies the federated multi-task learning formulation and the task-similarity regularizer that problem (P1) restates as a sheaf regularizer.","marker":"[23]"},{"why":"Supplies the convergence-analysis template, including the modified parameter vector with averaged shared parameters and the descent lemma that Theorem 1 follows.","marker":"[25]"},{"why":"Supplies the partially shared architecture idea that local models split into globally shared initial layers and client-specific heads.","marker":"[24]"},{"why":"Supplies the link blockage prediction dataset and problem setup used in Scenario I.","marker":"[2]"},{"why":"Supplies the DeepSense Drone dataset and the camera-plus-GPS beam prediction problem used in Scenario II.","marker":"[6]"},{"why":"Supplies the DMML-KD decentralized multimodal baseline against which the proposed algorithms are compared.","marker":"[27]"},{"why":"Supplies the cellular sheaf background, including stalks and restriction maps, that the key machinery relies on.","marker":"[21]"}],"fun_headline_variants":["Sheaf theory boosts decentralized multimodal learning","Decentralized multimodal learning gets a provable rate","Sheaf-DMFL: multimodal edge learning with convergence proof","Multimodal devices learn together via sheaf structure","Provable convergence for sheaf-based decentralized learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is Assumption 2: for every modality, the clients that possess it must form a connected subgraph of the communication network, because the convergence proof needs a strictly positive spectral gap for each modality-specific mixing matrix.","fun_headline_variants_meta":{"raw":{"variants":["Sheaf theory boosts decentralized multimodal learning","Decentralized multimodal learning gets a provable rate","Sheaf-DMFL: multimodal edge learning with convergence proof","Multimodal devices learn together via sheaf structure","Provable convergence for sheaf-based decentralized learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000143,"raw_usage":{"total_tokens":1212,"prompt_tokens":1024,"completion_tokens":188,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":128}},"tokens_in":640,"tokens_out":188,"duration_ms":2717,"temperature":1.0,"reasoning_tokens":128,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:05:55.422173+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Sheaf-DMFL-Att on a network where one modality is split across two disconnected clusters, with all other assumptions satisfied, and measure the average squared gradient norm of $\\Psi$ over $R$ rounds: if the bound still decays at $O(1/R)$, the connectivity assumption is not necessary; if the encoders drift apart or the bound fails, the assumption is doing the work.","supporting_citations":[{"cited_title":"Tackling feature and sample heterogeneity in decentralized multi-task learning: A sheaf- theoretic approach,","cited_arxiv_id":null,"evidence_quote":"Supplies the sheaf-theoretic decentralized multi-task learning machinery: the sheaf Laplacian, learnable restriction maps, and the update rules this paper adapts to multimodal clients."},{"cited_title":"Federated multi-task learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the federated multi-task learning formulation and the task-similarity regularizer that problem (P1) restates as a sheaf regularizer."},{"cited_title":"Distributed learning over networks with graph-attention-based personalization,","cited_arxiv_id":null,"evidence_quote":"Supplies the convergence-analysis template, including the modified parameter vector with averaged shared parameters and the descent lemma that Theorem 1 follows."},{"cited_title":"Exploiting shared representations for personalized federated learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the partially shared architecture idea that local models split into globally shared initial layers and client-specific heads."},{"cited_title":"Proactively predicting dynamic 6G link blockages using LiDAR and in-band signatures,","cited_arxiv_id":null,"evidence_quote":"Supplies the link blockage prediction dataset and problem setup used in Scenario I."},{"cited_title":"Towards real-world 6g drone communication: Position and camera aided beam prediction,","cited_arxiv_id":null,"evidence_quote":"Supplies the DeepSense Drone dataset and the camera-plus-GPS beam prediction problem used in Scenario II."},{"cited_title":"Knowledge distillation and training bal- ance for heterogeneous decentralized multi-modal learning over wireless networks,","cited_arxiv_id":null,"evidence_quote":"Supplies the DMML-KD decentralized multimodal baseline against which the proposed algorithms are compared."},{"cited_title":"Robinson, Topological signal processing","cited_arxiv_id":null,"evidence_quote":"Supplies the cellular sheaf background, including stalks and restriction maps, that the key machinery relies on."}],"review_version":1}