{"id":"e01e9ef5-9ef4-4fa5-ac1a-a1fd4ce7fc52","arxiv_id":"2502.01841","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A unified diffusion-model framework for generation and decision-making tasks in multi-antenna communications, with a beamforming case study showing gains over fully-connected and graph neural networks.","lead":"This paper proposes a unified framework for using diffusion models in multi-antenna wireless communications, splitting tasks into generation and decision-making, and supports it with a beamforming case study. The authors argue diffusion models are better than standard neural networks in noisy, interference-heavy settings because they model distributions and explore multiple good actions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No ablation isolates diffusion-specific exploration from stochastic or ensemble effects; Fig. 3's headline comparison may reflect training or sampling choices, not the claimed ability to escape local optima.","rationale":"I read the paper as a position/tutorial article whose quantitative support is a single synthetic beamforming case study. The central mechanistic claim is that diffusion models improve decision-making by sampling from a multimodal near-optimal action distribution, thereby escaping the equal-power averaging behavior of deterministic networks. This is plausible and consistent with diffusion-policy results in other domains, but the manuscript does not provide the controlled comparison needed to attribute Fig. 3's gains to this mechanism. The reader's weakest-assumption identification matches my own: reverse diffusion is assumed to produce samples from near-optimal modes rather than a blurred or averaged policy, and no experiment isolates exploration from distribution modeling, ensemble effects, or training/protocol differences. I therefore agree with the CONDITIONAL verdict; the concern is real but not disqualifying, because the core framework and taxonomy have independent value and the empirical claim is testable. My concrete test would settle whether the load-bearing mechanism actually operates as claimed.","tokens_in":8990,"tokens_out":3450,"duration_ms":36369,"concrete_test":"Re-run the Section IV-A beamforming experiment with controlled ablations: (1) train DM (FNN) and FNN with identical backbone capacity, optimizer, epochs, batch size, and training data, and set the number of reverse diffusion steps to a compute-matched value; (2) decode the DM deterministically by zeroing the added reverse noise at every step, or by using the conditional mean, and compare against stochastic sampling; (3) for one high-correlation channel (e.g., correlation 0.99), sample 100 actions from the trained DM and plot the power-allocation histogram against the four one-hot optima. The exploration claim is supported only if the deterministic DM still beats FNN and the histogram shows four well-separated modes; if all gain disappears when sampling noise is removed, or the DM samples cluster around the equal-power mean, the mechanism is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Sections II-C and IV-A) is that DMs outperform TAI in decision-making because reverse diffusion 'can jump out of local optimal solutions to explore better actions' and captures the multimodal p(action|state). The supporting evidence, Fig. 3, compares DM (GNN) vs GNN and DM (FNN) vs FNN, but each pair differs in several uncontrolled ways: diffusion training objective, noise schedule, inference-time stochastic sampling, and possibly sampling budget or network capacity. The manuscript never states how many samples are drawn per channel at inference, whether the best SE across samples is selected, or whether the same training epochs, data, optimizer, and loss are used for both architectures. If multiple stochastic samples are drawn and the best is chosen, the gain could be an ensemble or search effect that any stochastic policy would enjoy, rather than the claimed exploration mechanism. The narrative in Section IV-A -- four one-hot optima versus an equal-power average under a deterministic FNN -- is plausible but not connected to any measured action distribution. No histogram, diversity metric, or sample-level analysis of the DM's outputs is provided. Since the paper's practical conclusion rests on this mechanism, the missing isolation is load-bearing and currently unresolved.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a unified framework for applying diffusion models (DMs) to multiple-antenna communication tasks, dividing tasks into generation and decision-making categories. It argues that DMs outperform traditional AI and conventional optimization in strong-noise and strong-interference scenarios because they model complex conditional distributions and, in decision-making tasks, can explore better actions in non-convex problems. The paper describes supervised, model-based/model-free unsupervised, and reinforcement learning training methods; reviews generation tasks (channel estimation, signal detection) and decision-making tasks (beamforming, beam training, resource allocation); outlines applications to near-field, STAR-RIS, CAPA, and other emerging technologies; and presents a beamforming case study (Fig. 3) comparing DM(FNN) and DM(GNN) against FNN and GNN baselines.","tokens_in":9217,"tokens_out":6366,"duration_ms":57889,"significance":"The paper makes a useful organizing contribution: the generation/decision-making dichotomy is simple, and the mapping of channel estimation, signal detection, beamforming, beam training, and resource allocation onto that dichotomy is clear and pedagogically valuable. The case study uses a sensible metric (SE ratio relative to WMMSE) and includes both FNN and GNN backbones, with the GNN designed to exploit permutation symmetry, which is a good design choice. If the quantitative claims were properly supported, the paper would provide a valuable tutorial and a strong motivation for DM-based wireless policies. At present, however, the central mechanistic claim ('exploring better actions') is not isolated by the experiments, and the missing experimental details prevent reproducible verification. The paper's breadth is a strength, but the evidence bar for its central superiority claim should match the strength of the claim.","major_comments":[{"comment":"The headline comparison is not an isolated test of the paper's exploration claim. The DM and baseline pairs differ simultaneously in training objective (noise prediction versus direct mapping), inference-time stochastic sampling, noise schedule, and likely optimization hyperparameters and network capacity. The manuscript does not state how many reverse-diffusion samples are drawn per channel at inference, whether the best SE across samples is selected, or whether the same training epochs, learning rate, batch size, and data splits are used for all four architectures. If best-of-many samples are selected, the gain over a deterministic FNN/GNN could be an ensemble or search effect available to any stochastic policy. Please provide the full experimental protocol and an ablation that equalizes the sampling budget and isolates the diffusion-specific component (e.g., one sample per channel, or a stochastic MLP with the same number of forward passes).","section":"Section IV-A and Fig. 3"},{"comment":"The statement that 'By starting with noise and adding noise on each reverse step, the DM can jump out of local optimal solutions to explore better actions' is an assertion, not a demonstrated property. The reverse process is a sampler for the learned conditional distribution; without an additional argument or empirical measurement, there is no reason to expect samples to be better than the training targets or to escape local optima. The paper itself acknowledges in Section VI that theoretical understanding is limited. To support the mechanism, please provide sample-level evidence (e.g., histograms of actions, diversity metrics, or the distribution of output power allocations for the identical-user case) or a formal statement of the conditions under which reverse diffusion explores multiple modes of p(action|state).","section":"Section II-C and Section IV-A"},{"comment":"The quantitative demonstration lacks variability measures and scope. No error bars, confidence intervals, or multiple random trials are reported, and the only operating point is 8 BS antennas, 4 users, and 10 dB SNR. Since the abstract and Section II-C claim superiority 'especially in scenarios with strong interference and noise,' the simulation should include at least one sensitivity analysis (e.g., varying correlation, SNR, or number of users) with repeated trials, and the reported curves should indicate the spread across runs. This is necessary for the reader to assess whether the DM advantage is significant or within run-to-run variation.","section":"Section IV-A, Fig. 3"},{"comment":"The model-based unsupervised training method is too vague to reproduce. The sentence 'a better action is firstly updated with gradient descent' does not specify the initialization, the objective with respect to which the gradient is taken, or how the gradient is propagated through the DM's forward diffusion process. Since the beamforming case study uses this training paradigm, please give the exact loss function, the update rule for the action, and the constraint-handling procedure (e.g., normalization or the dual-problem formulation).","section":"Section II-A2"}],"minor_comments":[{"comment":"'variational encoder (V AE)' should be 'variational autoencoder (VAE)'.","section":"Section I"},{"comment":"Reference [8] lists 'NeuIPS'; the correct venue spelling is 'NeurIPS'.","section":"References"},{"comment":"Please use uppercase 'THz' instead of 'Thz' in 'Thz communications' and elsewhere.","section":"Section V-D"},{"comment":"The caption contains a non-standard full-width colon ('：') and a stray spacing; please format the caption consistently.","section":"Figure 1 caption"},{"comment":"Please define the channel correlation parameter precisely: the current text states that the channel vector is a weighted sum of a common and a per-user Rayleigh term, but does not specify how the 'weighting factor' maps to the reported correlation values (0.0, 0.3, ..., 0.99).","section":"Section IV-A"},{"comment":"It would help to include the WMMSE spectral efficiency as a horizontal reference line (100% in SE ratio) so the reader can see the absolute performance of all learned policies, not only their relative ordering.","section":"Fig. 3"}],"recommendation":"major_revision","confidential_remarks":"This is a magazine-style paper with a tutorial component and a new simulation. The main concern is the gap between the strength of the claims ('DMs are superior...') and the evidence (one uncontrolled simulation comparison). I recommend major revision rather than rejection because the framework and survey content are valuable and the claim could be fixed by either adding a proper ablation or softening the causal language. The editor may also want to check that the self-citations ([2], [5], [13]) are appropriate background citations; they are, in my reading, but the related-work coverage could be broadened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful organizing paper, but the load-bearing empirical claim is not actually demonstrated. The taxonomy — generation vs decision-making tasks — plus the unified training framework (supervised, model-based/model-free unsupervised, RL) is a solid contribution and will help people map the area. The beamforming case study is new: DM (GNN/FNN) against GNN/FNN with a channel-correlation sweep. That's worth some credit.\n\nNow the soft spot. The paper argues DMs 'explore better actions' and escape local optima because of reverse diffusion. The evidence is Fig. 3 — one synthetic scenario, no error bars, no training details (learning rate, epochs, diffusion steps, noise schedule, network widths), no code/data, and no information on how many samples are drawn at inference. If the DM is sampling multiple times and keeping the best SE, the gain could be an ensemble effect any stochastic policy would enjoy, not the diffusion mechanism. The stress-test note is right: there is no ablation isolating exploration from stochastic sampling, and no analysis of the action distribution to back the 'four peaks' story. The narrative is plausible, but it's not connected to what the model actually outputs.\n\nTo be fair: the qualitative argument is reasonable, and the paper is honest about its own limits — the final section explicitly says theoretical understanding is lacking. The citation pattern is unremarkable; the authors cite their own prior work but as background. This is a position/tutorial piece, so I'm not holding it to the bar of a systems paper. Still, if a referee asks for details and an ablation, the authors should be able to provide them or soften the claims.\n\nWho gets value: wireless researchers who want a bird's-eye view of where diffusion models could plug into beamforming, channel estimation, detection, and resource allocation. It's a good entry point, not a methods contribution. I'd send it to peer review for a magazine or tutorial track, but I'd want the case study either beefed up or reframed as an illustration rather than a demonstration. With that revision, I'd probably cite it.","headline":"A useful taxonomy and framework for diffusion models in multi-antenna systems, but the case study doesn't isolate the claimed exploration benefit.","tokens_in":9738,"tokens_out":2480,"would_cite":true,"duration_ms":23952,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that diffusion models outperform both traditional AI and conventional optimization for multiple-antenna tasks under strong noise and interference, and demonstrates this on a beamforming case study.","keywords":["diffusion model","multiple antenna communications","beamforming","generative AI","channel estimation","signal detection","non-convex optimization","resource allocation"],"falsifier":"Re-run the four-user beamforming case study at high channel correlation, such as 0.98, twice per architecture: once with normal reverse-process sampling and once with the injected noise removed at inference so the diffusion model outputs its mean action. If the SE ratio of the diffusion model drops to the level of the standalone FNN or GNN, the gain is caused by stochastic exploration; if it stays high, the gain comes from distribution modeling, which would contradict the paper's stated explanation.","tokens_in":8757,"feed_emoji":"📡","tokens_out":4156,"duration_ms":35268,"temperature":0.7,"pith_summary":"This paper argues that diffusion models—generative networks trained by adding and then removing noise—are the right tool for many multiple-antenna wireless tasks because those tasks are non-deterministic: the same observed state can have many good actions. It proposes a unified recipe: identify state and action, pick a training method (supervised, model-based or model-free unsupervised, or reinforcement learning), train the diffusion policy, and evaluate. Tasks are divided into generation tasks, such as channel estimation and signal detection, and decision-making tasks, such as beamforming, beam training, and resource allocation, with diffusion models claimed to help in both by modeling complex distributions, and additionally in decision tasks by escaping local optima through reverse-process noise. A beamforming case study with four users reports that diffusion-augmented networks beat their deterministic counterparts, especially when user channel vectors are strongly correlated. If the claim holds, a single generative architecture could handle estimation, detection, and optimization in interference-heavy 6G scenarios.","feed_headline":"Diffusion models beat standard AI for multi-antenna beamforming","feed_subtitle":"A beamforming case study shows diffusion-augmented networks win as channel correlation grows.","key_machinery":"The engine is the conditional diffusion model, which defines a forward process that adds Gaussian noise to an action and a reverse process in which a noise-prediction DNN gradually denoises, conditioned on the state, to sample from an approximation of $p(\\text{action} \\mid \\text{state})$. The trained object is therefore a stochastic policy rather than a deterministic mapping. The paper's task taxonomy is the other load-bearing piece: classifying tasks as generation or decision-making determines whether the value of the diffusion model comes from distribution modeling, from action exploration, or from both. In the beamforming case study, the noise-prediction DNN is instantiated as either a fully-connected network or a graph neural network.","core_discovery":"The paper's central claim is that a diffusion model trained to generate actions from states learns non-deterministic policies better than traditional AI, which outputs one action per state, and better than conventional optimization, which relies on simplified distributional assumptions. The mechanism is the reverse diffusion process: starting from noise and adding noise at each denoising step lets the model represent a full conditional distribution of good actions, and lets it jump out of local optima in non-convex problems such as multiuser beamforming under interference. The paper also claims that this is why diffusion models are well-suited to strong noise and interference, for both decision-making and generation tasks.","pith_inferences":["If the exploration explanation is right, inference-time behavior should show a diversity benefit: sampling multiple reverse trajectories for the same channel should produce multiple distinct high-performing beamforming matrices, whereas a deterministic network produces one; the paper does not report such a test.","The reported gain could be re-framed as a stochastic-sampling or ensemble effect rather than a fundamentally new optimization mechanism; a fair comparison would match the number of denoising steps against an ensemble of deterministic networks with the same backbone.","The same framework suggests a concrete recipe for other ill-posed inverse problems in wireless, such as channel prediction or DoA estimation with ambiguous phase, where the relevant object is a distribution of plausible reconstructions rather than a single estimate.","A testable extension is to measure whether exploration helps most at the correlation values where the action set is multi-modal, such as near-identical user channels, which would confirm or refute the paper's explanation."],"forward_implications":["Diffusion-based policies should outperform deterministic neural policies for beamforming, with the margin growing as user channels become more correlated.","In generation tasks, diffusion models should yield more accurate channel estimates and signal detection from limited or noisy pilots than traditional estimators.","In decision-making tasks, adding reverse-process noise is presented as an exploration mechanism that helps escape local optima in non-convex resource allocation and beamforming.","A unified training framework with four training methods covers supervised, model-based and model-free unsupervised, and reinforcement-learning settings for the same diffusion-model architecture.","Emerging multi-antenna technologies such as near-field communications, STAR-RIS, and continuous aperture arrays are identified as natural next targets, since they add non-determinism from phase ambiguity, channel non-stationarity, and compressed measurements."],"supporting_citations":[{"why":"Supplies the optimal multiuser beamforming solution structure that reduces the learned action to power allocation in the case study.","marker":"[11]"},{"why":"Defines the denoising diffusion probabilistic model training objective used for supervised diffusion-model training.","marker":"[8]"},{"why":"Provides the tutorial basis for using generative diffusion models in network optimization and deep reinforcement learning.","marker":"[2]"},{"why":"Supplies the dual-problem loss used for model-based unsupervised training of diffusion models.","marker":"[9]"},{"why":"Provides the learning-based MIMO signal detection approach that motivates diffusion-based denoising for beam training.","marker":"[6]"},{"why":"Describes EM property sensing via diffusion models, the basis for the CAPA sensing application.","marker":"[12]"},{"why":"Connects generative AI to integrated sensing and communication, supporting the DoA estimation application.","marker":"[7]"}],"fun_headline_variants":["Diffusion models beat standard AI for multi-antenna beamforming","Diffusion models escape local optima in beamforming","Diffusion models excel in noisy multi-antenna channels","Diffusion AI beats optimization in beamforming","Diffusion models handle interference better than classic AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the reverse diffusion process samples from the set of near-optimal actions rather than a blurred or averaged policy, and that the beamforming gains come from this exploration; no experiment isolates exploration from distribution modeling.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion models beat standard AI for multi-antenna beamforming","Diffusion models escape local optima in beamforming","Diffusion models excel in noisy multi-antenna channels","Diffusion AI beats optimization in beamforming","Diffusion models handle interference better than classic AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000846,"raw_usage":{"total_tokens":3614,"prompt_tokens":812,"completion_tokens":2802,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":2725}},"tokens_in":428,"tokens_out":2802,"duration_ms":17423,"temperature":1.0,"reasoning_tokens":2725,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T14:16:44.651759+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the four-user beamforming case study at high channel correlation, such as 0.98, twice per architecture: once with normal reverse-process sampling and once with the injected noise removed at inference so the diffusion model outputs its mean action. If the SE ratio of the diffusion model drops to the level of the standalone FNN or GNN, the gain is caused by stochastic exploration; if it stays high, the gain comes from distribution modeling, which would contradict the paper's stated explanation.","supporting_citations":[{"cited_title":"Optimal multiuser trans- mit beamforming: A difficult problem with a simple solution structure,","cited_arxiv_id":null,"evidence_quote":"Supplies the optimal multiuser beamforming solution structure that reduces the learned action to power allocation in the case study."},{"cited_title":"Denoising diffusion probabilistic models,","cited_arxiv_id":null,"evidence_quote":"Defines the denoising diffusion probabilistic model training objective used for supervised diffusion-model training."},{"cited_title":"Enhancing deep reinforcement learning: A tutorial on generative diffusion models in network optimization,","cited_arxiv_id":null,"evidence_quote":"Provides the tutorial basis for using generative diffusion models in network optimization and deep reinforcement learning."},{"cited_title":"Learning optimal resource allocations in wireless systems,","cited_arxiv_id":null,"evidence_quote":"Supplies the dual-problem loss used for model-based unsupervised training of diffusion models."},{"cited_title":"Learning-based signal detection for MIMO systems with unknown noise statistics,","cited_arxiv_id":null,"evidence_quote":"Provides the learning-based MIMO signal detection approach that motivates diffusion-based denoising for beam training."},{"cited_title":"Electromagnetic property sensing based on diffusion model in ISAC system,","cited_arxiv_id":null,"evidence_quote":"Describes EM property sensing via diffusion models, the basis for the CAPA sensing application."},{"cited_title":"Generative AI for integrated sensing and communication: Insights from the physical layer perspective,","cited_arxiv_id":null,"evidence_quote":"Connects generative AI to integrated sensing and communication, supporting the DoA estimation application."}],"review_version":1}