REVIEW 4 major objections 3 minor 1 cited by
GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks
T0 review · 4 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that a graph diffusion model trained only on suboptimal heuristic samples can generate near-optimal solutions to the multi-server multi-user computation offloading problem with high probability.
desk verdict Open-sourced graph diffusion solver for MSCO shows a real engineering gain, but its headline 'converges to optimal' claim rests on undisclosed large-scale test labels and an unverified unbiasedness assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on a solution-space parametrization: each feasible solution $y$ is assigned probability $p_\theta(y|x) \propto \exp(\sum_i y_i \theta_i)$, and the optimal solution corresponds to a parameter vector $\theta^*$. The model learns the distribution $p(\Theta_x)$ of suboptimal parameterizations, then repeatedly samples solutions $y \sim p_{\theta'}(y|x)$; a Chebyshev bound plus a finite-sampling hitting argument shows that enough samples make the probability of hitting the optimum near one. The generative engine is a graph neural network run as a two-head diffusion model—one head uses categorical (Bernoulli) transitions for binary offloading decisions, the other uses Gaussian transitions for resource-allocation ratios. Two engineering additions carry the experiments: a padding-edge gating flag that suppresses nonexistent edges in non-fully-connected graphs, and the observation that the two diffusion tasks' gradients are nearly orthogonal, so joint training does not suffer negative interference.
What would settle it
On a fixed instance, measure the average of the model's output parameter vectors over many parallel samples; if that average is far from the optimal solution's parameterization (for example, when the training heuristic always favors one server), the Exceed ratio should stay well above 1, contradicting the claim that sampling converges to optimal. The paper's own variance measurement (σ² ≈ 10⁻⁴) is only on outputs, not on the bias E(Θ_x) − θ*, so measuring that bias directly would settle the question.
Extended reading notes
Core claim
GDSG is a multi-task graph diffusion model that takes a network graph as input and generates both discrete offloading decisions and continuous resource-allocation ratios. The paper's central claim is that training this model on a dataset of suboptimal solutions—produced cheaply by a randomized minimum-cost-maximum-flow heuristic—is enough for it to learn the distribution of high-quality solutions, so that parallel sampling from the learned distribution converges to the optimal solution with high probability. On ten test scales, the generated solutions' costs sit within about 10% of the ground-truth optimal costs (Exceed ratios mostly below 1.1), which beats the heuristic generator, a same-architecture discriminative GNN, and prior diffusion-based combinatorial solvers. The reported savings reach up to 56.62% of the target cost when trained on an optimal dataset and 41.06% when trained on a suboptimal dataset, compared with discriminative baselines.
Load-bearing premise
The suboptimal solutions' parameterizations must be centered on the optimal parameterization, so that random mistakes cancel out rather than pushing the learned distribution toward a systematically wrong solution.
Editorial extensions
If this is right
- A model trained purely on suboptimal MCMF-generated data achieves Exceed ratios below 1.1 on most tested scales, so near-optimal decisions can be obtained without ground-truth labels.
- GDSG outperforms the heuristic that generated its training data, showing the model generalizes beyond its data generator.
- GDSG beats a same-architecture discriminative GNN on both suboptimal and ground-truth training datasets, with substantial cost savings.
- The discrete and continuous diffusion tasks share parameters with near-100% gradient orthogonality, so multi-task training avoids negative interference; the paper attributes this to the diffusion loss rather than the GNN architecture.
- Padding-edge gating lets one trained model generalize across graph sizes, substantially improving the Exceed ratio over a vanilla GNN backbone.
Reading between the lines
- If the heuristic is systematically biased, the paper's convergence proof does not guarantee optimality, but the diffusion sampler might still concentrate probability mass near good solutions; measuring the gap between the sample-mean parameterization and the optimal one on a biased instance would test this.
- The same pipeline—cheap suboptimal samples, a parametrized solution space, and a multi-task diffusion model—could transfer to other graph-structured MEC problems such as routing, scheduling, or spectrum allocation, where ground-truth labels are equally scarce.
- The near-total gradient orthogonality between the discrete and continuous heads suggests that diffusion objectives can sidestep standard multi-task interference, which would make joint discrete-continuous optimization easier for other network problems.
- Because the heuristic generates an 80-node instance in about 0.6 seconds while GDSG inference is roughly 0.03–0.4 seconds, the practical bottleneck shifts to dataset coverage; randomizing the heuristic's initialization is a cheap way to broaden the distribution the model learns.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GDSG, a graph-diffusion-based generative model for the multi-server multi-user computation offloading (MSCO) problem in MEC networks. The central claim is that GDSG can be trained on suboptimal datasets, produced cheaply by an MCMF heuristic, and still generate solutions that converge to the optimal solution with high probability. The authors reformulate the optimization as solution-distribution learning, provide a theoretical argument based on Chebyshev's inequality that repeated sampling from a learned distribution near the optimal parametrization can hit the optimum, and introduce a multi-task GNN with a padding-edge gating mechanism. Experiments on synthetic datasets report Exceed ratios below 1.1 on most test scales, with cost savings up to 56.62% over a discriminative GNN on ground-truth training data and 41.06% on suboptimal training data. The paper also claims near-100% task orthogonality between discrete and continuous diffusion tasks. The code and datasets are open-sourced.
Significance. If the central claims hold, the work would be a practically valuable contribution: it offers a way to train an optimization solver from cheap suboptimal data, avoiding the need for expensive ground-truth labels, and it provides an open-source dataset and implementation. The empirical comparison against a same-architecture discriminative GNN is a strength, and the padding-edge gating mechanism appears to improve generalization on non-fully-connected graphs. However, the theoretical guarantee rests on an unverified unbiasedness assumption E(Theta_x)=theta*, and the evaluation on large test scales relies on labels whose provenance is not described. The orthogonality claim, as written, is internally inconsistent because orthogonality corresponds to cosine 0, not 1. These issues leave the main contributions currently under-supported.
major comments (4)
- [Section V-A, Table I] The test sets gt7s24u, gt7s27u, gt10s31u, gt10s36u, gt20s61u, and gt20s68u are named 'gt' (ground-truth), but Section V-A only describes exhaustive generation for small-scale ground-truth datasets and explicitly states that 'an optimal solver is not available.' The method used to produce the labels for the larger test scales is not disclosed. If those labels are generated by the same MCMF heuristic, possibly with more restarts, then the Exceed ratios in Table I measure agreement with the heuristic's choices rather than closeness to the true optimum, and the central claim that GDSG 'converges to the optimal solution' is not supported by the data. Please specify the exact label-generation procedure for every test scale, or rename these sets and adjust the claims accordingly.
- [Section III-B, Eq. (9)] The convergence argument assumes E(Theta_x)=mu=theta*, i.e., that the mean of the suboptimal parameterizations equals the optimal parameterization. This assumption is not derived from the MCMF heuristic and is not tested in Section V-B; the empirical measurement of the model output variance (on the order of 1e-4) addresses spread but not bias. If the heuristic is systematically biased toward certain allocation patterns, Chebyshev's inequality with respect to theta* fails, and the repeated-sampling argument for hitting y* collapses. Please verify the mean condition on the actual generated datasets, or replace the guarantee with a bias-aware bound involving |E(Theta_x)-theta*|.
- [Section V-C, Fig. 4] Orthogonal gradient vectors have cosine 0, yet the text and the figure caption state that 'cosine values approaching 1' and 'orthogonal gradient proportions approaching 100%' indicate good orthogonality, and DiGNN is criticized for having cosine values 'far away from 1.' As written, the reported numbers are internally contradictory: a cosine near 1 means the gradients are nearly parallel, not orthogonal. Please clarify the metric actually computed (e.g., 1 - |cos|, or the angle itself) and re-state the claim; otherwise, the task-orthogonality contribution is not interpretable.
- [Section III-B, Eqs. (11)-(13)] The derivation of the lower bound on the number of samples n is not rigorous. The quantity p_epsilon is a probability determined by the distribution, not a free parameter, and replacing Chebyshev's '>=' with '>' on the grounds that the distribution is not three-point does not provide a uniform margin, so the passage from Eq. (11) to Eq. (12) is not justified. Moreover, Eq. (12) depends on the unknown p_epsilon, making it not a computable lower bound. The paper should either provide a clean derivation (e.g., using a bias-aware or union-bound argument) or present this as a heuristic scaling argument rather than a proof.
minor comments (3)
- [Abstract] The phrase 'converging to the optimal solution large probably' in the abstract should read 'with high probability' to match the full text.
- [Section V-A] The sentence 'we use exhaustive methods to generate small-scale ground-truth datasets but the time complexity is very high (over 30sec for a single instance of 80 nodes)' is confusing because 80 nodes is not a small scale relative to the listed small-scale sets (e.g., gt3s6u has 9 nodes). Please clarify which scales were solved exhaustively and what '80 nodes' refers to.
- [Section III-A] The parametrization of the discrete solution space with theta*_i0 and theta*_i1 is inconsistent with the earlier definition of theta in R^N; please unify the notation.
Circularity Check
No demonstrated circularity: the convergence proof is explicitly conditional on an unverified unbiasedness assumption, and the large-scale ground-truth label generation is undescribed, but no prediction reduces to a fitted input or to a self-citation by the paper's own equations.
full rationale
The paper's central theoretical claim in Sec. III-B is conditional rather than circular: it states 'we assume that the perturbations caused by suboptimal solutions are smoothly random and the mean E(Theta_x) = mu = theta* holds' and then applies Chebyshev's inequality (Eq. 9). This is an unverified assumption about the heuristic data distribution, not a derivation of that distribution from the model; the proof would fail if the assumption is false, but it does not redefine the prediction as its input. The empirical claim of near-optimal performance is measured against externally named 'gt' test sets; small-scale ground truth is said to come from exhaustive search ('we use exhaustive methods to generate small-scale ground-truth datasets'), while the paper also states 'an optimal solver is not available' and does not describe how the gt7s24u through gt20s68u labels were produced. That is a provenance gap that weakens support for the convergence claim, but the paper does not state that those labels are generated by the same heuristic, so a circular reduction cannot be exhibited from the text. Self-citations ([25], [38], [46]) and the DiffSG reference appear in related-work and baseline roles and are not load-bearing for the derivation. Accordingly, no circular step meets the quotation-and-reduction bar; the correct finding is no significant circularity, with the unbiasedness assumption and large-scale label provenance flagged as correctness risks.
Assumptions & free parameters
free parameters (4)
- epsilon (neighborhood radius) =
0.1
- gamma (Gaussian sampling std) =
0.04
- delta (tolerable gap) =
0.1
- a (probability slack) =
a -> 0
assumptions (5)
- domain assumption The mean of the suboptimal parametrization equals the optimal parametrization: E(Theta_x)=mu=theta*.
- ad hoc to paper p(Theta_x) is not a three-point distribution, so Chebyshev's inequality becomes strict.
- domain assumption The components of theta are independent across dimensions for the hitting probability calculation.
- domain assumption Ground-truth labels for all ten test scales are exact optima.
- domain assumption The trained GNN can represent the conditional solution distribution p_Phi(theta|G) well enough for the denoising objective.
Cite this review
Pith. "Pith review of GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks." pith.science (2026). https://pith.science/paper/DPHSG553
@misc{pith2026241208296,
author = {Pith},
title = {Pith review of: GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/DPHSG553}},
note = {Machine review of arXiv:2412.08296}
}
read the original abstract
Optimization is crucial for MEC networks to function efficiently and reliably, most of which are NP-hard and lack efficient approximation algorithms. This leads to a paucity of optimal solution, constraining the effectiveness of conventional deep learning approaches. Most existing learning-based methods necessitate extensive optimal data and fail to exploit the potential benefits of suboptimal data that can be obtained with greater efficiency and effectiveness. Taking the multi-server multi-user computation offloading (MSCO) problem, which is widely observed in systems like Internet-of-Vehicles (IoV) and Unmanned Aerial Vehicle (UAV) networks, as a concrete scenario, we present a Graph Diffusion-based Solution Generation (GDSG) method. This approach is designed to work with suboptimal datasets while converging to the optimal solution large probably. We transform the optimization issue into distribution-learning and offer a clear explanation of learning from suboptimal training datasets. We build GDSG as a multi-task diffusion model utilizing a Graph Neural Network (GNN) to acquire the distribution of high-quality solutions. We use a simple and efficient heuristic approach to obtain a sufficient amount of training data composed entirely of suboptimal solutions. In our implementation, we enhance the backbone GNN and achieve improved generalization. GDSG also reaches nearly 100\% task orthogonality, ensuring no interference between the discrete and continuous generation tasks. We further reveal that this orthogonality arises from the diffusion-related training loss, rather than the neural network architecture itself. The experiments demonstrate that GDSG surpasses other benchmark methods on both the optimal and suboptimal training datasets. The MSCO datasets has open-sourced at this http URL, as well as the GDSG algorithm codes at https://github.com/qiyu3816/GDSG.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences
Guiding a pretrained topology-diffusion generator with human-preference reward classifiers is claimed to suppress floating-material and boundary-violation failure modes without retraining the generator.
Reference graph
Works this paper leans on
-
[1]
A Survey on Mobile Edge Computing: The Communication Perspective,
Y . Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief, “A Survey on Mobile Edge Computing: The Communication Perspective,” IEEE Communications Surveys & Tutorials , vol. 19, no. 4, pp. 2322–2358, 2017
work page 2017
-
[2]
Mobile Edge Computing: A Survey on Archi- tecture and Computation Offloading,
P. Mach and Z. Becvar, “Mobile Edge Computing: A Survey on Archi- tecture and Computation Offloading,” IEEE Communications Surveys & Tutorials, vol. 19, no. 3, pp. 1628–1656, 2017
work page 2017
-
[3]
A2-UA V: Application-Aware Content and Network Optimization of Edge-Assisted UA V Systems,
A. Coletta, F. Giorgi, G. Maselli, M. Prata, D. Silvestri, J. Ashdown, and F. Restuccia, “A2-UA V: Application-Aware Content and Network Optimization of Edge-Assisted UA V Systems,” in Proceedings of IEEE International Conference on Computer Communications , 2023, pp. 1– 10
work page 2023
-
[4]
Energy-Efficient Trajectory Optimization for Aerial Video Surveillance under QoS Constraints,
C. Zhan, H. Hu, S. Mao, and J. Wang, “Energy-Efficient Trajectory Optimization for Aerial Video Surveillance under QoS Constraints,” in Proceedings of IEEE International Conference on Computer Communi- cations, 2022, pp. 1559–1568
work page 2022
-
[5]
S. Chai and V . K. N. Lau, “Multi-UA V Trajectory and Power Opti- mization for Cached UA V Wireless Networks With Energy and Content Recharging-Demand Driven Deep Learning Approach,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 10, pp. 3208–3224, 2021
work page 2021
-
[6]
C. Zhao, J. Liu, M. Sheng, W. Teng, Y . Zheng, and J. Li, “Multi- UA V Trajectory Planning for Energy-Efficient Content Coverage: A Decentralized Learning-Based Approach,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 10, pp. 3193–3207, 2021
work page 2021
-
[7]
B. Yang, X. Cao, C. Yuen, and L. Qian, “Offloading Optimization in Edge Computing for Deep-Learning-Enabled Target Tracking by Internet of UA Vs,”IEEE Internet of Things Journal , vol. 8, no. 12, pp. 9878– 9893, 2021
work page 2021
-
[8]
A. Masaracchia, K. K. Nguyen, T. Q. Duong, and V . Sharma, Deep Reinforcement Learning for Reconfigurable Intelligent Surfaces and UAV Empowered Smart 6G Communications . The Institution of Engineering and Technology, 2024. [Online]. Available: https: //digital-library.theiet.org/doi/abs/10.1049/PBTE106E
Show all 74 references
-
[9]
Generative AI-Augmented Graph Reinforcement Learning for Adaptive UA V Swarm Optimization,
B. Hazarika, P. Singh, K. Singh, S. L. Cotton, H. Shin, O. A. Dobre, and T. Q. Duong, “Generative AI-Augmented Graph Reinforcement Learning for Adaptive UA V Swarm Optimization,” IEEE Internet of Things Journal, pp. 1–1, 2025
2025
-
[10]
Latency Optimization for Blockchain-Empowered Federated Learning in Multi-Server Edge Computing,
D. C. Nguyen, S. Hosseinalipour, D. J. Love, P. N. Pathirana, and C. G. Brinton, “Latency Optimization for Blockchain-Empowered Federated Learning in Multi-Server Edge Computing,” IEEE Journal on Selected Areas in Communications , vol. 40, no. 12, pp. 3373–3390, 2022
2022
-
[11]
Federated Edge Network Utility Maximization for a Multi-Server System: Algorithm and Convergence,
N. Karakoc ¸, A. Scaglione, M. Reisslein, and R. Wu, “Federated Edge Network Utility Maximization for a Multi-Server System: Algorithm and Convergence,” IEEE/ACM Transactions on Networking , vol. 30, no. 5, pp. 2002–2017, 2022
2002
-
[12]
Federated Spectrum Learning for Reconfigurable Intelligent Surfaces-Aided Wireless Edge Networks,
B. Yang, X. Cao, C. Huang, C. Yuen, M. Di Renzo, Y . L. Guan, D. Niyato, L. Qian, and M. Debbah, “Federated Spectrum Learning for Reconfigurable Intelligent Surfaces-Aided Wireless Edge Networks,” IEEE Transactions on Wireless Communications , vol. 21, no. 11, pp. 9610–9626, 2022
2022
-
[13]
Reconfigurable Intelligent Surface-Assisted Aerial-Terrestrial Communications via Multi-Task Learning,
X. Cao, B. Yang, C. Huang, C. Yuen, M. D. Renzo, D. Niyato, and Z. Han, “Reconfigurable Intelligent Surface-Assisted Aerial-Terrestrial Communications via Multi-Task Learning,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 10, pp. 3035–3050, 2021
2021
-
[14]
Hybrid Beamforming for Reconfigurable Intelligent Surface based Multi-User Communications: Achievable Rates With Limited Discrete Phase Shifts,
B. Di, H. Zhang, L. Song, Y . Li, Z. Han, and H. V . Poor, “Hybrid Beamforming for Reconfigurable Intelligent Surface based Multi-User Communications: Achievable Rates With Limited Discrete Phase Shifts,” IEEE Journal on Selected Areas in Communications , vol. 38, no. 8, pp. 1...
2020
-
[15]
Resource Allocation for Power Minimization in RIS-Assisted Multi- UA V Networks With NOMA,
W. Feng, J. Tang, Q. Wu, Y . Fu, X. Zhang, D. K. C. So, and K.-K. Wong, “Resource Allocation for Power Minimization in RIS-Assisted Multi- UA V Networks With NOMA,”IEEE Transactions on Communications , vol. 71, no. 11, pp. 6662–6676, 2023
2023
-
[16]
Joint Base Station and IRS Deployment for En- hancing Network Coverage: A Graph-Based Modeling and Optimization Approach,
W. Mei and R. Zhang, “Joint Base Station and IRS Deployment for En- hancing Network Coverage: A Graph-Based Modeling and Optimization Approach,” IEEE Transactions on Wireless Communications , vol. 22, no. 11, pp. 8200–8213, 2023
2023
-
[17]
Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks: Design Optimization and Analysis,
X. Zhang, B. Yang, Z. Yu, X. Cao, G. C. Alexandropoulos, Y . Zhang, M. Debbah, and C. Yuen, “Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks: Design Optimization and Analysis,” IEEE Transactions on Intelligent Trans- portation Sys...
2024
-
[18]
Computation Offloading in MEC-Enabled IoV Networks: Average Energy Efficiency Analysis and Learning-Based Maximization,
T. Z. H. Ernest and A. S. Madhukumar, “Computation Offloading in MEC-Enabled IoV Networks: Average Energy Efficiency Analysis and Learning-Based Maximization,” IEEE Transactions on Mobile Comput- ing, vol. 23, no. 5, pp. 6074–6087, 2024
2024
-
[19]
Asynchronous Deep Reinforcement Learning for Data-Driven Task Offloading in MEC- Empowered Vehicular Networks,
P. Dai, K. Hu, X. Wu, H. Xing, and Z. Yu, “Asynchronous Deep Reinforcement Learning for Data-Driven Task Offloading in MEC- Empowered Vehicular Networks,” in Proceedings of IEEE International Conference on Computer Communications , 2021, pp. 1–10
2021
-
[20]
Edge Intelligence for Autonomous Driving in 6G Wireless System: Design Challenges and Solutions,
B. Yang, X. Cao, K. Xiong, C. Yuen, Y . L. Guan, S. Leng, L. Qian, and Z. Han, “Edge Intelligence for Autonomous Driving in 6G Wireless System: Design Challenges and Solutions,” IEEE Wireless Communica- tions, vol. 28, no. 2, pp. 40–47, 2021
2021
-
[21]
DeepScheduler: Enabling Flow-Aware Scheduling in Time-Sensitive Networking,
X. He, X. Zhuge, F. Dang, W. Xu, and Z. Yang, “DeepScheduler: Enabling Flow-Aware Scheduling in Time-Sensitive Networking,” in Proceedings of IEEE International Conference on Computer Commu- nications, 2023, pp. 1–10
2023
-
[22]
RouteNet: Leveraging graph neural networks for network modeling and optimization in SDN,
K. Rusek, J. Su ´arez-Varela, P. Almasan, P. Barlet-Ros, and A. Cabellos- Aparicio, “RouteNet: Leveraging graph neural networks for network modeling and optimization in SDN,” IEEE Journal on Selected Areas in Communications, vol. 38, no. 10, pp. 2260–2270, 2020
2020
-
[23]
A Joint Energy and Latency Framework for Transfer Learning Over 5G Industrial Edge Networks,
B. Yang, O. Fagbohungbe, X. Cao, C. Yuen, L. Qian, D. Niyato, and Y . Zhang, “A Joint Energy and Latency Framework for Transfer Learning Over 5G Industrial Edge Networks,” IEEE Transactions on Industrial Informatics, vol. 18, no. 1, pp. 531–541, 2022
2022
-
[24]
Deep- learning-based joint resource scheduling algorithms for hybrid MEC networks,
F. Jiang, K. Wang, L. Dong, C. Pan, W. Xu, and K. Yang, “Deep- learning-based joint resource scheduling algorithms for hybrid MEC networks,” IEEE Internet of Things Journal , vol. 7, no. 7, pp. 6252– 6265, 2019
2019
-
[25]
A multi-head ensemble multi-task learning approach for dynamical com- putation offloading,
R. Liang, B. Yang, Z. Yu, X. Cao, D. W. K. Ng, and C. Yuen, “A multi-head ensemble multi-task learning approach for dynamical com- putation offloading,” in Proceedings of IEEE Global Communications Conference, 2023, pp. 6079–6084
2023
-
[26]
GNN-Based Power Allocation and User Association in Digital Twin Network for the Terahertz Band,
H. Zhang, X. Ma, X. Liu, L. Li, and K. Sun, “GNN-Based Power Allocation and User Association in Digital Twin Network for the Terahertz Band,” IEEE Journal on Selected Areas in Communications , 2023
2023
-
[27]
Edge-Assisted Multi-Layer Offloading Optimization of LEO Satellite-Terrestrial Integrated Networks,
X. Cao, B. Yang, Y . Shen, C. Yuen, Y . Zhang, Z. Han, H. V . Poor, and L. Hanzo, “Edge-Assisted Multi-Layer Offloading Optimization of LEO Satellite-Terrestrial Integrated Networks,” IEEE Journal on Selected Areas in Communications , vol. 41, no. 2, pp. 381–398, 2023
2023
-
[28]
STaR: self-taught reasoner bootstrapping reasoning with reasoning,
E. Zelikman, Y . Wu, J. Mu, and N. D. Goodman, “STaR: self-taught reasoner bootstrapping reasoning with reasoning,” in Proceedings of Neural Information Processing Systems , 2024
2024
-
[29]
LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery,
P. Ma, T.-H. Wang, M. Guo, Z. Sun, J. B. Tenenbaum, D. Rus, C. Gan, and W. Matusik, “LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery,” in Proceedings of International Conference on Machine Learning , vol. 235, 21–27 Jul 2024, p...
2024
-
[30]
Dream the impossible: outlier imag- ination with diffusion models,
X. Du, Y . Sun, X. Zhu, and Y . Li, “Dream the impossible: outlier imag- ination with diffusion models,” in Proceedings of Neural Information Processing Systems, 2024
2024
-
[31]
Adding Conditional Control to Text-to-Image Diffusion Models,
L. Zhang, A. Rao, and M. Agrawala, “Adding Conditional Control to Text-to-Image Diffusion Models,” in Proceedings of International Conference on Computer Vision , 2023, pp. 3813–3824
2023
-
[32]
Gurobi Optimizer Reference Manual,
Gurobi Optimization, LLC, “Gurobi Optimizer Reference Manual,”
-
[33]
ApS, The MOSEK optimization toolbox for MATLAB manual
M. ApS, The MOSEK optimization toolbox for MATLAB manual. Version 10.1., 2024. [Online]. Available: http://docs.mosek.com/latest/ toolbox/index.html
2024
-
[34]
IBM ILOG CPLEX Optimization Studio,
IBM, “IBM ILOG CPLEX Optimization Studio,” 2024. [Online]. Available: https://www.ibm.com/docs/en/icos/22.1.1
2024
-
[35]
GEKKO Optimization Suite,
L. Beal, D. Hill, R. Martin, and J. Hedengren, “GEKKO Optimization Suite,” Processes, vol. 6, no. 8, p. 106, 2018
2018
-
[36]
A Survey on Generative Diffusion Models,
H. Cao, C. Tan, Z. Gao, Y . Xu, G. Chen, P.-A. Heng, and S. Z. Li, “A Survey on Generative Diffusion Models,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 7, pp. 2814–2830, 2024
2024
-
[37]
A GNN- based supervised learning framework for resource allocation in wireless IoT networks,
T. Chen, X. Zhang, M. You, G. Zheng, and S. Lambotharan, “A GNN- based supervised learning framework for resource allocation in wireless IoT networks,” IEEE Internet of Things Journal, vol. 9, no. 3, pp. 1712– 1724, 2021
2021
-
[38]
Computation offloading in multi-access edge computing: A multi-task learning approach,
B. Yang, X. Cao, J. Bassey, X. Li, and L. Qian, “Computation offloading in multi-access edge computing: A multi-task learning approach,” IEEE Transactions on Mobile Computing, vol. 20, no. 9, pp. 2745–2762, 2020
2020
-
[39]
Denoising Diffusion Probabilistic Models,
J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” in Proceedings of Neural Information Processing Systems , vol. 33, 2020, pp. 6840–6851
2020
-
[40]
Classifier-free diffusion guidance,
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598, 2022
2022 arXiv
-
[41]
Diffusion Models Beat GANs on Image Synthesis,
P. Dhariwal and A. Nichol, “Diffusion Models Beat GANs on Image Synthesis,” in Proceedings of Neural Information Processing Systems , vol. 34, 2021, pp. 8780–8794
2021
-
[42]
Structured Denoising Diffusion Models in Discrete State-Spaces,
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. van den Berg, “Structured Denoising Diffusion Models in Discrete State-Spaces,” in Proceedings of Neural Information Processing Systems , vol. 34, 2021, pp. 17 981–17 993
2021
-
[43]
DIFUSCO: Graph-based Diffusion Solvers for Combinatorial Optimization,
Z. Sun and Y . Yang, “DIFUSCO: Graph-based Diffusion Solvers for Combinatorial Optimization,” in Proceedings of Neural Information Processing Systems, vol. 36, 2023, pp. 3706–3731
2023
-
[44]
T2T: From Distribution Learning in Training to Gradient Search in Testing for Combinatorial Optimization,
Y . Li, J. Guo, R. Wang, and J. Yan, “T2T: From Distribution Learning in Training to Gradient Search in Testing for Combinatorial Optimization,” in Proceedings of Neural Information Processing Systems, vol. 36, 2023, pp. 50 020–50 040
2023
-
[45]
Enhancing Deep Reinforcement Learning: A Tutorial on Generative Diffusion Models in Network Optimization,
H. Du, R. Zhang, Y . Liu, J. Wang, Y . Lin, Z. Li, D. Niyato, J. Kang, Z. Xiong, S. Cui, B. Ai, H. Zhou, and D. I. Kim, “Enhancing Deep Reinforcement Learning: A Tutorial on Generative Diffusion Models in Network Optimization,” IEEE Communications Surveys & Tutorials, pp. 1–1, 2024
2024
-
[46]
DiffSG: A Generative Solver for Network Optimization with Diffusion Model,
R. Liang, B. Yang, Z. Yu, B. Guo, X. Cao, M. Debbah, H. V . Poor, and Y . Chau, “DiffSG: A Generative Solver for Network Optimization with Diffusion Model,” 2024. [Online]. Available: https://arxiv.org/abs/2408.06701
2024 arXiv
-
[47]
Generative AI based Secure Wireless Sensing for ISAC Networks,
J. Wang, H. Du, Y . Liu, G. Sun, D. Niyato, S. Mao, D. I. Kim, and X. Shen, “Generative AI based Secure Wireless Sensing for ISAC Networks,” 2024. [Online]. Available: https://arxiv.org/abs/2408.11398
2024 arXiv
-
[48]
Generative AI for Deep Reinforcement Learning: Framework, Analysis, and Use Cases,
G. Sun, W. Xie, D. Niyato, F. Mei, J. Kang, H. Du, and S. Mao, “Generative AI for Deep Reinforcement Learning: Framework, Analysis, and Use Cases,” 2024. [Online]. Available: https://arxiv.org/ abs/2405.20568
2024 arXiv
-
[49]
Diffusion-Based Reinforcement Learning for Edge-Enabled AI-Generated Content Services,
H. Du, Z. Li, D. Niyato, J. Kang, Z. Xiong, H. Huang, and S. Mao, “Diffusion-Based Reinforcement Learning for Edge-Enabled AI-Generated Content Services,” IEEE Transactions on Mobile Com- puting, vol. 23, no. 9, pp. 8902–8918, 2024
2024
-
[50]
Deep Generative Model and Its Applications in Efficient Wireless Network Management: A Tutorial and Case Study,
Y . Liu, H. Du, D. Niyato, J. Kang, Z. Xiong, D. I. Kim, and A. Ja- malipour, “Deep Generative Model and Its Applications in Efficient Wireless Network Management: A Tutorial and Case Study,” IEEE Wireless Communications, vol. 31, no. 4, pp. 199–207, 2024
2024
-
[51]
ADMM for Mobile Edge Intelligence: A Survey,
A. He, H. Pan, Y . Dai, X. Si, C. Yuen, and Y . Zhang, “ADMM for Mobile Edge Intelligence: A Survey,” IEEE Communications Surveys & Tutorials, pp. 1–1, 2024
2024
-
[52]
Online Distributed ADMM Algorithm With RLS-Based Multitask Graph Filter Models,
Y . Lai, F. Chen, M. Feng, and J. Kurths, “Online Distributed ADMM Algorithm With RLS-Based Multitask Graph Filter Models,” IEEE Transactions on Network Science and Engineering , vol. 9, no. 6, pp. 4115–4128, 2022
2022
-
[53]
Hybridized MA- DRL for Serving xURLLC With Cognizable RIS and UA V Integration,
A. Paul, R. Allu, K. Singh, C.-P. Li, and T. Q. Duong, “Hybridized MA- DRL for Serving xURLLC With Cognizable RIS and UA V Integration,” IEEE Transactions on Wireless Communications , vol. 23, no. 10, pp. 15 507–15 524, 2024
2024
-
[54]
A survey on uplink resource allocation in OFDMA wireless networks,
E. Yaacoub and Z. Dawy, “A survey on uplink resource allocation in OFDMA wireless networks,” IEEE Communications Surveys & Tutorials, vol. 14, no. 2, pp. 322–337, 2011
2011
-
[55]
An overview of radio resource management in relay-enhanced OFDMA-based networks,
M. Salem, A. Adinoyi, M. Rahman, H. Yanikomeroglu, D. Falconer, Y .-D. Kim, E. Kim, and Y .-C. Cheong, “An overview of radio resource management in relay-enhanced OFDMA-based networks,” IEEE Com- munications Surveys & Tutorials , vol. 12, no. 3, pp. 422–438, 2010
2010
-
[56]
Decentralized computation offloading game for mobile cloud computing,
X. Chen, “Decentralized computation offloading game for mobile cloud computing,” IEEE Transactions on Parallel and Distributed Systems , vol. 26, no. 4, pp. 974–983, 2014
2014
-
[57]
Processor design for portable systems,
T. D. Burd and R. W. Brodersen, “Processor design for portable systems,” Journal of VLSI signal processing systems for signal, image and video technology , vol. 13, no. 2, pp. 203–221, 1996
1996
-
[58]
DIMES: A Differentiable Meta Solver for Combinatorial Optimization Problems,
R. Qiu, Z. Sun, and Y . Yang, “DIMES: A Differentiable Meta Solver for Combinatorial Optimization Problems,” in Proceedings of Neural Information Processing Systems , vol. 35, 2022, pp. 25 531–25 546
2022
-
[59]
Chebyshev inequality with estimated mean and variance,
J. G. Saw, M. C. Yang, and T. C. Mo, “Chebyshev inequality with estimated mean and variance,” The American Statistician, vol. 38, no. 2, pp. 130–132, 1984
1984
-
[60]
Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,
D. Kingma and R. Gao, “Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,” in Proceedings of Neural Information Processing Systems , vol. 36, 2023, pp. 65 484–65 516
2023
-
[61]
Variational Diffusion Models,
D. Kingma, T. Salimans, B. Poole, and J. Ho, “Variational Diffusion Models,” in Proceedings of Neural Information Processing Systems , vol. 34, 2021, pp. 21 696–21 707
2021
-
[62]
Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions,
E. Hoogeboom, D. Nielsen, P. Jaini, P. Forr ´e, and M. Welling, “Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions,” in Proceedings of Neural Information Processing Systems, vol. 34, 2021, pp. 12 454–12 465
2021
-
[63]
Analog bits: Generating discrete data using diffusion models with self-conditioning,
T. Chen, R. Zhang, and G. Hinton, “Analog bits: Generating discrete data using diffusion models with self-conditioning,” arXiv preprint arXiv:2208.04202, 2022
2022 arXiv
-
[64]
Learning the travelling salesperson problem requires rethinking generalization,
C. K. Joshi, Q. Cappart, L.-M. Rousseau, and T. Laurent, “Learning the travelling salesperson problem requires rethinking generalization,” Constraints, vol. 27, no. 1, pp. 70–98, 2022
2022
-
[65]
Benchmarking graph neural networks,
V . P. Dwivedi, C. K. Joshi, A. T. Luu, T. Laurent, Y . Bengio, and X. Bresson, “Benchmarking graph neural networks,” MIT Press Journal of Machine Learning Research , vol. 24, no. 43, pp. 1–48, 2023
2023
-
[66]
A systematic survey on deep generative models for graph generation,
X. Guo and L. Zhao, “A systematic survey on deep generative models for graph generation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 5, pp. 5370–5390, 2022
2022
-
[67]
Towards Impartial Multi-task Learning,
L. Liu, Y . Li, Z. Kuang, J.-H. Xue, Y . Chen, W. Yang, Q. Liao, and W. Zhang, “Towards Impartial Multi-task Learning,” in Proceedings of International Conference on Learning Representations , 2021
2021
-
[68]
Gradient Surgery for Multi-Task Learning,
T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn, “Gradient Surgery for Multi-Task Learning,” in Proceedings of Neural Information Processing Systems , vol. 33, 2020, pp. 5824–5836
2020
-
[69]
Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout,
Z. Chen, J. Ngiam, Y . Huang, T. Luong, H. Kretzschmar, Y . Chai, and D. Anguelov, “Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout,” in Proceedings of Neural Information Processing Systems, vol. 33, 2020, pp. 2039–2050
2020
-
[70]
Addressing Negative Transfer in Diffusion Models,
H. Go, , Y . Lee, S. Lee, S. Oh, H. Moon, and S. Choi, “Addressing Negative Transfer in Diffusion Models,” in Proceedings of Neural Information Processing Systems , vol. 36, 2023, pp. 27 199–27 222
2023
-
[71]
DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data,
H. Ye and D. Xu, “DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data,” in Proceedings of IEEE/CVF Computer Vision and Pattern Recognition Conference, 2024, pp. 27 960–27 969
2024
-
[72]
Diffusion Model is an Effective Planner and Data Synthesizer for Multi- Task Reinforcement Learning,
H. He, C. Bai, K. Xu, Z. Yang, W. Zhang, D. Wang, B. Zhao, and X. Li, “Diffusion Model is an Effective Planner and Data Synthesizer for Multi- Task Reinforcement Learning,” in Proceedings of Neural Information Processing Systems, vol. 36, 2023, pp. 64 896–64 917
2023
-
[73]
Denoising Diffusion Implicit Models,
J. Song, C. Meng, and S. Ermon, “Denoising Diffusion Implicit Models,” in Proceedings of International Conference on Learning Representa- tions, 2021
2021
-
[2024]
Available: https://www.gurobi.com
[Online]. Available: https://www.gurobi.com
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.