Pith. sign in

REVIEW 4 major objections 3 minor 1 cited by

GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks

T0 review · 4 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that a graph diffusion model trained only on suboptimal heuristic samples can generate near-optimal solutions to the multi-server multi-user computation offloading problem with high probability.

desk verdict Open-sourced graph diffusion solver for MSCO shows a real engineering gain, but its headline 'converges to optimal' claim rests on undisclosed large-scale test labels and an unverified unbiasedness assumption. read the letter →

arxiv 2412.08296 v3 pith:DPHSG553 submitted 2024-12-11 cs.NI cs.LG

classification cs.NIcs.LG
keywords graphdiffusioncomputationoffloadingmulti-accessedgecomputingsuboptimaltrainingdatagenerativemodelneuralnetworkmulti-tasklearningsolutiondistribution
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Many optimization problems in mobile edge computing are NP-hard and have no efficient approximation algorithm, so labeled optimal solutions are scarce. The paper tries to show that this scarcity need not block learning: a graph diffusion model trained only on cheap suboptimal solutions can still converge to the optimal solution with high probability. The key move is to treat optimization as distribution learning—parameterize the solution space, learn the distribution of suboptimal parameterizations, and sample repeatedly. If correct, GDSG would let network operators train near-optimal solvers from data that is easy to generate, and the paper reports that it beats both the heuristic that produced its training data and a same-architecture discriminative GNN.

What carries the argument

The argument rests on a solution-space parametrization: each feasible solution $y$ is assigned probability $p_\theta(y|x) \propto \exp(\sum_i y_i \theta_i)$, and the optimal solution corresponds to a parameter vector $\theta^*$. The model learns the distribution $p(\Theta_x)$ of suboptimal parameterizations, then repeatedly samples solutions $y \sim p_{\theta'}(y|x)$; a Chebyshev bound plus a finite-sampling hitting argument shows that enough samples make the probability of hitting the optimum near one. The generative engine is a graph neural network run as a two-head diffusion model—one head uses categorical (Bernoulli) transitions for binary offloading decisions, the other uses Gaussian transitions for resource-allocation ratios. Two engineering additions carry the experiments: a padding-edge gating flag that suppresses nonexistent edges in non-fully-connected graphs, and the observation that the two diffusion tasks' gradients are nearly orthogonal, so joint training does not suffer negative interference.

What would settle it

On a fixed instance, measure the average of the model's output parameter vectors over many parallel samples; if that average is far from the optimal solution's parameterization (for example, when the training heuristic always favors one server), the Exceed ratio should stay well above 1, contradicting the claim that sampling converges to optimal. The paper's own variance measurement (σ² ≈ 10⁻⁴) is only on outputs, not on the bias E(Θ_x) − θ*, so measuring that bias directly would settle the question.

Watch

Extended reading notes

Core claim

GDSG is a multi-task graph diffusion model that takes a network graph as input and generates both discrete offloading decisions and continuous resource-allocation ratios. The paper's central claim is that training this model on a dataset of suboptimal solutions—produced cheaply by a randomized minimum-cost-maximum-flow heuristic—is enough for it to learn the distribution of high-quality solutions, so that parallel sampling from the learned distribution converges to the optimal solution with high probability. On ten test scales, the generated solutions' costs sit within about 10% of the ground-truth optimal costs (Exceed ratios mostly below 1.1), which beats the heuristic generator, a same-architecture discriminative GNN, and prior diffusion-based combinatorial solvers. The reported savings reach up to 56.62% of the target cost when trained on an optimal dataset and 41.06% when trained on a suboptimal dataset, compared with discriminative baselines.

Load-bearing premise

The suboptimal solutions' parameterizations must be centered on the optimal parameterization, so that random mistakes cancel out rather than pushing the learned distribution toward a systematically wrong solution.

Editorial extensions

If this is right

  • A model trained purely on suboptimal MCMF-generated data achieves Exceed ratios below 1.1 on most tested scales, so near-optimal decisions can be obtained without ground-truth labels.
  • GDSG outperforms the heuristic that generated its training data, showing the model generalizes beyond its data generator.
  • GDSG beats a same-architecture discriminative GNN on both suboptimal and ground-truth training datasets, with substantial cost savings.
  • The discrete and continuous diffusion tasks share parameters with near-100% gradient orthogonality, so multi-task training avoids negative interference; the paper attributes this to the diffusion loss rather than the GNN architecture.
  • Padding-edge gating lets one trained model generalize across graph sizes, substantially improving the Exceed ratio over a vanilla GNN backbone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the heuristic is systematically biased, the paper's convergence proof does not guarantee optimality, but the diffusion sampler might still concentrate probability mass near good solutions; measuring the gap between the sample-mean parameterization and the optimal one on a biased instance would test this.
  • The same pipeline—cheap suboptimal samples, a parametrized solution space, and a multi-task diffusion model—could transfer to other graph-structured MEC problems such as routing, scheduling, or spectrum allocation, where ground-truth labels are equally scarce.
  • The near-total gradient orthogonality between the discrete and continuous heads suggests that diffusion objectives can sidestep standard multi-task interference, which would make joint discrete-continuous optimization easier for other network problems.
  • Because the heuristic generates an 80-node instance in about 0.6 seconds while GDSG inference is roughly 0.03–0.4 seconds, the practical bottleneck shifts to dataset coverage; randomizing the heuristic's initialization is a cheap way to broaden the distribution the model learns.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper proposes GDSG, a graph-diffusion-based generative model for the multi-server multi-user computation offloading (MSCO) problem in MEC networks. The central claim is that GDSG can be trained on suboptimal datasets, produced cheaply by an MCMF heuristic, and still generate solutions that converge to the optimal solution with high probability. The authors reformulate the optimization as solution-distribution learning, provide a theoretical argument based on Chebyshev's inequality that repeated sampling from a learned distribution near the optimal parametrization can hit the optimum, and introduce a multi-task GNN with a padding-edge gating mechanism. Experiments on synthetic datasets report Exceed ratios below 1.1 on most test scales, with cost savings up to 56.62% over a discriminative GNN on ground-truth training data and 41.06% on suboptimal training data. The paper also claims near-100% task orthogonality between discrete and continuous diffusion tasks. The code and datasets are open-sourced.

Significance. If the central claims hold, the work would be a practically valuable contribution: it offers a way to train an optimization solver from cheap suboptimal data, avoiding the need for expensive ground-truth labels, and it provides an open-source dataset and implementation. The empirical comparison against a same-architecture discriminative GNN is a strength, and the padding-edge gating mechanism appears to improve generalization on non-fully-connected graphs. However, the theoretical guarantee rests on an unverified unbiasedness assumption E(Theta_x)=theta*, and the evaluation on large test scales relies on labels whose provenance is not described. The orthogonality claim, as written, is internally inconsistent because orthogonality corresponds to cosine 0, not 1. These issues leave the main contributions currently under-supported.

major comments (4)
  1. [Section V-A, Table I] The test sets gt7s24u, gt7s27u, gt10s31u, gt10s36u, gt20s61u, and gt20s68u are named 'gt' (ground-truth), but Section V-A only describes exhaustive generation for small-scale ground-truth datasets and explicitly states that 'an optimal solver is not available.' The method used to produce the labels for the larger test scales is not disclosed. If those labels are generated by the same MCMF heuristic, possibly with more restarts, then the Exceed ratios in Table I measure agreement with the heuristic's choices rather than closeness to the true optimum, and the central claim that GDSG 'converges to the optimal solution' is not supported by the data. Please specify the exact label-generation procedure for every test scale, or rename these sets and adjust the claims accordingly.
  2. [Section III-B, Eq. (9)] The convergence argument assumes E(Theta_x)=mu=theta*, i.e., that the mean of the suboptimal parameterizations equals the optimal parameterization. This assumption is not derived from the MCMF heuristic and is not tested in Section V-B; the empirical measurement of the model output variance (on the order of 1e-4) addresses spread but not bias. If the heuristic is systematically biased toward certain allocation patterns, Chebyshev's inequality with respect to theta* fails, and the repeated-sampling argument for hitting y* collapses. Please verify the mean condition on the actual generated datasets, or replace the guarantee with a bias-aware bound involving |E(Theta_x)-theta*|.
  3. [Section V-C, Fig. 4] Orthogonal gradient vectors have cosine 0, yet the text and the figure caption state that 'cosine values approaching 1' and 'orthogonal gradient proportions approaching 100%' indicate good orthogonality, and DiGNN is criticized for having cosine values 'far away from 1.' As written, the reported numbers are internally contradictory: a cosine near 1 means the gradients are nearly parallel, not orthogonal. Please clarify the metric actually computed (e.g., 1 - |cos|, or the angle itself) and re-state the claim; otherwise, the task-orthogonality contribution is not interpretable.
  4. [Section III-B, Eqs. (11)-(13)] The derivation of the lower bound on the number of samples n is not rigorous. The quantity p_epsilon is a probability determined by the distribution, not a free parameter, and replacing Chebyshev's '>=' with '>' on the grounds that the distribution is not three-point does not provide a uniform margin, so the passage from Eq. (11) to Eq. (12) is not justified. Moreover, Eq. (12) depends on the unknown p_epsilon, making it not a computable lower bound. The paper should either provide a clean derivation (e.g., using a bias-aware or union-bound argument) or present this as a heuristic scaling argument rather than a proof.
minor comments (3)
  1. [Abstract] The phrase 'converging to the optimal solution large probably' in the abstract should read 'with high probability' to match the full text.
  2. [Section V-A] The sentence 'we use exhaustive methods to generate small-scale ground-truth datasets but the time complexity is very high (over 30sec for a single instance of 80 nodes)' is confusing because 80 nodes is not a small scale relative to the listed small-scale sets (e.g., gt3s6u has 9 nodes). Please clarify which scales were solved exhaustively and what '80 nodes' refers to.
  3. [Section III-A] The parametrization of the discrete solution space with theta*_i0 and theta*_i1 is inconsistent with the earlier definition of theta in R^N; please unify the notation.

Circularity Check

0 steps flagged · score 2.0 of 10

No demonstrated circularity: the convergence proof is explicitly conditional on an unverified unbiasedness assumption, and the large-scale ground-truth label generation is undescribed, but no prediction reduces to a fitted input or to a self-citation by the paper's own equations.

full rationale

The paper's central theoretical claim in Sec. III-B is conditional rather than circular: it states 'we assume that the perturbations caused by suboptimal solutions are smoothly random and the mean E(Theta_x) = mu = theta* holds' and then applies Chebyshev's inequality (Eq. 9). This is an unverified assumption about the heuristic data distribution, not a derivation of that distribution from the model; the proof would fail if the assumption is false, but it does not redefine the prediction as its input. The empirical claim of near-optimal performance is measured against externally named 'gt' test sets; small-scale ground truth is said to come from exhaustive search ('we use exhaustive methods to generate small-scale ground-truth datasets'), while the paper also states 'an optimal solver is not available' and does not describe how the gt7s24u through gt20s68u labels were produced. That is a provenance gap that weakens support for the convergence claim, but the paper does not state that those labels are generated by the same heuristic, so a circular reduction cannot be exhibited from the text. Self-citations ([25], [38], [46]) and the DiffSG reference appear in related-work and baseline roles and are not load-bearing for the derivation. Accordingly, no circular step meets the quotation-and-reduction bar; the correct finding is no significant circularity, with the unbiasedness assumption and large-scale label provenance flagged as correctness risks.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The ledger shows that the 'convergence to optimum' result is conditional on several unverified premises, chiefly the mean assumption E(Theta_x)=theta* and the exactness of ground-truth labels for large test sets. The hand-chosen parameters epsilon, gamma, delta only enter illustrative bounds, not the algorithm itself.

free parameters (4)
  • epsilon (neighborhood radius) = 0.1
    Chosen by hand in Sec. V-B for the hitting-probability illustration; not fitted to data.
  • gamma (Gaussian sampling std) = 0.04
    Chosen in Sec. V-B for continuous solution sampling in Eq. (10).
  • delta (tolerable gap) = 0.1
    Chosen in Sec. V-B to define the acceptable interval in Eq. (10).
  • a (probability slack) = a -> 0
    Introduced in Eq. (11) as an arbitrarily small positive constant; never specified numerically.
assumptions (5)
  • domain assumption The mean of the suboptimal parametrization equals the optimal parametrization: E(Theta_x)=mu=theta*.
    Stated in Sec. III-B as an assumption; without it the convergence guarantee in Eqs. (9)-(13) does not follow.
  • ad hoc to paper p(Theta_x) is not a three-point distribution, so Chebyshev's inequality becomes strict.
    Invoked in Sec. III-B.3 to replace '>=' with '>' in Eq. (11); no empirical justification.
  • domain assumption The components of theta are independent across dimensions for the hitting probability calculation.
    Used in Sec. III-B.2, 'with independent variables in each dimension'; likely violated by correlated offloading decisions.
  • domain assumption Ground-truth labels for all ten test scales are exact optima.
    The dataset section explicitly describes exhaustive generation only for small-scale datasets; the method for large test sets is not stated, yet the Exceed ratio metric assumes these labels are optimal.
  • domain assumption The trained GNN can represent the conditional solution distribution p_Phi(theta|G) well enough for the denoising objective.
    Implicit in Sec. IV-B; no capacity or approximation guarantees are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks." pith.science (2026). https://pith.science/paper/DPHSG553

@misc{pith2026241208296,
  author       = {Pith},
  title        = {Pith review of: GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DPHSG553}},
  note         = {Machine review of arXiv:2412.08296}
}
read the original abstract

Optimization is crucial for MEC networks to function efficiently and reliably, most of which are NP-hard and lack efficient approximation algorithms. This leads to a paucity of optimal solution, constraining the effectiveness of conventional deep learning approaches. Most existing learning-based methods necessitate extensive optimal data and fail to exploit the potential benefits of suboptimal data that can be obtained with greater efficiency and effectiveness. Taking the multi-server multi-user computation offloading (MSCO) problem, which is widely observed in systems like Internet-of-Vehicles (IoV) and Unmanned Aerial Vehicle (UAV) networks, as a concrete scenario, we present a Graph Diffusion-based Solution Generation (GDSG) method. This approach is designed to work with suboptimal datasets while converging to the optimal solution large probably. We transform the optimization issue into distribution-learning and offer a clear explanation of learning from suboptimal training datasets. We build GDSG as a multi-task diffusion model utilizing a Graph Neural Network (GNN) to acquire the distribution of high-quality solutions. We use a simple and efficient heuristic approach to obtain a sufficient amount of training data composed entirely of suboptimal solutions. In our implementation, we enhance the backbone GNN and achieve improved generalization. GDSG also reaches nearly 100\% task orthogonality, ensuring no interference between the discrete and continuous generation tasks. We further reveal that this orthogonality arises from the diffusion-related training loss, rather than the neural network architecture itself. The experiments demonstrate that GDSG surpasses other benchmark methods on both the optimal and suboptimal training datasets. The MSCO datasets has open-sourced at this http URL, as well as the GDSG algorithm codes at https://github.com/qiyu3816/GDSG.

Figures

Figures reproduced from arXiv: 2412.08296 by the authors.

Figure 1
Figure 1. System model of multi-server and multi-user computation offloading. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Framework for training the diffusion generative model with suboptimal dataset to achieve the optimal solution generation. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The expectation of hitting y ∗ for different number of samples n and variable dimension N. The dotted lines represent the lower bounds on n, where the red and black dotted lines coincide. 3s8u 4s12u 7s27u 10s36u 20s68u gt3s8u gt4s12u Training dataset 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Cosine of the average angle GDSG DiGNN 3s8u 4s12u 7s27u 10s36u 20s68u gt3s8u gt4s12u Training dataset 0 20 40 60 80 100 Proportion near orth… view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: The performance Exceed ratio of GDSG (GNN padding edge not handled). gt3s6u gt3s8u gt4s10u gt4s12u gt7s24u gt7s27u gt10s31u gt10s36u gt20s61u gt20s68u Evaluation dataset lq3s6u lq3s8u lq4s10u lq4s12u lq7s24u lq7s27u lq10s31u lq10s36u lq20s61u lq20s68u Training dataset …
Figure 8
Figure 8. Figure 8: Performance comparison of three trained models: GDSG trained on a [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    Guiding a pretrained topology-diffusion generator with human-preference reward classifiers is claimed to suppress floating-material and boundary-violation failure modes without retraining the generator.

Reference graph

Works this paper leans on

74 extracted references · 62 canonical work pages · cited by 1 Pith paper

  1. [1]

    A Survey on Mobile Edge Computing: The Communication Perspective,

    Y . Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief, “A Survey on Mobile Edge Computing: The Communication Perspective,” IEEE Communications Surveys & Tutorials , vol. 19, no. 4, pp. 2322–2358, 2017

  2. [2]

    Mobile Edge Computing: A Survey on Archi- tecture and Computation Offloading,

    P. Mach and Z. Becvar, “Mobile Edge Computing: A Survey on Archi- tecture and Computation Offloading,” IEEE Communications Surveys & Tutorials, vol. 19, no. 3, pp. 1628–1656, 2017

  3. [3]

    A2-UA V: Application-Aware Content and Network Optimization of Edge-Assisted UA V Systems,

    A. Coletta, F. Giorgi, G. Maselli, M. Prata, D. Silvestri, J. Ashdown, and F. Restuccia, “A2-UA V: Application-Aware Content and Network Optimization of Edge-Assisted UA V Systems,” in Proceedings of IEEE International Conference on Computer Communications , 2023, pp. 1– 10

  4. [4]

    Energy-Efficient Trajectory Optimization for Aerial Video Surveillance under QoS Constraints,

    C. Zhan, H. Hu, S. Mao, and J. Wang, “Energy-Efficient Trajectory Optimization for Aerial Video Surveillance under QoS Constraints,” in Proceedings of IEEE International Conference on Computer Communi- cations, 2022, pp. 1559–1568

  5. [5]

    Multi-UA V Trajectory and Power Opti- mization for Cached UA V Wireless Networks With Energy and Content Recharging-Demand Driven Deep Learning Approach,

    S. Chai and V . K. N. Lau, “Multi-UA V Trajectory and Power Opti- mization for Cached UA V Wireless Networks With Energy and Content Recharging-Demand Driven Deep Learning Approach,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 10, pp. 3208–3224, 2021

  6. [6]

    Multi- UA V Trajectory Planning for Energy-Efficient Content Coverage: A Decentralized Learning-Based Approach,

    C. Zhao, J. Liu, M. Sheng, W. Teng, Y . Zheng, and J. Li, “Multi- UA V Trajectory Planning for Energy-Efficient Content Coverage: A Decentralized Learning-Based Approach,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 10, pp. 3193–3207, 2021

  7. [7]

    Offloading Optimization in Edge Computing for Deep-Learning-Enabled Target Tracking by Internet of UA Vs,

    B. Yang, X. Cao, C. Yuen, and L. Qian, “Offloading Optimization in Edge Computing for Deep-Learning-Enabled Target Tracking by Internet of UA Vs,”IEEE Internet of Things Journal , vol. 8, no. 12, pp. 9878– 9893, 2021

  8. [8]

    Masaracchia, K

    A. Masaracchia, K. K. Nguyen, T. Q. Duong, and V . Sharma, Deep Reinforcement Learning for Reconfigurable Intelligent Surfaces and UAV Empowered Smart 6G Communications . The Institution of Engineering and Technology, 2024. [Online]. Available: https: //digital-library.theiet.org/doi/abs/10.1049/PBTE106E

Show all 74 references
  1. [9]

    Generative AI-Augmented Graph Reinforcement Learning for Adaptive UA V Swarm Optimization,

    B. Hazarika, P. Singh, K. Singh, S. L. Cotton, H. Shin, O. A. Dobre, and T. Q. Duong, “Generative AI-Augmented Graph Reinforcement Learning for Adaptive UA V Swarm Optimization,” IEEE Internet of Things Journal, pp. 1–1, 2025

  2. [10]

    Latency Optimization for Blockchain-Empowered Federated Learning in Multi-Server Edge Computing,

    D. C. Nguyen, S. Hosseinalipour, D. J. Love, P. N. Pathirana, and C. G. Brinton, “Latency Optimization for Blockchain-Empowered Federated Learning in Multi-Server Edge Computing,” IEEE Journal on Selected Areas in Communications , vol. 40, no. 12, pp. 3373–3390, 2022

  3. [11]

    Federated Edge Network Utility Maximization for a Multi-Server System: Algorithm and Convergence,

    N. Karakoc ¸, A. Scaglione, M. Reisslein, and R. Wu, “Federated Edge Network Utility Maximization for a Multi-Server System: Algorithm and Convergence,” IEEE/ACM Transactions on Networking , vol. 30, no. 5, pp. 2002–2017, 2022

  4. [12]

    Federated Spectrum Learning for Reconfigurable Intelligent Surfaces-Aided Wireless Edge Networks,

    B. Yang, X. Cao, C. Huang, C. Yuen, M. Di Renzo, Y . L. Guan, D. Niyato, L. Qian, and M. Debbah, “Federated Spectrum Learning for Reconfigurable Intelligent Surfaces-Aided Wireless Edge Networks,” IEEE Transactions on Wireless Communications , vol. 21, no. 11, pp. 9610–9626, 2022

  5. [13]

    Reconfigurable Intelligent Surface-Assisted Aerial-Terrestrial Communications via Multi-Task Learning,

    X. Cao, B. Yang, C. Huang, C. Yuen, M. D. Renzo, D. Niyato, and Z. Han, “Reconfigurable Intelligent Surface-Assisted Aerial-Terrestrial Communications via Multi-Task Learning,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 10, pp. 3035–3050, 2021

  6. [14]

    Hybrid Beamforming for Reconfigurable Intelligent Surface based Multi-User Communications: Achievable Rates With Limited Discrete Phase Shifts,

    B. Di, H. Zhang, L. Song, Y . Li, Z. Han, and H. V . Poor, “Hybrid Beamforming for Reconfigurable Intelligent Surface based Multi-User Communications: Achievable Rates With Limited Discrete Phase Shifts,” IEEE Journal on Selected Areas in Communications , vol. 38, no. 8, pp. 1...

  7. [15]

    Resource Allocation for Power Minimization in RIS-Assisted Multi- UA V Networks With NOMA,

    W. Feng, J. Tang, Q. Wu, Y . Fu, X. Zhang, D. K. C. So, and K.-K. Wong, “Resource Allocation for Power Minimization in RIS-Assisted Multi- UA V Networks With NOMA,”IEEE Transactions on Communications , vol. 71, no. 11, pp. 6662–6676, 2023

  8. [16]

    Joint Base Station and IRS Deployment for En- hancing Network Coverage: A Graph-Based Modeling and Optimization Approach,

    W. Mei and R. Zhang, “Joint Base Station and IRS Deployment for En- hancing Network Coverage: A Graph-Based Modeling and Optimization Approach,” IEEE Transactions on Wireless Communications , vol. 22, no. 11, pp. 8200–8213, 2023

  9. [17]

    Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks: Design Optimization and Analysis,

    X. Zhang, B. Yang, Z. Yu, X. Cao, G. C. Alexandropoulos, Y . Zhang, M. Debbah, and C. Yuen, “Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks: Design Optimization and Analysis,” IEEE Transactions on Intelligent Trans- portation Sys...

  10. [18]

    Computation Offloading in MEC-Enabled IoV Networks: Average Energy Efficiency Analysis and Learning-Based Maximization,

    T. Z. H. Ernest and A. S. Madhukumar, “Computation Offloading in MEC-Enabled IoV Networks: Average Energy Efficiency Analysis and Learning-Based Maximization,” IEEE Transactions on Mobile Comput- ing, vol. 23, no. 5, pp. 6074–6087, 2024

  11. [19]

    Asynchronous Deep Reinforcement Learning for Data-Driven Task Offloading in MEC- Empowered Vehicular Networks,

    P. Dai, K. Hu, X. Wu, H. Xing, and Z. Yu, “Asynchronous Deep Reinforcement Learning for Data-Driven Task Offloading in MEC- Empowered Vehicular Networks,” in Proceedings of IEEE International Conference on Computer Communications , 2021, pp. 1–10

  12. [20]

    Edge Intelligence for Autonomous Driving in 6G Wireless System: Design Challenges and Solutions,

    B. Yang, X. Cao, K. Xiong, C. Yuen, Y . L. Guan, S. Leng, L. Qian, and Z. Han, “Edge Intelligence for Autonomous Driving in 6G Wireless System: Design Challenges and Solutions,” IEEE Wireless Communica- tions, vol. 28, no. 2, pp. 40–47, 2021

  13. [21]

    DeepScheduler: Enabling Flow-Aware Scheduling in Time-Sensitive Networking,

    X. He, X. Zhuge, F. Dang, W. Xu, and Z. Yang, “DeepScheduler: Enabling Flow-Aware Scheduling in Time-Sensitive Networking,” in Proceedings of IEEE International Conference on Computer Commu- nications, 2023, pp. 1–10

  14. [22]

    RouteNet: Leveraging graph neural networks for network modeling and optimization in SDN,

    K. Rusek, J. Su ´arez-Varela, P. Almasan, P. Barlet-Ros, and A. Cabellos- Aparicio, “RouteNet: Leveraging graph neural networks for network modeling and optimization in SDN,” IEEE Journal on Selected Areas in Communications, vol. 38, no. 10, pp. 2260–2270, 2020

  15. [23]

    A Joint Energy and Latency Framework for Transfer Learning Over 5G Industrial Edge Networks,

    B. Yang, O. Fagbohungbe, X. Cao, C. Yuen, L. Qian, D. Niyato, and Y . Zhang, “A Joint Energy and Latency Framework for Transfer Learning Over 5G Industrial Edge Networks,” IEEE Transactions on Industrial Informatics, vol. 18, no. 1, pp. 531–541, 2022

  16. [24]

    Deep- learning-based joint resource scheduling algorithms for hybrid MEC networks,

    F. Jiang, K. Wang, L. Dong, C. Pan, W. Xu, and K. Yang, “Deep- learning-based joint resource scheduling algorithms for hybrid MEC networks,” IEEE Internet of Things Journal , vol. 7, no. 7, pp. 6252– 6265, 2019

  17. [25]

    A multi-head ensemble multi-task learning approach for dynamical com- putation offloading,

    R. Liang, B. Yang, Z. Yu, X. Cao, D. W. K. Ng, and C. Yuen, “A multi-head ensemble multi-task learning approach for dynamical com- putation offloading,” in Proceedings of IEEE Global Communications Conference, 2023, pp. 6079–6084

  18. [26]

    GNN-Based Power Allocation and User Association in Digital Twin Network for the Terahertz Band,

    H. Zhang, X. Ma, X. Liu, L. Li, and K. Sun, “GNN-Based Power Allocation and User Association in Digital Twin Network for the Terahertz Band,” IEEE Journal on Selected Areas in Communications , 2023

  19. [27]

    Edge-Assisted Multi-Layer Offloading Optimization of LEO Satellite-Terrestrial Integrated Networks,

    X. Cao, B. Yang, Y . Shen, C. Yuen, Y . Zhang, Z. Han, H. V . Poor, and L. Hanzo, “Edge-Assisted Multi-Layer Offloading Optimization of LEO Satellite-Terrestrial Integrated Networks,” IEEE Journal on Selected Areas in Communications , vol. 41, no. 2, pp. 381–398, 2023

  20. [28]

    STaR: self-taught reasoner bootstrapping reasoning with reasoning,

    E. Zelikman, Y . Wu, J. Mu, and N. D. Goodman, “STaR: self-taught reasoner bootstrapping reasoning with reasoning,” in Proceedings of Neural Information Processing Systems , 2024

  21. [29]

    LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery,

    P. Ma, T.-H. Wang, M. Guo, Z. Sun, J. B. Tenenbaum, D. Rus, C. Gan, and W. Matusik, “LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery,” in Proceedings of International Conference on Machine Learning , vol. 235, 21–27 Jul 2024, p...

  22. [30]

    Dream the impossible: outlier imag- ination with diffusion models,

    X. Du, Y . Sun, X. Zhu, and Y . Li, “Dream the impossible: outlier imag- ination with diffusion models,” in Proceedings of Neural Information Processing Systems, 2024

  23. [31]

    Adding Conditional Control to Text-to-Image Diffusion Models,

    L. Zhang, A. Rao, and M. Agrawala, “Adding Conditional Control to Text-to-Image Diffusion Models,” in Proceedings of International Conference on Computer Vision , 2023, pp. 3813–3824

  24. [32]

    Gurobi Optimizer Reference Manual,

    Gurobi Optimization, LLC, “Gurobi Optimizer Reference Manual,”

  25. [33]

    ApS, The MOSEK optimization toolbox for MATLAB manual

    M. ApS, The MOSEK optimization toolbox for MATLAB manual. Version 10.1., 2024. [Online]. Available: http://docs.mosek.com/latest/ toolbox/index.html

  26. [34]

    IBM ILOG CPLEX Optimization Studio,

    IBM, “IBM ILOG CPLEX Optimization Studio,” 2024. [Online]. Available: https://www.ibm.com/docs/en/icos/22.1.1

  27. [35]

    GEKKO Optimization Suite,

    L. Beal, D. Hill, R. Martin, and J. Hedengren, “GEKKO Optimization Suite,” Processes, vol. 6, no. 8, p. 106, 2018

  28. [36]

    A Survey on Generative Diffusion Models,

    H. Cao, C. Tan, Z. Gao, Y . Xu, G. Chen, P.-A. Heng, and S. Z. Li, “A Survey on Generative Diffusion Models,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 7, pp. 2814–2830, 2024

  29. [37]

    A GNN- based supervised learning framework for resource allocation in wireless IoT networks,

    T. Chen, X. Zhang, M. You, G. Zheng, and S. Lambotharan, “A GNN- based supervised learning framework for resource allocation in wireless IoT networks,” IEEE Internet of Things Journal, vol. 9, no. 3, pp. 1712– 1724, 2021

  30. [38]

    Computation offloading in multi-access edge computing: A multi-task learning approach,

    B. Yang, X. Cao, J. Bassey, X. Li, and L. Qian, “Computation offloading in multi-access edge computing: A multi-task learning approach,” IEEE Transactions on Mobile Computing, vol. 20, no. 9, pp. 2745–2762, 2020

  31. [39]

    Denoising Diffusion Probabilistic Models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” in Proceedings of Neural Information Processing Systems , vol. 33, 2020, pp. 6840–6851

  32. [40]

    Classifier-free diffusion guidance,

    J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598, 2022

  33. [41]

    Diffusion Models Beat GANs on Image Synthesis,

    P. Dhariwal and A. Nichol, “Diffusion Models Beat GANs on Image Synthesis,” in Proceedings of Neural Information Processing Systems , vol. 34, 2021, pp. 8780–8794

  34. [42]

    Structured Denoising Diffusion Models in Discrete State-Spaces,

    J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. van den Berg, “Structured Denoising Diffusion Models in Discrete State-Spaces,” in Proceedings of Neural Information Processing Systems , vol. 34, 2021, pp. 17 981–17 993

  35. [43]

    DIFUSCO: Graph-based Diffusion Solvers for Combinatorial Optimization,

    Z. Sun and Y . Yang, “DIFUSCO: Graph-based Diffusion Solvers for Combinatorial Optimization,” in Proceedings of Neural Information Processing Systems, vol. 36, 2023, pp. 3706–3731

  36. [44]

    T2T: From Distribution Learning in Training to Gradient Search in Testing for Combinatorial Optimization,

    Y . Li, J. Guo, R. Wang, and J. Yan, “T2T: From Distribution Learning in Training to Gradient Search in Testing for Combinatorial Optimization,” in Proceedings of Neural Information Processing Systems, vol. 36, 2023, pp. 50 020–50 040

  37. [45]

    Enhancing Deep Reinforcement Learning: A Tutorial on Generative Diffusion Models in Network Optimization,

    H. Du, R. Zhang, Y . Liu, J. Wang, Y . Lin, Z. Li, D. Niyato, J. Kang, Z. Xiong, S. Cui, B. Ai, H. Zhou, and D. I. Kim, “Enhancing Deep Reinforcement Learning: A Tutorial on Generative Diffusion Models in Network Optimization,” IEEE Communications Surveys & Tutorials, pp. 1–1, 2024

  38. [46]

    DiffSG: A Generative Solver for Network Optimization with Diffusion Model,

    R. Liang, B. Yang, Z. Yu, B. Guo, X. Cao, M. Debbah, H. V . Poor, and Y . Chau, “DiffSG: A Generative Solver for Network Optimization with Diffusion Model,” 2024. [Online]. Available: https://arxiv.org/abs/2408.06701

  39. [47]

    Generative AI based Secure Wireless Sensing for ISAC Networks,

    J. Wang, H. Du, Y . Liu, G. Sun, D. Niyato, S. Mao, D. I. Kim, and X. Shen, “Generative AI based Secure Wireless Sensing for ISAC Networks,” 2024. [Online]. Available: https://arxiv.org/abs/2408.11398

  40. [48]

    Generative AI for Deep Reinforcement Learning: Framework, Analysis, and Use Cases,

    G. Sun, W. Xie, D. Niyato, F. Mei, J. Kang, H. Du, and S. Mao, “Generative AI for Deep Reinforcement Learning: Framework, Analysis, and Use Cases,” 2024. [Online]. Available: https://arxiv.org/ abs/2405.20568

  41. [49]

    Diffusion-Based Reinforcement Learning for Edge-Enabled AI-Generated Content Services,

    H. Du, Z. Li, D. Niyato, J. Kang, Z. Xiong, H. Huang, and S. Mao, “Diffusion-Based Reinforcement Learning for Edge-Enabled AI-Generated Content Services,” IEEE Transactions on Mobile Com- puting, vol. 23, no. 9, pp. 8902–8918, 2024

  42. [50]

    Deep Generative Model and Its Applications in Efficient Wireless Network Management: A Tutorial and Case Study,

    Y . Liu, H. Du, D. Niyato, J. Kang, Z. Xiong, D. I. Kim, and A. Ja- malipour, “Deep Generative Model and Its Applications in Efficient Wireless Network Management: A Tutorial and Case Study,” IEEE Wireless Communications, vol. 31, no. 4, pp. 199–207, 2024

  43. [51]

    ADMM for Mobile Edge Intelligence: A Survey,

    A. He, H. Pan, Y . Dai, X. Si, C. Yuen, and Y . Zhang, “ADMM for Mobile Edge Intelligence: A Survey,” IEEE Communications Surveys & Tutorials, pp. 1–1, 2024

  44. [52]

    Online Distributed ADMM Algorithm With RLS-Based Multitask Graph Filter Models,

    Y . Lai, F. Chen, M. Feng, and J. Kurths, “Online Distributed ADMM Algorithm With RLS-Based Multitask Graph Filter Models,” IEEE Transactions on Network Science and Engineering , vol. 9, no. 6, pp. 4115–4128, 2022

  45. [53]

    Hybridized MA- DRL for Serving xURLLC With Cognizable RIS and UA V Integration,

    A. Paul, R. Allu, K. Singh, C.-P. Li, and T. Q. Duong, “Hybridized MA- DRL for Serving xURLLC With Cognizable RIS and UA V Integration,” IEEE Transactions on Wireless Communications , vol. 23, no. 10, pp. 15 507–15 524, 2024

  46. [54]

    A survey on uplink resource allocation in OFDMA wireless networks,

    E. Yaacoub and Z. Dawy, “A survey on uplink resource allocation in OFDMA wireless networks,” IEEE Communications Surveys & Tutorials, vol. 14, no. 2, pp. 322–337, 2011

  47. [55]

    An overview of radio resource management in relay-enhanced OFDMA-based networks,

    M. Salem, A. Adinoyi, M. Rahman, H. Yanikomeroglu, D. Falconer, Y .-D. Kim, E. Kim, and Y .-C. Cheong, “An overview of radio resource management in relay-enhanced OFDMA-based networks,” IEEE Com- munications Surveys & Tutorials , vol. 12, no. 3, pp. 422–438, 2010

  48. [56]

    Decentralized computation offloading game for mobile cloud computing,

    X. Chen, “Decentralized computation offloading game for mobile cloud computing,” IEEE Transactions on Parallel and Distributed Systems , vol. 26, no. 4, pp. 974–983, 2014

  49. [57]

    Processor design for portable systems,

    T. D. Burd and R. W. Brodersen, “Processor design for portable systems,” Journal of VLSI signal processing systems for signal, image and video technology , vol. 13, no. 2, pp. 203–221, 1996

  50. [58]

    DIMES: A Differentiable Meta Solver for Combinatorial Optimization Problems,

    R. Qiu, Z. Sun, and Y . Yang, “DIMES: A Differentiable Meta Solver for Combinatorial Optimization Problems,” in Proceedings of Neural Information Processing Systems , vol. 35, 2022, pp. 25 531–25 546

  51. [59]

    Chebyshev inequality with estimated mean and variance,

    J. G. Saw, M. C. Yang, and T. C. Mo, “Chebyshev inequality with estimated mean and variance,” The American Statistician, vol. 38, no. 2, pp. 130–132, 1984

  52. [60]

    Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,

    D. Kingma and R. Gao, “Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,” in Proceedings of Neural Information Processing Systems , vol. 36, 2023, pp. 65 484–65 516

  53. [61]

    Variational Diffusion Models,

    D. Kingma, T. Salimans, B. Poole, and J. Ho, “Variational Diffusion Models,” in Proceedings of Neural Information Processing Systems , vol. 34, 2021, pp. 21 696–21 707

  54. [62]

    Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions,

    E. Hoogeboom, D. Nielsen, P. Jaini, P. Forr ´e, and M. Welling, “Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions,” in Proceedings of Neural Information Processing Systems, vol. 34, 2021, pp. 12 454–12 465

  55. [63]

    Analog bits: Generating discrete data using diffusion models with self-conditioning,

    T. Chen, R. Zhang, and G. Hinton, “Analog bits: Generating discrete data using diffusion models with self-conditioning,” arXiv preprint arXiv:2208.04202, 2022

  56. [64]

    Learning the travelling salesperson problem requires rethinking generalization,

    C. K. Joshi, Q. Cappart, L.-M. Rousseau, and T. Laurent, “Learning the travelling salesperson problem requires rethinking generalization,” Constraints, vol. 27, no. 1, pp. 70–98, 2022

  57. [65]

    Benchmarking graph neural networks,

    V . P. Dwivedi, C. K. Joshi, A. T. Luu, T. Laurent, Y . Bengio, and X. Bresson, “Benchmarking graph neural networks,” MIT Press Journal of Machine Learning Research , vol. 24, no. 43, pp. 1–48, 2023

  58. [66]

    A systematic survey on deep generative models for graph generation,

    X. Guo and L. Zhao, “A systematic survey on deep generative models for graph generation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 5, pp. 5370–5390, 2022

  59. [67]

    Towards Impartial Multi-task Learning,

    L. Liu, Y . Li, Z. Kuang, J.-H. Xue, Y . Chen, W. Yang, Q. Liao, and W. Zhang, “Towards Impartial Multi-task Learning,” in Proceedings of International Conference on Learning Representations , 2021

  60. [68]

    Gradient Surgery for Multi-Task Learning,

    T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn, “Gradient Surgery for Multi-Task Learning,” in Proceedings of Neural Information Processing Systems , vol. 33, 2020, pp. 5824–5836

  61. [69]

    Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout,

    Z. Chen, J. Ngiam, Y . Huang, T. Luong, H. Kretzschmar, Y . Chai, and D. Anguelov, “Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout,” in Proceedings of Neural Information Processing Systems, vol. 33, 2020, pp. 2039–2050

  62. [70]

    Addressing Negative Transfer in Diffusion Models,

    H. Go, , Y . Lee, S. Lee, S. Oh, H. Moon, and S. Choi, “Addressing Negative Transfer in Diffusion Models,” in Proceedings of Neural Information Processing Systems , vol. 36, 2023, pp. 27 199–27 222

  63. [71]

    DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data,

    H. Ye and D. Xu, “DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data,” in Proceedings of IEEE/CVF Computer Vision and Pattern Recognition Conference, 2024, pp. 27 960–27 969

  64. [72]

    Diffusion Model is an Effective Planner and Data Synthesizer for Multi- Task Reinforcement Learning,

    H. He, C. Bai, K. Xu, Z. Yang, W. Zhang, D. Wang, B. Zhao, and X. Li, “Diffusion Model is an Effective Planner and Data Synthesizer for Multi- Task Reinforcement Learning,” in Proceedings of Neural Information Processing Systems, vol. 36, 2023, pp. 64 896–64 917

  65. [73]

    Denoising Diffusion Implicit Models,

    J. Song, C. Meng, and S. Ermon, “Denoising Diffusion Implicit Models,” in Proceedings of International Conference on Learning Representa- tions, 2021

  66. [2024]

    Available: https://www.gurobi.com

    [Online]. Available: https://www.gurobi.com

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.