REVIEW 3 major objections 5 minor 41 references
A Lyapunov-Guided Diffusion-Based Reinforcement Learning Approach for UAV-Assisted Vehicular Networks with Delayed CSI Feedback
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A diffusion-model actor, steered by a Lyapunov energy queue, can jointly assign channels, set transmit powers, and choose drone altitude under delayed channel feedback to raise vehicle-to-drone sum rates beyond what conventional deep…
desk verdict Competent application of the authors' earlier D3PG framework to UAV vehicular networks, but the delayed-CSI model in Eqs. (11)-(12) is algebraically wrong and the reported gains over D3PG-WCSI may not survive a corrected simulator. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a denoising diffusion policy network wrapped in a deterministic actor-critic loop. A virtual queue is updated as $Q(t+1)=\max\{Q(t)+P(t)\Delta-E_U^{th},0\}$, and the per-slot problem minimizes $Q(t)(P(t)\Delta-E_U^{th})-V\sum_m R_m^U(t)$; the same expression, negated and with an outage penalty, is the reward. The actor is the diffusion denoiser $\hat{\epsilon}_\theta(\pi_i(t),i,s(t))$, which over $I$ denoising steps reconstructs the action vector $\pi_0(t)=\{x(t),p(t),\Delta H(t)\}$ from Gaussian noise, conditioned on the state $s(t)$ that includes delayed V2V channel gains and the queue. This conditional iterative generation is what the paper credits for robust decisions under CSI delay, and the action-amender step enforces the discrete channel-allocation and power constraints.
What would settle it
Re-run the experiment with the Gauss-Markov process simulated directly: draw $g(t)=\rho\hat{g}(t)+\delta$ with $\rho=J_0(2\pi f_c s_{rel} T_{delay}/c)$, compute $|g(t)|^2$ exactly, and compare D3PG against D3PG-WCSI; if the 6.39% sum-rate gap shrinks or disappears, the reported advantage came from the dropped cross term in Eq. (12) rather than from the diffusion policy.
Extended reading notes
Core claim
The paper's central claim is that diffusion-model-based action generation, rather than an ordinary multilayer-perceptron actor, is what lets a deterministic policy gradient agent cope with the exploration-exploitation trade-off and with outdated CSI in a UAV-assisted vehicular network. Concretely, the agent's reward is the negative of the Lyapunov drift-plus-penalty objective: the sum rate minus the virtual energy queue times the excess energy draw, so maximizing reward both improves throughput and keeps long-term propulsion energy under the threshold. At inference, the actor samples Gaussian noise and iteratively denoises it, conditioned on the current channel state and queue, to produce the channel-allocation, power, and altitude actions; the diffusion pass provides stochastic refinement that a single forward pass lacks. In the reported simulations the result is a higher converged reward and a higher V2U sum rate than the three benchmarks, with the gap over the no-delay baseline widening as the CSI feedback delay grows.
Load-bearing premise
The whole delayed-CSI simulation rests on Eq. (12), which assumes the squared magnitude of a correlated complex Gaussian channel equals the correlation coefficient squared times the squared magnitude of the delayed estimate plus an independent squared-error term, dropping the cross term; if that approximation misrepresents the actual Gauss-Markov fading process, the reported D3PG advantage over D3PG-WCSI may be an artifact of the simulator rather than a real algorithmic gain.
Editorial extensions
If this is right
- If the reported gains hold, diffusion-model actors are a viable replacement for MLP actors in continuous-action wireless resource allocation, not just in image generation.
- The Lyapunov decoupling means the long-term battery constraint can be enforced online without future knowledge, so the same per-slot reward design can be ported to other UAV control problems with energy limits.
- Explicitly modeling CSI delay matters most when the feedback delay is large: the reported gap between D3PG and D3PG-WCSI widens as the Bessel correlation $J_0$ falls.
- The method's per-slot inference cost grows only linearly in denoising steps and network layers, so the accuracy gain is bought with a modest running-time increase (about 3.34 ms per slot versus 0.66 ms for DDPG at $K=10$).
Reading between the lines
- A direct extension the authors leave implicit: the same diffusion actor could be applied to multi-UAV coordination, with the denoising process conditioned on neighboring UAV states and the Lyapunov queue becoming per-UAV energy debt.
- Because the reward is already the negative drift-plus-penalty objective, one could test whether freezing the diffusion noise seed at inference collapses the D3PG advantage; if it does, the gain comes from stochastic exploration rather than representation power.
- The delayed-CSI modeling assumption in Eq. (12) is the most fragile point: simulating the Gauss-Markov process directly and squaring it, instead of using the approximation, would reveal whether the reported gains are robust or partly a simulator artifact.
- The reported plateau in sum rate as the Lyapunov weight $V$ grows suggests a performance ceiling set by the channel and interference structure; a neighboring question is whether the same ceiling appears when the energy budget is tightened, which would expose the true energy-throughput trade-off frontier.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies joint channel allocation, power control, and flight-altitude adjustment in a single-UAV-assisted vehicular network with delayed CSI. It formulates a long-term MINLP that maximizes the V2U communication sum rate subject to V2V outage constraints (C7) and a UAV long-term energy constraint (C8). The authors apply Lyapunov drift-plus-penalty optimization to decompose the long-term problem into per-slot deterministic subproblems, then propose D3PG, a DDPG-style algorithm whose actor is implemented as a diffusion-model denoiser and whose reward encodes the Lyapunov objective plus an outage-violation penalty. Simulations with OpenStreetMap/SUMO vehicle traces compare D3PG with DDPG, H-DDQN, and a no-CSI-delay variant (D3PG-WCSI), reporting V2U sum-rate gains of up to 6.39%, 12.55%, and 23.25% at K=10 while keeping the UAV's moving-average energy below the threshold.
Significance. If the results hold, the paper would demonstrate a useful integration of diffusion-based policy parameterization with Lyapunov-guided rewards for a realistic networking problem. Strengths: the Lyapunov derivation is standard and the appendices supply the drift bound and the sample-path argument for constraint C8; the simulation uses a real-world road layout and SUMO mobility traces; three baselines are compared; and the paper reports the Lyapunov-weight tradeoff (Fig. 11) and per-slot inference times (Table III). However, the empirical claim is weakened by an internally inconsistent delayed-channel model in Eqs. (11)-(12), an incompletely specified gradient path for the stochastic and discrete components of the actor, and the absence of any statistical significance information. These issues are fixable within the manuscript's scope, but they must be addressed before the reported percentage gains can be trusted.
major comments (3)
- [III-C, Eqs. (10)-(12)] The delayed-fading power model in Eqs. (11)-(12) is not a valid consequence of the Gauss-Markov process in Eq. (10). For g(t)=ρĝ(t)+δ with δ~CN(0,1-ρ²), the exact squared magnitude is |g|²=ρ²|ĝ|²+|δ|²+2ρ Re(ĝ*δ); the cross term is dropped in Eqs. (11)-(12). The approximation replaces an exponential |g|² (mean 1, variance 1) with a hypoexponential mixture whose variance is 1-2ρ²+2ρ⁴, so the V2V SINR statistics and the C7 outage probabilities that enter the reward (30) and the constraints are not those of the stated model. Since Section VII-C4 attributes the widening D3PG-versus-D3PG-WCSI gap in Fig. 9 to the Gauss-Markov delay structure, the quantitative gains (e.g., 6.39% at K=10 in Fig. 8) may be artifacts of a mis-specified simulator. Please either simulate the exact |ρĝ+δ|² or justify the approximation explicitly (e.g., as a conditional-mean model) and rerun the experiments to confirm the conclusions.
- [VI-B, Eq. (33) and the action amender] The learning rule in Eq. (33) is the deterministic policy gradient, but the diffusion actor's output is stochastic: the reverse process (27) samples fresh noise at each of the I denoising steps, so η_θ(s) is not a deterministic function of the state. The paper should specify the gradient estimator (e.g., reparameterized pathwise gradients that treat the sampled noise as fixed in each backward pass) and state why it is unbiased for the stochastic policy. Moreover, the action amender resolves the channel allocation by an argmax over the K×M preference scores, which is non-differentiable; the gradient of the critic with respect to the pre-argmax scores is zero almost everywhere, so the channel-allocation component of the actor is not trained by (33) as written. Please state the mechanism used for the discrete component (e.g., a straight-through estimator or a REINFORCE term); without this, the algorithm is not fully specified and the learning behavior reported in Fig. 7 is hard to reproduce.
- [VII-C, Figs. 8-10 and the percentage gains in the text] The performance comparisons are reported as averages over five seeds without confidence intervals, error bars, or significance tests. The margins against D3PG-WCSI in Fig. 8 are 4.37%–6.39% at the quoted operating points, and the claims in the text are point estimates; with only five seeds and no variance information, the reader cannot assess whether the differences are statistically meaningful. Please report the per-seed spread (e.g., shaded regions or error bars) and, ideally, a paired test across the five seeds for the headline numbers in Sections VII-C3 and VII-C4.
minor comments (5)
- [VI-A, Action Space] The action space is said to contain 2K+M+1 elements, but the raw channel allocation x̃(t) is described in VI-B as a K×M preference matrix; the raw action dimension would then be KM+M+K+1. Please reconcile these statements.
- [VII-B, D3PG-WCSI benchmark] Please clarify whether D3PG-WCSI observes the delayed CSI and ignores the delay (so that its state is mismatched to the reward) or observes undelayed CSI; the current text is ambiguous, and the two readings have opposite implications for interpreting the gap in Fig. 8.
- [V, Remarks 2 and 3] The paper correctly acknowledges that the forward diffusion process is omitted and that training uses the RL objective rather than the standard diffusion loss; accordingly, the diffusion model functions as a stochastic policy parameterization. Positioning the novelty relative to existing diffusion-policy RL methods (e.g., Diffuser, DIPO, and the authors' prior work [27]) would make the contribution clearer.
- [IV, Eq. (19)] The term ½(P(t)Δ-E_U^th)² is dropped in P2 as a 'constant,' but it depends on the action through P(t); since the action space is bounded, it is more accurate to say that the term is bounded by a constant upper bound and that P2 minimizes the remaining part of the upper bound.
- [III-C, Eqs. (11)-(12)] For complex discrepancy terms δ, the notation δ² should read |δ|².
Circularity Check
No significant circularity: the Lyapunov decomposition, diffusion-based actor, and baseline comparisons are self-contained; the authors' self-citations are motivational and do not force the numerical results.
full rationale
The paper's derivation chain is not circular. The Lyapunov drift-plus-penalty bound in Appendix B follows directly from the queue update (15) via the standard inequality (max{a+b-c,0})^2 <= (a+b-c)^2, and the per-slot problem P2 in (20) is exactly the negative of the reward (30) used to train the D3PG agent. Thus the algorithm optimizes the claimed objective rather than fitting a parameter and then renaming the fit as a prediction. The delayed-CSI model in (10)-(12) is taken from the external reference [31]; while Eqs. (11)-(12) drop the cross term of |rho*g_hat + delta|^2 and may misrepresent the Gauss-Markov statistics, this is a modeling-accuracy issue, not a circular reduction. The numerical claims are obtained by comparing D3PG against independent baselines (DDPG, H-DDQN, and the D3PG-WCSI ablation) using the same environment and reward, so the reported gains are not forced by construction. The self-citations [10] and [27] are used to motivate diffusion-based DRL and to justify the assumption of ample cloud training resources, but the central Lyapunov algebra and the simulation results do not reduce to those citations. Remark 2 explicitly omits the forward process because the optimal solution is unavailable, but the reverse process is trained through RL against the reward (30) rather than by feeding the target solution back into the derivation, so no self-definitional loop is present. Overall, no load-bearing circular step was found.
Assumptions & free parameters
free parameters (5)
- Lyapunov weight V =
100 (swept in Fig. 11)
- Number of denoising steps I =
4
- Reward penalty Gamma_pen =
10
- Discount factor omega =
0.99
- Target network update rate tau =
0.005
assumptions (5)
- standard math Lyapunov drift-plus-penalty optimization enforces the long-term energy constraint through a virtual queue.
- domain assumption The delayed CSI is modeled by a first-order Gauss-Markov process with correlation J0(2*pi*f_c*s_rel*T_delay/c).
- domain assumption The LoS probability formula in Eq. (6) and the path-loss models in Eqs. (3)-(5) describe the air-to-ground channel.
- domain assumption The UAV propulsion power model in Eq. (13) with the parameters in Table II captures the dominant flight energy consumption.
- standard math P1 is NP-hard because it is an MINLP.
invented entities (1)
-
Virtual queue Q(t)
Cite this review
Pith. "Pith review of A Lyapunov-Guided Diffusion-Based Reinforcement Learning Approach for UAV-Assisted Vehicular Networks with Delayed CSI Feedback." pith.science (2026). https://pith.science/paper/7DFSUZKI
@misc{pith2026250720524,
author = {Pith},
title = {Pith review of: A Lyapunov-Guided Diffusion-Based Reinforcement Learning Approach for UAV-Assisted Vehicular Networks with Delayed CSI Feedback},
year = {2026},
howpublished = {\url{https://pith.science/paper/7DFSUZKI}},
note = {Machine review of arXiv:2507.20524}
}
read the original abstract
Low altitude uncrewed aerial vehicles (UAVs) are expected to facilitate the development of aerial-ground integrated intelligent transportation systems and unlocking the potential of the emerging low-altitude economy. However, several critical challenges persist, including the dynamic optimization of network resources and UAV trajectories, limited UAV endurance, and imperfect channel state information (CSI). In this paper, we offer new insights into low-altitude economy networking by exploring intelligent UAV-assisted vehicle-to-everything communication strategies aligned with UAV energy efficiency. Particularly, we formulate an optimization problem of joint channel allocation, power control, and flight altitude adjustment in UAV-assisted vehicular networks. Taking CSI feedback delay into account, our objective is to maximize the vehicle-to-UAV communication sum rate while satisfying the UAV's long-term energy constraint. To this end, we first leverage Lyapunov optimization to decompose the original long-term problem into a series of per-slot deterministic subproblems. We then propose a diffusion-based deep deterministic policy gradient (D3PG) algorithm, which innovatively integrates diffusion models to determine optimal channel allocation, power control, and flight altitude adjustment decisions. Through extensive simulations using real-world vehicle mobility traces, we demonstrate the superior performance of the proposed D3PG algorithm compared to existing benchmark solutions.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[27]
Z. Liu, H. Du, J. Lin, Z. Gao, L. Huang, S. Hosseinalipour, and D. Niyato, “DNN partitioning, task offloading, and resource allocation in dynamic vehicular networks: A Lyapunov-guided diffusion-based reinforcement learning approach,” IEEE Trans. Mobile Comput. , vol. 24, no. 3, pp. 1945–1962, 2025
work page 1945
-
[1]
Named data networking in vehicular ad hoc networks: State-of-the-art and challenges,
H. Khelifi, S. Luo, B. Nour, H. Moungla, Y . Faheem, R. Hussain, and A. Ksentini, “Named data networking in vehicular ad hoc networks: State-of-the-art and challenges,” IEEE Commun. Surveys & Tut. , vol. 22, no. 1, pp. 320–351, 2019
work page 2019
-
[2]
A comprehensive simulation platform for space-air-ground integrated network,
N. Cheng, W. Quan, W. Shi, H. Wu, Q. Ye, H. Zhou, W. Zhuang, X. Shen, and B. Bai, “A comprehensive simulation platform for space-air-ground integrated network,” IEEE Wireless Commun., vol. 27, no. 1, pp. 178–185, 2020
work page 2020
-
[3]
Intelligent task offloading in vehicular edge computing networks,
H. Guo, J. Liu, J. Ren, and Y . Zhang, “Intelligent task offloading in vehicular edge computing networks,” IEEE Wireless Commun., vol. 27, no. 4, pp. 126–132, 2020
work page 2020
-
[4]
Aerial-ground integrated vehicular networks: A UA V-vehicle collaboration perspective,
Y . He, D. Wang, F. Huang, R. Zhang, and L. Min, “Aerial-ground integrated vehicular networks: A UA V-vehicle collaboration perspective,” IEEE Trans. Intell. Transp. Syst. , vol. 25, no. 6, pp. 5154–5169, 2023
work page 2023
-
[5]
UA V-assisted task offloading in vehicular edge computing networks,
X. Dai, Z. Xiao, H. Jiang, and J. C. Lui, “UA V-assisted task offloading in vehicular edge computing networks,” IEEE Trans. Mobile Comput. , vol. 23, no. 4, pp. 2520–2534, 2023
work page 2023
-
[6]
M. Samir, D. Ebrahimi, C. Assi, S. Sharafeddine, and A. Ghrayeb, “Leveraging UA Vs for coverage in cell-free vehicular networks: A deep reinforcement learning approach,” IEEE Trans. Mobile Comput. , vol. 20, no. 9, pp. 2835–2847, 2020
work page 2020
-
[7]
RFID: Towards low latency and reliable DAG task scheduling over dynamic vehicular clouds,
Z. Liu, M. Liwang, S. Hosseinalipour, H. Dai, Z. Gao, and L. Huang, “RFID: Towards low latency and reliable DAG task scheduling over dynamic vehicular clouds,” IEEE Trans. Veh. Technol., vol. 72, no. 9, pp. 12 139–12 153, 2023
work page 2023
Show all 41 references
-
[8]
Energy-efficient resource allocation for UA V-assisted vehicular networks with spectrum sharing,
W. Qi, Q. Song, L. Guo, and A. Jamalipour, “Energy-efficient resource allocation for UA V-assisted vehicular networks with spectrum sharing,” IEEE Trans. Veh. Technol., vol. 71, no. 7, pp. 7691–7702, 2022
2022
-
[9]
Joint optimization of relay selection and transmission scheduling for UA V- aided mmwave vehicular networks,
J. Li, Y . Niu, H. Wu, B. Ai, R. He, N. Wang, and S. Chen, “Joint optimization of relay selection and transmission scheduling for UA V- aided mmwave vehicular networks,” IEEE Trans. Veh. Technol., vol. 72, no. 5, pp. 6322–6334, 2023
2023
-
[10]
Generative AI for Lyapunov optimization theory in UA V-based low- altitude economy networking,
Z. Liu, D. Niyato, J. Wang, G. Sun, L. Huang, Z. Gao, and X. Wang, “Generative AI for Lyapunov optimization theory in UA V-based low- altitude economy networking,” arXiv preprint arXiv:2501.15928 , 2025
2025 arXiv
-
[11]
Relaying data with joint optimization of energy and delay in cluster- based UA V-assisted vanets,
S. Mokhtari, N. Nouri, J. Abouei, A. Avokh, and K. N. Plataniotis, “Relaying data with joint optimization of energy and delay in cluster- based UA V-assisted vanets,”IEEE Internet Things J. , vol. 9, no. 23, pp. 24 541–24 559, 2022
2022
-
[12]
Space-air-ground integrated vehicular network for connected and automated vehicles: Challenges and solutions,
Z. Niu, X. S. Shen, Q. Zhang, and Y . Tang, “Space-air-ground integrated vehicular network for connected and automated vehicles: Challenges and solutions,” Intell. Converged Networks, vol. 1, no. 2, pp. 142–169, 2020
2020
-
[13]
Energy-aware 3D-deployment of UA V for IoV with highway interchange,
Z. Liao, Y . Ma, J. Huang, and J. Wang, “Energy-aware 3D-deployment of UA V for IoV with highway interchange,” IEEE Trans. Commun., vol. 71, no. 3, pp. 1536–1548, 2022
2022
-
[14]
Performance analysis and 3D position deployment for V2V-assisted UA V communications in vehicular networks,
B. Zhang, Z. He, Y . Feng, and Z. Han, “Performance analysis and 3D position deployment for V2V-assisted UA V communications in vehicular networks,” IEEE Trans. Veh. Technol., 2024
2024
-
[15]
GA- DRL: Graph neural network-augmented deep reinforcement learning for DAG task scheduling over dynamic vehicular clouds,
Z. Liu, L. Huang, Z. Gao, M. Luo, S. Hosseinalipour, and H. Dai, “GA- DRL: Graph neural network-augmented deep reinforcement learning for DAG task scheduling over dynamic vehicular clouds,” IEEE Trans. Netw. Service Manage., 2024
2024
-
[16]
OpenStreetMap: User-generated street maps,
M. Haklay and P. Weber, “OpenStreetMap: User-generated street maps,” IEEE Pervasive Comput., vol. 7, no. 4, pp. 12–18, 2008
2008
-
[17]
Microscopic traffic simulation using SUMO,
P. A. Lopez, M. Behrisch, L. Bieker-Walz, J. Erdmann, Y .-P. Fl ¨otter¨od, R. Hilbrich, L. L ¨ucken, J. Rummel, P. Wagner, and E. Wießner, “Microscopic traffic simulation using SUMO,” in Proc. Int. Conf. Intell. Transp. Syst. Ieee, 2018, pp. 2575–2582
2018
-
[18]
Joint resources allocation and 3D trajectory optimization for UA V-enabled space-air-ground integrated networks,
Z. Hu, F. Zeng, Z. Xiao, B. Fu, H. Jiang, H. Xiong, Y . Zhu, and M. Alazab, “Joint resources allocation and 3D trajectory optimization for UA V-enabled space-air-ground integrated networks,”IEEE Trans. Veh. Technol., vol. 72, no. 11, pp. 14 214–14 229, 2023
2023
-
[19]
Optimizing number, placement, and backhaul connectivity of multi-UA V networks,
J. Sabzehali, V . K. Shah, Q. Fan, B. Choudhury, L. Liu, and J. H. Reed, “Optimizing number, placement, and backhaul connectivity of multi-UA V networks,” IEEE Internet Things J. , vol. 9, no. 21, pp. 21 548–21 560, 2022
2022
-
[20]
Deployment and association of multiple UA Vs in UA V-assisted cellular networks with the knowledge of statistical user position,
L. Wang, H. Zhang, S. Guo, and D. Yuan, “Deployment and association of multiple UA Vs in UA V-assisted cellular networks with the knowledge of statistical user position,” IEEE Trans. Wireless Commun. , vol. 21, no. 8, pp. 6553–6567, 2022
2022
-
[21]
Multi-UA V trajectory and power optimization for cached UA V wireless networks with energy and content recharging- demand driven deep learning approach,
S. Chai and V . K. N. Lau, “Multi-UA V trajectory and power optimization for cached UA V wireless networks with energy and content recharging- demand driven deep learning approach,” IEEE J. Sel. Areas Commun. , vol. 39, no. 10, pp. 3208–3224, 2021
2021
-
[22]
A UA V-enabled data dissemination protocol with proactive caching and file sharing in V2X networks,
R. Zhang, R. Lu, X. Cheng, N. Wang, and L. Yang, “A UA V-enabled data dissemination protocol with proactive caching and file sharing in V2X networks,” IEEE Trans. Commun. , vol. 69, no. 6, pp. 3930–3942, 2021
2021
-
[23]
DRL-UTPS: DRL-based trajectory planning for unmanned aerial vehicles for data collection in dynamic IoT network,
R. Liu, Z. Qu, G. Huang, M. Dong, T. Wang, S. Zhang, and A. Liu, “DRL-UTPS: DRL-based trajectory planning for unmanned aerial vehicles for data collection in dynamic IoT network,” IEEE Trans. Intell. Vehicles, vol. 8, no. 2, pp. 1204–1218, 2023
2023
-
[24]
Path planning for cellular-connected UA V: A DRL solution with quantum-inspired experience replay,
Y . Li, A. H. Aghvami, and D. Dong, “Path planning for cellular-connected UA V: A DRL solution with quantum-inspired experience replay,” IEEE Trans. Wireless Commun., vol. 21, no. 10, pp. 7897–7912, 2022
2022
-
[25]
Deep reinforcement learning-based resource allocation in cooperative UA V-assisted wireless networks,
P. Luong, F. Gagnon, L.-N. Tran, and F. Labeau, “Deep reinforcement learning-based resource allocation in cooperative UA V-assisted wireless networks,” IEEE Trans. Wireless Commun., vol. 20, no. 11, pp. 7610– 7625, 2021
2021
-
[26]
UA V-assisted wireless cooperative communication and coded caching: A multiagent two-timescale DRL approach,
B. Tian, L. Wang, L. Xu, W. Pan, H. Wu, L. Li, and Z. Han, “UA V-assisted wireless cooperative communication and coded caching: A multiagent two-timescale DRL approach,” IEEE Trans. Mobile Comput. , vol. 23, no. 5, pp. 4389–4404, 2024
2024
-
[28]
Multiagent deep-reinforcement- learning-based resource allocation for heterogeneous QoS guarantees for vehicular networks,
J. Tian, Q. Liu, H. Zhang, and D. Wu, “Multiagent deep-reinforcement- learning-based resource allocation for heterogeneous QoS guarantees for vehicular networks,” IEEE Internet Things J. , vol. 9, no. 3, pp. 1683–1695, 2021
2021
-
[29]
A comprehensive survey on UA V communication channel modeling,
C. Yan, L. Fu, J. Zhang, and J. Wang, “A comprehensive survey on UA V communication channel modeling,” IEEE Access, vol. 7, pp. 107 769– 107 792, 2019
2019
-
[30]
Deep-reinforcement-learning- based mode selection and resource allocation for cellular V2X com- munications,
X. Zhang, M. Peng, S. Yan, and Y . Sun, “Deep-reinforcement-learning- based mode selection and resource allocation for cellular V2X com- munications,” IEEE Internet Things J. , vol. 7, no. 7, pp. 6380–6391, 2019
2019
-
[31]
Spectrum and power allocation for vehicular communications with delayed CSI feedback,
L. Liang, J. Kim, S. C. Jha, K. Sivanesan, and G. Y . Li, “Spectrum and power allocation for vehicular communications with delayed CSI feedback,” IEEE Wireless Commun. Letters , vol. 6, no. 4, pp. 458–461, 2017
2017
-
[32]
UA V-assisted content delivery in intelligent transportation systems-joint trajectory planning and cache management,
A. Al-Hilo, M. Samir, C. Assi, S. Sharafeddine, and D. Ebrahimi, “UA V-assisted content delivery in intelligent transportation systems-joint trajectory planning and cache management,” IEEE Trans. Intell. Transp. Syst., vol. 22, no. 8, pp. 5155–5167, 2020
2020
-
[33]
Resource allocation and 3D trajectory design for power-efficient IRS-assisted UA V- NOMA communications,
Y . Cai, Z. Wei, S. Hu, C. Liu, D. W. K. Ng, and J. Yuan, “Resource allocation and 3D trajectory design for power-efficient IRS-assisted UA V- NOMA communications,” IEEE Trans. Wireless Commun., vol. 21, no. 12, pp. 10 315–10 334, 2022
2022
-
[34]
Enhancing deep reinforcement learning: A tutorial on generative diffusion models in network optimization,
H. Du, R. Zhang, Y . Liu, J. Wang, Y . Lin, Z. Li, D. Niyato, J. Kang, Z. Xiong, S. Cui et al. , “Enhancing deep reinforcement learning: A tutorial on generative diffusion models in network optimization,” IEEE Commun. Surveys & Tut. , 2024
2024
-
[35]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. Int. Conf. Neural Inf. Process. Syst. , vol. 33, 2020, pp. 6840– 6851
2020
-
[36]
Improved DDPG based two-timescale multi-dimensional resource allocation for multi-access edge computing networks,
Q. Liu, H. Zhang, X. Zhang, and D. Yuan, “Improved DDPG based two-timescale multi-dimensional resource allocation for multi-access edge computing networks,” IEEE Trans. Veh. Technol., vol. 73, no. 6, pp. 9153–9158, 2024
2024
-
[37]
Adam: A method for stochastic optimization,
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980 , 2014
2014 arXiv
-
[38]
Online trajectory and resource optimization for stochastic UA V-enabled MEC systems,
Z. Yang, S. Bi, and Y .-J. A. Zhang, “Online trajectory and resource optimization for stochastic UA V-enabled MEC systems,” IEEE Trans. Wireless Commun., vol. 21, no. 7, pp. 5629–5643, 2022
2022
-
[39]
Continuous control with deep reinforcement learning,
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” arXiv preprint arXiv:1509.02971 , 2015
2015 arXiv
-
[40]
Deep reinforcement learning with double Q-learning,
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double Q-learning,” in Proc. Conf. Artif. Intell., vol. 30, no. 1, 2016
2016
-
[41]
A Lyapunov-based approach to joint optimization of resource allocation and 3-D trajectory for solar- powered UA V MEC systems,
X.-H. Lin, S. Bi, G. Su, and Y .-J. A. Zhang, “A Lyapunov-based approach to joint optimization of resource allocation and 3-D trajectory for solar- powered UA V MEC systems,” IEEE Internet Things J. , vol. 11, no. 11, pp. 20 797–20 815, 2024
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.