REVIEW 4 major objections 5 minor 29 references
Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read JUROR's central claim is that routing and UAV flight in a delay-tolerant network should be learned jointly, because each heading choice reshapes the contact graph that forwarding decisions face at the next step.
desk verdict A plausible joint UAV-routing RL framework whose headline comparison against routing-only baselines is confounded by control authority; worth reviewing, but the abstract needs to be softened until a UAV-enabled baseline is tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a factored partially observable Markov decision process with sequential motion–routing coupling: at each step the UAVs move first, then routing actions execute under the contact matrix $C^{(t)}$ their motion created, then TTLs decay and contacts rebuild. The mechanism carrying the argument is a shared CTDE–PPO policy with decentralized actors and a training-time centralized critic; each actor outputs a masked choice over $K+1$ forwarding candidates (or idle), and each UAV actor also selects one of eight heading bins, all trained by the per-step team reward. An optional multi-horizon LSTM hotspot predictor and hotspot-guided alignment shaping are auxiliary modules the paper finds traffic-dependent, and the default stack disables them.
What would settle it
Run the same simulator with a fixed heuristic UAV flight policy such as heading toward the current congestion centroid while using PRoPHET or MaxProp for forwarding; if that configuration matches JUROR's held-out delivery ratios, then the jointly learned heading-and-routing policy is not the source of the claimed gains.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that joint UAV–routing control can be cast as a factored, partially observable cooperative decision process and solved under centralized training with decentralized execution. The environment couples the two subsystems in both directions: UAV headings determine which forwarding contacts exist at the next step, and routing outcomes change the buffer-pressure stress fields that steer UAV headings. JUROR factorizes the policy into per-node routing heads and per-UAV heading heads sharing one PPO update, and it trains them with a per-step team reward that rewards deliveries and penalizes expiry, drops, congestion, heading switches, and idle replication. The optional multi-horizon LSTM hotspot predictor and hotspot-guided alignment shaping are auxiliary modules that the paper finds traffic-dependent; the default configuration disables both. The reported result is that this joint policy, evaluated on held-out simulators, delivers substantially more messages than PRoPHET, MaxProp, Fan DPUVR, and ICC Q-learning variants under all four traffic modes, while a deployment-near-real variant using only contact-limited inputs retains most of the peak benefit.
Load-bearing premise
The comparison assumes the partially reimplemented routing-only baselines are fair yardsticks, yet none of them controls UAV headings and all omit the full simulator transceiver stack, so part of the margin could come from extra control authority or from implementation differences rather than from the joint learning algorithm.
Editorial extensions
If this is right
- UAV trajectory planning should be treated as an action of the forwarding system: every heading choice enlarges or shrinks the set of contacts available to routing, so fixed flight paths leave delivery gains on the table.
- Decentralized execution is compatible with the joint policy: the deployment-near-real variant, which uses contact-limited UAV context instead of network-wide statistics, retains most of the peak delivery benefit, so the approach does not require a central controller at run time.
- The optional hotspot modules are not always helpful: adding the LSTM auxiliary or HGA shaping improves some bursty multi-source modes but hurts single-source and sustained-congestion modes, which is why the paper's default configuration disables both.
- Removing the controllable UAV fleet from the learned policy cuts peak held-out delivery from roughly 57.6 to 13.2 messages on M1, which the paper reads as evidence that joint motion–routing control is a structural prerequisite, not a marginal add-on.
- The advantage over routing-only baselines widens under heavier sustained load, where encounter-based scoring methods saturate their buffers.
Reading between the lines
- An implication the paper leaves implicit is that its reported margins over routing-only baselines may partly reflect the extra control authority JUROR has, since none of the baselines controls UAV headings; rerunning the same simulator with a fixed or heuristic UAV flight policy under PRoPHET or MaxProp routing would isolate the learning contribution.
- The evaluation uses one seed, one map, a disk contact model, and hand-tuned reward weights; whether the gains survive across seeds, maps, non-disk radio effects, and reward-weight perturbations is not established by the paper.
- The jointness result suggests a broader design principle: in any relay network where nodes control their own trajectories, route utility and motion control should share one objective and one credit signal, so applying the same factored CTDE recipe to surface or underwater mobile relays is a natural next test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes JUROR, a CTDE–PPO framework that jointly learns per-node opportunistic forwarding decisions and discrete UAV heading actions in a simulated delay-tolerant network. The system model is a factored partially observable MDP with a hand-crafted per-step team reward, decentralized actors, a centralized critic, and optional LSTM-based hotspot prediction and HGA shaping. Experiments on four traffic modes compare JUROR against PRoPHET, MaxProp, Fan DPUVR, and two ICC Q-learning variants, reporting higher held-out delivery ratios for JUROR. The paper also includes ablations over UAV presence, CTDE observation scope, LSTM auxiliary loss, and HGA.
Significance. If the central claim is established, the paper would be a useful contribution to UAV-assisted DTN routing: it formalizes joint motion–routing control as a factored POMDP, preserves decentralized execution, and evaluates on a non-trivial simulator with four traffic modes. The work is also honest about the optional nature of its hotspot auxiliaries, showing that they are traffic-dependent. However, the headline comparison is currently confounded because the baselines are routing-only and cannot control UAV headings, so the reported margin may come from extra control authority rather than from the learned joint policy. The paper's own ablation data show that removing UAVs drastically reduces delivery, which strengthens this concern rather than resolving it. The central claim is defensible but not yet evidenced.
major comments (4)
- [Sec. VI-C, Table VI] The comparison lacks a control for UAV heading authority. PRoPHET, MaxProp, Fan DPUVR, and the ICC variants are routing-only and do not make UAV heading decisions, while JUROR simultaneously controls both forwarding and discrete headings. Since the base−UAV ablation in Table IV shows that removing UAVs collapses delivery (M1 peak 13.2 vs 57.6), the margin in Table VI could be due to granting JUROR an extra control channel rather than to the learned joint policy. Add baselines that pair PRoPHET and MaxProp with fixed waypoints, random headings, and a simple stress-following heuristic for the same UAV fleet, and report delivery ratios for those controls.
- [Sec. VI-A, Sec. VI-C, Table VI] The abstract's claim of 'effective gains over PRoPHET and MaxProp' rests on Table VI, in which the JUROR rows are flagged as selected peak held-out references 'not directly comparable to the routing-only protocol,' while each baseline is a single trial without error bars. Peak-over-training values are not a stable estimator of policy quality. Report terminal or fixed-checkpoint metrics, averaged over multiple training seeds and evaluation episodes, for both JUROR and the baselines under the same episode distribution, and use those numbers for the headline claim.
- [Sec. VI-C] The baseline implementations are explicitly partial reimplementations: they omit The ONE transceiver stack and substitute simplified proxy metrics, yet no calibration against the original implementations or published results is provided. Lower baseline performance could therefore reflect implementation quality rather than protocol weakness. Validate the reimplementations on a standard benchmark or against published numerical results, and discuss what the omitted transceiver stack could change in the evaluated scenarios.
- [Sec. VI-D.1] The statement that the base−UAV ablation 'verifies joint routing-UA V optimization as the core source of JUROR's performance gains' overreaches: that ablation removes UAVs entirely, so it demonstrates the value of aerial relays, not the value of learned heading control. Separately test a fixed or heuristic heading policy with the same routing mechanisms to support the claimed decomposition.
minor comments (5)
- [Abstract] The sentence introducing JUROR has an unbalanced parenthesis: 'JUROR (Joint UAV flight and Opportunistic Routing, based on the proximal policy optimization (PPO) framework.' should close the first parenthesis.
- [Eq. (18)] In the delivery objective, α_s is described as a per-step idle penalty, but as written it is a constant added at every step; clarify how this term incentivizes routing activity and how its magnitude interacts with α_h I_h.
- [Sec. VI-A] All training runs use seed 42 only; reporting at least a small number of seeds for the default configuration would help assess PPO variance, especially because Table VI reports peak held-out values.
- [References] Reference [19] is listed as an author manuscript without a venue or DOI; if it is not published, mark it as a preprint and consider whether it should be a primary citation.
- [Sec. VI-A and Table III] The traffic-mode description says M2 injects about 4300 messages per episode while Sec. VI-C states 4321; align these numbers for consistency.
Circularity Check
No constructional circularity: JUROR's reward is explicitly engineered and evaluation is on held-out simulator episodes; baseline confounds are experimental, not circular.
full rationale
JUROR does not derive its claimed gains from fitted parameters or from a self-citation chain. The per-step team reward in Eq. (32) is an explicitly engineered linear combination of delivery, expiry, drop, buffer, and UAV-placement terms, and the held-out metrics (D_test, E_test, R_test) are measured on separate evaluation environments after training rather than being constructed from the reward itself. No parameter is fit to the evaluation quantity and then renamed as a prediction. The baseline comparison to PRoPHET and MaxProp is external: those protocols are independently described in the literature, and the paper candidly states that the baselines are 'partially reimplemented' and that selected JUROR peak rows are 'not directly comparable to the routing-only protocol.' Any concern that JUROR's margin partly reflects extra control authority over UAV headings is a fairness or correctness confound, not a circularity of the derivation chain. The paper also does not invoke a uniqueness theorem from the authors' prior work, and there are no author self-citations carrying a load-bearing premise. The optional hotspot modules are ablated and are explicitly disabled in the default stack, and the LSTM auxiliary loss is trained against replay-derived targets that are not the same as the final delivery objective. Overall, the derivation from MDP formulation to PPO objective to simulator evaluation is self-contained, and the reported gains are empirical claims subject to experimental scrutiny rather than reductions of the conclusion into the premises.
Assumptions & free parameters
free parameters (7)
- Delivery/queue reward weights alpha_d, alpha_e, alpha_dr, alpha_b, alpha_s, alpha_h =
4.0, -1.5, -0.5, -0.1, -0.03, -0.005
- UAV placement reward weights alpha_rho, alpha_c, alpha_x, alpha_k, alpha_f =
0.05, 0.015, 0.006, 0.003, 0.01
- HGA reward weights alpha_a, alpha_ad, alpha_ar =
default 0,0,0; ablation 0.003,0.012,0.004
- State-dependent gating thresholds g_s, g_m, g_e =
g_s=1 if D=0 else 0.3; g_m=1 if D=0 else 0.5; g_e=min(1,E/2)
- K (max routing candidates per node) =
not stated
- K_sigma (top stress neighbors in UAV navigation vector) =
not stated
- omega_ex (neighbor exchange decay factor) =
not stated
assumptions (5)
- domain assumption Ground mobility follows an exogenous road-constrained process independent of the policy.
- domain assumption Communication is a disk model with no interference, channel errors, or energy constraints, and GNSS localization is always available.
- domain assumption Each node can receive at most one replicated message per step and buffer overflow drops the oldest message.
- ad hoc to paper The partial reimplementations of PRoPHET, MaxProp, Fan DPUVR, and ICC Q-learning faithfully represent the original protocols' forwarding logic.
- ad hoc to paper PPO with a centralized critic and factored policy converges to a useful joint policy under the chosen reward and observation features.
Cite this review
Pith. "Pith review of Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks." pith.science (2026). https://pith.science/paper/MRYEZNMG
@misc{pith2026260804590,
author = {Pith},
title = {Pith review of: Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/MRYEZNMG}},
note = {Machine review of arXiv:2608.04590}
}
read the original abstract
The growing deployment of delay-tolerant networks (DTNs) has made store-carry-forward (SCF) communication indispensable under sparse connectivity. However, intermittent contacts, finite buffers, and limited message time-to-live (TTL) often give rise to sparse delivery and congestion, leading to substantial end-to-end performance degradation. To address this challenge, this study explores the joint optimization of decentralized opportunistic routing and controllable unmanned aerial vehicle (UAV) flight, aiming to enlarge future contacts through discrete UAV headings while enabling per-node replication under contact-limited observations. Building upon this architecture, we study cooperative factored routing--UAV control under centralized training and decentralized execution (CTDE) and propose JUROR (Joint UAV flight and Opportunistic Routing, based on the proximal policy optimization (PPO) framework. In our design, we first cast the problem as a factored partially observable Markov decision process with sequential motion--routing coupling and a per-step team reward; subsequently, decentralized actors act on local observations while a training-time critic uses global statistics, and an optional multi-horizon hotspot predictor provides auxiliary supervision. Simulation results over four traffic modes demonstrate effective gains over PRoPHET and MaxProp, while retaining contact-limited decentralized execution.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Routing Protocols in FANET for Disaster Area Networks: A Review,
Badawi, Salma, Ahmad, Norulhusna, Dziyauddin, Rudzidatul Akmam, Mohamed, Norliza, and Sam, Suriani Mohd, “Routing Protocols in FANET for Disaster Area Networks: A Review,”ASEAN Engineering Journal, vol. 15, no. 3, pp. 81–100, 2025, doi: 10.11113/aej.V15.22954
-
[2]
Probabilistic Rout- ing in Intermittently Connected Networks,
Lindgren, Anders, Doria, Avri, and Schel ´en, Olov, “Probabilistic Rout- ing in Intermittently Connected Networks,”ACM SIGMOBILE Mobile Computing and Communications Review, vol. 7, no. 3, pp. 19–26, 2003, doi: 10.1145/961268.961272
-
[3]
A Routing Protocol for UA V-Assisted Vehicular Delay Tolerant Networks,
Du, Zhaoyang, Wu, Celimuge, Yoshinaga, Tsutomu, Chen, Xianfu, Wang, Xiaoyan, Yau, Kok-Lim Alvin, and Ji, Yusheng, “A Routing Protocol for UA V-Assisted Vehicular Delay Tolerant Networks,”IEEE Open Journal of the Computer Society, vol. 2, pp. 85–98, 2021, doi: 10.1109/OJCS.2021.3054759
arXiv 2021
-
[4]
A Novel UA V-assisted V ANET Routing Protocol for Post-Disaster Emergency Communica- tions,
Fan, Zhijie, Zhang, Mansi, Cao, Yue, Liu, Zilong, Kaiwartya, Om- prakash, Javed, Yasir, and Hussain, Faisal Bashir, “A Novel UA V-assisted V ANET Routing Protocol for Post-Disaster Emergency Communica- tions,”IEEE Transactions on Network Science and Engineering, vol. 13, pp. 4863–4882, 2026, doi: 10.1109/TNSE.2025.3644432
-
[5]
Reinforcement Learning Based Routing in Net- works: Review and Classification of Approaches,
Mammeri, Zoubir, “Reinforcement Learning Based Routing in Net- works: Review and Classification of Approaches,”IEEE Access, vol. 7, pp. 55916–55950, 2019, doi: 10.1109/ACCESS.2019.2913776
-
[6]
Social and Geographical Routing for Vehicular Delay-Tolerant Networks,
Fernandes, In ˆes and Pereira, Paulo Rog ´erio, “Social and Geographical Routing for Vehicular Delay-Tolerant Networks,”Proc. Int. Young Engi- neers Forum Electr. Comput. Eng. (YEF-ECE), 2025, doi: 10.1109/YEF- ECE66503.2025.11117524
-
[7]
Sammou, El Mastapha, “PF-DTN: Predictive Routing for Intelligent Delay Tolerant Networks Using RNN-LSTM Deep Learning With Monte Carlo Dropout Uncertainty Estimation and Hybrid Deterministic, Probabilistic, and Uncertain Routing Strategies,”IEEE Access, vol. 14, pp. 10841–10859, 2026, doi: 10.1109/ACCESS.2026.3655507
-
[8]
Enhancing DTN Routing Strategies with Deep Reinforcement Learning in Disaster Recovery Networks,
Wen, Xin and Tan, Long, “Enhancing DTN Routing Strategies with Deep Reinforcement Learning in Disaster Recovery Networks,”Proc. Int. Conf. Frontier Technol. Inf. Comput. (ICFTIC), pp. 450–456, 2024, doi: 10.1109/ICFTIC64248.2024.10913424
arXiv 2024
Show all 29 references
-
[9]
A Noble Route Discovery Technique in UA V Based Delay Tolerant Network for Disaster Management,
Chakrabarti, Chandrima, “A Noble Route Discovery Technique in UA V Based Delay Tolerant Network for Disaster Management,”Proc. Int. Conf. Res. Methodol. Knowl. Manag., Artif. Intell. Telecommun. Eng. (RMKMATE), 2025, doi: 10.1109/RMKMATE64874.2025.11042440
2025
-
[10]
Reinforcement Learning-Based Routing Protocol for Opportunistic Networks,
Dhurandher, Sanjay K., Singh, Jagdeep, Woungang, Isaac, Srivastava, Sanjeev, and Rodrigues, Joel J. P. C., “Reinforcement Learning-Based Routing Protocol for Opportunistic Networks,”Proc. IEEE Int. Conf. Commun. (ICC), 2020, doi: 10.1109/ICC40277.2020.9149039
2020
-
[11]
An Improved Spray and Wait Algorithm Based on Q-learning in Delay Tolerant Network,
Yao, Lei, Bai, Xiangyu, and Zhou, Kexin, “An Improved Spray and Wait Algorithm Based on Q-learning in Delay Tolerant Network,” Proc. Int. Joint Conf. Neural Netw. (IJCNN), pp. 1–8, 2024, doi: 10.1109/IJCNN60899.2024.10650780
2024
-
[12]
Performance Evaluation for Q-Learning Based Anycast Routing Protocol in Unmanned Aerial Vehicle Networks with Multiple Base Stations,
Xiang, Yuhong, Gao, Shuai, Wang, Hongchao, Yang, Dong, Zhang, Yuming, and Zhang, Hongke, “Performance Evaluation for Q-Learning Based Anycast Routing Protocol in Unmanned Aerial Vehicle Networks with Multiple Base Stations,”Ad Hoc Networks, vol. 168, pp. 103719, 2025, doi: 10....
2025
-
[13]
Delay-Tolerant Multi-Agent DRL for Trajectory Plan- ning and Transmission Control in UA V-Assisted Wireless Networks,
Fan, Zesong, Gong, Shimin, Long, Yusi, Li, Lanhua, Gu, Bo, and Luong, Nguyen Cong, “Delay-Tolerant Multi-Agent DRL for Trajectory Plan- ning and Transmission Control in UA V-Assisted Wireless Networks,” Proc. IEEE Veh. Technol. Conf. (VTC Spring), pp. 1–5, 2024, doi: 10.1109/V...
2024
-
[14]
Multi-Agent Actor-Critic for Mixed Cooperative- Competitive Environments,
Lowe, Ryan, Wu, Yi, Tamar, Aviv, Harb, Jean, Abbeel, Pieter, and Mordatch, Igor, “Multi-Agent Actor-Critic for Mixed Cooperative- Competitive Environments,”Advances in Neural Information Processing Systems, vol. 30, 2017
2017
-
[15]
Counterfactual Multi-Agent Policy Gradients,
Foerster, Jakob N., Farquhar, Gregory, Afouras, Triantafyllos, Nardelli, Nantas, and Whiteson, Shimon, “Counterfactual Multi-Agent Policy Gradients,”Proceedings of the AAAI Conference on Artificial Intelli- gence, vol. 32, no. 1, 2018
2018
-
[16]
Proximal Policy Optimization Algorithms,
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg, “Proximal Policy Optimization Algorithms,”arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[17]
Routing for Disruption Tolerant Networks: Taxonomy and Design,
Spyropoulos, Thrasyvoulos, Psounis, Konstantinos, and Raghavendra, C. S., “Routing for Disruption Tolerant Networks: Taxonomy and Design,”Wireless Communications, vol. 16, no. 6, pp. 39–46, 2009, doi: 10.1109/MWC.2009.5302297
2009
-
[18]
MaxProp: Routing for Vehicle-Based Disruption-Tolerant Networks,
Burgess, John, Gallagher, Brian, Jensen, David, and Levine, Brian Neil, “MaxProp: Routing for Vehicle-Based Disruption-Tolerant Networks,” Proc. IEEE INFOCOM, 2006, doi: 10.1109/INFOCOM.2006.228
2006 doi
-
[19]
Epidemic Routing Optimization Algorithm Based on Node Cache Status,
Zhang, Yu, “Epidemic Routing Optimization Algorithm Based on Node Cache Status,”EBNCS flooding-family variant with cache-state classifi- cation and ACK cleanup; author manuscript, 2025
2025
-
[20]
SR-SAAD: A Social Rank-Based Routing Protocol for Enhanced Efficiency in Delay Tolerant Networks,
Ullah, Saif, Muhammad, Asif, Ali, Zulfiqar, Waqar, Muhammad, and Kim, Ajung, “SR-SAAD: A Social Rank-Based Routing Protocol for Enhanced Efficiency in Delay Tolerant Networks,”IET Communications, vol. 19, no. 1, pp. e70099, 2025, doi: 10.1049/cmu2.70099
2025 doi
-
[21]
Comparing Statistical, Analytical, and Learning-Based Rout- ing Approaches for Delay-Tolerant Networks,
D’Argenio, Pedro R., Fraire, Juan, Hartmanns, Arnd, and Raverta, Fernando, “Comparing Statistical, Analytical, and Learning-Based Rout- ing Approaches for Delay-Tolerant Networks,”ACM Transactions on Modeling and Computer Simulation, vol. 35, no. 2, pp. 1–26, 2025, doi: 10.114...
2025 doi
-
[22]
A Constructive Analysis on Machine Learning Integration in High Delay Tolerant Networking (HDTN),
Salam, Mohammad Abdus, Saif, A. F. M. Saifuddin, Katroju, Priya Hamsa, and Kassouf-Short, Robert, “A Constructive Analysis on Machine Learning Integration in High Delay Tolerant Networking (HDTN),”Proc. Int. Conf. Adv. Commun. Technol. Netw. (CommNet), pp. 1–6, 2024, doi: 10.1...
2024
-
[23]
A Review of Research on W ANG AND Y ANG: JUROR FOR DELAY-TOLERANT NETWORKS 16 Routing Protocols for Unmanned Aerial Vehicle Cluster Self-Organizing Networks,
Zhou, Xuan, Lin, Haitao, and Chen, Jin, “A Review of Research on W ANG AND Y ANG: JUROR FOR DELAY-TOLERANT NETWORKS 16 Routing Protocols for Unmanned Aerial Vehicle Cluster Self-Organizing Networks,”Proc. Int. Conf. Intell. Syst., Commun. Comput. Netw. (ISCCN), pp. 246–255, 20...
2025
-
[24]
Distributed Routing and Data Scheduling in IPNs With GNN-Based Multiagent DRL,
Zhou, Xixuan, Xu, Ruoyu, Tian, Xiaojian, Zhang, Yueyue, Liang, Yu, Chen, Xiaoliang, and Zhu, Zuqing, “Distributed Routing and Data Scheduling in IPNs With GNN-Based Multiagent DRL,”IEEE Internet of Things Journal, vol. 12, no. 12, pp. 21565–21576, 2025, doi: 10.1109/JIOT.2025.3547341
2025
-
[25]
On the Scaling of Reliable Interplanetary Networks with Deep Reinforce- ment Learning,
Tian, Xiaojian, Chen, Xiaoliang, Zhou, Xixuan, and Zhu, Zuqing, “On the Scaling of Reliable Interplanetary Networks with Deep Reinforce- ment Learning,”Proc. Int. Conf. Design Rel. Commun. Netw. (DRCN), 2025, doi: 10.1109/DRCN65040.2025.11046163
2025
-
[26]
Rein- forcement Learning with Unsupervised Auxiliary Tasks,
Jaderberg, Max, Mnih, V olodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z., Silver, David, and Kavukcuoglu, Koray, “Rein- forcement Learning with Unsupervised Auxiliary Tasks,”International Conference on Learning Representations, 2017
2017
-
[27]
OpenAI Gym,
Brockman, Greg, Cheung, Vicki, Petrov, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech, “OpenAI Gym,” arXiv preprint arXiv:1606.01540, 2016
2016 arXiv
-
[28]
Tianshou: A Highly Mod- ularized Deep Reinforcement Learning Library,
Weng, Jiayi, Chen, Huan, Yan, Dong, You, Kaichao, Duburcq, Antoine, Zhang, Minghao, Su, Hang, and Zhu, Jun, “Tianshou: A Highly Mod- ularized Deep Reinforcement Learning Library,”Journal of Machine Learning Research, vol. 22, no. 267, pp. 1–6, 2021
2021
-
[29]
The ONE Simulator for DTN Protocol Evaluation,
Ker ¨anen, Ari, Ott, J ¨org, and K ¨arkk¨ainen, Teemu, “The ONE Simulator for DTN Protocol Evaluation,”Proceedings of the 2nd International Conference on Simulation Tools and Techniques, pp. 1–10, 2009
2009
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.