REVIEW 3 major objections 1 minor 60 references
Achieving Fair-Effective Communications and Robustness in Underwater Acoustic Sensor Networks: A Semi-Cooperative Approach
T0 review · 3 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims SECOPA, a distributed multi-agent reinforcement learning approach, lets each underwater node choose transmission power to meet its own QoS and improve global fair-effective performance, robust to time-varying channels and n
desk verdict As submitted, this arXiv paper is two different documents: the abstract describes SECOPA, an underwater MARL power-allocation method, while the body is a demographic forecasting paper about Estonia—so the claimed work cannot be reviewed at all. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is SECOPA, a distributed multi-agent reinforcement learning (MARL) approach to transmission-power allocation. The mechanism it proposes: each node independently chooses its transmit power to optimize a reward that couples its own QoS satisfaction with a global fair-effective communication objective, and the training environment is intentionally made imperfect (time-varying channels, node failures) so the learned policies become robust. The supplied text does not present the underlying equations, state space, reward formulation, or training algorithm.
What would settle it
Deploy the learned power-allocation policies in a high-fidelity underwater acoustic simulator (or sea trial) whose channel model is statistically matched to measured ocean environments, induce a random node failure, and check whether per-node QoS and the global fair-effectiveness metric remain within the ranges claimed numerically; a significant degradation would falsify the robustness claim.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that a semi-cooperative power-allocation policy (SECOPA), learned by distributed multi-agent reinforcement learning, can make each underwater acoustic sensor node meet its Quality-of-Service requirements while the network as a whole achieves fair-effective communication, and that training in deliberately imperfect environments—time-varying acoustic channels and unexpected node failures—produces policies robust to those imperfections. The abstract asserts numerical validation of this behavior. The supplied full text, however, is an unrelated paper on probabilistic population forecasting, so the claimed method, equations, simulations, and resu
Load-bearing premise
The robustness claim rests on the premise that the simulated 'imperfect environments' used in training faithfully represent real underwater acoustic channels and unexpected node failures; if the simulator diverges from the field, the learned policies' fair-effective behavior need not transfer to deployed networks.
Editorial extensions
If this is right
- If SECOPA works as claimed, each underwater node can set its own transmit power without a central controller, preserving its QoS while the network as a whole stays fair and effective.
- Policies trained in imperfect environments would keep underwater networks functional when acoustic channels change rapidly or when some nodes suddenly fail.
- The semi-cooperative formulation offers a middle path between fully cooperative and fully selfish power control, which could be exported to other wireless systems with conflicting individual and network objectives.
- The claimed numerical validation, if reproducible, would give network designers a practical way to choose transmission powers under uncertainty.
Reading between the lines
- My inference: the 'imperfect environments' phrase points to training with simulated channel non-stationarity and injected node faults; if so, the robustness guarantee is bounded by how well those simulated faults match real failure modes, a testable modeling question.
- My inference: the paper's abstract does not define its fairness metric; without a formal measure connecting per-node QoS to a global fair-effective index, the claim is hard to quantify across different network topologies.
- My inference: because the supplied full text does not contain the method, a reader seeking to verify the claim should look for the actual version of this paper's technical sections or supplementary code.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript under review presents an abstract claiming a distributed multi-agent reinforcement learning power allocation approach (SECOPA) for underwater acoustic sensor networks, with two objectives (individual QoS and global fair-effective communication) and robustness via training in imperfect environments, and states that numerical results validate the approach. However, the body of the manuscript (Sections 1–5 and the Appendix) is an unrelated demographic forecasting paper, 'A new approach to probabilistic population forecasting with an application to Estonia' by Swanson and Tayman. None of the key concepts from the abstract—SECOPA, underwater acoustic sensor networks, MARL, power allocation, QoS, node failures—appear in the body. Thus, as submitted, the manuscript provides no algorithm, no model, no simulation, and no numerical results supporting its abstract.
Significance. If the claimed SECOPA contribution were present, it could be significant for underwater acoustic networks; the idea of semi-cooperative distributed power allocation under imperfect channels and node failures is a plausible research direction. However, the submitted artifact contains none of the claimed work. There is no derivable contribution, no testable prediction, and no reproducible code or proofs to evaluate. The demographic forecasting text in the body is a self-contained paper on a different topic, but it does not substantiate the abstract and is outside the scope of the claimed networking contribution. Consequently, the significance of the claimed result cannot be assessed, and the manuscript in its current form has no scientific content matching its abstract.
major comments (3)
- [Full text (Sections 1–5 and Appendix)] The body is an entirely different manuscript. The abstract's central claim—that 'this paper presents a SEmi-COoperative Power Allocation approach (SECOPA)'—is unsubstantiated because the body never defines SECOPA, no MARL formulation is given (state/action/reward design), and no power-allocation algorithm or equations appear. The only 'approach' described is the Espenshade–Tayman method for translating ARIMA confidence intervals onto cohort-component population forecasts (Sections 2–3). This is a load-bearing absence: the claimed central contribution is missing.
- [Abstract (validation claim)] The abstract states 'Numerical results are presented to validate our proposed approach.' No such results are in the submitted text. The only numerical tables (Tables 1.A–1.F and the Appendix) contain Estonian population forecasts and ARIMA diagnostics, not underwater network simulations. There is no comparison against baselines, no fairness/QoS metrics, and no evaluation under time-varying acoustic channels or node failures. The stated validation is therefore unsupported.
- [Abstract (robustness claim)] The second objective—'advanced training algorithms are developed to provide imperfect environments for training robust models'—cannot be inspected because no environment description, channel model (propagation loss, multipath, Doppler), node-failure model, or training procedure is provided anywhere in the manuscript. The robustness claim is load-bearing for the paper's contribution, and its complete absence is a separate deficiency from the missing algorithm and results.
minor comments (1)
- [References] Several reference entries contain garbled characters (e.g., 'Alkema ������ ������' appears multiple times), suggesting OCR corruption in the submitted PDF; these should be corrected in any version. This is a presentation issue independent of the central mismatch.
Circularity Check
No circularity identified; the submitted full text does not contain the claimed SECOPA derivation, so there is no derivation chain that could reduce to its own inputs.
full rationale
The abstract claims a SEmi-COoperative Power Allocation approach (SECOPA) for underwater acoustic sensor networks, validated by numerical results. However, the full text is a demographic forecasting paper titled 'A new approach to probabilistic population forecasting with an application to Estonia' by Swanson and Tayman. It contains no equations, algorithm descriptions, simulation setup, or numerical results related to SECOPA, MARL, power allocation, QoS, underwater acoustics, or node failures. Consequently, there is no derivation chain to inspect for circularity. The central claim is unsupported as submitted, but unsupported is not the same as circular. Under the hard rules, circularity can be flagged only when the paper itself exhibits a specific reduction (e.g., Eq. X = Eq. Y by construction, or a fitted parameter renamed as a prediction). No such reduction can be quoted because the claimed paper is absent. The potential same-distribution circularity of training and validating on the same simulated 'imperfect environments' cannot be confirmed without the methods text, and speculation about it is not permitted. Therefore, the circularity score is 0, with the caveat that the manuscript fails for other reasons (content mismatch, missing evidence).
Assumptions & free parameters
assumptions (3)
- domain assumption Underwater acoustic sensor networks are imperfect and energy-constrained, with time-varying channels.
- ad hoc to paper A per-node 'semi-cooperative' power allocation objective can simultaneously satisfy individual QoS and global fair-effective communication.
- ad hoc to paper Training in simulated imperfect environments produces policies robust to real time-varying acoustic channels and unexpected node failures.
invented entities (1)
-
SECOPA (SEmi-COoperative Power Allocation approach)
Cite this review
Pith. "Pith review of Achieving Fair-Effective Communications and Robustness in Underwater Acoustic Sensor Networks: A Semi-Cooperative Approach." pith.science (2026). https://pith.science/paper/TAAW3RYX
@misc{pith2026250807578,
author = {Pith},
title = {Pith review of: Achieving Fair-Effective Communications and Robustness in Underwater Acoustic Sensor Networks: A Semi-Cooperative Approach},
year = {2026},
howpublished = {\url{https://pith.science/paper/TAAW3RYX}},
note = {Machine review of arXiv:2508.07578}
}
read the original abstract
This paper investigates the fair-effective communication and robustness in imperfect and energy-constrained underwater acoustic sensor networks (IC-UASNs). Specifically, we investigate the impact of unexpected node malfunctions on the network performance under the time-varying acoustic channels. Each node is expected to satisfy Quality of Service (QoS) requirements. However, achieving individual QoS requirements may interfere with other concurrent communications. Underwater nodes rely excessively on the rationality of other underwater nodes when guided by fully cooperative approaches, making it difficult to seek a trade-off between individual QoS and global fair-effective communications under imperfect conditions. Therefore, this paper presents a SEmi-COoperative Power Allocation approach (SECOPA) that achieves fair-effective communication and robustness in IC-UASNs. The approach is distributed multi-agent reinforcement learning (MARL)-based, and the objectives are twofold. On the one hand, each intelligent node individually decides the transmission power to simultaneously optimize individual and global performance. On the other hand, advanced training algorithms are developed to provide imperfect environments for training robust models that can adapt to the time-varying acoustic channels and handle unexpected node failures in the network. Numerical results are presented to validate our proposed approach.
Reference graph
Works this paper leans on
-
[1]
11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEauthordefaulttextstyle \@IEEEcompsocnotconfonly \@IEEEauthor...
arXiv 2018
-
[2]
J. Liu, X. Du, J. Cui, M. Pan, and D. Wei, ``Task-oriented intelligent networking architecture for the space--air--ground--aqua integrated network,'' IEEE Internet of Things Journal, vol. 7, no. 6, pp. 5345--5358, 2020
work page 2020
-
[3]
Z. Mohammadi, M. Soleimanpour-Moghadam, M. Askarizadeh, and S. Talebi, ``Increasing the lifetime of underwater acoustic sensor networks: Difference convex approach,'' IEEE Systems Journal, vol. 14, no. 3, pp. 3214--3224, 2020
work page 2020
-
[4]
Stojanovic and J
M. Stojanovic and J. Preisig, ``Underwater acoustic communication channels: Propagation models and statistical characterization,'' IEEE communications magazine, vol. 47, no. 1, pp. 84--89, 2009
2009
-
[5]
Y. Noh, U. Lee, S. Han, P. Wang, D. Torres, J. Kim, and M. Gerla, ``Dots: A propagation delay-aware opportunistic mac protocol for mobile underwater networks,'' IEEE Transactions on Mobile Computing, vol. 13, no. 4, pp. 766--782, 2014
work page 2014
-
[6]
Gupta and P
P. Gupta and P. R. Kumar, ``The capacity of wireless networks,'' IEEE Transactions on information theory, vol. 46, no. 2, p. 22, 2000
2000
-
[7]
A. A. Syed, W. Ye, J. Heidemann, and B. Krishnamachari, ``Understanding spatio-temporal uncertainty in medium access with aloha protocols,'' in Proceedings of the second workshop on Underwater networks, 2007, pp. 41--48
work page 2007
-
[8]
Z. Guan, T. Melodia, and D. Yuan, ``Stochastic channel access for underwater acoustic networks with spatial and temporal interference uncertainty,'' in Proceedings of the Seventh ACM International Conference on Underwater Networks and Systems, 2012, pp. 1--8
work page 2012
Show all 60 references
-
[9]
C. Li, Y. Xu, C. Xu, Z. An, B. Diao, and X. Li, ``Dtmac: A delay tolerant mac protocol for underwater wireless sensor networks,'' IEEE Sensors Journal, vol. 16, no. 11, pp. 4137--4146, 2015
2015
-
[10]
Zhong, F
X. Zhong, F. Ji, F. Chen, Q. Guan, and H. Yu, ``A new acoustic channel interference model for 3-d underwater acoustic sensor networks and throughput analysis,'' IEEE Internet of Things Journal, vol. 7, no. 10, pp. 9930--9942, 2020
2020
-
[11]
Z. Guan, H. Kulhandjian, and T. Melodia, ``Stochastic channel access in underwater networks with statistical interference modeling,'' IEEE Transactions on Mobile Computing, vol. 20, no. 10, pp. 3020--3033, 2020
2020
-
[12]
Y. Shi, Y. T. Hou, J. Liu, and S. Kompella, ``Bridging the gap between protocol and physical models for wireless networks,'' IEEE Transactions on Mobile Computing, vol. 12, no. 7, pp. 1404--1416, 2012
2012
-
[13]
Y. Su, Y. Zhu, H. Mo, J.-H. Cui, and Z. Jin, ``A joint power control and rate adaptation mac protocol for underwater sensor networks,'' Ad Hoc Networks, vol. 26, pp. 36--49, 2015
2015
-
[14]
L. Qian, S. Zhang, M. Liu, and Q. Zhang, ``A maca-based power control mac protocol for underwater wireless sensor networks,'' in 2016 IEEE/OES China Ocean Acoustics (COA). 1em plus 0.5em minus 0.4em IEEE, 2016, pp. 1--8
2016
-
[15]
C. Zhen, Y. Feng, D. Nie, J. Zhang, and G. Ning, ``Transmission power allocation for underwater acoustic multicarrier-cdma communication networks based on genetic algorithm,'' in OCEANS 2016-Shanghai. 1em plus 0.5em minus 0.4em IEEE, 2016, pp. 1--4
2016
-
[16]
Mohsan, H
S. Mohsan, H. Amjad, A. Mazinani, S. Shahzad, M. Khan, A. Islam, A. Mahmood, and A. Soban, ``Investigating transmission power control strategy for underwater wireless sensor networks,'' Int. J. Adv. Comput. Sci. Appl, vol. 11, pp. 281--285, 2020
2020
-
[17]
Zhang, Y
T. Zhang, Y. Gou, J. Liu, T. Yang, and J.-H. Cui, ``Udarmf: An underwater distributed and adaptive resource management framework,'' IEEE Internet of Things Journal, 2021
2021
-
[18]
Y. Gou, T. Zhang, J. Liu, T. Yang, S. Song, and J.-H. Cui, ``Achieving time-sharing and spatial-reuse underwater wireless sensor networks with communication fairness: A distributed deep multi-agent reinforcement learning approach,'' in The 15th International Conference on Unde...
2021
-
[19]
Y. Gou, T. Zhang, T. Yang, J. Liu, S. Song, and J.-H. Cui, ``A deep marl-based power management strategy for improving the fair reuse of uwsns,'' IEEE Internet of Things Journal, 2022
2022
-
[20]
J. M. Jornet and M. Stojanovic, ``Distributed power control for underwater acoustic networks,'' in OCEANS 2008. 1em plus 0.5em minus 0.4em IEEE, 2008, pp. 1--7
2008
-
[21]
Sathiaseelan and G
A. Sathiaseelan and G. Fairhurst, ``Multimedia congestion control for broadband wireless networks,'' in 2007 16th IST Mobile and Wireless Communications Summit. 1em plus 0.5em minus 0.4em IEEE, 2007, pp. 1--5
2007
-
[22]
Huaizhou, R
S. Huaizhou, R. V. Prasad, E. Onur, and I. Niemegeers, ``Fairness in wireless networks: Issues, measures and challenges,'' IEEE Communications Surveys & Tutorials, vol. 16, no. 1, pp. 5--24, 2013
2013
-
[23]
Diamant, P
R. Diamant, P. Casari, F. Campagnaro, and M. Zorzi, ``Leveraging the near--far effect for improved spatial-reuse scheduling in underwater acoustic networks,'' IEEE Transactions on Wireless Communications, vol. 16, no. 3, pp. 1480--1493, 2016
2016
-
[24]
Y. Luo, L. Pu, Z. Peng, Z. Zhou, and J.-H. Cui, ``An efficient mac protocol for underwater multi-user uplink communication networks,'' Ad Hoc Networks, vol. 34, pp. 75--91, 2015
2015
-
[25]
Y. Xu, V. Mahendran, W. Guo, and S. Radhakrishnan, ``Fairness in fog networks: Achieving fair throughput performance in mqtt-based iots,'' in 2017 14th IEEE Annual Consumer Communications & Networking Conference (CCNC). 1em plus 0.5em minus 0.4em IEEE, 2017, pp. 191--196
2017
-
[26]
J. Zhao, Y. Zhao, W. Wang, M. Yang, X. Hu, W. Zhou, J. Hao, and H. Li, ``Coach-assisted multi-agent reinforcement learning framework for unexpected crashed agents,'' Frontiers of Information Technology & Electronic Engineering, vol. 23, no. 7, pp. 1032--1042, 2022
2022
-
[27]
Young and R
M. Young and R. Boutaba, ``Overcoming adversaries in sensor networks: A survey of theoretical models and algorithmic approaches for tolerating malicious interference,'' IEEE Communications Surveys & Tutorials, vol. 13, no. 4, pp. 617--641, 2011
2011
-
[28]
ElBatt and A
T. ElBatt and A. Ephremides, ``Joint scheduling and power control for wireless ad hoc networks,'' IEEE Transactions on Wireless communications, vol. 3, no. 1, pp. 74--85, 2004
2004
-
[29]
Jain, The art of computer systems performance analysis
R. Jain, The art of computer systems performance analysis. 1em plus 0.5em minus 0.4em john wiley & sons, 2008
2008
-
[30]
Xie and J.-H
P. Xie and J.-H. Cui, ``R-mac: An energy-efficient mac protocol for underwater sensor networks,'' in International Conference on Wireless Algorithms, Systems and Applications (WASA 2007). 1em plus 0.5em minus 0.4em IEEE, 2007, pp. 187--198
2007
-
[31]
A. A. Syed, W. Ye, and J. Heidemann, ``T-lohi: A new class of mac protocols for underwater acoustic sensor networks,'' in IEEE INFOCOM 2008-The 27th Conference on Computer Communications. 1em plus 0.5em minus 0.4em IEEE, 2008, pp. 231--235
2008
-
[32]
Bouabdallah, C
F. Bouabdallah, C. Zidi, and R. Boutaba, ``Joint routing and energy management in underwater acoustic sensor networks,'' IEEE Transactions on Network and Service Management, vol. 14, no. 2, pp. 456--471, 2017
2017
-
[33]
H. Wang, Y. Li, and J. Qian, ``Self-adaptive resource allocation in underwater acoustic interference channel: A reinforcement learning approach,'' IEEE Internet of Things Journal, vol. 7, no. 4, pp. 2816--2827, 2019
2019
-
[34]
X. Ye, Y. Yu, and L. Fu, ``Deep reinforcement learning based mac protocol for underwater acoustic networks,'' IEEE Transactions on Mobile Computing, vol. 21, no. 5, pp. 1625--1638, 2020
2020
-
[35]
Zhang, Scaling multi-agent learning in complex environments
C. Zhang, Scaling multi-agent learning in complex environments. 1em plus 0.5em minus 0.4em University of Massachusetts Amherst, 2011
2011
-
[36]
J. Lin, K. Dzeparoska, S. Q. Zhang, A. Leon-Garcia, and N. Papernot, ``On the robustness of cooperative multi-agent reinforcement learning,'' in 2020 IEEE Security and Privacy Workshops (SPW). 1em plus 0.5em minus 0.4em IEEE, 2020, pp. 62--68
2020
-
[37]
Chefrour, ``One-way delay measurement from traditional networks to sdn: A survey,'' ACM Computing Surveys (CSUR), vol
D. Chefrour, ``One-way delay measurement from traditional networks to sdn: A survey,'' ACM Computing Surveys (CSUR), vol. 54, no. 7, pp. 1--35, 2021
2021
-
[38]
Song and P.-Y
Y. Song and P.-Y. Kong, ``Optimizing design and performance of underwater acoustic sensor networks with 3d topology,'' IEEE Transactions on Mobile Computing, vol. 19, no. 7, pp. 1689--1701, 2019
2019
-
[39]
Y. He, G. Han, J. Jiang, H. Wang, and M. Martinez-Garcia, ``A trust update mechanism based on reinforcement learning in underwater acoustic sensor networks,'' IEEE Transactions on Mobile Computing, 2020
2020
-
[40]
Stojanovic, ``On the relationship between capacity and distance in an underwater acoustic communication channel,'' ACM SIGMOBILE Mobile Computing and Communications Review, vol
M. Stojanovic, ``On the relationship between capacity and distance in an underwater acoustic communication channel,'' ACM SIGMOBILE Mobile Computing and Communications Review, vol. 11, no. 4, pp. 34--43, 2007
2007
-
[41]
Wills, W
J. Wills, W. Ye, and J. Heidemann, ``Low-power acoustic modem for dense underwater sensor networks,'' in Proceedings of the 1st International Workshop on Underwater Networks, 2006, pp. 79--85
2006
-
[42]
R. J. Urick, ``Principles of underwater sound-2,'' 1975
1975
-
[43]
Pompili, T
D. Pompili, T. Melodia, and I. F. Akyildiz, ``A cdma-based medium access control for underwater acoustic sensor networks,'' IEEE Transactions on Wireless Communications, vol. 8, no. 4, pp. 1899--1909, 2009
1909
-
[44]
W. Bai, H. Wang, X. Shen, and R. Zhao, ``Link scheduling method for underwater acoustic sensor networks based on correlation matrix,'' IEEE Sensors Journal, vol. 16, no. 11, pp. 4015--4022, 2015
2015
-
[45]
H. Yang, B. Liu, F. Ren, H. Wen, and C. Lin, ``Optimization of energy efficient transmission in underwater sensor networks,'' in GLOBECOM 2009-2009 IEEE Global Telecommunications Conference. 1em plus 0.5em minus 0.4em IEEE, 2009, pp. 1--6
2009
-
[46]
Guerra, P
F. Guerra, P. Casari, and M. Zorzi, ``A performance comparison of mac protocols for underwater networks using a realistic channel simulator,'' in OCEANS 2009. 1em plus 0.5em minus 0.4em IEEE, 2009, pp. 1--8
2009
-
[47]
Morozs, W
N. Morozs, W. Gorma, B. T. Henson, L. Shen, P. D. Mitchell, and Y. V. Zakharov, ``Channel modeling for underwater acoustic network simulation,'' IEEE Access, vol. 8, pp. 136\,151--136\,175, 2020
2020
-
[48]
Y. Zhu, Z. Peng, J.-H. Cui, and H. Chen, ``Toward practical mac design for underwater acoustic networks,'' IEEE Transactions on Mobile Computing, vol. 14, no. 4, pp. 872--886, 2014
2014
-
[49]
M. B. Porter, ``The bellhop manual and user’s guide: Preliminary draft,'' Heat, Light, and Sound Research, Inc., La Jolla, CA, USA, Tech. Rep, vol. 260, 2011
2011
-
[50]
L. Jing, C. He, J. Huang, and Z. Ding, ``Energy management and power allocation for underwater acoustic sensor network,'' IEEE Sensors Journal, vol. 17, no. 19, pp. 6451--6462, 2017
2017
-
[51]
Berger-Sabbatel, A
G. Berger-Sabbatel, A. Duda, M. Heusse, and F. Rousseau, ``Short-term fairness of 802.11 networks with several hosts,'' in IFIP International Conference on Mobile and Wireless Communication Networks. 1em plus 0.5em minus 0.4em Springer, 2004, pp. 263--274
2004
-
[52]
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein, ``The complexity of decentralized control of markov decision processes,'' Mathematics of operations research, vol. 27, no. 4, pp. 819--840, 2002
2002
-
[53]
W. Zhao, J. P. Queralta, and T. Westerlund, ``Sim-to-real transfer in deep reinforcement learning for robotics: a survey,'' in 2020 IEEE symposium series on computational intelligence (SSCI). 1em plus 0.5em minus 0.4em IEEE, 2020, pp. 737--744
2020
-
[54]
Foerster, I
J. Foerster, I. A. Assael, N. De Freitas, and S. Whiteson, ``Learning to communicate with deep multi-agent reinforcement learning,'' Advances in neural information processing systems, vol. 29, 2016
2016
-
[55]
Sunehag, G
P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. F. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls et al., ``Value-decomposition networks for cooperative multi-agent learning based on team reward,'' in AAMAS, 2018
2018
-
[56]
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., ``Human-level control through deep reinforcement learning,'' nature, vol. 518, no. 7540, pp. 529--533, 2015
2015
-
[57]
Hausknecht and P
M. Hausknecht and P. Stone, ``Deep recurrent q-learning for partially observable mdps,'' in 2015 aaai fall symposium series, 2015
2015
-
[58]
Di Valerio, F
V. Di Valerio, F. L. Presti, C. Petrioli, L. Picari, D. Spaccini, and S. Basagni, ``Carma: Channel-aware reinforcement learning-based multi-path adaptive routing for underwater wireless sensor networks,'' IEEE Journal on Selected Areas in Communications, vol. 37, no. 11, pp. 2...
2019
-
[59]
Y. Ye, Y. Tang, H. Wang, X.-P. Zhang, and G. Strbac, ``A scalable privacy-preserving multi-agent deep reinforcement learning approach for large-scale peer-to-peer transactive energy trading,'' IEEE transactions on smart grid, vol. 12, no. 6, pp. 5185--5200, 2021
2021
-
[60]
W. Yu, Y. Chen, Y. Tang, and X. Xu, ``Power allocation for underwater source nodes in uwa cooperative networks,'' in 2018 IEEE international conference on signal processing, communications and computing (ICSPCC). 1em plus 0.5em minus 0.4em IEEE, 2018, pp. 1--6
2018
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.