REVIEW 3 major objections 5 minor 60 references
PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Under sustained non-stationarity, shared-parameter cooperative MARL loses learning capacity as neurons go dormant; PRIME restores it by resetting only neurons that are simultaneously forward-dormant and backward-silent, gaining 24.9% IQM re
desk verdict Solid method paper with honest empirical work, but Theorem 1 overclaims by counting neurons where the reset modifies whole parameter slices. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The silent neuron intersection S_l = D_l ∩ G_l: D_l is the forward-dormant set (normalized mean absolute activation ≤ τ_d, aggregated over the full B×n team batch), and G_l is the backward-silent set (normalized mean absolute gradient of the live PPO training loss ≤ τ_g). Only neurons in this intersection are reinitialized, via Kaiming-uniform incoming weights, zeroed bias, zeroed outgoing weight column, and cleared Adam state. This machinery carries the argument by guaranteeing that every reset targets capacity that is genuinely expendable, which keeps the perturbation term of the regret bound proportional to the silent-subspace dimension rather than the full parameter count.
What would settle it
Run PRIME against plain MAPPO on standard cooperative MARL benchmarks with larger networks (more hidden units and agents) and a wider seed population, and measure the per-layer forward-dormant versus backward-silent gap. If the 24.9% IQM advantage does not persist, or if the gap vanishes so the intersection collapses to a forward-only criterion, the claim that the bidirectional team-aggregated gate carries the gain would be contradicted. A second targeted check: on the paper's own setup, replace the live PPO gradient with an output-sensitivity proxy; the paper reports this proxy stalls on one
Extended reading notes
Core claim
The paper claims that in shared-parameter cooperative MAPPO under sustained non-stationarity, the only safe neuron-reset set is the intersection of forward-dormant and backward-silent neurons, aggregated over the full team batch. Forward dormancy alone is misleading: in three of four hidden layers, the forward-dormant set is substantially larger than the backward-silent set, so a forward-only reset would discard neurons the optimizer is still actively steering. PRIME's criterion reads the backward signal from the gradient the PPO training loss has already deposited, and reinitializes only the intersection set with an output-preserving reset. On a phase-switching UAV emergency communication s
Load-bearing premise
The decisive premise is that the single custom simulator configuration (three UAVs, twenty users, 32 hidden units, one cyclic phase schedule, two seeds) fairly represents shared-parameter cooperative MARL generally; the regret bound additionally assumes the clipped PPO surrogate is locally strongly convex inside its trust region.
Editorial extensions
If this is right
- PRIME can be added to any shared-parameter MAPPO/CTDE pipeline without changing the network architecture; detection is purely periodic and needs no external signal about when the environment changes.
- In the paper's phase-switching UAV emergency communication simulator, PRIME improves IQM return by 24.9% over MAPPO (72.626 vs 58.135) and keeps the dormant neuron fraction at 10-20% versus MAPPO's 40-45%.
- Ablations locate the performance gain in the backward signal being the live PPO training gradient and in team-level B×n aggregation, not in the reset operator: a matched-cadence stochastic perturbation of the same silent set performs on par, while restricting detection to a single agent's slice costs 8.1 IQM points.
- The dynamic regret bound shows the cost of PRIME's resets scales with the time-averaged silent-subspace dimension rather than the full parameter count, so the improvement over blanket perturbation is expected to grow with network size.
- The forward-only baseline is worse than plain MAPPO in change mode (39.410 IQM vs 58.135), supporting the paper's claim that resetting dormant-but-gradient-active neurons destroys learning in progress.
Reading between the lines
- Editorial extension: the measured decoupling between forward dormancy and training-gradient silence may be a generic signature of distributional shift in shared-parameter MARL; one could instrument other cooperative domains for the same gap and use it as an early-warning indicator before phase changes.
- Editorial extension: the exchangeability of the reset operator suggests future work can focus on detection quality and trigger schedule rather than the intervention itself; making the reset period adaptive to online dormancy or feature-rank monitoring is a natural next step.
- Editorial extension: a direct testable extension is to measure whether the same intersection criterion helps in standard cooperative MARL benchmarks and larger actor networks; if the forward/backward decoupling gap shrinks there, the gain attributed to the gradient gate would be expected to shrink correspondingly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies plasticity loss in shared-parameter cooperative MAPPO under sustained non-stationarity, in the context of UAV-assisted emergency communication networks. It reports diagnostic measurements (dormancy persistence, forward/backward decoupling, cross-agent dormancy disagreement), proposes PRIME, a periodic neuron-level reset method that resets only neurons that are both forward-dormant and backward-silent, aggregates detection statistics over the full team batch, and reinitializes with an output-preserving reset. The authors report a 24.9% IQM improvement over vanilla MAPPO on a custom phase-switching UAV-ECN simulator, extensive ablations along seven design axes, and a dynamic regret bound (Theorem 1) intended to show that the perturbation cost scales with the silent-subspace dimension rather than the full parameter count.
Significance. If the results hold, PRIME is a practical, low-overhead plasticity-maintenance method for a realistic and under-studied regime: shared-parameter cooperative MARL under non-stationarity. The paper's strengths include a systematic motivation study, a clearly specified algorithm, and a broad ablation matrix that isolates the training-gradient signal and team-level aggregation as the key design choices. The authors are also unusually transparent about limitations (single simulator, two seeds, deferred benchmarks). However, the theoretical pillar as stated has a load-bearing dimensional flaw, and the main experimental configuration appears to include a phase-triggered reset not present in the algorithm description. These issues can likely be repaired, but they require changes to the paper's claims and presentation.
major comments (3)
- [Section IV-E, Assumption 4, Eq. (17), Eqs. (12)-(14)] Theorem 1's perturbation accounting conflates neurons with scalar parameters. Assumption 4 sets d*_t = |S_t| (number of silent neurons) and models the reset as noise projected onto a subspace of dimension d*_t. But the reset in Eqs. (12)-(14) modifies, for each silent neuron i, the d_in incoming weights, the bias, and the entire outgoing column [W_{l+1}]_{:,i}; in parameter space the modified subspace has dimension |S_t|(d_in+1+d_out), not |S_t|. Since the theorem tracks parameter error e_t = w_t - w*_t, zeroing an outgoing column changes e_t by the current column norm and cannot be discarded as 'output continuity'. Appendix C's justification counts only the Kaiming incoming energy and ignores outgoing weights; it also does not reconcile the Gaussian proxy's support with the actual reset support. With H=32 and Section V-C reporting Policy Layer 0 silent fractions around 20-30%, even a co
- [Section V-D, Fig. 10(b), Algorithm 1] There is a mismatch between the described algorithm and the evaluated configuration. Algorithm 1 specifies a purely periodic reset (c mod F=0) with no external phase information. However, Section V-D states that 'the main experimental runs (variant A in Fig. 10(b)) additionally trigger a full reset sweep at each boundary.' The headline 24.9% IQM gain (72.626 vs 58.135) and the results in Section V-B/Table II therefore appear to be produced by variant A, not by the pure periodic procedure that the abstract and contributions claim. Variant B, which is exactly Algorithm 1, yields an even higher IQM (75.301), so the design claim is empirically salvageable, but the paper must report the clean-algorithm configuration as primary and clearly state which variant generated each figure and table. As written, readers cannot determine whether the main results validate the paper's central claim of pha
- [Section V-A, V-B] The empirical evidence is confined to one custom simulator with U=3, N=20, H=32, and two random seeds. The paper explicitly defers larger teams/networks, standard MARL benchmarks (SMAC, MPE), and bootstrap confidence intervals to future work. Two seeds is a weak basis for the quantitative claims, especially because the AuxGrad ablation shows a 15.5-point cross-seed spread, and several figures report only min-max bands without confidence intervals. This does not invalidate the within-simulator comparison, but it sharply limits the generalizability claims in the abstract and introduction ('first method in multi-agent communication systems'). At minimum, the authors should add at least one more seed or report bootstrap CIs, and either temper the general claims or include a standard benchmark with larger networks.
minor comments (5)
- [Supplementary Appendix B] Typo: 'Substituting the update rule the update rule of the main paper' and the reference to 'Section IV-F' should be 'Section IV-E'.
- [Abstract / throughout] 'UA V' should be 'UAV' for readability.
- [Figure S10] The legend uses 'LD', 'LZG', 'LDI' and the formula 'LDI = LD LZG' without defining the intersection symbol; the notation is clear only after reading the main text.
- [Section V-D, AuxGrad paragraph] The text says 'The sweep spans 47 IQM points' and then lists four values; this should be '4'.
- [Table II] The notation 'τ_aux^g' appears in the text but not in Table II, and the threshold values are not fully annotated in the table header.
Circularity Check
Theorem 1's perturbation-cost conclusion is Assumption 4 restated: the regret bound's d*_t term is put in by construction, not derived from the reset operator.
-
self definitional
[Section IV-E, Assumption 4; Theorem 1 Eq. (17); Supplementary Appendix C]
"Assumption 4 (Selective Reset Energy). At each reset event, parameters in the silent set S_t = ∪_l S_l of cardinality d*_t = |S_t| are reinitialized. The perturbation is modeled as a projection-restricted noise injection: E_t = ηγ_r Π_t ζ_t with ζ_t ∼ N(0, I_d), where Π_t is the orthogonal projection onto the silent subspace of dimension d*_t, and γ_r > 0 controls the perturbation strength. The expected perturbation energy satisfies E‖E_t‖² = η²γ²_r d*_t, with d*_t ≤ d [19]."
The abstract and Section IV-E advertise the theorem as showing that 'the perturbation cost scales with the small silent-subspace dimension rather than the full parameter count.' But the perturbation term in Eq. (17), (2ηγ²_r/μ)·(1/T)Σ d*_t, is exactly Assumption 4's E‖E_t‖² = η²γ²_r d*_t inserted into a standard tracking bound. The actual reset (12)–(14) reinitializes each silent neuron's incoming weights, bias, and the entire outgoing column of W_{l+1}, so in parameter space the modified subspace has dimension |S_t|·(d_in + 1 + d_out), not |S_t|. The proof never derives d*_t from the reset operations; the 'small silent-subspace dimension' conclusion is an input of the model, not a consequence derived from PRIME. What remains empirical—that |S_t| is small—is measured, but the theoretical c
full rationale
The paper's empirical contribution is not circular: MAPPO, PRIME-Forward, and PRIME are compared in the same simulator with the same backbone, and the ablations (Table II) vary detection signal, aggregation, and reset operator independently, so the 24.9% IQM gain and the attribution to gradient signal and team-level aggregation are self-contained findings. The self-citation to ReSiN [19] is genuine prior work and the tracking proof is reproduced in Appendix B, so I do not count the self-citation itself as load-bearing. However, the central theoretical pillar—the dynamic regret bound showing that reset cost scales with the silent-subspace dimension—reduces by construction to Assumption 4, which defines the perturbation energy to be η²γ²_r d*_t. The reset equations touch per-neuron incoming weights, bias, and the entire outgoing column, yet the analysis assigns each silent neuron cost 1 in parameter tracking error; no bridge is supplied from Eqs. (12)–(14) to the projection model. Thus the headline theoretical claim is equivalent to its modeling assumption, while the empirical claim remains independent. Score 6: partial circularity in the theoretical validation, not in the experiments.
Assumptions & free parameters
free parameters (4)
- τ_d (dormancy threshold) =
0.5
- τ_g (gradient-silence threshold) =
0.08
- F (reset period) =
200 mini-batch steps
- τ_aux^g (AuxGrad ablation threshold) =
0.15 (best of {0.03, 0.08, 0.15, 0.30} on seed 42)
assumptions (7)
- domain assumption Assumptions 1–2: team-averaged PPO surrogate loss L_t is L-smooth and locally µ-strongly convex within the PPO trust region (Sec. IV-E; proof in App. B).
- domain assumption Assumption 3: bounded cumulative path length of the phase-conditional optimum, P_T = Σ ||w*_{t+1} − w*_t|| < ∞.
- ad hoc to paper Assumption 4: resets modeled as projection-restricted Gaussian noise E_t = ηγ_r Π_t ζ_t with E||Π_t ζ_t||² = d*_t.
- domain assumption ReSiN equivalence: a neuron is silent iff both its mean absolute activation and its activation-gradient magnitude vanish (Definition 2, 'under the boundedness and non-degeneracy conditions of [19]').
- domain assumption The clipped-PPO loss gradient accumulated over mini-batches is a faithful indicator that the optimizer has abandoned a neuron.
- domain assumption Simulator fidelity: closed-form LoS approximation (Eq. 2), 3GPP path loss, rotary-wing energy model (Eq. 4), RPGM mobility, and the engineered {0,1,2} phase cycle reproduce relevant post-disaster UAV-ECN non-stationarity.
- standard math Appendix B proof steps: Young's inequality with α = µη/(2−µη), telescoping of the recursion, and Cauchy–Schwarz for the path-length sum.
Cite this review
Pith. "Pith review of PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks." pith.science (2026). https://pith.science/paper/TBLOENV2
@misc{pith2026260717922,
author = {Pith},
title = {Pith review of: PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/TBLOENV2}},
note = {Machine review of arXiv:2607.17922}
}
read the original abstract
Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state unexamined. We show that sustained non-stationarity damages this internal state directly: as objectives shift, neurons progressively fall dormant and the shared policy loses the capacity to learn. The obvious remedy, resetting dormant neurons, is unsafe under shared-parameter multi-agent training: many neurons that appear inactive are still receiving strong training gradients, and whether a neuron appears dormant depends on which agent's observations it processes. PRIME (Plasticity Recovery In Multi-agent Environments) therefore verifies both directions before intervening. Extending the bidirectional Silent Neuron framework to cooperative multi-agent reinforcement learning, it aggregates activation and gradient statistics over the full team batch, reads the backward signal from the gradient the training loss has already deposited , not from a hand-crafted proxy, and reinitializes only neurons that are simultaneously activation-dormant and gradient-silent. Useful representations are preserved while learning capacity is restored. On a phase-switching UAV emergency communication simulator, PRIME improves interquartile mean return by 24.9\% over MAPPO and holds dormant neuron fractions at 10--20\% versus 40--45\%; ablations attribute the gains to the gradient signal and team-level aggregation rather than to the specific reset operator. A dynamic regret bound shows that the perturbation cost scales with the small silent-subspace dimension rather than the full parameter count.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Richard and Wong, Kai-Kit , journal=
Zhao, Nan and Lu, Weidang and Sheng, Min and Chen, Yunfei and Tang, Jie and Yu, F. Richard and Wong, Kai-Kit , journal=. UAV-Assisted Emergency Networks in Disasters , year=
-
[2]
A Tutorial on UAVs for Wireless Networks: Applications, Challenges, and Open Problems , year=
Mozaffari, Mohammad and Saad, Walid and Bennis, Mehdi and Nam, Young-Han and Debbah, Mérouane , journal=. A Tutorial on UAVs for Wireless Networks: Applications, Challenges, and Open Problems , year=
-
[3]
UAV Communications for 5G and Beyond: Recent Advances and Future Trends , year=
Li, Bin and Fei, Zesong and Zhang, Yan , journal=. UAV Communications for 5G and Beyond: Recent Advances and Future Trends , year=
-
[4]
Cellular-Connected UAV: Potential, Challenges, and Promising Technologies , year=
Zeng, Yong and Lyu, Jiangbin and Zhang, Rui , journal=. Cellular-Connected UAV: Potential, Challenges, and Promising Technologies , year=
-
[5]
IEEE Trans
Joint trajectory and communication design for multi-UAV enabled wireless networks , author=. IEEE Trans. Wireless Commun. , volume=. 2018 , publisher=
2018
-
[6]
Liu, Chi Harold and Chen, Zheyu and Tang, Jian and Xu, Jie and Piao, Chengzhe , title =. IEEE J.Sel. A. Commun. , month = sep, pages =. 2018 , issue_date =. doi:10.1109/JSAC.2018.2864373 , abstract =
arXiv 2018
-
[7]
Deep Reinforcement Learning Based Dynamic Trajectory Control for UAV-Assisted Mobile Edge Computing , year=
Wang, Liang and Wang, Kezhi and Pan, Cunhua and Xu, Wei and Aslam, Nauman and Nallanathan, Arumugam , journal=. Deep Reinforcement Learning Based Dynamic Trajectory Control for UAV-Assisted Mobile Edge Computing , year=
-
[8]
Jiang, Ye and Zhai, Daosen and Yang, Mengke and Lin, Zheng and Li, Yuanzhan , title =. Proc. ACM MobiCom Workshop Drone Assisted Wireless Commun. 5G Beyond , pages =. 2022 , isbn =. doi:10.1145/3555661.3560866 , abstract =
arXiv 2022
Show all 60 references
-
[9]
IEEE Trans
Dynamic spectrum access with QoS and interference temperature constraints , author=. IEEE Trans. Mobile Comput. , volume=. 2007 , publisher=
2007
-
[10]
IEEE Trans
Trajectory planning and resource allocation for multi-UAV cooperative computation , author=. IEEE Trans. Commun. , volume=. 2024 , publisher=
2024
-
[11]
2016 , eprint=
Prioritized Experience Replay , author=. 2016 , eprint=
2016
-
[12]
Finn, Chelsea and Abbeel, Pieter and Levine, Sergey , title =. Proc. Int. Conf. Mach. Learn. , pages =. 2017 , publisher =
2017
-
[13]
Loss of plasticity in continual deep reinforcement learning , author=. Proc. Conf. Lifelong Learn. Agents , pages=. 2023 , organization=
2023
-
[14]
and Adams, Ryan P
Ash, Jordan T. and Adams, Ryan P. , title =. Proc. Adv. Neural Inf. Process. Syst. , articleno =. 2020 , isbn =
2020
-
[15]
The primacy bias in deep reinforcement learning , author=. Proc. Int. Conf. Mach. Learn. , pages=. 2022 , organization=
2022
-
[16]
arXiv preprint arXiv:2505.01584 , year=
Understanding and Exploiting Plasticity for Non-stationary Network Resource Adaptation , author=. arXiv preprint arXiv:2505.01584 , year=
-
[17]
A Survey on DRL-Based UAV Communications and Networking: DRL Fundamentals, Applications and Implementations , year=
Zhao, Wei and Cui, Shaoxin and Qiu, Wen and He, Zhiqiang and Liu, Zhi and Zheng, Xiao and Mao, Bomin and Kato, Nei , journal=. A Survey on DRL-Based UAV Communications and Networking: DRL Fundamentals, Applications and Implementations , year=
-
[18]
2023 , issue_date =
Sun, Jie and Sheng, Zhichao and Nasir, Ali Arshad and Huang, Zhiyu and Yu, Hongwen and Fang, Yong , title =. 2023 , issue_date =. doi:10.1016/j.phycom.2023.102200 , journal =
2023
-
[19]
, author=
Computing Challenges of UAV Networks: A Comprehensive Survey. , author=. Comput. Mater. Continua , volume=
-
[20]
2024 , eprint=
An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning , author=. 2024 , eprint=
2024
-
[21]
and Farquhar, Gregory and Afouras, Triantafyllos and Nardelli, Nantas and Whiteson, Shimon , title =
Foerster, Jakob N. and Farquhar, Gregory and Afouras, Triantafyllos and Nardelli, Nantas and Whiteson, Shimon , title =. Proc. AAAI Conf. Artif. Intell. , articleno =. 2018 , isbn =
2018
-
[22]
Yu, Chao and Velu, Akash and Vinitsky, Eugene and Gao, Jiaxuan and Wang, Yu and Bayen, Alexandre and Wu, Yi , title =. Proc. Adv. Neural Inf. Process. Syst. , articleno =. 2022 , isbn =
2022
-
[23]
Fernando and Lan, Qingfeng and Rahman, Parash and Mahmood, A
Dohare, Shibhansh and Hernandez-Garcia, J. Fernando and Lan, Qingfeng and Rahman, Parash and Mahmood, A. Rupam and Sutton, Richard S. , title =. Nature , year =
-
[24]
Sokar, Ghada and Agarwal, Rishabh and Castro, Pablo Samuel and Evci, Utku , title =. Proc. Int. Conf. Mach. Learn. , articleno =. 2023 , publisher =
2023
-
[25]
Lyle, Clare and Zheng, Zeyu and Nikishin, Evgenii and Pires, Bernardo Avila and Pascanu, Razvan and Dabney, Will , title =. Proc. Int. Conf. Mach. Learn. , articleno =. 2023 , publisher =
2023
-
[26]
IEEE Trans
He, Zhiqiang and Liu, Zhenyu , title =. IEEE Trans. Mobile Comput. , year =. doi:10.1109/TMC.2025.3546230 , note =
2025
-
[27]
2026 , eprint=
Plasticity-Enhanced Multi-Agent Mixture of Experts for Dynamic Objective Adaptation in UAVs-Assisted Emergency Communication Networks , author=. 2026 , eprint=
2026
-
[28]
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer , author=. Proc. Int. Conf. Learn. Represent. , year=
-
[29]
Fedus, William and Zoph, Barret and Shazeer, Noam , title =. J. Mach. Learn. Res. , month = jan, articleno =. 2022 , issue_date =
2022
-
[30]
2017 , eprint=
Proximal Policy Optimization Algorithms , author=. 2017 , eprint=
2017
-
[31]
2018 , eprint=
High-Dimensional Continuous Control Using Generalized Advantage Estimation , author=. 2018 , eprint=
2018
-
[32]
Study on Channel Model for Frequencies from 0.5 to 100\,
-
[33]
Wireless Commun
Camp, Tracy and Boleng, Jeff and Davies, Vanessa , title =. Wireless Commun. Mobile Comput. , volume =. doi:https://doi.org/10.1002/wcm.72 , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1002/wcm.72 , abstract =
-
[34]
1999 , isbn =
Hong, Xiaoyan and Gerla, Mario and Pei, Guangyu and Chiang, Ching-Chuan , title =. 1999 , isbn =. doi:10.1145/313237.313248 , booktitle =
1999
-
[35]
New Energy Consumption Model for Rotary-Wing UAV Propulsion , year=
Yan, Hua and Chen, Yunfei and Yang, Shuang-Hua , journal=. New Energy Consumption Model for Rotary-Wing UAV Propulsion , year=
-
[36]
Energy Minimization for Wireless Communication With Rotary-Wing UAV , year=
Zeng, Yong and Xu, Jie and Zhang, Rui , journal=. Energy Minimization for Wireless Communication With Rotary-Wing UAV , year=
-
[37]
2024 , eprint=
Disentangling the Causes of Plasticity Loss in Neural Networks , author=. 2024 , eprint=
2024
-
[38]
2024 , eprint=
Weight Clipping for Deep Continual and Reinforcement Learning , author=. 2024 , eprint=
2024
-
[39]
2024 , eprint=
Maintaining Plasticity in Continual Learning via Regenerative Regularization , author=. 2024 , eprint=
2024
-
[40]
Deep reinforcement learning with plasticity injection , year =
Nikishin, Evgenii and Oh, Junhyuk and Ostrovski, Georg and Lyle, Clare and Pascanu, Razvan and Dabney, Will and Barreto, Andr\'. Deep reinforcement learning with plasticity injection , year =. Proc. Adv. Neural Inf. Process. Syst. , articleno =
-
[41]
The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value Factorization , url =
Qin, Haoyuan and Ma, Chennan and Deng, Mian and Liu, Zhengzhu and Mei, Songzhu and Liu, Xinwang and Wang, Cheng and Shen, Siqi , booktitle =. The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value Factorization , url =. doi:10.52202/079017-1127 , editor =
-
[42]
arXiv preprint arXiv:1906.04737 , year=
Dealing with non-stationarity in multi-agent deep reinforcement learning , author=. arXiv preprint arXiv:1906.04737 , year=
1906 arXiv
-
[43]
Multiagent Continual Coordination via Progressive Task Contextualization , year=
Yuan, Lei and Li, Lihe and Zhang, Ziqian and Zhang, Fuxiang and Guan, Cong and Yu, Yang , journal=. Multiagent Continual Coordination via Progressive Task Contextualization , year=
-
[44]
Foerster, Jakob and Nardelli, Nantas and Farquhar, Gregory and Afouras, Triantafyllos and Torr, Philip H. S. and Kohli, Pushmeet and Whiteson, Shimon , title =. Proc. Int. Conf. Mach. Learn. , pages =. 2017 , publisher =
2017
-
[45]
and Al-Shedivat, Maruan and Whiteson, Shimon and Abbeel, Pieter and Mordatch, Igor , title =
Foerster, Jakob and Chen, Richard Y. and Al-Shedivat, Maruan and Whiteson, Shimon and Abbeel, Pieter and Mordatch, Igor , title =. Proc. Int. Conf. Auton. Agents Multiagent Syst. , pages =. 2018 , publisher =
2018
-
[46]
Zhao, Stephen and Lu, Chris and Grosse, Roger and Foerster, Jakob , title =. Proc. Adv. Neural Inf. Process. Syst. , articleno =. 2022 , isbn =
2022
-
[47]
2024 , eprint=
Mixture of Experts in a Mixture of RL settings , author=. 2024 , eprint=
2024
-
[48]
and Dai, Andrew and Chen, Zhifeng and Le, Quoc and Laudon, James , title =
Zhou, Yanqi and Lei, Tao and Liu, Hanxiao and Du, Nan and Huang, Yanping and Zhao, Vincent Y. and Dai, Andrew and Chen, Zhifeng and Le, Quoc and Laudon, James , title =. Proc. Adv. Neural Inf. Process. Syst. , articleno =. 2022 , isbn =
2022
-
[49]
, title =
Agarwal, Rishabh and Schwarzer, Max and Castro, Pablo Samuel and Courville, Aaron and Bellemare, Marc G. , title =. Proc. Adv. Neural Inf. Process. Syst. , articleno =. 2021 , isbn =
2021
-
[50]
Lyle, Clare and Zheng, Zeyu and Khetarpal, Khimya and Martens, James and van Hasselt, Hado and Pascanu, Razvan and Dabney, Will , title =. Proc. Adv. Neural Inf. Process. Syst. , articleno =. 2024 , isbn =
2024
-
[51]
Spectrum-Aware Mobile Edge Computing for UAVs Using Reinforcement Learning , year=
Badnava, Babak and Kim, Taejoon and Cheung, Kenny and Ali, Zaheer and Hashemi, Morteza , booktitle=. Spectrum-Aware Mobile Edge Computing for UAVs Using Reinforcement Learning , year=
-
[52]
2026 , eprint=
PPO in the Fisher-Rao geometry , author=. 2026 , eprint=
2026
-
[53]
2025 , eprint=
Plasticity-Aware Mixture of Experts for Learning Under QoE Shifts in Adaptive Video Streaming , author=. 2025 , eprint=
2025
-
[54]
, title =
Juliani, Arthur and Ash, Jordan T. , title =. Proc. Adv. Neural Inf. Process. Syst. , articleno =. 2024 , isbn =
2024
-
[55]
Samvelyan, Mikayel and Rashid, Tabish and Schroeder de Witt, Christian and Farquhar, Gregory and Nardelli, Nantas and Rudner, Tim G. J. and Hung, Chia-Man and Torr, Philip H. S. and Foerster, Jakob and Whiteson, Shimon , title =. Proc. Int. Conf. Auton. Agents Multiagent Syst....
2019
-
[56]
Lowe, Ryan and Wu, Yi and Tamar, Aviv and Harb, Jean and Abbeel, Pieter and Mordatch, Igor , title =. Proc. Adv. Neural Inf. Process. Syst. , pages =. 2017 , isbn =
2017
-
[57]
Continuous Adaptation via Meta-Learning in Nonstationary and Competitive Environments , author=. Proc. Int. Conf. Learn. Represent. , year=
-
[58]
Mathematics , VOLUME =
Wang, Suyu and Yue, Quan and Xu, Zhenlei and Qiao, Peihong and Lyu, Zhentao and Gao, Feng , TITLE =. Mathematics , VOLUME =. 2025 , NUMBER =
2025
-
[59]
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset , url =
Galashov, Alexandre and Titsias, Michalis and Gy\". Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset , url =. Proc. Adv. Neural Inf. Process. Syst. , doi =
-
[60]
Cheng, Xiang and Huang, Ziwei and Bai, Lu , title =. Commun. Surveys Tuts. , month = jul, pages =. 2022 , issue_date =. doi:10.1109/COMST.2022.3184049 , abstract =
2022
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.