REVIEW 3 major objections 5 minor 59 references
Mind the Gap: Towards Generalizable Autonomous Penetration Testing via Domain Randomization and Meta-Reinforcement Learning
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read GAP, a Real-to-Sim-to-Real framework combining domain randomization and meta-RL, claims to make autonomous pentesting agents that transfer zero-shot to similar vulnerabilities and adapt quickly to new ones.
desk verdict A solid, well-engineered first step for generalizable RL pentesting, but the headline zero-shot numbers are measured in a way that oversells transfer to real hosts. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the Real-to-Sim-to-Real pipeline, specifically its LLM-powered domain randomization step and its MAML-based meta-training loop. Domain randomization here is not about rendering settings but about text-form host configurations: the LLM is prompted with the official vulnerability description and the original JSON simulation to produce semantically coherent variants with altered ports, versions, operating system, and fingerprints. MAML then treats each synthetic variant as a task, performs an inner-loop policy-gradient update (PPO) from the shared initialization, and updates that initialization in the outer loop with the meta-gradient. The whole design is what lets one trained agent recognize a changed host configuration as another instance of the same vulnerability instead of an out-of-distribution observation.
What would settle it
Run the zero-shot policy of the five-environment GAP variant on authorized real hosts found through a public search engine that have the same vulnerability but non-default ports, versions, and fingerprints, and compare the compromise success rate with the reported 0.92; if it falls toward the 0.14 baseline level, the claimed real-world generalization is not established.
Extended reading notes
Core claim
The paper's central claim is that the generalization gap in RL-based autonomous pentesting can be bridged by combining domain randomization with meta-reinforcement learning, and that this can be done while learning end-to-end in unknown emulated environments. GAP first trains a PPO agent in a vulnerable virtual machine, records the discovered host configuration as a JSON simulation, and treats that simulation as a faithful digital copy of the real environment. It then prompts a large language model with the official vulnerability description plus the JSON example to generate randomized synthetic variants—changing ports, service versions, operating system versions, and web fingerprints—and uses those variants as meta-training environments for MAML. The learned initialization is claimed to transfer zero-shot to similar hosts (same vulnerability, different configuration) and to adapt quickly to dissimilar environments (different vulnerability). The experiments report a zero-shot generalization gap of 51.61 and success rate 0.92 for the five-environment GAP variant, versus 990.21 and 0.14 for the comparison framework, and a reduction of average adaptation training time from 4045 seconds to 2429 seconds compared with learning from scratch.
Load-bearing premise
The load-bearing premise is that the vulnerable virtual machines and the LLM-generated synthetic variants capture the range of configurations a pentesting agent would actually meet on real networks, including effects of network protocols and honeypots; the paper explicitly concedes in its discussion that the vulnerability environments are idealized and overlook such effects.
Editorial extensions
If this is right
- A policy meta-trained on about five LLM-generated variants of one host can compromise same-vulnerability testing hosts with changed ports, versions, and fingerprints at a 0.92 success rate without any further training.
- The generalization gap drops from 990.21 (baseline) to 51.61, so the deployment performance of the transferred policy is close to its training-performance ceiling.
- When transferred to a host with a different vulnerability, the meta-learned initialization reaches a usable policy roughly 40 percent faster than training from scratch, and about 22 percent faster than the transfer baseline.
- Because the pipeline only needs text-based scan observations, the same recipe can in principle apply to any tool-driven security task where host state is reported as text.
Reading between the lines
- Editorial inference: The same LLM-powered randomization recipe should extend to text-observable security tasks beyond remote code execution, such as privilege escalation or web misconfiguration, because the pipeline only assumes that host state arrives as text.
- Editorial inference: The reported near-perfect negative correlation between jumpstart and training time suggests the main benefit of meta-training is a better starting policy; an ablation of MAML against plain PPO on the same five synthetic environments would test that mechanism directly.
- Editorial inference: The admitted gap between idealized vulnerability environments and real networks could be probed by evaluating the zero-shot policy against honeypot or network-middlebox configurations; the paper's own discussion suggests this is the next decisive test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GAP, an autonomous penetration-testing framework that combines end-to-end RL policy learning on vulnerable virtual machines with an LLM-powered domain-randomization step that generates synthetic environment variants, followed by MAML-based meta-reinforcement learning. The claimed contributions are a Real-to-Sim-to-Real pipeline that reduces the training-environment dilemma and improves generalization, measured through zero-shot transfer to similar environments (same vulnerability, different host configurations) and few-shot adaptation to dissimilar environments (different vulnerabilities). Experiments across 20 Vulhub-based CVEs report that 5-GAP achieves a generalization gap of 51.61 and a success rate of 0.92, versus 990.21 and 0.14 for the April baseline, and that GAP-Transfer reduces training time by roughly 40% relative to learning from scratch.
Significance. If the generalization results hold for genuinely independent test distributions, the work would be a meaningful step in applying meta-RL to realistic pentesting: it is among the first to combine domain randomization and meta-RL in this domain, it ships an open-source implementation, and it evaluates across a fairly broad set of real vulnerable products with three seeds per condition. The reported zero-shot gains and faster adaptation are potentially valuable for the autonomous-pentesting community. However, the central external-validity assumption—that the synthetic and test environments capture real-world host-configuration variability—is not established, and the test environments appear to be drawn from the same randomization recipe used for meta-training. The paper's own Section 7 concedes that the environments are idealized and ignore network-protocol and honeypot effects, which directly bears on the zero-shot generalization claim.
major comments (3)
- [§5.5.2 and §4.3] The zero-shot test environments for RQ2 are generated by changing host configurations (ports, services, OS versions, website fingerprints), as described in Section 5.5.2, and these are exactly the same categories of variation that the LLM-powered domain randomization in Section 4.3 uses to construct meta-training environments. The paper does not state that the test variants come from an independent source or distribution, so the GenGap values in Table 3 measure robustness within the authors' randomization recipe rather than generalization across the true distribution of real host configurations. This is load-bearing for the central claim that GAP 'bridges the generalization gap.' Section 7's concession that the environments are idealized and overlook network-protocol and honeypot effects reinforces, rather than resolves, this concern. To support the claim, the authors should evaluate on test environments whose generation is independent of the LLM-based randomization process, for example manually configured hosts or real-world scan data.
- [§6.2, Table 3] The number of meta-training environments n is selected on the test set: the paper reports results for n ∈ {3, 5, 8} and then states that 5-GAP 'serves as the default setting in the subsequent experiments' based on the observed generalization performance. Because the reported headline GenGap of 51.61 corresponds to a model chosen after inspecting test-set performance, the result is subject to selection bias. The authors should either pre-specify n, perform a nested validation split, or report the full set of results without treating the test set as a model-selection criterion.
- [§4.2.2] The claim that 'the agent interacting with this simulated environment equates to interacting with a real environment' is too strong. The JSON simulation is constructed from observed scan feedback and therefore can only capture aspects of the environment that the agent has already queried; it cannot represent unobserved service behavior, protocol-level dynamics, or adversary interactions such as honeypots. Section 7 later acknowledges that the environments are idealized, which is inconsistent with the unconditional phrasing in Section 4.2.2. This overstatement affects the credibility of the Real-to-Sim stage, and the claim should be qualified to state that the simulation matches the observed features, not the full real environment.
minor comments (5)
- [§3.3, Eq. (3)] The definition of GenGap uses expected cumulative reward, but the experiments appear to compute empirical average rewards over episodes; the paper should state explicitly how the expectation in Eq. (3) is estimated in the reported tables.
- [§3.2, Eq. (2)] The symbol G(τ) is used in Eq. (2) before it is defined; please define it immediately before or after the equation.
- [§6.3, Table 4 and Fig. 11] The Pearson correlation of -0.99 is computed from only four points (one per method); this is a very small sample for a quantitative claim. Either report a significance test, present it as descriptive, or add error bars and more methods.
- [§5.5.2] The success rate metric is described as the proportion of hosts compromised, but it is not clear whether the denominator is the number of test variants, the number of seeds, or both; please define the aggregation procedure precisely.
- [References] Reference [56] contains a broken author field ('Team GLM, :,'); this should be cleaned up before publication.
Circularity Check
No significant circularity: the zero-shot evaluation is an empirical held-out comparison; the overlapping host-configuration axes are an external-validity limitation, not a constructional reduction.
full rationale
GAP's derivation chain is empirical rather than definitional. Section 4.3 constructs meta-training environments by LLM-powered domain randomization over host-configuration axes, while Section 5.5.2 creates held-out zero-shot test variants by changing the same kinds of axes; this overlap is a real limitation for the broader 'real-world' generalization claim, and Section 7 explicitly concedes that the environments are idealized and overlook network-protocol and honeypot effects. However, the paper measures GenGap via Eq. 3 and success rates on test environments that are not used during meta-training, and the comparison with April and PPO is an empirical benchmark; the reported results could have failed and are not forced by the definitions. Choosing n=5 after comparing n=3, 5, and 8 on the test set is a model-selection weakness rather than a constructional equivalence. Reference [13] (April) is a self-citation, but it serves only as a baseline and as a setting to follow, not as the load-bearing justification of the generalization claim. Therefore there is no significant circularity; the main risks are external validity and test-set selection, not circularity.
Assumptions & free parameters
free parameters (3)
- Reward function weights (compromise, information, penalty) =
+1000, +100, -10
- Number of meta-training environments n =
5 (default, best of {3,5,8})
- MAML learning rates (alpha, beta) and PPO hyperparameters =
not reported in text
assumptions (4)
- standard math Gradient-based meta-learning (MAML) yields a good initialization for tasks sampled from a fixed task distribution.
- domain assumption Interactions with the JSON-simulated environment are equivalent to interactions with the real environment.
- domain assumption LLM-generated synthetic environments, validated by human experts, are representative of the real-world distribution of hosts with the same vulnerability.
- ad hoc to paper Test variants formed by changing host configurations are the correct operationalization of 'similar environments'.
Cite this review
Pith. "Pith review of Mind the Gap: Towards Generalizable Autonomous Penetration Testing via Domain Randomization and Meta-Reinforcement Learning." pith.science (2026). https://pith.science/paper/QNBTOTUC
@misc{pith2026241204078,
author = {Pith},
title = {Pith review of: Mind the Gap: Towards Generalizable Autonomous Penetration Testing via Domain Randomization and Meta-Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/QNBTOTUC}},
note = {Machine review of arXiv:2412.04078}
}
read the original abstract
With increasing numbers of vulnerabilities exposed on the internet, autonomous penetration testing (pentesting) has emerged as a promising research area. Reinforcement learning (RL) is a natural fit for studying this topic. However, two key challenges limit the applicability of RL-based autonomous pentesting in real-world scenarios: (a) training environment dilemma -- training agents in simulated environments is sample-efficient while ensuring their realism remains challenging; (b) poor generalization ability -- agents' policies often perform poorly when transferred to unseen scenarios, with even slight changes potentially causing significant generalization gap. To this end, we propose GAP, a generalizable autonomous pentesting framework that aims to realizes efficient policy training in realistic environments and train generalizable agents capable of drawing inferences about other cases from one instance. GAP introduces a Real-to-Sim-to-Real pipeline that (a) enables end-to-end policy learning in unknown real environments while constructing realistic simulations; (b) improves agents' generalization ability by leveraging domain randomization and meta-RL learning.Specially, we are among the first to apply domain randomization in autonomous pentesting and propose a large language model-powered domain randomization method for synthetic environment generation. We further apply meta-RL to improve agents' generalization ability in unseen environments by leveraging synthetic environments. The combination of two methods effectively bridges the generalization gap and improves agents' policy adaptation performance.Experiments are conducted on various vulnerable virtual machines, with results showing that GAP can enable policy learning in various realistic environments, achieve zero-shot policy transfer in similar environments, and realize rapid policy adaptation in dissimilar environments.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Holm, Lore a red team emulation tool, IEEE Trans
H. Holm, Lore a red team emulation tool, IEEE Trans. De- pendable Secur. Comput. 20 (2) (2023) 1596–1608. doi:10. 1109/TDSC.2022.3160792
arXiv 2023
-
[2]
X. Chen, J. Hu, C. Jin, L. Li, L. Wang, Understanding domain randomization for sim-to-real transfer, in: The Tenth Interna- tional Conference on Learning Representations, ICLR 2022, Vir- tual Event, April 25-29, 2022, OpenReview.net, 2022
work page 2022
-
[3]
J. Schwartz, H. Kurniawati, Autonomous penetration testing using reinforcement learning, arXiv preprint arXiv:1905.05965 (2019)
arXiv 2019
-
[4]
F. M. Zennaro, L. Erdodi, Modeling penetration testing with re- inforcement learning using capture-the-flag challenges and tab- ular q-learning, arXiv preprint arXiv:2005.12632 (2020)
work page Pith review arXiv 2020
-
[5]
Z. Hu, R. Beuran, Y. Tan, Automated penetration testing using deep reinforcement learning, 2020 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) (2020) 2–10
work page 2020
-
[6]
K. Tran, A. Akella, M. Standen, J. Kim, D. Bowman, T. J. Richer, C.-T. L. I. One, I. Two, Deep hierarchical rein- forcement agents for automated penetration testing, ArXiv abs/2109.06449 (2021)
arXiv 2021
-
[7]
Y. Yang, X. Liu, Behaviour-diverse automatic penetration testing: A curiosity-driven multi-objective deep reinforcement learning approach, CoRR abs/2202.10630 (2022). arXiv:2202. 10630
work page Pith review arXiv 2022
-
[8]
J. Chen, S. Hu, H. Zheng, C. Xing, G. Zhang, Gail-pt: An intel- ligent penetration testing framework with generative adversarial imitation learning, Comput. Secur. 126 (2023) 103055
work page 2023
Show all 59 references
-
[9]
Y. Yang, M. Chen, H. Fu, X. Liu, Settron: Towards better gen- eralisation in penetration testing with reinforcement learning, in: IEEE Global Communications Conference, GLOBECOM 2023, Kuala Lumpur, Malaysia, December 4-8, 2023, 2023, pp. 4662–4667. doi:10.1109/GLOBECOM54140.20...
2023
-
[10]
Q. Li, R. Wang, D. Li, F. Shi, M. Zhang, A. Chattopadhyay, Y. Shen, Y. Li, Dynpen: Automated penetration testing in dynamic network scenarios using deep reinforcement learning, IEEE Transactions on Information Forensics and Security 19 (2024) 8966–8981. doi:10.1109/TIFS.2024.3461950
2024
-
[11]
Schwartz, H
J. Schwartz, H. Kurniawatti, Nasim: Network attack simulator (2019). URL https://networkattacksimulator.readthedocs.io/
2019
-
[12]
M. D. R. Team., Cyberbattlesim, created by Christian Seifert, Michael Betser, William Blum, James Bono, Kate Farris, Emily Goren, Justin Grana, Kristian Holsheimer, Brandon Marken, Joshua Neil, Nicole Nichols, Jugal Parikh, Haoran Wei. (2021). URL https://github.com/microsoft/...
2021
-
[13]
S. Zhou, J. Liu, Y. Lu, J. Yang, D. Hou, Y. Zhang, S. Hu, April: towards scalable and transferable autonomous penetra- tion testing in large action space via action embedding, IEEE Transactions on Dependable and Secure Computing (2024) 1– 17doi:10.1109/TDSC.2024.3518500
2024
-
[14]
K. Wang, B. Kang, J. Shao, J. Feng, Improving generaliza- tion in reinforcement learning with mixture regularization, in: H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, H. Lin (Eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Informat...
2020
-
[15]
X. Song, Y. Jiang, S. Tu, Y. Du, B. Neyshabur, Observational overfitting in reinforcement learning, in: 8th International Con- ference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020
2020
-
[16]
R. Kirk, A. Zhang, E. Grefenstette, T. Rockt¨ aschel, A survey of zero-shot generalisation in deep reinforcement learning, J. Artif. Intell. Res. 76 (2023) 201–264. doi:10.1613/JAIR.1.14174
2023 doi
-
[17]
Cobbe, O
K. Cobbe, O. Klimov, C. Hesse, T. Kim, J. Schulman, Quantify- ing generalization in reinforcement learning, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, ...
2019
-
[18]
Tobin, R
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, P. Abbeel, Domain randomization for transferring deep neural networks from simulation to the real world, in: 2017 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems, IROS 2017, Vancouver, BC, Canada, Septe...
2017
-
[19]
Horv´ ath, G
D. Horv´ ath, G. Erd¨ os, Z. Istenes, T. Horv´ ath, S. F¨ oldi, Ob- ject detection using sim2real domain randomization for robotic applications, IEEE Trans. Robotics 39 (2) (2023) 1225–1243. doi:10.1109/TRO.2022.3207619
2023
-
[20]
Z. Li, H. Zhu, Z. Lu, M. Yin, Synthetic data generation with large language models for text classification: Potential and limi- tations, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023,...
2023 doi
- [21]
-
[22]
T. M. Hospedales, A. Antoniou, P. Micaelli, A. J. Storkey, Meta-learning in neural networks: A survey, IEEE Trans. Pat- tern Anal. Mach. Intell. 44 (9) (2022) 5149–5169. doi:10.1109/ TPAMI.2021.3079209
2022
-
[23]
D. Ye, T. Zhu, K. Gao, W. Zhou, Defending against label- only attacks via meta-reinforcement learning, IEEE Trans. Inf. Forensics Secur. 19 (2024) 3295–3308. doi:10.1109/TIFS.2024. 3357292
2024 doi
- [24]
-
[25]
C. Finn, P. Abbeel, S. Levine, Model-agnostic meta-learning for fast adaptation of deep networks, in: D. Precup, Y. W. Teh (Eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6- 11 August 2017, Vol. 70 of Proceedings...
2017
-
[26]
M. M. Botvinick, S. Ritter, J. X. Wang, Z. Kurth-Nelson, C. Blundell, D. Hassabis, Reinforcement learning, fast and slow, Trends in Cognitive Sciences 23 (2019) 408–422
2019
-
[27]
Schneier, Attack trees, Dr
B. Schneier, Attack trees, Dr. Dobb’s journal 24 (12) (1999) 21–29
1999
-
[28]
Sheyner, J
O. Sheyner, J. W. Haines, S. Jha, R. Lippmann, J. M. Wing, Automated generation and analysis of attack graphs, in: 2002 IEEE Symposium on Security and Privacy, Berkeley, California, USA, May 12-15, 2002, IEEE Computer Society, 2002, pp. 273–
2002
-
[29]
H. Jmal, F. B. Hmida, N. Basta, M. Ikram, M. A. Kˆ aafar, A. Walker, SPGNN-API: A transferable graph neural net- work for attack paths identification and autonomous mitiga- tion, IEEE Trans. Inf. Forensics Secur. 19 (2024) 1601–1613. doi:10.1109/TIFS.2023.3338965
2024
-
[30]
Vinyals, I
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, 14 A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Daliba...
2019
-
[31]
S. Feng, H. Sun, X. Yan, H. Zhu, Z. Zou, S. Shen, H. X. Liu, Dense reinforcement learning for safety validation of au- tonomous vehicles, Nature 615 (2023) 620 – 627
2023
-
[32]
L. Bo, T. Zhang, H. Zhang, J. Hong, M. Liu, C. Zhang, B. Liu, 3d UA V path planning in unknown environment: A transfer re- inforcement learning method based on low-rank adaption, Adv. Eng. Informatics 62 (2024) 102920. doi:10.1016/J.AEI.2024. 102920
2024 doi
-
[33]
L. Li, Y. Luo, J. Yang, L. Pu, Reinforcement learning enabled intelligent energy attack in green iot networks, IEEE Trans. Inf. Forensics Secur. 17 (2022) 644–658. doi:10.1109/TIFS.2022. 3149148
2022 doi
-
[34]
N. Ilic, D. Dasic, M. Vucetic, A. Makarov, R. Petrovic, Dis- tributed web hacking by adaptive consensus-based reinforce- ment learning, Artif. Intell. 326 (2024) 104032. doi:10.1016/ J.ARTINT.2023.104032
2024
-
[35]
Y. YANG, L. CHEN, S. LIU, L. W ANG, H. FU, X. LIU, Z. CHEN, Behaviour-diverse automatic penetration testing: a coverage-based deep reinforcement learning approach, Fron- tiers of Computer Science 19 (3) (2025) 193309. doi:10.1007/ s11704-024-3380-1
2025
-
[36]
Takaesu, Deepexploit (2018)
I. Takaesu, Deepexploit (2018). URL https://www.mbsd.jp/blog/20180228.html
2018
-
[37]
Maeda, M
R. Maeda, M. Mimura, Automating post-exploitation with deep reinforcement learning, Comput. Secur. 100 (2021) 102108
2021
-
[38]
Tremblay, A
J. Tremblay, A. Prakash, D. Acuna, M. Brophy, V. Jampani, C. Anil, T. To, E. Cameracci, S. Boochoon, S. Birchfield, Train- ing deep networks with synthetic data: Bridging the reality gap by domain randomization, in: 2018 IEEE/CVF Confer- ence on Computer Vision and Pattern Rec...
2018 doi
-
[39]
X. B. Peng, M. Andrychowicz, W. Zaremba, P. Abbeel, Sim-to- real transfer of robotic control with dynamics randomization, in: 2018 IEEE International Conference on Robotics and Au- tomation, ICRA 2018, Brisbane, Australia, May 21-25, 2018, IEEE, 2018, pp. 1–8. doi:10.1109/ICRA...
2018
-
[40]
Tiboni, A
G. Tiboni, A. Protopapa, T. Tommasi, G. Averta, Domain randomization for robust, affordable and effective closed-loop control of soft robots, in: IROS, 2023, pp. 612–619. doi: 10.1109/IROS55552.2023.10342537. URL https://doi.org/10.1109/IROS55552.2023.10342537
2023
-
[41]
J. Chen, D. Tam, C. Raffel, M. Bansal, D. Yang, An empirical survey of data augmentation for limited data learning in NLP, Trans. Assoc. Comput. Linguistics 11 (2023) 191–211. doi: 10.1162/TACL\_A\_00542
2023 doi
- [42]
- [43]
- [44]
- [45]
-
[46]
H. Ju, R. Juan, R. Gomez, K. Nakamura, G. Li, Transferring policy of deep reinforcement learning from simulation to reality for robotics, Nat. Mac. Intell. 4 (12) (2022) 1077–1087. doi: 10.1038/S42256-022-00573-6
2022 doi
-
[47]
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, P. Abbeel, Rl 2: Fast reinforcement learning via slow rein- forcement learning, CoRR abs/1611.02779 (2016). arXiv:1611. 02779
2016 arXiv
-
[48]
Arndt, M
K. Arndt, M. Hazara, A. Ghadirzadeh, V. Kyrki, Meta rein- forcement learning for sim-to-real domain adaptation, in: 2020 IEEE International Conference on Robotics and Automation, ICRA 2020, Paris, France, May 31 - August 31, 2020, IEEE, 2020, pp. 2725–2731. doi:10.1109/ICRA409...
2020
-
[49]
R. S. Sutton, A. G. Barto, Reinforcement learning: An intro- duction, MIT press, 2018
2018
-
[50]
Z. Bing, D. Lerch, K. Huang, A. C. Knoll, Meta-reinforcement learning in non-stationary and dynamic environments, IEEE Trans. Pattern Anal. Mach. Intell. 45 (3) (2023) 3476–3491. doi:10.1109/TPAMI.2022.3185549
2023
-
[51]
C. Lyle, M. Rowland, W. Dabney, M. Kwiatkowska, Y. Gal, Learning dynamics and generalization in deep reinforcement learning, in: K. Chaudhuri, S. Jegelka, L. Song, C. Szepesv´ ari, G. Niu, S. Sabato (Eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 20...
2022
-
[52]
A. A. Ta ¨ ıga, R. Agarwal, J. Farebrother, A. C. Courville, M. G. Bellemare, Investigating multi-task pretraining and generaliza- tion in reinforcement learning, in: The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, O...
2023
-
[53]
W. Zhao, J. P. Queralta, T. Westerlund, Sim-to-real transfer in deep reinforcement learning for robotics: a survey, in: 2020 IEEE Symposium Series on Computational Intelligence, SSCI 2020, Canberra, Australia, December 1-4, 2020, IEEE, 2020, pp. 737–744. doi:10.1109/SSCI47803....
2020
-
[54]
K. Wang, N. Reimers, I. Gurevych, Tsdae: Using transformer- based sequential denoising auto-encoderfor unsupervised sen- tence embedding learning, arXiv preprint arXiv:2104.06979 (4 2021)
2021 arXiv
-
[55]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms, CoRR abs/1707.06347 (2017). arXiv:1707.06347
2017 arXiv
- [56]
-
[57]
Z. Zhu, K. Lin, A. K. Jain, J. Zhou, Transfer learning in deep reinforcement learning: A survey, IEEE Trans. Pattern Anal. Mach. Intell. 45 (11) (2023) 13344–13362. doi:10.1109/TPAMI. 2023.3292075
2023
-
[58]
A. M. Metelli, Recent advancements in inverse reinforcement learning, in: M. J. Wooldridge, J. G. Dy, S. Natarajan (Eds.), Thirty-Eighth AAAI Conference on Artificial Intel- ligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI...
2024 doi
-
[284]
doi:10.1109/SECPRI.2002.1004377
2002 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.