REVIEW 3 major objections 5 minor 2 cited by
A Survey of Reinforcement Learning for Software Engineering
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This survey sets out to provide the first systematic map of reinforcement learning applied to software engineering, based on 115 peer-reviewed studies from 22 top venues since 2015.
desk verdict Useful first systematic map of RL across the SE lifecycle, but the search string typo and arXiv-preprint miscounting undercut the comprehensiveness claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the systematic mapping-study protocol: a fixed search string over title, keyword, and abstract fields in three major scholarly databases, restricted to 22 premier software-engineering venues, followed by two-author abstract screening, snowballing, and card-sorting-based keywording. This machinery turns a scattered literature into distributions over venues, SE activities, task types, algorithms, data sources, evaluation metrics, and replicability that answer the paper's five research questions.
What would settle it
Re-run the same search over the same 22 venues with a corrected string that adds 'reward' and removes 'award', then apply the same inclusion criteria; if the corrected search retrieves additional relevant papers beyond the 115, the survey's claim that snowballing found no missed studies is contradicted.
Extended reading notes
Core claim
The central discovery is a quantitative landscape of RL-for-SE: among 115 studies, 72 percent target software quality assurance, and test generation alone accounts for 49 papers. Task-wise, 74 percent of the underlying SE tasks are generation tasks, while ranking, classification, and regression are rare. Value-based algorithms appear in 52 percent of papers, led by Q-learning (27 papers) and DQN (22 papers), with PPO the most common policy-based method (17 papers). Only 5.2 percent of papers focus solely on refining core RL concepts, while 82.6 percent treat RL as a tool with task-specific tuning. Evaluation is effectiveness-heavy, and 41 percent of papers are non-replicable because source code or complete packages are missing. The paper frames these results as the first systematic mapping covering the whole software-engineering lifecycle, in contrast to earlier reviews limited to testing or to machine learning and deep learning broadly.
Load-bearing premise
The survey's entire map depends on the assumption that its search and screening procedure found all relevant RL-for-SE papers in the chosen venues; the search string omits the common term 'reward' and instead contains the misspelled term 'award', so papers that only mention reward-based learning without the word 'reinforcement' in title, keywords, or abstract could be missing from the 115.
Editorial extensions
If this is right
- Software quality assurance is by far the most explored activity (72 percent of 115 papers), with test generation alone accounting for 49 papers, so other lifecycle activities such as requirements and management have almost no RL work.
- Because 74 percent of RL-for-SE tasks are generation tasks, RL is currently used mainly to produce artifacts like tests, code, and comments rather than for classification, ranking, or regression.
- Q-learning and DQN dominate the algorithm landscape, while more recent methods such as SAC, TD3, meta-RL, and RLHF appear in only a handful of studies, indicating room for the field to adopt newer RL techniques.
- Only 5.2 percent of papers innovate on core RL concepts, and 41 percent are non-replicable, which together limit cumulative progress and fair comparison across studies.
- Integration of LLMs with RL has emerged since 2024 in about 10 of the surveyed studies, pointing to a rapidly growing direction for RL-for-SE.
Reading between the lines
- Because the search string omits the term 'reward', the reported 115-paper corpus likely undercounts studies that describe their learning signal without using the word 'reinforcement'; a corrected search could shift the activity and algorithm distributions.
- The restriction to 22 premier venues and papers of at least 8 pages means the sharp post-2022 growth curve probably underestimates the total volume of RL-for-SE work appearing in workshops, short papers, and adjacent venues.
- The survey's own distributions imply a concrete next step it does not take: assembling a shared benchmark environment suite for test generation and repair, since a 41 percent non-replicability rate makes cross-study comparison currently unreliable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a systematic mapping study of reinforcement learning applied to software engineering. It claims to be the first systematic map of RL-for-SE, based on 115 peer-reviewed studies from 22 premier SE venues between 2015 and May 2025. The authors formulate five research questions and report quantitative distributions: 72% of studies address software quality assurance, 74% address generation tasks, Q-learning and DQN are the most used algorithms, and 41% of studies are non-replicable. The paper also identifies challenges and opportunities for RL-for-SE. The central claim depends on a supposedly comprehensive retrieval and screening procedure described in Sections 4.2 and 4.3.
Significance. If the underlying population of studies is correctly identified, the survey would be a useful and timely reference: it provides a structured taxonomy of SE activities, task types, and RL algorithms, along with trend data and a public artifact repository. The paper gives detailed tables (e.g., Table 6) and explicit quantitative summaries that are internally consistent with its own tables (e.g., 85/115 generation = 74%). The authors also make a welcome effort to report replicability rates and to discuss evaluation practices. However, the significance of the contribution is conditional on the completeness and purity of the 115-paper population, and the retrieval and screening methodology has load-bearing weaknesses that prevent accepting the current version as a comprehensive systematic map.
major comments (3)
- [§4.2 and §4.3] The search string is defined as (“reinforcement” OR “Q-learning” or “Q-Network” or “award”) over titles, keywords, and abstracts. This string omits the RL core term “reward” and common algorithm terms such as “policy gradient”, “PPO”, “actor-critic”, and “bandit”. If “award” is a typo for “reward”, the intended term is misspelled and will not match; if it is taken literally, it adds irrelevant matches and still misses reward-based papers. Section 4.3 then states that snowballing yielded no additional papers and concludes that no relevant studies were missed, but snowballing only inspects reference lists of already selected papers and cannot recover papers omitted by the flawed initial search. This undermines every RQ1–RQ4 distribution and the abstract's claim of a comprehensive map. The search must be rerun with an expanded, correctly spelled keyword set and the screening repeated, with the reported 115-count and all derived percentages recomputed.
- [§4.3 and references [60], [105], [137]] The abstract and Section 1 describe the surveyed set as “115 peer-reviewed studies published across 22 premier SE venues”, and Section 4.3 explicitly excludes “non-published manuscripts”. However, at least three retained references are arXiv preprints rather than venue publications: Kim et al. [60] (arXiv:2411.07098), Sanchez-Stern et al. [105] (arXiv:2408.09237), and Wang et al. [137] (arXiv:2407.19487). These papers appear in the analyzed corpus (e.g., [60] is listed in Table 6 under Multi-Agent Q-Learning) and their inclusion contradicts the stated inclusion criteria. The authors should replace these with their peer-reviewed versions if they exist, or explicitly reclassify the corpus. This materially affects the claim of surveying only peer-reviewed venue publications.
- [§4.3] The screening process is described as a manual review of abstracts by two authors, but no inter-rater agreement metric (e.g., Cohen's kappa) is reported and no explicit procedure for resolving screening disagreements is given. Since the selection of the 115 papers is the foundation of all subsequent statistics, the absence of reliability information weakens the claim that the screening is reproducible and that papers were not arbitrarily excluded. The authors should report screening agreement and conflict resolution, and ideally apply the inclusion criteria to full texts where abstracts are ambiguous.
minor comments (5)
- [Table 2] Row 5 appears to be a typographical merge: “International Symposium on Testing and Analysis Working Conference on Mining Software Repositories” should be two separate entries, ISSTA and MSR. Also, “Information and Software Systems” should likely be “Information and Software Technology”, and “Automated Engineering” should likely be “Automated Software Engineering”.
- [Table 6] The abbreviation “TPRO” is used for Trust Region Policy Optimization; the standard abbreviation is TRPO, and the text elsewhere uses TRPO-context terminology.
- [§5, Figure 4] The text says “The cumulative plot in Figure 4 also illustrates…”, but the cumulative plot appears in Figure 3(b); Figure 4 instead shows venue distributions.
- [§1 and Figure 5] There are minor typos: “alogorithms” in Section 1 and “Saftey Improvement” in Figure 5 should be “algorithms” and “Safety Improvement”, respectively.
- [§8, Replicability paragraph] The sentence “journal papers exhibit lower applicability” should read “lower replicability”, matching the subsequent discussion and Table 9.
Circularity Check
No circularity found: the survey is an empirical literature-mapping claim with no fitted parameters, predictive equations, or load-bearing self-citation chain.
full rationale
This paper is a systematic literature review, not a derivation. Its central claim—that it offers the first systematic mapping of RL-for-SE based on 115 peer-reviewed studies from 22 venues—rests on a retrieval and screening process, not on equations that reduce to inputs. There are no fitted parameters renamed as predictions, no ansatz smuggled in via citation, and no uniqueness theorem imported from the authors' prior work. The RL background sections (e.g., Bellman equations, policy gradient theorem) are expository and self-contained. The authors do cite their own prior thread in a few places (e.g., references [18], [19], and [148]), and two of their own RL-for-SE papers appear in the surveyed corpus, but the corpus-level statistics and qualitative classifications do not depend on those specific papers; the survey would stand unchanged if those self-authored entries were removed. The search string omits 'reward' and includes the typo 'award', and some retained references are arXiv preprints rather than peer-reviewed venue publications; these are genuine threats to the validity and completeness of the mapping, but they are coverage/methodology concerns, not circular reasoning. The snowballing claim that no additional papers were found is also an evidentiary weakness, but it is not an instance of a result being equivalent to its input by definition. Accordingly, no circular step can be quoted and reduced as required by the criteria, and the honest finding is no significant circularity (score 0).
Assumptions & free parameters
assumptions (4)
- domain assumption The 22 selected venues and the search string are sufficient to recover the population of RL-for-SE studies.
- domain assumption Abstract-level screening by two authors reliably separates RL-for-SE studies from non-RL software testing work.
- domain assumption All 115 retained papers are peer-reviewed full papers.
- domain assumption The SWEBOK six-activity taxonomy is an appropriate and consistently applied categorization for RL-for-SE papers.
Cite this review
Pith. "Pith review of A Survey of Reinforcement Learning for Software Engineering." pith.science (2026). https://pith.science/paper/4VODNYYX
@misc{pith2026250712483,
author = {Pith},
title = {Pith review of: A Survey of Reinforcement Learning for Software Engineering},
year = {2026},
howpublished = {\url{https://pith.science/paper/4VODNYYX}},
note = {Machine review of arXiv:2507.12483}
}
read the original abstract
Reinforcement Learning (RL) has emerged as a powerful paradigm for sequential decision-making and has attracted growing interest across various domains, particularly following the advent of Deep Reinforcement Learning (DRL) in 2015. Simultaneously, the rapid advancement of Large Language Models (LLMs) has further fueled interest in integrating RL with LLMs to enable more adaptive and intelligent systems. In the field of software engineering (SE), the increasing complexity of systems and the rising demand for automation have motivated researchers to apply RL to a broad range of tasks, from software design and development to quality assurance and maintenance. Despite growing research in RL-for-SE, there remains a lack of a comprehensive and systematic survey of this evolving field. To address this gap, we reviewed 115 peer-reviewed studies published across 22 premier SE venues since the introduction of DRL. We conducted a comprehensive analysis of publication trends, categorized SE topics and RL algorithms, and examined key factors such as dataset usage, model design and optimization, and evaluation practices. Furthermore, we identified open challenges and proposed future research directions to guide and inspire ongoing work in this evolving area. To summarize, this survey offers the first systematic mapping of RL applications in software engineering, aiming to support both researchers and practitioners in navigating the current landscape and advancing the field. Our artifacts are publicly available: https://github.com/KaiWei-Lin-lanina/RL4SE.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 2 Pith papers
-
Evaluating Fuzz Testing for Reinforcement Learning Agents
Under unified budgets, MDPFuzz leads crash count and speed; SeqDivFuzz leads diversity; fuzz crashes improve robustness and train cross-fuzzer safety monitors.
-
Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation
CATGen improves LLM unit-test reliability by combining structured project-context retrieval, deterministic test-class skeletons, and static-analysis repair, beating six baselines on compilation success, coverage, and cost.
Reference graph
Works this paper leans on
-
[60]
Myeongsoo Kim, Tyler Stennett, Saurabh Sinha, and Alessandro Orso. 2024. A Multi-Agent Approach for REST API Testing with Semantic Graphs and LLM-Driven Inputs. arXiv preprint arXiv:2411.07098 (2024)
arXiv 2024
-
[105]
Alex Sanchez-Stern, Abhishek Varghese, Zhanna Kaufman, Dylan Zhang, Talia Ringer, and Yuriy Brun. 2024. QEDCartographer: Automating formal verification using reward-free reinforcement learning. arXiv preprint arXiv:2408.09237 (2024)
work page Pith review arXiv 2024
-
[137]
Yanlin Wang, Yanli Wang, Daya Guo, Jiachi Chen, Ruikai Zhang, Yuchi Ma, and Zibin Zheng. 2024. Rlcoder: Reinforcement learning for repository-level code completion. arXiv preprint arXiv:2407.19487 (2024)
arXiv 2024
-
[1]
Maryam Nooraei Abadeh. 2024. Knowledge-enhanced software refinement: leveraging reinforcement learning for search-based quality engineering. Automated Software Engineering 31, 2 (2024), 57
2024
-
[2]
Amr Abo-eleneen, Ahammed Palliyali, and Cagatay Catal. 2023. The role of Reinforcement Learning in software testing. Information and Software Technology 164 (2023), 107325
2023
-
[3]
Hamidreza Ahmadi, Mehrdad Ashtiani, Mohammad Abdollahi Azgomi, and Raana Saheb-Nassagh. 2022. A DQN-based agent for automatic software refactoring. Information and Software Technology 147 (2022), 106893
2022
-
[4]
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 (2022)
arXiv 2022
-
[5]
Takumi Akazaki, Shuang Liu, Yoriyuki Yamagata, Yihai Duan, and Jianye Hao. 2018. Falsification of cyber-physical systems using deep reinforcement learning. In Formal Methods: 22nd International Symposium, FM 2018, Held as Part of the Federated Logic Conference, FloC 2018, Oxford, UK, July 15-17, 2018, Proceedings 22 . Springer, 456–465
2018
Show all 159 references
-
[6]
Hussein Almulla and Gregory Gay. 2022. Learning how to search: generating effective test cases through adaptive fitness function selection. Empirical Software Engineering 27, 2 (2022), 38
2022
-
[7]
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. 2017. Hindsight Experience Replay. In Advances in Neural Information Processing Systems , Vol. 30
2017
-
[8]
Mojtaba Bagherzadeh, Nafiseh Kahani, and Lionel Briand. 2021. Reinforcement learning for test case prioritization. IEEE Transactions on Software Engineering 48, 8 (2021), 2836–2856
2021
-
[9]
Richard Ernest Bellman. 2003. Dynamic Programming. Dover Publications, Inc., New York, NY, USA
2003
-
[10]
Antonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, and Stefano Russo. 2020. Learning-to-rank vs ranking-to-learn: Strategies for regression testing in continuous integration. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engin...
2020
-
[11]
Andrea Borgarelli, Constantin Enea, Rupak Majumdar, and Srinidhi Nagendra. 2024. Reward Augmentation in Reinforcement Learning for Testing Distributed Systems. Proceedings of the ACM on Programming Languages 8, OOPSLA2 (2024), 1928–1954
2024
-
[12]
Pierre Bourque, Robert Dupuis, Alain Abran, James W Moore, and Leonard Tripp. 2002. The guide to the software engineering body of knowledge. IEEE software 16, 6 (2002), 35–44
2002
-
[13]
Anthony Canino, Yu David Liu, and Hidehiko Masuhara. 2018. Stochastic energy optimization for mobile GPS applications. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 703–713
2018
-
[14]
Nicolás Cardozo and Ivana Dusparic. 2023. Auto-COP: Adaptation generation in context-oriented programming using reinforcement learning options. Information and Software Technology 164 (2023), 107308
2023
-
[15]
Santo Carino and James H Andrews. 2015. Dynamically testing GUIs using ant colony optimization (t). In 2015 30th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 138–148. Manuscript submitted to ACM 40
2015
-
[16]
Partha Chakraborty, Mahmoud Alfadel, and Meiyappan Nagappan. 2024. Rlocator: Reinforcement learning for bug localization. IEEE Transactions on Software Engineering (2024)
2024
-
[17]
Chao Chen, Wenrui Diao, Yingpei Zeng, Shanqing Guo, and Chengyu Hu. 2018. DRLgencert: Deep learning-based automated testing of certificate verification in SSL/TLS implementations. In 2018 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 48–58
2018
-
[18]
Junjie Chen, Haoyang Ma, and Lingming Zhang. 2020. Enhanced compiler bug isolation via memoized search. In Proceedings of the 35th IEEE/ACM international conference on automated software engineering . 78–89
2020
-
[19]
Junjie Chen, Chenyao Suo, Jiajun Jiang, Peiqi Chen, and Xingjian Li. 2023. Compiler test-program generation via memoized configuration search. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 2035–2047
2023
-
[20]
Jia Chen, Jiayi Wei, Yu Feng, Osbert Bastani, and Isil Dillig. 2019. Relational verification using reinforcement learning. Proceedings of the ACM on Programming Languages 3, OOPSLA (2019), 1–30
2019
-
[21]
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. 2021. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems 34 (2021), 15084–15097
2021
-
[22]
Po-Wei Chou, Daniel Maturana, and Sebastian Scherer. 2017. Improving Stochastic Policy Gradients in Continuous Control with Deep Reinforcement Learning using the Beta Distribution. In Proceedings of the 34th International Conference on Machine Learning . 834–843
2017
-
[23]
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems 30 (2017)
2017
-
[24]
Davide Corradini, Zeno Montolli, Michele Pasqua, and Mariano Ceccato. 2024. DeepREST: Automated Test Case Generation for REST APIs Exploiting Deep Reinforcement Learning. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 1383–1394
2024
-
[25]
Stevão Alves de Andrade, Fatima LS Nunes, and Márcio Eduardo Delamaro. 2023. Exploiting deep reinforcement learning and metamorphic testing to automatically test virtual reality applications. Software Testing, Verification and Reliability 33, 8 (2023), e1863
2023
-
[26]
Christian Degott, Nataniel P Borges Jr, and Andreas Zeller. 2019. Learning user interface element interactions. In Proceedings of the 28th ACM SIGSOFT international symposium on software testing and analysis . 296–306
2019
-
[27]
Ruomeng Ding, Chaoyun Zhang, Lu Wang, Yong Xu, Minghua Ma, Xiaomin Wu, Meng Zhang, Qingjun Chen, Xin Gao, Xuedong Gao, et al. 2023. Tracediag: Adaptive, interpretable, and efficient root cause analysis on large-scale microservice systems. In Proceedings of the 31st ACM Joint E...
2023
-
[28]
Yanru Ding, Yanmei Zhang, Guan Yuan, Shujuan Jiang, Wei Dai, and Luciano Baresi. 2025. Optimizing Class Integration Testing with Criticality- Driven Test Order Generation. In 2025 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 276–286
2025
-
[29]
Yanru Ding, Yanmei Zhang, Guan Yuan, Shujuan Jiang, Wei Dai, and Yinghui Zhang. 2023. Integration test order generation based on reinforcement learning considering class importance. Journal of Systems and Software 205 (2023), 111823
2023
-
[30]
Davide Domini, Filippo Cavallari, Gianluca Aguzzi, and Mirko Viroli. 2024. Scarlib: Towards a hybrid toolchain for aggregate computing and many-agent reinforcement learning. Science of Computer Programming 238 (2024), 103176
2024
-
[31]
Andréa Doreste, Matteo Biagiola, and Paolo Tonella. 2024. Adversarial testing with reinforcement learning: A case study on autonomous driving. In 2024 IEEE Conference on Software Testing, Verification and Validation (ICST) . IEEE, 293–304
2024
-
[32]
Zachary Eberhart and Collin McMillan. 2021. Dialogue management for interactive api search. In 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 274–285
2021
-
[33]
Seyedeh Sepideh Emam and James Miller. 2015. Test case prioritization using extended digraphs. ACM Transactions on Software Engineering and Methodology (TOSEM) 25, 1 (2015), 1–41
2015
-
[34]
Seyedeh Sepideh Emam and James Miller. 2018. Inferring extended probabilistic finite-state automaton models from software executions. ACM Transactions on Software Engineering and Methodology (TOSEM) 27, 1 (2018), 1–39
2018
-
[35]
Jueon Eom, Seyeon Jeong, and Taekyoung Kwon. 2024. Fuzzing JavaScript Interpreters with Coverage-Guided Reinforcement Learning for LLM-Based Mutation. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis . 1656–1668
2024
-
[36]
Yujia Fan, Sinan Wang, Zebang Fei, Yao Qin, Huaxuan Li, and Yepang Liu. 2024. Can Cooperative Multi-Agent Reinforcement Learning Boost Automatic Web Testing? An Exploratory Study. InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 14–26
2024
-
[37]
Raihana Ferdous, Fitsum Kifetew, Davide Prandi, and Angelo Susi. 2022. Towards agent-based testing of 3D games using reinforcement learning. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering . 1–8
2022
-
[38]
Stefano Ferretti, Silvia Mirri, Catia Prandi, and Paola Salomoni. 2016. Automatic web content personalization through reinforcement learning. Journal of Systems and Software 121 (2016), 157–169
2016
-
[39]
Xiaoqin Fu, Haipeng Cai, Wen Li, and Li Li. 2020. Seads: Scalable and cost-effective dynamic dependence analysis of distributed systems via reinforcement learning. ACM Transactions on Software Engineering and Methodology (TOSEM) 30, 1 (2020), 1–45
2020
-
[40]
Xinyu Gao, Yun Xiong, Deze Wang, Zhenhan Guan, Zejian Shi, Haofen Wang, and Shanshan Li. 2024. Preference-Guided Refactored Tuning for Retrieval Augmented Code Generation. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 65–77
2024
-
[41]
Saeedeh Sadat Sajjadi Ghaemmaghami, Seyedeh Sepideh Emam, and James Miller. 2022. Automatically inferring user behavior models in large-scale web applications. Information and Software Technology 141 (2022), 106704. Manuscript submitted to ACM A Survey of Reinforcement Learnin...
2022
-
[42]
Luca Giamattei, Matteo Biagiola, Roberto Pietrantuono, Stefano Russo, and Paolo Tonella. 2025. Reinforcement learning for online testing of autonomous driving systems: a replication and extension study. Empirical Software Engineering 30, 1 (2025), 19
2025
-
[43]
Rong Gu, Peter G Jensen, Cristina Seceleanu, Eduard Enoiu, and Kristina Lundqvist. 2022. Correctness-guaranteed strategy synthesis and compression for multi-agent autonomous systems. Science of Computer Programming 224 (2022), 102894
2022
-
[44]
Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, and Alois Knoll. 2024. A review of safe reinforcement learning: Methods, theories and applications. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[45]
Tianxiao Gu, Chun Cao, Tianchi Liu, Chengnian Sun, Jing Deng, Xiaoxing Ma, and Jian Lü. 2017. Aimdroid: Activity-insulated multi-level automated testing for android applications. In 2017 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 103–114
2017
-
[46]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al . 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025)
2025 arXiv
-
[47]
Hanyang Guo, Yingye Chen, Xiangping Chen, Yuan Huang, and Zibin Zheng. 2024. Smart contract code repair recommendation based on reinforcement learning and multi-metric optimization. ACM Transactions on Software Engineering and Methodology 33, 4 (2024), 1–31
2024
-
[48]
Hui Guo, Ting Su, Xiaoqiang Liu, Siyi Gu, and Jingling Sun. 2023. Effectively finding ICC-related bugs in android apps via reinforcement learning. In 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE) . IEEE, 403–414
2023
-
[49]
Wunan Guo, Zhen Dong, Liwei Shen, Daihong Zhou, Bin Hu, Chen Zhang, and Hai Xue. 2025. Effectively Modeling UI Transition Graphs for Android Apps Via Reinforcement Learning. In 2025 IEEE/ACM 33rd International Conference on Program Comprehension (ICPC) . IEEE Computer Society, 13–24
2025
-
[50]
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning . 1861–1870
2018
-
[51]
Carol Hanna, Aymeric Blot, and Justyna Petke. 2025. Reinforcement learning for mutation operator selection in automated program repair. Automated Software Engineering 32, 2 (2025), 1–33
2025
-
[52]
Fitash Ul Haq, Donghwan Shin, and Lionel C Briand. 2023. Many-objective reinforcement learning for online testing of dnn-enabled systems. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1814–1826
2023
-
[53]
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–79
2024
-
[54]
Yuan Huang, Shaohao Huang, Huanchao Chen, Xiangping Chen, Zibin Zheng, Xiapu Luo, Nan Jia, Xinyu Hu, and Xiaocong Zhou. 2020. Towards automatically generating block comments for code snippets. Information and Software Technology 127 (2020), 106373
2020
-
[55]
Yuchao Huang, Junjie Wang, Zhe Liu, Yawen Wang, Song Wang, Chunyang Chen, Yuanzhe Hu, and Qing Wang. 2024. Crashtranslator: Automatically reproducing mobile application crashes directly from stack trace. In Proceedings of the 46th ieee/acm international conference on software ...
2024
-
[56]
Dmytro Humeniuk, Foutse Khomh, and Giuliano Antoniol. 2024. Reinforcement learning informed evolutionary search for autonomous systems testing. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–45
2024
-
[57]
Rinkesh Joshi and Nafiseh Kahani. 2024. Comparative Study of Reinforcement Learning in GitHub Pull Request Outcome Predictions. In 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 489–500
2024
-
[58]
Shuting Kang, Qian Dong, Yunzhi Xue, and W Yanjun. 2024. MACS: Multi-Agent Adversarial Reinforcement Learning for Finding Diverse Critical Driving Scenarios. In 2024 IEEE Conference on Software Testing, Verification and Validation (ICST) . IEEE, 1–12
2024
-
[59]
Myeongsoo Kim, Saurabh Sinha, and Alessandro Orso. [n. d.]. Adaptive rest api testing with reinforcement learning. In 2023 38th IEEE. In ACM International Conference on Automated Software Engineering (ASE) . 446–458
2023
-
[61]
Jinkyu Koo, Charitha Saumya, Milind Kulkarni, and Saurabh Bagchi. 2019. Pyse: Automatic worst-case test generation by reinforcement learning. In 2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST) . IEEE, 136–147
2019
-
[62]
Yavuz Koroglu and Alper Sen. 2021. Functional test generation from UI test scenarios using reinforcement learning for android applications. Software Testing, Verification and Reliability 31, 3 (2021), e1752
2021
-
[63]
Yavuz Koroglu, Alper Sen, Ozlem Muslu, Yunus Mete, Ceyda Ulker, Tolga Tanriverdi, and Yunus Donmez. 2018. Qbe: Qlearning-based exploration of android applications. In 2018 IEEE 11th International Conference on Software Testing, Verification and Validation (ICST) . IEEE, 105–115
2018
-
[64]
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. 2021. Offline reinforcement learning with implicit q-learning. arXiv preprint arXiv:2110.06169 (2021)
2021 arXiv
-
[65]
Christian Krupitzer, Christian Gruhl, Bernhard Sick, and Sven Tomforde. 2022. Proactive hybrid learning and optimisation in self-adaptive systems: The swarm-fleet infrastructure scenario. Information and Software Technology 145 (2022), 106826
2022
-
[66]
Abdullah Lakhan, Mazin Abed Mohammed, Omar Ibrahim Obaid, Chinmay Chakraborty, Karrar Hameed Abdulkareem, and Seifedine Kadry. 2022. Efficient deep-reinforcement learning aware resource allocation in SDN-enabled fog paradigm. Automated Software Engineering 29 (2022), 1–25
2022
-
[67]
Yuanhong Lan, Yifei Lu, Zhong Li, Minxue Pan, Wenhua Yang, Tian Zhang, and Xuandong Li. 2024. Deeply reinforcing android gui testing with deep reinforcement learning. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13
2024
-
[68]
Sascha Lange, Thomas Gabel, and Martin Riedmiller. 2012. Batch Reinforcement Learning. 45–73. Manuscript submitted to ACM 42
2012
-
[69]
Paul-Antoine Le Tolguenec, Emmanuel Rachelson, Yann Besse, Florent Teichteil-Koenigsbuch, Nicolas Schneider, Hélène Waeselynck, and Dennis Wilson. 2024. Exploration-Driven Reinforcement Learning for Avionic System Fault Detection (Experience Paper). In Proceedings of the 33rd ...
2024
-
[70]
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. nature 521, 7553 (2015), 436–444
2015
-
[71]
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. CoRR abs/2005.01643 (2020). arXiv:2005.01643
2020 arXiv
-
[72]
Bolun Li, Zhihong Sun, Tao Huang, Hongyu Zhang, Yao Wan, Ge Li, Zhi Jin, and Chen Lyu. 2024. Ircoco: Immediate rewards-guided deep reinforcement learning for code completion. Proceedings of the ACM on Software Engineering 1, FSE (2024), 182–203
2024
-
[73]
Yun Li, Yanmei Zhang, Yanru Ding, Shujuan Jiang, and Guan Yuan. 2024. A class integration test order generation approach based on Sarsa algorithm. Automated Software Engineering 31, 1 (2024), 7
2024
-
[74]
Zikun Li, Jinjun Peng, Yixuan Mei, Sina Lin, Yi Wu, Oded Padon, and Zhihao Jia. 2024. Quarl: A learning-based quantum circuit optimizer. Proceedings of the ACM on Programming Languages 8, OOPSLA1 (2024), 555–582
2024
-
[75]
Linfeng Liang, Yao Deng, Kye Morton, Valtteri Kallinen, Alice James, Avishkar Seth, Endrowednes Kuantama, Subhas Mukhopadhyay, Richard Han, and Xi Zheng. 2025. GARL: Genetic Algorithm-Augmented Reinforcement Learning to Detect Violations in Marker-Based Autonomous Landing Syst...
2025
-
[76]
Long-Ji Lin. 1992. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning 8, 3 (1992), 293–321
1992
-
[77]
Junrui Liu, Yanju Chen, Bryan Tan, Isil Dillig, and Yu Feng. 2022. Learning contract invariants using reinforcement learning. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering . 1–11
2022
-
[78]
Jiawei Liu, Yuheng Huang, Zhijie Wang, Lei Ma, Chunrong Fang, Mingzheng Gu, Xufan Zhang, and Zhenyu Chen. 2023. Generation-based differential fuzzing for deep learning libraries. ACM Transactions on Software Engineering and Methodology 33, 2 (2023), 1–28
2023
-
[79]
Yujie Liu, Mingxuan Zhu, Jinhao Dong, Junzhe Yu, and Dan Hao. 2024. Compiler Bug Isolation via Enhanced Test Program Mutation. InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 819–830
2024
-
[80]
Zhongxin Liu, Xin Xia, Christoph Treude, David Lo, and Shanping Li. 2019. Automatic generation of pull request descriptions. In 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 176–188
2019
-
[81]
Chengjie Lu, Yize Shi, Huihui Zhang, Man Zhang, Tiexin Wang, Tao Yue, and Shaukat Ali. 2022. Learning configurations of operating environment of autonomous vehicles to maximize their collisions. IEEE Transactions on Software Engineering 49, 1 (2022), 384–402
2022
-
[82]
Paulina Stevia Nouwou Mindom, Amin Nikanjam, and Foutse Khomh. 2025. Harnessing pre-trained generalist agents for software engineering tasks. Empirical Software Engineering 30, 1 (2025), 1–53
2025
-
[83]
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing Atari with Deep Reinforcement Learning. In cite arxiv:1312.5602Comment: NIPS Deep Learning Workshop 2013
2013 arXiv
-
[84]
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015. Human-level control through deep reinforcement learning. nature 518, 7540 (2015), 529–533
2015
-
[85]
Mahshid Helali Moghadam, Mehrdad Saadatmand, Markus Borg, Markus Bohlin, and Björn Lisper. 2021. An autonomous performance testing framework using self-adaptive fuzzy reinforcement learning. Software quality journal (2021), 1–33
2021
-
[86]
Christopher Molloy, Jeremy Banks, Steven HH Ding, Furkan Alaca, Philippe Charland, and Andrew Walenstein. 2025. Mecha: A Neural-Symbolic Open-Set Homogeneous Decision Fusion Approach for Zero-day Malware Similarity Detection. IEEE Transactions on Software Engineering (2025)
2025
-
[87]
Suvam Mukherjee, Pantazis Deligiannis, Arpita Biswas, and Akash Lal. 2020. Learning-based controlled concurrency testing. Proceedings of the ACM on Programming Languages 4, OOPSLA (2020), 1–31
2020
-
[88]
Mona Nashaat and James Miller. 2024. Towards efficient fine-tuning of language models with organizational data for automated software review. IEEE Transactions on Software Engineering (2024)
2024
-
[89]
Vu Nguyen and Bach Le. 2021. Rltcp: A reinforcement learning approach to prioritizing automated user interface tests. Information and Software Technology 136 (2021), 106574
2021
-
[90]
Paulina Stevia Nouwou Mindom, Amin Nikanjam, and Foutse Khomh. 2023. A comparison of reinforcement learning frameworks for software testing tasks. Empirical Software Engineering 28, 5 (2023), 111
2023
-
[91]
Ciprian Păduraru, Rareş Cristea, and Alin Stefanescu. 2024. End-to-end RPA-like testing using reinforcement learning. In 2024 IEEE Conference on Software Testing, Verification and Validation (ICST). IEEE, 419–429
2024
-
[92]
Minxue Pan, An Huang, Guoxin Wang, Tian Zhang, and Xuandong Li. 2020. Reinforcement learning based curiosity-driven testing of android applications. In Proceedings of the 29th ACM SIGSOFT international symposium on software testing and analysis . 153–164
2020
-
[93]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu ...
2019
-
[94]
Kai Petersen, Robert Feldt, Shahid Mujtaba, and Michael Mattsson. 2008. Systematic Mapping Studies in Software Engineering. In Proceedings of the 12th International Conference on Evaluation and Assessment in Software Engineering (EASE’08) . 68–77
2008
-
[95]
Kai Petersen, Robert Feldt, Shahid Mujtaba, and Michael Mattsson. 2008. Systematic mapping studies in software engineering. In 12th international conference on evaluation and assessment in software engineering (EASE) . BCS Learning & Development. Manuscript submitted to ACM A ...
2008
-
[96]
Puterman
Martin L. Puterman. 1994. Markov Decision Processes: Discrete Stochastic Dynamic Programming (1st ed.). John Wiley & Sons, Inc., New York, NY, USA
1994
-
[97]
Martin L Puterman. 2014. Markov decision processes: discrete stochastic dynamic programming . John Wiley & Sons
2014
-
[98]
Zhongsheng Qian, Qingyuan Yu, Hui Zhu, Jinping Liu, and Tingfeng Fu. 2025. Reinforcement learning for test case prioritization based on LLEed K-means clustering and dynamic priority factor. Information and Software Technology 179 (2025), 107654
2025
-
[99]
Dezhi Ran, Hao Wang, Wenyu Wang, and Tao Xie. 2023. Badge: prioritizing UI events with hierarchical multi-armed bandits for automated UI testing. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 894–905
2023
-
[100]
Sameer Reddy, Caroline Lemieux, Rohan Padhye, and Koushik Sen. 2020. Quickly generating diverse valid test inputs with reinforcement learning. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering . 1410–1421
2020
-
[101]
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. 2022. A generalist agent. arXiv preprint arXiv:2205.06175 (2022)
2022 arXiv
-
[102]
Zilong Ren, Xiaolin Ju, Xiang Chen, and Hao Shen. 2024. ProRLearn: boosting prompt tuning-based vulnerability detection by reinforcement learning. Automated Software Engineering 31, 2 (2024), 38
2024
-
[103]
Andrea Romdhana, Mariano Ceccato, Alessio Merlo, and Paolo Tonella. 2022. Ifrit: Focused testing through deep reinforcement learning. In 2022 IEEE Conference on Software Testing, Verification and Validation (ICST) . IEEE, 24–34
2022
-
[104]
Andrea Romdhana, Alessio Merlo, Mariano Ceccato, and Paolo Tonella. 2022. Deep reinforcement learning for black-box testing of android apps. ACM Transactions on Software Engineering and Methodology (TOSEM) 31, 4 (2022), 1–29
2022
-
[106]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. arXiv:1707.06347 (2017). https://arxiv.org/pdf/1707.06347.pdf
2017 arXiv
-
[107]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al . 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300 (2024)
2024 arXiv
-
[108]
Ashish Sharma, Sanjiv Tokekar, and Sunita Varma. 2022. Actor-critic architecture based probabilistic meta-reinforcement learning for load balancing of controllers in software defined networks. Automated Software Engineering 29, 2 (2022), 59
2022
-
[109]
Salman Sherin, Asmar Muqeet, Muhammad Uzair Khan, and Muhammad Zohaib Iqbal. 2023. QExplore: An exploration strategy for dynamic web applications using guided search. Journal of Systems and Software 195 (2023), 111512
2023
-
[110]
Chaochen Shi, Yong Xiang, Jiangshan Yu, Keshav Sood, and Longxiang Gao. 2023. Machine translation-based fine-grained comments generation for solidity smart contracts. Information and Software Technology 153 (2023), 107065
2023
-
[111]
Ting Shu, Cuiping Wu, and Zuohua Ding. 2023. Boosting input data sequences generation for testing EFSM-specified systems using deep reinforcement learning. Information and Software Technology 155 (2023), 107114
2023
-
[112]
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madele...
2016
-
[113]
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. 2017. Ma...
2017
-
[114]
Jiayang Song, Xuan Xie, and Lei Ma. 2023. SIEGE: A Semantics-Guided Safety Enhancement Framework for AI-Enabled Cyber-Physical Systems. IEEE Transactions on Software Engineering 49, 8 (2023), 4058–4080
2023
-
[115]
Donna Spencer. 2009. Card sorting: Designing usable categories . Rosenfeld Media
2009
-
[116]
Helge Spieker and Arnaud Gotlieb. 2020. Adaptive metamorphic testing with contextual bandits. Journal of Systems and Software 165 (2020), 110574
2020
-
[117]
Helge Spieker, Arnaud Gotlieb, Dusica Marijan, and Morten Mossige. 2017. Reinforcement learning for automatic test case prioritization and selection in continuous integration. In Proceedings of the 26th ACM SIGSOFT international symposium on software testing and analysis . 12–22
2017
-
[118]
Jianzhong Su, Hong-Ning Dai, Lingjun Zhao, Zibin Zheng, and Xiapu Luo. 2022. Effectively generating vulnerable transaction sequences in smart contracts with reinforcement learning-guided fuzzing. In Proceedings of the 37th IEEE/ACM International Conference on Automated Softwar...
2022
-
[119]
Richard S Sutton. 1988. Learning to predict by the methods of temporal differences. Machine learning 3 (1988), 9–44
1988
-
[120]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. A Bradford Book, Cambridge, MA, USA
2018
-
[121]
Richard S Sutton, Andrew G Barto, et al. 1998. Reinforcement learning: An introduction. Vol. 1. MIT press Cambridge
1998
-
[122]
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999. Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems 12 (1999)
1999
-
[123]
Sutton, David McAllester, Satinder Singh, and Yishay Mansour
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999. Policy Gradient Methods for Reinforcement Learning with Function Approximation. In Advances in Neural Information Processing Systems (NIPS) . 1057–1063. Manuscript submitted to ACM 44
1999
-
[124]
Wannita Takerngsaksiri, Rujikorn Charakorn, Chakkrit Tantithamthavorn, and Yuan-Fang Li. 2025. Pytester: Deep reinforcement learning for text-to-testcase generation. Journal of Systems and Software 224 (2025), 112381
2025
-
[125]
Lizhuang Tan, Amjad Aldweesh, Ning Chen, Jian Wang, Jianyong Zhang, Yi Zhang, Konstantin Igorevich Kostromitin, and Peiying Zhang. 2024. Energy efficient resource allocation based on virtual network embedding for IoT data generation. Automated Software Engineering 31, 2 (2024), 66
2024
-
[126]
Haoxin Tu, Zhide Zhou, He Jiang, Imam Nur Bani Yusuf, Yuxian Li, and Lingxiao Jiang. 2024. Isolating compiler bugs by generating effective witness programs with large language models. IEEE Transactions on Software Engineering (2024)
2024
-
[127]
Rosalia Tufano, Simone Scalabrino, Luca Pascarella, Emad Aghajani, Rocco Oliveto, and Gabriele Bavota. 2022. Using reinforcement learning for load testing of video games. In Proceedings of the 44th international conference on software engineering . 2303–2314
2022
-
[128]
Uraz Cengiz Türker, Robert M Hierons, Khaled El-Fakih, Mohammad Reza Mousavi, and Ivan Y Tyukin. 2024. Accelerating finite state machine-based testing using reinforcement learning. IEEE Transactions on Software Engineering 50, 3 (2024), 574–597
2024
-
[129]
Uraz Cengiz Türker, Robert M Hierons, Mohammad Reza Mousavi, and Ivan Y Tyukin. 2021. Efficient state synchronisation in model-based testing through reinforcement learning. In 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 368–380
2021
-
[130]
Hado Van Hasselt, Arthur Guez, and David Silver. 2016. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI conference on artificial intelligence, Vol. 30
2016
-
[131]
Yao Wan, Zhou Zhao, Min Yang, Guandong Xu, Haochao Ying, Jian Wu, and Philip S Yu. 2018. Improving automatic source code summarization via deep reinforcement learning. In Proceedings of the 33rd ACM/IEEE international conference on automated software engineering . 397–407
2018
-
[132]
Dong Wang, Yuki Ueda, Raula Gaikovina Kula, Takashi Ishio, and Kenichi Matsumoto. 2021. Can we benchmark code review studies? a systematic mapping study of methodology, dataset, and metric. Journal of Systems and Software 180 (2021), 111009
2021
-
[133]
Jingbo Wang and Chao Wang. 2022. Learning to synthesize relational invariants. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. 1–12
2022
-
[134]
Simin Wang, Liguo Huang, Amiao Gao, Jidong Ge, Tengfei Zhang, Haitao Feng, Ishna Satyarth, Ming Li, He Zhang, and Vincent Ng. 2022. Machine/deep learning for software engineering: A systematic literature review. IEEE Transactions on Software Engineering 49, 3 (2022), 1188–1231
2022
-
[135]
Wenhua Wang, Yuqun Zhang, Yulei Sui, Yao Wan, Zhou Zhao, Jian Wu, Philip S Yu, and Guandong Xu. 2020. Reinforcement-learning-guided source code summarization using hierarchical attention. IEEE Transactions on software Engineering 48, 1 (2020), 102–119
2020
-
[136]
Yilun Wang, Pengfei Chen, Hui Dou, Yiwen Zhang, Guangba Yu, Zilong He, and Haiyu Huang. 2024. FaaSConf: QoS-aware Hybrid Resources Configuration for Serverless Workflows. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 957–969
2024
-
[138]
Yidan Wang, Zhouruixing Zhu, Qiuai Fu, Yuchi Ma, and Pinjia He. 2024. MRCA: Metric-level Root Cause Analysis for Microservices via Multi-Modal Data. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 1057–1068
2024
-
[139]
Christopher J C H Watkins and Peter Dayan. 1992. Q-learning. Machine learning 8, 3-4 (1992), 279–292
1992
-
[140]
Cody Watson, Nathan Cooper, David Nader Palacio, Kevin Moran, and Denys Poshyvanyk. 2022. A systematic literature review on the use of deep learning in software engineering research. ACM Transactions on Software Engineering and Methodology (TOSEM) 31, 2 (2022), 1–58
2022
-
[141]
Williams
Ronald J. Williams. 1992. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Machine Learning (May 1992), 229âĂŞ256
1992
-
[142]
Zhengkai Wu, Evan Johnson, Wei Yang, Osbert Bastani, Dawn Song, Jian Peng, and Tao Xie. 2019. REINAM: reinforcement learning for input- grammar inference. In Proceedings of the 2019 27th acm joint meeting on european software engineering conference and symposium on the foundat...
2019
-
[143]
Mingrui Yang and Dalin Zhang. 2023. Deep reinforcement learning guided decision tree learning for program synthesis. In 2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 925–932
2023
-
[144]
Yang Yang, Zheng Li, Liuliu He, and Ruilian Zhao. 2020. A systematic study of reward for reinforcement learning based continuous integration testing. Journal of Systems and Software 170 (2020), 110787
2020
-
[145]
Yang Yang, Zheng Li, Ying Shang, and Qianyu Li. 2023. Sparse reward for reinforcement learning-based continuous integration testing. Journal of Software: Evolution and Process 35, 6 (2023), e2409
2023
-
[146]
Yanming Yang, Xin Xia, David Lo, and John Grundy. 2022. A survey on deep learning for software engineering. ACM Computing Surveys (CSUR) 54, 10s (2022), 1–73
2022
-
[147]
Kaichun Yao, Hao Wang, Chuan Qin, Hengshu Zhu, Yanjun Wu, and Libo Zhang. 2024. CARL: Unsupervised Code-Based Adversarial Attacks for Programming Language Models via Reinforcement Learning. ACM Transactions on Software Engineering and Methodology 34, 1 (2024), 1–32
2024
-
[148]
Hanmo You, Zan Wang, Bin Lin, and Junjie Chen. 2025. Navigating the Testing of Evolving Deep Learning Systems: An Exploratory Interview Study. In ICSE. IEEE, 2726–2738
2025
-
[149]
Shengcheng Yu, Chunrong Fang, Xin Li, Yuchen Ling, Zhenyu Chen, and Zhendong Su. 2024. Effective, Platform-Independent GUI Testing via Image Embedding and Reinforcement Learning. ACM Transactions on Software Engineering and Methodology 33, 7 (2024), 1–27
2024
-
[150]
Shiwen Yu, Ting Wang, and Ji Wang. 2023. Loop invariant inference through smt solving enhanced reinforcement learning. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis . 175–187
2023
-
[151]
Xinglin Yu, Hongliang Liang, and Chunlin Wang. 2024. Multiple Targets Directed Greybox Fuzzing: From Reachable to Exploited. In 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 907–917. Manuscript submitted to ACM A Survey of...
2024
-
[152]
Shaokun Zhang, Hanwen Lei, Yuanpeng Wang, Ding Li, Yao Guo, and Xiangqun Chen. 2023. How Android Apps Break the Data Minimization Principle: An Empirical Study. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 1238–1250
2023
-
[153]
Shaohua Zhang, Shuang Liu, Jun Sun, Yuqi Chen, Wenzhi Huang, Jinyi Liu, Jian Liu, and Jianye Hao. 2021. Figcps: Effective failure-inducing input generation for cyber-physical systems with deep reinforcement learning. In 2021 36th IEEE/ACM International Conference on Automated ...
2021
-
[154]
Shaokun Zhang, Linna Wu, Yuanchun Li, Ziqi Zhang, Hanwen Lei, Ding Li, Yao Guo, and Xiangqun Chen. 2023. ReSPlay: Improving Cross-Platform Record-and-Replay with GUI Sequence Matching. In 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE) . IEEE...
2023
-
[155]
Weiwei Zhang, Shengjian Guo, Hongyu Zhang, Yulei Sui, Yinxing Xue, and Yun Xu. 2023. Challenging machine learning-based clone detectors via semantic-preserving code transformations. IEEE Transactions on Software Engineering 49, 5 (2023), 3052–3070
2023
-
[156]
Zhaoxu Zhang, Robert Winn, Yu Zhao, Tingting Yu, and William GJ Halfond. 2023. Automatically reproducing android bug reports using natural language processing and reinforcement learning. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Ana...
2023
-
[157]
Yan Zheng, Yi Liu, Xiaofei Xie, Yepang Liu, Lei Ma, Jianye Hao, and Yang Liu. 2021. Automatic web testing using curiosity-driven reinforcement learning. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 423–435
2021
-
[158]
Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan. 2019. Wuji: Automatic online combat game testing using evolutionary deep reinforcement learning. In 2019 34th IEEE/ACM International Conference on Automa...
2019
-
[159]
Lingwei Zhu, Haseeb Shah, Han Wang, Yukie Nagai, and Martha White. 2025. q-exponential policy optimization. In International Conference on Learning Representations (ICLR). Manuscript submitted to ACM
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.