REVIEW 4 major objections 5 minor 1 cited by
Adversarial Reasoning for Repair Based on Inferred Program Intent
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that automated repair succeeds by adversarially inferring several program intents and generating intent-specific tests, correctly repairing 77 Defects4J 2.0 bugs and 105 HumanEval-Java bugs under realistic fault…
desk verdict AdverIntent-Agent is a genuine conceptual step for LLM repair—adversarial intents plus in-loop test generation—but the headline counts rest on a single non-deterministic run and an under-reported correctness-label breakdown. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the adversarial program intent: a natural-language specification of a function's expected behavior that is intentionally constructed to conflict with previously inferred intents. The paper measures the conflict with an adversarial score, defined as the fraction of oracle tests (same inputs, intent-dependent expected outputs) on which two intents disagree, and requires each new intent to score above $100\%/K$ with the first intent. The second load-bearing piece is dynamic precise prompting in the repair agent, which first asks for the top three root causes of the bug under one intent and then issues a separate patch-generation prompt per root cause. Together these mechanisms convert the untestable question “what did the developer mean?” into a space of testable hypotheses, and they force the patch pool to be diverse by construction.
What would settle it
Run each generated adversarial test against the ground-truth fixed program for a random sample of repaired bugs and count how many assert the wrong expected output; a material share of wrong assertions would mean the tests can reject correct patches, undermining the correctness counts. Alternatively, measure intent-alignment coverage on a fresh benchmark and check whether it stays near the reported 81.7 percent, since that coverage is the method's upper bound.
Extended reading notes
Core claim
The paper's central claim is that intent diversity, made concrete through adversarial test oracles, is a repair mechanism in its own right. For a buggy function, the reasoning agent produces an initial natural-language statement of expected behavior and then, prompted with “what if the previous intents are incorrect,” produces two more intents that are deliberately distinct. The test agent turns each intent into executable tests, reusing the same inputs but changing expected outputs per intent, and measures an adversarial score as the fraction of tests whose expected outputs differ between intents; intents below a threshold of 100 percent divided by K are regenerated. The repair agent then asks for the top three root causes consistent with each intent and generates a patch per root cause, accepting only patches that pass both the original tests and the intent-specific adversarial tests. The paper reports that this pipeline correctly repaired 77 Defects4J 2.0 bugs and 105 HumanEval-Java bugs under realistic fault localization, and that on 300 sampled Defects4J bugs at least one of the three inferred intents was judged aligned with the ground-truth intent in 81.7 percent of cases.
Load-bearing premise
The load-bearing premise is that at least one of the few LLM-inferred adversarial intents matches the developer's true intent: the paper's own RQ2 finds this in 81.7% of 300 sampled Defects4J bugs, leaving 18.3% where the intent-driven pipeline cannot produce a correct patch through its intended mechanism.
Editorial extensions
If this is right
- With intent inference built into repair, correct repairs no longer require a single lucky first patch: the paper's alignment study shows three intents cover the true intent in 81.7% of sampled bugs versus 62.0% for the first intent alone.
- Generated adversarial tests act as a filter during repair rather than after it, and in the evaluation they removed likely-overfitting patches for 12 Defects4J bugs and 7 HumanEval-Java bugs.
- Adversarial intent exploration also helps locate the bug, improving fault-localization precision by 13.8% and patch-generation success by 19.4% in the ablation study.
- The developer-facing output changes from a single patch to a set of inferred intents, tests, and patches, so a human can judge intent in natural language instead of reading a code diff only.
Reading between the lines
- The method's ceiling is set by intent coverage: if the LLM never proposes the true intent among the K candidates, a correct patch cannot emerge through the intended mechanism, so any way to widen or sharpen the intent search (better prompts, retrieval from issue reports, more candidates) should directly raise the repair ceiling.
- The adversarial-score threshold is a diversity heuristic, not a correctness guarantee; one could test whether requiring higher pairwise disagreement between intents actually increases correct-repair yield or instead pushes the model toward implausible intents.
- Because the pipeline outputs intent descriptions alongside patches, it offers a natural experiment on developer acceptance: if developers can reliably select the aligned intent, then intent-selection accuracy could become a repair metric in its own right.
- The same three-agent loop should transfer to other languages and defect types, but the bottleneck will likely be oracle quality in the generated tests, since incorrect expected outputs can reject correct patches; measuring oracle accuracy on a held-out set of ground-truth fixes would quantify that risk.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes AdverIntent-Agent, a multi-agent LLM-based program repair system. The system uses a reasoning agent to infer multiple deliberately adversarial program intents and locate faulty statements, a test agent to generate tests that differentiate these intents (using the same inputs with different expected outputs), and a repair agent to generate patches that satisfy each inferred intent. The authors evaluate on Defects4J 2.0 and HumanEval-Java, reporting 77 and 105 correct repairs under realistic fault localization, respectively, and claim state-of-the-art performance compared with prior APR tools. The paper also includes an ablation study (RQ2) examining the contribution of adversarial intents to fault localization and patch generation, an analysis (RQ3) of how generated tests filter overfitting patches, and a token-cost analysis (RQ4).
Significance. If the reported results hold, this is a valuable contribution to the APR literature. The paper is one of the first to shift the focus of LLM-based repair from generating diverse patches to generating diverse, adversarial program-intent hypotheses, using tests as a way to measure the degree of adversarial disagreement. The evaluation on two standard benchmarks with multiple baselines is appropriate in scope, and the authors make several praiseworthy efforts: they use exact-match and manual semantic-equivalence checks in addition to LLM-based assessment, they explicitly acknowledge data leakage and non-determinism in Section 6, and they state that patches and interaction logs will be publicly released. The idea of supplying developers with inferred intents and tests, rather than only patches, is a novel and potentially useful paradigm. However, the central empirical claim (the 77/105 counts) rests on a patch-correctness labeling procedure that is partly circular, and the reported margins over prior work are small enough that labeling and protocol differences could change the conclusion.
major comments (4)
- [§4.1.2, Table 3, Table 6]
- [§4.2, Table 3]
- [§4.3, Table 5, RQ2]
- [§4.1, §6, Table 3]
minor comments (5)
- [Table 3]
- [§3.2, equation for adversarial score]
- [§4.1.2]
- [§6]
- [§2, Figure 1]
Circularity Check
Minor self-referential correctness metric; central repair claims rest on external benchmarks and manual review.
-
self definitional
[Section 4.1.2, Patch Quality Evaluation, Top@N bullet.]
"• 𝑇𝑜𝑝@𝑁: A metric that evaluates whether at least one of the top-𝑛 generated patches is correct. A patch is considered correct if it passes both the original test cases and the automatically generated test cases."
The automatically generated test cases are produced by Agenttest from an inferred program intent, and Agentrepair is explicitly tasked with generating patches that ensure 'both the original and adversarial test cases pass' (Section 3.3). Thus a patch that passes its intended test case satisfies the very condition it was prompted to satisfy; labeling that as correct is self-confirming rather than independent evidence that the patch matches developer intent.
full rationale
The central claim is empirical and benchmarked against external corpora (Defects4J 2.0 and HumanEval-Java) and prior tools, so the repair counts are not forced by construction. The core mechanism, inferring multiple adversarial intents and generating patches per intent, is a genuine generative pipeline rather than a rename of the evaluation data. The main caveat is the Top@N correctness definition, which equates correctness with passing automatically generated tests that were themselves derived from the same inferred intent guiding patch generation; this is circular if used as the final correctness label. The paper mitigates this by requiring exact-match or two-author manual semantic equivalence for the reported correct counts, and the manual review is a standard APR practice. Minor self-citations exist (ITER [77] and SelfAPR [74] are used as baselines, and [77] justifies a fault-localization evaluation convention), but they are not load-bearing: the headline comparison does not reduce to those citations. Overall, the circularity is partial and confined to an evaluation metric, not the derivation of the method's outputs.
Assumptions & free parameters
free parameters (5)
- K (number of inferred intents) =
3
- adversarial threshold =
33.3%
- test filtering ratio =
70%
- patch refinement rounds =
3
- LLM temperature =
1
assumptions (4)
- domain assumption GPT-4o can reliably infer program intent, localize faults, generate correct test oracles, and produce correct patches from natural-language intents.
- domain assumption The two benchmarks' original test suites encode sufficient developer intent to judge patch plausibility.
- domain assumption LLM-extracted ground-truth intent from human-written patches is a reliable oracle for measuring intent alignment in RQ2.
- domain assumption Adversarial intents generated sequentially with criticism prompts are sufficiently diverse to cover the true intent.
Cite this review
Pith. "Pith review of Adversarial Reasoning for Repair Based on Inferred Program Intent." pith.science (2026). https://pith.science/paper/YN352QLZ
@misc{pith2026250513008,
author = {Pith},
title = {Pith review of: Adversarial Reasoning for Repair Based on Inferred Program Intent},
year = {2026},
howpublished = {\url{https://pith.science/paper/YN352QLZ}},
note = {Machine review of arXiv:2505.13008}
}
read the original abstract
Automated program repair (APR) has shown promising results, particularly with the use of neural networks. Currently, most APR tools focus on code transformations specified by test suites, rather than reasoning about the program intent and the high-level bug specification. Without a proper understanding of program intent, these tools tend to generate patches that overfit incomplete test suites and fail to reflect the developers intentions. However, reasoning about program intent is challenging. In our work, we propose an approach called AdverIntent-Agent, based on critique and adversarial reasoning. Our approach is novel to shift the focus from generating multiple APR patches to inferring multiple potential program intents. Ideally, we aim to infer intents that are, to some extent, adversarial to each other, maximizing the probability that at least one aligns closely with the developers original intent. AdverIntent-Agent is a multi-agent approach consisting of three agents: a reasoning agent, a test agent, and a repair agent. First, the reasoning agent generates adversarial program intents along with the corresponding faulty statements. Next, the test agent produces adversarial test cases that align with each inferred intent, constructing oracles that use the same inputs but have different expected outputs. Finally, the repair agent uses dynamic and precise LLM prompts to generate patches that satisfy both the inferred program intent and the generated tests. AdverIntent-Agent was evaluated on two benchmarks: Defects4J 2.0 and HumanEval-Java. AdverIntent-Agent correctly repaired 77 and 105 bugs in both benchmarks, respectively.
Figures
Forward citations
Cited by 1 Pith paper
-
Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing
GUIRepair, a cross-modal LLM pipeline that converts issue screenshots into reproduction code and rendered patch screenshots into validation feedback, resolves 157/517 SWE-bench M instances with GPT-4o and 175 with o4-mini.
Reference graph
Works this paper leans on
-
[1]
E. T. Barr, M. Harman, P. McMinn, M. Shahbaz, and S. Yoo. 2015. The Oracle Problem in Software Testing: A Survey. IEEE Transactions on Software Engineering 41, 05 (may 2015), 507–525. https://doi.org/10.1109/TSE.2014.2372785
arXiv 2015
-
[2]
Clark Barrett, Roberto Sebastiani, Sanjit Seshia, and Cesare Tinelli. 2009. Satisfiability modulo theories. (2009), 1–885
2009
-
[3]
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024. Graph of Thoughts: Solving Elaborate Problems with Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence 38, 16 (March 2024), 17682–176...
-
[4]
Islem Bouzenia, Premkumar Devanbu, and Michael Pradel. 2024. RepairAgent: An Autonomous, LLM-Based Agent for Program Repair. arXiv:2403.17134 [cs.SE]
arXiv 2024
-
[5]
S. Chakraborty, Y. Ding, M. Allamanis, and B. Ray. 2020. CODIT: Code Editing with Tree-Based Neural Models. IEEE Transactions on Software Engineering (2020). https://doi.org/10.1109/TSE.2020.3020502
arXiv 2020
-
[6]
Saikat Chakraborty and Baishakhi Ray. 2021. On Multi-Modal Learning of Editing Source Code. In2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE)
2021
-
[7]
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2024. Teaching Large Language Models to Self-Debug. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=KuPixIqPiq
2024
-
[8]
Z. Chen, S. J. Kommrusch, M. Tufano, L. Pouchet, D. Poshyvanyk, and M. Monperrus. 2019. SEQUENCER: Sequence- to-Sequence Learning for End-to-End Program Repair. IEEE Transactions on Software Engineering (2019)
2019
Show all 84 references
-
[9]
Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan. 2023. Automated Repair of Programs from Large Language Models. In Proceedings of the 45th International Conference on Software Engineering (ICSE ’23)
2023
-
[11]
Xiang Gao, Sergey Mechtaev, and Abhik Roychoudhury. 2019. Crash-avoiding program repair. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis . 8–18
2019
-
[12]
Luca Gazzola, Daniela Micucci, and Leonardo Mariani. 2017. Automatic Software Repair: A Survey. IEEE Transactions on Software Engineering (2017)
2017
-
[14]
Google. 2024. Large sequence models for software development activities. Google (2024)
2024
-
[15]
Dávid Hidvégi, Khashayar Etemadi, Sofia Bobadilla, and Martin Monperrus. 2024. CigaR: Cost-efficient Program Repair with LLMs. arXiv:2402.06598 [cs.SE]
2024 arXiv
-
[16]
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou
-
[17]
Elkhan Ismayilzada, Md Mazba Ur Rahman, Dongsun Kim, and Jooyong Yi. 2023. Poracle: Testing Patches under Preservation Conditions to Combat the Overfitting Problem of Program Repair. ACM Trans. Softw. Eng. Methodol. 33, 2, Article 44 (dec 2023), 39 pages. https://doi.org/10.11...
2023 doi
-
[18]
Yue Jia and Mark Harman. 2010. An analysis and survey of the development of mutation testing. IEEE transactions on software engineering 37, 5 (2010), 649–678
2010
-
[19]
Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan. 2023. Impact of Code Language Models on Automated Program Repair. arXiv:2302.05020 [cs.SE]
2023 arXiv
-
[20]
Nan Jiang, Thibaud Lutellier, Yiling Lou, Lin Tan, Dan Goldwasser, and Xiangyu Zhang. 2023. KNOD: Domain Knowledge Distilled Tree Decoder for Automated Program Repair. In Proceedings of the 45th International Conference on Software Engineering (ICSE 2023)
2023
-
[21]
Nan Jiang, Thibaud Lutellier, and Lin Tan. 2021. CURE: Code-Aware Neural Machine Translation for Automatic Program Repair. In Proceedings of the ACM/IEEE 43rd International Conference on Software Engineering
2021
-
[22]
Rene Just, Darioush Jalali, and Michael D Ernst. 2014. Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of the 2014 International Symposium on Software Testing and Analysis . ACM, 437–440
2014
-
[23]
YoungJae Kim, Seungheon Han, Askar Yeltayuly Khamit, and Jooyong Yi. 2023. Automated Program Repair from Fuzzing Perspective. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (Seattle, WA, USA) (ISSTA 2023). Association for Comput...
2023
-
[24]
Jiaolong Kong, Mingfei Cheng, Xiaofei Xie, Shangqing Liu, Xiaoning Du, and Qi Guo. 2024. ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs. arXiv:2403.01971 [cs.SE]
2024
-
[25]
Le, Duc-Hiep Chu, David Lo, Claire Le Goues, and Willem Visser
Xuan-Bach D. Le, Duc-Hiep Chu, David Lo, Claire Le Goues, and Willem Visser. 2017. JFIX: Semantics-Based Repair of Java Programs via Symbolic PathFinder. In Proceedings of the 26th ACM SIGSOFT International Symposium on Software Testing and Analysis (Santa Barbara, CA, USA) (I...
2017
-
[26]
Claire Le Goues, ThanhVu Nguyen, Stephanie Forrest, and Westley Weimer. 2012. GenProg: A generic method for automatic software repair. Software Engineering, IEEE Transactions on 38, 1 (2012), 54–72. https://doi.org/10.1109/TSE. 2011.104
2012 doi
-
[27]
Claire Le Goues, ThanhVu Nguyen, Stephanie Forrest, and Westley Weimer. 2012. GenProg: A Generic Method for Automatic Software Repair. IEEE Transactions on Software Engineering 38, 1 (Jan. 2012), 54–72. https://doi.org/10. 1109/TSE.2011.104
2012
-
[28]
Cheryl Lee, Chunqiu Steven Xia, Jen tse Huang, Zhouruixin Zhu, Lingming Zhang, and Michael R. Lyu. 2024. A Unified Debugging Approach via LLM-Based Multi-Agent Synergy. arXiv:2404.17153 [cs.SE]
2024
-
[29]
Berger, and Stephen N
Kyla Levin, Nicolas van Kempen, Emery D. Berger, and Stephen N. Freund. 2024. ChatDBG: An AI-Powered Debugging Assistant. arXiv:2403.16354 [cs.SE]
2024 arXiv
-
[30]
Yi Li, Shaohua Wang, and Tien N. Nguyen. 2020. DLFix: Context-Based Code Transformation Learning for Automated Program Repair. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering (Seoul, South Korea) (ICSE ’20). 602–614. https://doi.org/10.1145...
2020
-
[31]
Changshu Liu, Pelin Cetin, Yogesh Patodia, Baishakhi Ray, Saikat Chakraborty, and Yangruibo Ding. 2024. Automated Code Editing with Search-Generate-Modify. IEEE Transactions on Software Engineering (2024), 1–12. https://doi.org/ 10.1109/TSE.2024.3376387
2024
-
[32]
Bissyandé
Kui Liu, Anil Koyuncu, Dongsun Kim, and Tegawendé F. Bissyandé. 2019. TBar: Revisiting Template-based Automated Program Repair. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis . ACM, 31–42. https://doi.org/10.1145/3293882.3330577
2019
-
[33]
Liu and H
X. Liu and H. Zhong. 2018. Mining stackoverflow for program repair. In 2018 IEEE 25th International Conference on Software Analysis, Evolution and Reengineering (SANER)
2018
-
[34]
Fan Long and Martin Rinard. 2015. Staged program repair with condition synthesis. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering . Bergamo Italy, 166–178. https://doi.org/10.1145/2786805.2786811
2015
-
[35]
Thibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li, Moshi Wei, and Lin Tan. 2020. CoCoNuT: Combining Context-Aware Neural Translation Models Using Ensemble for Program Repair (ISSTA 2020). , Vol. 1, No. 1, Article . Publication date: August 2025. Adversarial Reasoning...
2020
-
[36]
Marginean, J
A. Marginean, J. Bader, S. Chandra, M. Harman, Y. Jia, K. Mao, A. Mols, and A. Scott. 2019. SapFix: automated end-to-end repair at scale. In Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice (Montreal, Quebec, Canada) (I...
2019
-
[37]
Matias Martinez and Martin Monperrus. 2016. ASTOR: A Program Repair Library for Java. In Proceedings of ISSTA
2016
-
[38]
Matias Martinez and Martin Monperrus. 2016. Astor: A program repair library for java. In Proceedings of the 25th International Symposium on Software Testing and Analysis . 441–444
2016
-
[39]
Sergey Mechtaev, Xiang Gao, Shin Hwei Tan, and Abhik Roychoudhury. 2018. Test-Equivalence Analysis for Automatic Patch Generation. ACM Trans. Softw. Eng. Methodol.27, 4, Article 15 (oct 2018), 37 pages. https://doi.org/10.1145/3241980
2018 doi
-
[40]
Mechtaev, J
S. Mechtaev, J. Yi, and A. Roychoudhury. 2015. DirectFix: Looking for Simple Program Repairs. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering , Vol. 1. 448–458. https://doi.org/10.1109/ICSE.2015.63
2015 doi
-
[41]
Sergey Mechtaev, Jooyong Yi, and Abhik Roychoudhury. 2016. Angelix: Scalable Multiline Program Patch Synthesis via Symbolic Analysis. In 2016 IEEE/ACM 38th International Conference on Software Engineering (ICSE)
2016
-
[42]
Martin Monperrus. 2017. Automatic Software Repair: a Bibliography. ACM Computing Surveys 51 (2017), 1–24. https://doi.org/10.1145/3105906
2017 doi
-
[43]
Hoang Duong Thien Nguyen, Dawei Qi, Abhik Roychoudhury, and Satish Chandra. 2013. SemFix: Program repair via semantic analysis. In 2013 35th International Conference on Software Engineering (ICSE) . 772–781. https://doi.org/10. 1109/ICSE.2013.6606623
2013
-
[44]
Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama
Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama. 2024. Is Self-Repair a Silver Bullet for Code Generation?. In International Conference on Learning Representations (ICLR)
2024
-
[45]
Zichao Qi, Fan Long, Sara Achour, and Martin Rinard. 2015. An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems. In Proceedings of the 2015 International Symposium on Software Testing and Analysis (Baltimore, MD, USA) (ISSTA 2015...
2015
-
[46]
Haifeng Ruan, Yuntong Zhang, and Abhik Roychoudhury. 2025. SpecRover: Code Intent Extraction via LLMs. In Proceedings of the 47th International Conference on Software Engineering (ICSE 2025) . arXiv:2408.02232 https://arxiv. org/abs/2408.02232
2025 arXiv
-
[47]
Seemanta Saha et al. 2019. Harnessing evolution for multi-hunk program repair. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) . IEEE, 13–24
2019
-
[48]
Ridwan Shariffdeen, Yannic Noller, Lars Grunske, and Abhik Roychoudhury. 2021. Concolic Program Repair. In 42nd ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI)
2021
-
[49]
Ridwan Shariffdeen, Shin Hwei Tan, Mingyuan Gao, and Abhik Roychoudhury. 2021. Automated Patch Transplantation. In ACM Transactions on Software Engineering and Methodology (TOSEM) . 1–36
2021
-
[50]
Smith, Earl T
Edward K. Smith, Earl T. Barr, Claire Le Goues, and Yuriy Brun. 2015. Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering (ESEC/FSE 2015)
2015
-
[51]
Shin Hwei Tan and Abhik Roychoudhury. 2015. relifix: Automated Repair of Software Regressions. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering , Vol. 1. 471–482. https://doi.org/10.1109/ICSE.2015.65
2015 doi
-
[52]
Prasad, and Abhik Roychoudhury
Shin Hwei Tan, Hiroaki Yoshida, Mukul R. Prasad, and Abhik Roychoudhury. 2016. Anti-patterns in Search-based Program Repair. In Proceedings of the 2016 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering (Seattle, WA, USA) (FSE 2016)
2016
-
[53]
Bissyande
Haoye Tian, Yinghua Li, Weiguo Pian, Abdoul Kader Kaboré, Kui Liu, Jacques Klein, and Tegawendé F. Bissyande
-
[54]
Bissyandé
Haoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu, Li Li, Jacques Klein, and Tegawendé F. Bissyandé. 2020. Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair. In ASE. IEEE, 981–992. https://doi.org/10.1145/3324884.3416532
2020
-
[55]
Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk. 2019. On Learning Meaningful Code Changes via Neural Machine Translation. In Proceedings of the 41st International Conference on Software Engineering (Montreal, Quebec, Canada) (ICSE ’19...
2019
-
[56]
Bing Wang, Yan Gao, Zhoujun Li, and Jian-Guang Lou. 2023. Know What I don’t Know: Handling Ambiguous and Unknown Questions for Text-to-SQL. In Findings of the Association for Computational Linguistics: ACL 2023 , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Asso...
2023 doi
-
[57]
Shangwen Wang, Ming Wen, Bo Lin, Hongjun Wu, Yihao Qin, Deqing Zou, Xiaoguang Mao, and Hai Jin. 2020. Automated Patch Correctness Assessment: How Far are We?. InASE. ACM
2020
-
[58]
Yuxiang Wei, Chunqiu Steven Xia, and Lingming Zhang. 2023. Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations...
2023
-
[59]
Westley Weimer, ThanhVu Nguyen, Claire Le Goues, and Stephanie Forrest. 2009. Automatically finding patches using genetic programming. In 2009 IEEE 31st International Conference on Software Engineering . 364–374. https: //doi.org/10.1109/ICSE.2009.5070536
2009
-
[60]
Chu-Pan Wong, Priscila Santiesteban, Christian Kästner, and Claire Le Goues. 2021. VarFix: Balancing Edit Ex- pressiveness and Search Effectiveness in Automated Program Repair. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposi...
2021
-
[61]
Chunqiu Steven Xia and Lingming Zhang. 2022. Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-Shot Learning. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering...
2022
-
[62]
Chunqiu Steven Xia and Lingming Zhang. 2024. Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA)
2024
-
[63]
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal. 2024. SORRY-Bench: Systematically Evaluating Large Language Model Saf...
2024 arXiv
-
[64]
Qi Xin and Steven Reiss. 2019. Better Code Search and Reuse for Better Program Repair. In2019 IEEE/ACM International Workshop on Genetic Improvement (GI). 10–17. https://doi.org/10.1109/GI.2019.00012
2019
-
[65]
Qi Xin and Steven P. Reiss. 2017. Identifying Test-Suite-Overfitted Patches through Test Case Generation. In ISSTA
2017
-
[66]
Xin and S
Q. Xin and S. P. Reiss. 2017. Leveraging syntax-related code for automated program repair. In 2017 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE)
2017
-
[67]
Yingfei Xiong, Xinyuan Liu, Muhan Zeng, Lu Zhang, and Gang Huang. 2018. Identifying patch correctness in test-based program repair. In Proceedings of the 40th International Conference on Software Engineering
2018
-
[68]
Yingfei Xiong, Jie Wang, Runfa Yan, Jiachen Zhang, Shi Han, Gang Huang, and Lu Zhang. 2017. Precise Condition Synthesis for Program Repair. In Proceedings of the 39th International Conference on Software Engineering (Buenos Aires, Argentina) (ICSE ’17). IEEE Press, 416–426. ht...
2017 doi
-
[69]
Jifeng Xuan, Matias Martinez, Favio Demarco, Maxime Clément, Sebastian Lamelas, Thomas Durieux, Daniel Le Berre, and Martin Monperrus. 2016. Nopol: Automatic Repair of Conditional Statement Bugs in Java Programs. IEEE Transactions on Software Engineering (2016)
2016
-
[70]
Guang Yang, Yu Zhou, Xiang Chen, Xiangyu Zhang, Terry Yue Zhuo, and Taolue Chen. 2023. Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models. arXiv:2312.05562 [cs.SE]
2023 arXiv
-
[71]
Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. SWE-agent: Agent Computer Interfaces Enable Software Engineering Language Models
2024
-
[72]
Griffiths, Yuan Cao, and Karthik Narasimhan
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 [cs.CL]
2023 arXiv
-
[73]
He Ye, Jian Gu, Matias Martinez, Thomas Durieux, and Martin Monperrus. 2021. Automated Classification of Overfitting Patches with Statically Extracted Code Features. IEEE Transactions on Software Engineering (2021). https://doi.org/10. 1109/tse.2021.3071750
2021
-
[74]
He Ye, Matias Martinez, Xiapu Luo, Tao Zhang, and Martin Monperrus. 2022. SelfAPR: Self-supervised Program Repair with Test Execution Diagnostics. arXiv preprint arXiv:2203.12755 (2022)
2022 arXiv
-
[75]
He Ye, Matias Martinez, and Martin Monperrus. 2021. Automated patch assessment for program repair at scale. Empirical Software Engineering 26, 2 (2021), 20. https://doi.org/10.1007/s10664-020-09920-w
2021 doi
-
[76]
He Ye, Matias Martinez, and Martin Monperrus. 2022. Neural Program Repair with Execution-based Backpropagation. In Proceedings of the ACM/IEEE 44th International Conference on Software Engineering
2022
-
[77]
He Ye and Martin Monperrus. 2024. ITER: Iterative Neural Repair for Multi-Location Patches. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE ’24) . Article 10, 13 pages. https://doi.org/10. 1145/3597503.3623337
2024
-
[78]
Burak Yetiştiren, Işık Özsoy, Miray Ayerdem, and Eray Tüzün. 2023. Evaluating the code quality of ai-assisted code generation tools: An empirical study on github copilot, amazon codewhisperer, and chatgpt. arXiv preprint arXiv:2304.10778 (2023)
2023 arXiv
-
[79]
Yuan Yuan and Wolfgang Banzhaf. 2018. ARJA: Automated Repair of Java Programs via Multi-Objective Genetic Programming. In IEEE Transactions on Software Engineering
2018
-
[80]
Yuan Yuan and Wolfgang Banzhaf. 2020. Toward Better Evolutionary Program Repair: An Integrated Approach. ACM Trans. Softw. Eng. Methodol. 29, 1, Article 5 (jan 2020), 53 pages. https://doi.org/10.1145/3360004 , Vol. 1, No. 1, Article . Publication date: August 2025. Adversaria...
2020 doi
-
[81]
Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. 2024. CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges. arXiv:2401.07339 [cs.SE]
2024 arXiv
-
[82]
Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. 2024. AutoCodeRover: Autonomous Program Improvement. arXiv:2404.05427 [cs.SE]
2024 arXiv
-
[83]
Qihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang, Kang Yuan, Yingfei Xiong, and Lu Zhang. 2021. A Syntax-Guided Edit Decoder for Neural Program Repair. InProceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of So...
2021
-
[84]
Armin Zirak and Hadi Hemmati. 2024. Improving Automated Program Repair with Domain Adaptation. ACM Trans. Softw. Eng. Methodol. 33, 3, Article 65 (mar 2024), 43 pages. https://doi.org/10.1145/3631972 , Vol. 1, No. 1, Article . Publication date: August 2025
2024 doi
-
[2022]
ACM Trans
Checking Patch Behaviour against Test Specification. ACM Trans. Softw. Eng. Methodol. (2022)
2022
-
[2024]
In The Twelfth International Conference on Learning Representations
Large Language Models Cannot Self-Correct Reasoning Yet. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=IkmD3fKBPQ
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.