REVIEW 3 major objections 6 minor 74 references
Multi-hunk bug repair improves when a coordinator schedules hunk order and selects among candidate patches, fixing 326 of 835 benchmark bugs and 420 with a stronger base model.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:43 UTC pith:HCCM4OAQ
load-bearing objection A genuine engineering step for multi-hunk APR, but the headline numbers rest on oracle fault locations and loosely matched baselines, so the strong claims are conditional until non-oracle and matched-budget results are shown. the 3 major comments →
MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the bottleneck for multi-location repair is scheduling and selection, not patch generation alone. The framework's key mechanism is a Coordinator-Proposer loop: the Coordinator builds a hunk-dependency graph, selects the next hunk to repair based on the current partial patch and failure relevance, solicits several independent candidate edits from Proposers, groups similar candidates, scores each by cluster support and consistency with context, previous patches, and the dependency graph, and accepts the winner before moving on. Accepted patches can later be revisited if the trajectory proves inconsistent. Two-stage refinement then compiles the assembled patch and iter
What carries the argument
The load-bearing object is the Coordinator-Proposer architecture combined with the hunk-dependency graph. The Coordinator treats repair as an autoregressive sequence over hunks: it selects the next hunk via a conditional policy, Proposers (assigned diverse sampling temperatures) generate a candidate set, and the Coordinator selects by maximizing a score of cluster agreement plus context, state, and dependency consistency. The dependency graph, with edge types for symmetric fixes, stepwise prerequisites, and change propagation, converts the repair-order problem into a scheduling policy; the revisit action lets the trajectory be corrected. This machinery operationalizes the generation-recognit
Load-bearing premise
The framework's headline numbers assume the benchmark provides the exact function-level buggy locations; without that oracle, the fix counts would likely fall, and no non-oracle measurement is reported for the main 835-bug benchmark.
What would settle it
Take a random sample of the 835-bug Java benchmark, run the same framework but replace the oracle buggy locations with a standard fault localizer's top-k output, and count correct fixes; if the count drops to baseline levels, the coordinator-proposer advantage depends on the oracle rather than on scheduling. A cheaper check: perturb the oracle by injecting one extra or one shifted hunk and observe whether the coordinator still produces a correct patch.
If this is right
- Repair order matters: dynamic scheduling by a coordinator fixes more multi-hunk bugs than repairing hunks in the order they appear in the developer patch or in random order, across every base model tested.
- Candidate diversity plus selection beats self-correction: the proposer pool with varied temperatures and coordinator selection outperforms generating one patch and iteratively asking the same model to fix it.
- The advantage concentrates on complex bugs: the largest gaps over baselines are in multi-method and multi-file categories, with smaller gaps on single-hunk bugs.
- Syntax refinement and test refinement both add correct fixes, with the larger share coming from compile-error fixes, suggesting many near-miss patches are syntactically broken.
- The same mechanism transfers to vulnerability repair: the framework repairs multi-hunk vulnerabilities in Java and other languages under the same base model, including several fixed after the model's release.
Where Pith is reading between the lines
- Because the paper only reports non-oracle results on the vulnerability benchmark, the full 835-bug Java benchmark number's sensitivity to localization errors remains unknown; an end-to-end evaluation would likely lower the headline count when oracle buggy locations are replaced by a fault localizer.
- The coordinator's selection is scoring by consistency rather than by executing tests, so its upper bound depends on how accurately an LLM can judge patch correctness; replacing that judge with a semantic or execution-based checker is a natural extension.
- The scheduling insight—repair state-producing hunks before state-consuming hunks—is language-agnostic; one testable extension is applying the dependency-graph scheduler to non-Java projects using lightweight static analysis to infer producer/consumer relations.
- If generation-recognition asymmetry is the real driver, then even single-hunk repair might benefit from coordinator-style selection when candidate diversity is high—something the paper does not directly test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MultiFixer (styled Multi2Fixer in the body), a Coordinator-Proposer multi-agent framework for repairing multi-hunk bugs. The framework operates in four stages: tool-augmented bug analysis, repair-context construction, iterative patch generation in which a Coordinator schedules hunk repairs and selects among candidates proposed by multiple Proposers, and two-stage syntactic/test refinement. The authors evaluate on 835 Defects4J bugs and three vulnerability benchmarks (VUL4J, SEC-bench, PatchEval), reporting 326 correct fixes on Defects4J with GPT-3.5, 420 correct fixes with Claude-3.5-Sonnet, and 24 VUL4J fixes. They claim these results outperform prior APR baselines and establish a new state of the art on Defects4J.
Significance. If the headline results held under controlled conditions, the paper would make a meaningful contribution: multi-hunk repair is a recognized weakness of LLM-based APR, and the Coordinator-Proposer design is a plausible mechanism for inter-hunk scheduling and candidate selection. The paper has several strengths: it ships an implementation on Zenodo, includes a detailed ablation study across five base models, reports a time-separated data-leakage analysis, and includes a trajectory case study that makes the scheduling mechanism concrete. However, the significance is currently conditional on two load-bearing assumptions that are not adequately controlled: all Defects4J numbers use oracle fault locations, and the baseline comparisons are drawn from prior papers with different scopes and patch budgets. These issues must be resolved before the empirical claims can be taken at face value.
major comments (3)
- [§4.5, Tables 2, 4, 9] The central empirical claim—326 Defects4J fixes with GPT-3.5 and 420 with Claude-3.5-Sonnet—rests on the sentence in §4.5: 'we use oracle function-level buggy locations provided by the benchmark setting as input to the repair framework.' The framework therefore never performs fault localization on Defects4J; the exact set of buggy hunks is supplied. The paper's own non-oracle experiment on VUL4J (Multi2Fixer* in Table 9) shows a drop from 24 to 22 total fixes and from 5 to 3 multi-hunk fixes. No comparable non-oracle experiment is reported on Defects4J, so the magnitude of the degradation on the headline benchmark is unknown. Because the Coordinator schedules and selects hunks from the provided set, an incomplete or incorrect hunk set would change both the scheduling and the candidate-selection behavior. The claims should either be re-scoped to 'repair given oracle fault locations' or ac
- [§4.3, Table 2] The comparison with prior APR baselines is not a matched comparison. Baseline results are 'taken directly from the original papers' (§4.3), and Table 2 itself shows that baselines were evaluated on different scopes (476, 483, 337, 835, 372 bugs) and with very different patch budgets (ranging from 1 to 1000). On the full-scope comparison, the only full-scope baselines are RepairAgent, MultiMend, and PReMM, but RepairAgent and MultiMend use different base models/settings and much larger patch budgets. The abstract and §1 state that MultiFixer 'outperforms prior APR baselines in the reported comparisons' and 'establishes a new state of the art on Defects4J'; those statements outrun the evidence. To support the claim, the authors should either run a representative set of baselines under the same oracle settings, patch budget, and base model, or explicitly restrict the claim to 'outperforms t
- [§5.3 RQ3.2, Table 6] The line-level context length of 20 is selected by grid search on the 372 multi-hunk Defects4J subset, which is part of the same benchmark used for the headline results. In particular, RQ3.2 states: 'we experiment with different lengths ranging from 5 to 50 in increments of 5. By comparing the results, we find that a context length of 20 achieves the best trade-off.' Other hyperparameters—number of Proposers K, temperatures, maximum refinement iterations, repair rounds, tool-invocation budget—appear to be fixed after similar development on Defects4J. This is not circularity, but it creates a risk of overfitting to Defects4J that is not quantified by the reported ablation. The paper should either report a validation/holdout split, or evaluate the chosen hyperparameters on the vulnerability benchmarks as out-of-sample tests.
minor comments (6)
- [Title/Abstract and §1] The framework name is inconsistent: the title and abstract use 'MultiFixer,' while the body consistently uses 'Multi2Fixer.' This should be unified.
- [Table 2] Several rows have '-' entries that are hard to interpret; e.g., RepairLLaMA and ContrastRepair show odd column alignments. A note explaining which entries are inapplicable versus unreported would improve clarity.
- [Eq. (1)] The notation has spacing artifacts (e.g., '𝜋𝛾(𝑃𝑖|𝐵 𝑖,𝐻𝑖)') and the decomposition assumes hunks are ordered; the definition of the ordering for a multi-hunk bug is introduced only later. Clarify the notation.
- [§3.3.2] The three context granularities are described, but the choice of 'buggy hunks' plus line-level context of 20 is not justified at that point; the justification appears only in the ablations. Consider moving or referencing the ablation justification earlier.
- [§6.1] The time-separated analysis in Table 12 is a useful check, but it only reports counts; the total number of post-2022 vulnerabilities in each benchmark is not given, so the reader cannot assess the baseline rates. Reporting denominators would strengthen the argument.
- [§7] The internal-validity paragraph mentions stochasticity mitigation via five repair rounds, but the sensitivity experiment in RQ3.3 shows 3–10% performance variation when K changes. The paper should state whether the reported numbers are single runs or averaged over seeds, and if averaged, over how many.
Circularity Check
No circularity: empirical claims rest on external benchmarks; the oracle fault-location setting is explicitly disclosed and not an equivalence.
full rationale
The paper is an empirical evaluation of a new APR pipeline against external benchmarks. Its load-bearing claims (326 Defects4J fixes, 420 with Claude-3.5-Sonnet, 24 VUL4J fixes) are counts of bugs for which generated patches passed tests and manual semantic checks; they are not derived from the paper's equations or from a self-citation. The oracle fault-location setting is explicitly disclosed in §4.5 ('we use oracle function-level buggy locations provided by the benchmark setting as input to the repair framework'), and the no-oracle variant (Multi2Fixer*) is separately reported in Table 9, showing a decline from 24 to 22 VUL4J fixes. This is a scope/validity limitation, not a circularity: the gold locations define the multi-hunk task, but the generated patches are not equivalent to those locations by construction. The Coordinator-Proposer formalization (Eqs. 1–7) describes the method rather than deriving the outcome; no parameter is solved for from the headline results. Hyperparameter choices such as line-level context length 20 and proposer count 3 are tuned on subsets of Defects4J, which is a test-set-tuning threat to external validity, but it does not make the headline fix counts equal to the fitted values. Self-citations (e.g., Refs. 15–18, 37–38, 62–68) are background or prior-framework papers and are not invoked as load-bearing evidence for the central comparative claims. No uniqueness theorem is imported, and no ansatz is smuggled via self-citation. I therefore find no circular step.
Axiom & Free-Parameter Ledger
free parameters (7)
- Number of Proposers K =
3
- Line-level context length =
20
- Proposer temperatures =
0, 0.5, 1.0
- Maximum hunk-repair steps per round =
20
- Syntax and test refinement iterations =
3 and 3
- Repair rounds per bug =
5
- Maximum tool invocations in BugAnalyzer =
20
axioms (5)
- domain assumption Generation-Recognition Asymmetry: selecting a correct patch from candidates is easier than generating it directly.
- domain assumption Oracle function-level fault locations are a reasonable substitute for practical repair input.
- domain assumption Defects4J hunk-dependency patterns (Symmetric Fixing, Stepwise Fixing, Change Propagation) generalize to other languages and benchmarks.
- domain assumption Manual semantic-equivalence checking is a reliable ground truth for Correct Fix labels.
- domain assumption JavaParser-based static-analysis tools accurately provide repository structure, types, and method bodies.
read the original abstract
Automated Program Repair (APR) has benefited greatly from Large Language Models (LLMs), but existing LLM-based APR methods still struggle with multi-hunk bugs that require coordinated changes across multiple locations. These bugs demand repository-level context understanding, repair-order scheduling, and effective hunk-level patch generation and selection. To address these challenges, we propose MultiFixer, a novel Coordinator-Proposer based multi-agent framework for multi-hunk repair. MultiFixer performs tool-augmented bug analysis, constructs fine-grained repair context, iteratively generates patches through a Coordinator-Proposer architecture, and applies two-stage patch refinement for syntactic and semantic correctness. We evaluate MultiFixer on 835 bugs from Defects4J and three vulnerability benchmarks. On Defects4J, MultiFixer fixes 326 bugs, including 62 multi-method and 27 multi-file bugs, and outperforms prior APR baselines in the reported comparisons with the same base model. Moreover, MultiFixer also fixes 46 multi-hunk bugs among 95 unique fixes. When combined with Claude-3.5-Sonnet, MultiFixer repairs 420 bugs, establishing a new state of the art on Defects4J. On VUL4J, MultiFixer repairs 24 real-world vulnerabilities, including 5 multi-hunk cases. On the multi-hunk subsets of SEC-bench and PatchEval, MultiFixer fixes 11 and 19 vulnerabilities, respectively, outperforming all compared baselines under GPT-3.5. These results demonstrate the effectiveness of MultiFixer for multi-hunk repair.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report.arXiv preprint arXiv:2303.08774 (2023)
Pith/arXiv arXiv 2023
-
[2]
Islem Bouzenia, Premkumar Devanbu, and Michael Pradel. 2024. Repairagent: An autonomous, llm-based agent for program repair.arXiv preprint arXiv:2403.17134 (2024)
Pith/arXiv arXiv 2024
-
[3]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural information processing systems33 (2020), 1877–1901
2020
-
[4]
Quang-Cuong Bui, Ranindya Paramitha, Duc-Ly Vu, Fabio Massacci, and Riccardo Scandariato. 2024. APR4Vul: an empirical study of automatic program repair techniques on real-world Java vulnerabilities.Empirical software engineering29, 1 (2024), 18
2024
-
[5]
Quang-Cuong Bui, Riccardo Scandariato, and Nicolás E Díaz Ferreyra. 2022. Vul4j: A dataset of reproducible java vulnerabilities geared towards the study of program repair techniques. InProceedings of the 19th International Conference on Mining Software Repositories. 464–468
2022
-
[6]
Viola Campos, Ridwan Shariffdeen, Adrian Ulges, and Yannic Noller. 2025. Empir- ical Evaluation of Generalizable Automated Program Repair with Large Language Models.arXiv preprint arXiv:2506.03283(2025)
arXiv 2025
-
[7]
Jialun Cao, Meiziniu Li, Ming Wen, and Shing-chi Cheung. 2025. A study on prompt design, advantages and limitations of chatgpt for deep learning program repair.Automated Software Engineering32, 1 (2025), 1–29
2025
-
[8]
Dawn Drain, Chen Wu, Alexey Svyatkovskiy, and Neel Sundaresan. 2021. Gen- erating bug-fixes using pretrained transformers. InProceedings of the 5th ACM SIGPLAN international symposium on machine programming. 1–8
2021
-
[9]
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2024. Evaluating large language models in class-level code generation. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13
2024
-
[10]
Zhiyu Fan, Haifeng Ruan, Sergey Mechtaev, and Abhik Roychoudhury. 2024. Oracle-guided Program Selection from Large Language Models. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 628–640
2024
-
[11]
Mahdi Farzandway and Fatemeh Ghassemi. 2025. Automated repair of c programs using large language models.arXiv preprint arXiv:2509.01947(2025)
Pith/arXiv arXiv 2025
-
[12]
Reza Gharibi, Mohammad Hadi Sadreddini, and Seyed Mostafa Fakhrahmad
-
[13]
Sichong Hao, Xianjun Shi, Hongwei Liu, and Yanjun Shu. 2023. Enhancing code language models for program repair by curricular fine-tuning framework. In2023 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 136–146
2023
-
[14]
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John C. Grundy, and Haoyu Wang. 2023. Large Language Models for Software Engineering: A Systematic Literature Review.CoRRabs/2308.10620 (2023), arXiv–2308
Pith/arXiv arXiv 2023
-
[15]
Haichuan Hu, Ye Shang, Weifeng Sun, and Quanjun Zhang. 2025. TSAPR: A Tree Search Framework For Automated Program Repair.arXiv preprint arXiv:2507.01827(2025)
Pith/arXiv arXiv 2025
-
[16]
Haichuan Hu, Ye Shang, Guolin Xu, Congqing He, and Quanjun Zhang. 2025. Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs. In2025 IEEE/ACM International Workshop on Automated Program Repair (APR). 11–18. doi:10.1109/APR66717.2025.00007
arXiv 2025
-
[17]
Haichuan Hu, Guoqing Xie, Quanjun Zhang, Jiawei Liu, Shengcheng Yu, Chun- rong Fang, Zhenyu Chen, and Liang Xiao. 2026. EvoRepair: Enhancing Vulnera- bility Repair Agents Through Experience-Based Self-Evolution.arXiv preprint arXiv:2605.30105(2026)
Pith/arXiv arXiv 2026
-
[18]
Haichuan Hu, Xiaochen Xie, and Quanjun Zhang. 2025. Repair-r1: Better test before repair.arXiv preprint arXiv:2507.22853(2025)
Pith/arXiv arXiv 2025
-
[19]
Kai Huang, Jian Zhang, Xiangxin Meng, and Yang Liu. 2025. Template-Guided Program Repair in the Era of Large Language Models.. InICSE. 1895–1907
2025
-
[20]
Zhili Huang, Ling Xu, Chao Liu, Weifeng Sun, Xu Zhang, Yan Lei, Meng Yan, and Hongyu Zhang. 2025. DynaFix: Iterative Automated Program Repair Driven by Execution-Level Dynamic Information.arXiv preprint arXiv:2512.24635(2025)
Pith/arXiv arXiv 2025
-
[21]
Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2024. A survey on large language models for code generation.arXiv preprint arXiv:2406.00515 (2024)
Pith/arXiv arXiv 2024
-
[22]
Nan Jiang, Thibaud Lutellier, and Lin Tan. 2021. Cure: Code-aware neural machine translation for automatic program repair. In2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 1161–1173
2021
-
[23]
René Just, Darioush Jalali, and Michael D Ernst. 2014. Defects4J: A database of ex- isting faults to enable controlled testing studies for Java programs. InProceedings of the 2014 international symposium on software testing and analysis. 437–440
2014
-
[24]
Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du, and Qi Guo. 2025. Contrastrepair: Enhancing conversation-based automated program repair via contrastive test case pairs.ACM Transactions on Software Engineering and Methodology34, 8 (2025), 1–31
2025
-
[25]
Ummay Kulsum, Haotian Zhu, Bowen Xu, and Marcelo d’Amorim. 2024. A case study of llm for automated vulnerability repair: Assessing impact of reasoning and patch validation feedback. InProceedings of the 1st ACM International Conference on AI-Powered Software. 103–111
2024
-
[26]
Hwiwon Lee, Ziqi Zhang, Hanxiao Lu, and Lingming Zhang. 2025. SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks. arXiv preprint arXiv:2506.11791(2025)
arXiv 2025
-
[27]
Fengjie Li, Jiajun Jiang, Jiajun Sun, and Hongyu Zhang. 2025. Hybrid automated program repair by combining large language models and program analysis.ACM Transactions on Software Engineering and Methodology34, 7 (2025), 1–28
2025
-
[28]
Yizhou Liu, Pengfei Gao, Xinchen Wang, Jie Liu, Yexuan Shi, Zhao Zhang, and Chao Peng. 2024. Marscode agent: Ai-native automated bug fixing.arXiv preprint arXiv:2409.00899(2024)
Pith/arXiv arXiv 2024
-
[29]
Ehsan Mashhadi and Hadi Hemmati. 2021. Applying codebert for automated pro- gram repair of java simple bugs. In2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR). IEEE, 505–509
2021
-
[31]
Noor Nashid, Daniel Ding, Keheliya Gallaba, Ahmed E Hassan, and Ali Mesbah
-
[32]
arXiv preprint arXiv:2511.11012(2025)
Beyond Accuracy: Behavioral Dynamics of Agentic Multi-Hunk Repair. arXiv preprint arXiv:2511.11012(2025)
Pith/arXiv arXiv 2025
-
[33]
Anvith Pabba, Alex Mathai, Anindya Chakraborty, and Baishakhi Ray. 2025. SemAgent: A Semantics Aware Program Repair Agent.arXiv preprint arXiv:2506.16650(2025)
Pith/arXiv arXiv 2025
-
[34]
Characterizing Multi-Hunk Patches: Divergence, Proximity, and LLM Repair Challenges.arXiv preprint arXiv:2506.04418(2025)
arXiv 2025
-
[35]
Noor Nashid, Mifta Sintaha, and Ali Mesbah. 2023. Retrieval-based prompt selection for code-related few-shot learning. In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2450–2462
2023
-
[36]
Pat Rondon, Renyao Wei, José Cambronero, Jürgen Cito, Aaron Sun, Siddhant Sanyam, Michele Tufano, and Satish Chandra. 2025. Evaluating agent-based program repair at google.arXiv preprint arXiv:2501.07531(2025)
Pith/arXiv arXiv 2025
-
[37]
Romain Peyrichou. 2026. The Generation-Recognition Asymmetry: Six Dimen- sions of a Fundamental Divide in Formal Language Theory.arXiv preprint arXiv:2603.10139(2026)
Pith/arXiv arXiv 2026
-
[38]
Joseph Renzullo, Pemma Reiter, Westley Weimer, and Stephanie Forrest. 2025. Automated Program Repair: Emerging trends pose and expose problems for benchmarks.Comput. Surveys57, 8 (2025), 1–18
2025
-
[39]
André Silva, Sen Fang, and Martin Monperrus. 2025. Repairllama: Efficient representations and fine-tuned adapters for program repair.IEEE Transactions on Software Engineering(2025)
2025
-
[40]
Ye Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu, Jianyi Zhou, and Zhenyu Chen. 2025. A large-scale empirical study on fine-tuning large language models for unit testing.Proceedings of the ACM on Software Engineering2, ISSTA (2025), 1678–1700
2025
-
[41]
Ye Shang, Quanjun Zhang, Haichuan Hu, Chunrong Fang, Liang Xiao, and Zhenyu Chen. 2026. Breaking, Stale, or Missing? Benchmarking Coding Agents on Project- Level Test Evolution.arXiv preprint arXiv:2605.06125(2026)
Pith/arXiv arXiv 2026
-
[42]
Alex Wang and Kyunghyun Cho. 2019. BERT has a mouth, and it must speak: BERT as a Markov random field language model.arXiv preprint arXiv:1902.04094 (2019)
Pith/arXiv arXiv 2019
-
[43]
Dominik Sobania, Martin Briesch, Carol Hanna, and Justyna Petke. 2023. An analysis of the automatic bug fixing performance of chatgpt. In2023 IEEE/ACM International Workshop on Automated Program Repair (APR). IEEE, 23–30
2023
-
[44]
Felix Stahlberg. 2020. Neural machine translation: A review.Journal of Artificial Intelligence Research69 (2020), 343–418
2020
-
[45]
Zichao Wei, Jun Zeng, Ming Wen, Zeliang Yu, Kai Cheng, Yiding Zhu, Jingyi Guo, Shiqi Zhou, Le Yin, Xiaodong Su, and Zhechao Ma. 2025. PATCHEVAL: A New Multi2Fixer: ACoordinator-ProposerBased Multi-Agent Framework For Fixing Multi-Hunk Bugs ASE’26, October 12–16, 2026, Munich, Germany Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities.arX...
arXiv 2025
-
[46]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems35 (2022), 24824–24837
2022
-
[47]
Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, and Sida I Wang. 2025. Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution.arXiv preprint arXiv:2502.18449(2025)
Pith/arXiv arXiv 2025
-
[48]
Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang. 2023. Automated program repair in the era of large pre-trained language models. In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1482–1494
2023
-
[49]
Yi Wu, Nan Jiang, Hung Viet Pham, Thibaud Lutellier, Jordan Davis, Lin Tan, Petr Babkin, and Sameena Shah. 2023. How effective are neural networks for fixing security vulnerabilities. InProceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 1282–1294
2023
-
[50]
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang. 2024. Agentless: Demystifying llm-based software engineering agents.arXiv preprint arXiv:2407.01489(2024)
Pith/arXiv arXiv 2024
-
[51]
Linna Xie, Zhong Li, Yu Pei, Zhongzhen Wen, Kui Liu, Tian Zhang, and Xuandong Li. 2025. PReMM: LLM-Based Program Repair for Multi-method Bugs via Divide and Conquer.Proceedings of the ACM on Programming Languages9, OOPSLA2 (2025), 1316–1344
2025
-
[52]
Chunqiu Steven Xia and Lingming Zhang. 2024. Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 819–831
2024
-
[53]
Jiahong Xiang, Xiaoyang Xu, Fanchu Kong, Mingyuan Wu, Zizheng Zhang, Haotian Zhang, and Yuqun Zhang. 2024. How far can we go with practical function-level program repair?arXiv preprint arXiv:2404.12833(2024)
Pith/arXiv arXiv 2024
-
[54]
Aidan ZH Yang, Sophia Kolak, Vincent Hellendoorn, Ruben Martins, and Claire Le Goues. 2025. Revisiting unnaturalness for automated program repair in the era of large language models. In2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 2561–2573
2025
-
[55]
Qi Xin, Haojun Wu, Steven P Reiss, and Jifeng Xuan. 2024. Towards Practical and Useful Automated Program Repair for Debugging.arXiv preprint arXiv:2407.08958 (2024)
Pith/arXiv arXiv 2024
-
[56]
Junjielong Xu, Ying Fu, Shin Hwei Tan, and Pinjia He. 2025. Aligning the objective of llm-based program repair. In2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 2548–2560
2025
-
[57]
He Ye and Martin Monperrus. 2024. Iter: Iterative neural repair for multi-location patches. InProceedings of the 46th IEEE/ACM international conference on software engineering. 1–13
2024
-
[58]
Aidan ZH Yang, Sophia Kolak, Vincent J Hellendoorn, Ruben Martins, and Claire Le Goues. 2024. Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models.arXiv preprint arXiv:2404.15236(2024)
Pith/arXiv arXiv 2024
-
[59]
Boyang Yang, Zijian Cai, Fengling Liu, Bach Le, Lingming Zhang, Tegawendé F Bissyandé, Yang Liu, and Haoye Tian. 2025. A survey of LLM-based automated program repair: Taxonomies, design paradigms, and applications.arXiv preprint arXiv:2506.23749(2025)
arXiv 2025
-
[60]
Wei Yuan, Quanjun Zhang, Tieke He, Chunrong Fang, Nguyen Quoc Viet Hung, Xiaodong Hao, and Hongzhi Yin. 2022. CIRCLE: Continual repair across program- ming languages. InProceedings of the 31st ACM SIGSOFT international symposium on software testing and analysis. 678–690
2022
-
[61]
Xin Yin, Chao Ni, Shaohua Wang, Zhenhao Li, Limin Zeng, and Xiaohu Yang
-
[62]
Quanjun Zhang, Chunrong Fang, Siqi Gu, Ye Shang, Zhenyu Chen, and Liang Xiao. 2025. Large Language Models for Unit Testing: A Systematic Literature Review.arXiv preprint arXiv:2506.15227(2025)
Pith/arXiv arXiv 2025
-
[63]
Zheng Yu, Ziyi Guo, Yuhang Wu, Jiahao Yu, Meng Xu, Dongliang Mu, Yan Chen, and Xinyu Xing. 2025. PatchAgent: A practical program repair agent mimicking human expertise. InProceedings of the 34th USENIX Security Symposium (USENIX Security’25), Seattle, W A, USA
2025
-
[64]
Quanjun Zhang, Chunrong Fang, Yang Xie, YuXiang Ma, Weisong Sun, Yun Yang, and Zhenyu Chen. 2024. A systematic literature review on large language models for automated program repair.arXiv preprint arXiv:2405.01466(2024)
arXiv 2024
-
[65]
Jiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li, and Milos Glig- oric. 2022. Coditt5: Pretraining for source code and natural language editing. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. 1–12
2022
-
[66]
Quanjun Zhang, Haichuan Hu, Chunrong Fang, Ye Shang, Tao Zheng, Zhenyu Chen, Yun Yang, and Liang Xiao. 2026. On the Effectiveness of Code Representa- tion in Deep Learning-Based Automated Patch Correctness Assessment.arXiv preprint arXiv:2603.07520(2026)
Pith/arXiv arXiv 2026
-
[67]
Quanjun Zhang, Chunrong Fang, Yuxiang Ma, Weisong Sun, and Zhenyu Chen
-
[68]
Quanjun Zhang, Yi Zheng, Ye Shang, Weifeng Sun, Haichuan Hu, Chunrong Fang, Zhenyu Chen, and Liang Xiao. 2026. ReProAgent: Tool-Augmented Multi-Stage Agentic Generation of Bug Reproduction Tests from Issue Reports.arXiv preprint arXiv:2607.09123(2026)
Pith/arXiv arXiv 2026
-
[69]
Jiuang Zhao, Donghao Yang, Li Zhang, Xiaoli Lian, Zitian Yang, and Fang Liu
-
[70]
Quanjun Zhang, Chunrong Fang, Tongke Zhang, Bowen Yu, Weisong Sun, and Zhenyu Chen. 2023. Gamma: Revisiting template-based automated program repair via mask prediction. In2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 535–547
2023
-
[72]
Quanjun Zhang, Ye Shang, Haichuan Hu, Chunrong Fang, Zhenyu Chen, and Liang Xiao. 2026. ComPass: Contrastive Learning for Automated Patch Correct- ness Assessment in Program Repair.arXiv preprint arXiv:2602.07561(2026)
Pith/arXiv arXiv 2026
-
[75]
InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering
Enhancing Automated Program Repair with Solution Design. InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 1706–1718
-
[2023]
A survey of learning-based automated program repair.ACM Transactions on Software Engineering and Methodology33, 2 (2023), 1–69
2023
-
[2024]
InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis
Thinkrepair: Self-directed automated program repair. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 1274–1286
-
[2025]
MultiMend: Multilingual Program Repair with Context Augmentation and Multi-Hunk Patch Generation.arXiv preprint arXiv:2501.16044(2025)
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.