REVIEW 3 major objections 5 minor 67 references
The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Fine-tuning instruction-tuned LLMs on just 1,000 examples improves program repair by up to 78%, challenging the need for large APR datasets.
desk verdict A genuinely useful APR empirical sweep, but the headline FFT-1K gain needs a parse-failure check before you trust it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is an iterative repair pipeline with a hard cap of ten patches per bug. The pipeline prompts a model with bug-delimited code, validates the generated patch by compiling and running the benchmark tests, and feeds the failing test code, a timeout notice, or the compilation error back into the chat context for the next attempt. Seven generation strategies (A: 10×1, B: 8-2, C: 5×2, D: 6-2-2, E: 4-3-3, F: 2×5, G: 1×10) carve the ten-patch budget into different splits of outputs per iteration and iterations, so the paper can isolate the value of iteration from the value of sampling breadth. The fine-tuning axis pairs full fine-tuning (FFT) against LoRA at dataset sizes 1K, 30K, and 65K, giving the three instruction-tuned models a factorial comparison.
What would settle it
Run the full pipeline (all 12 model variants and seven strategies) and validate every plausible patch, not just a third of them, against a stronger oracle—for instance, additional hidden tests for HumanEval-Java and extra edge-case tests for Defects4J. If the overfitting rate is roughly constant across strategies, the paper's rankings survive; if iterative strategies or FFT-1K patches overfit more heavily than average, the 78% gain and the iteration-benefit findings lose their force. A second check is to rerun FFT-1K on several different random 1K subsets to see whether the reported improvement is stable across sampling seeds.
Extended reading notes
Core claim
The paper claims that for instruction-tuned LLMs doing automatic program repair, both fine-tune data size and iteration count show a non-monotonic relationship with the number of plausible patches produced. Full fine-tuning on 1K samples lifts CodeLlama from 60 to 107 solved HumanEval-Java problems, DeepSeek-Coder from 76 to 129, and Llama3.1 from 68 to 108—gains of 59–78%—while 65K samples give no further gain and often regress, which the authors attribute to overfitting. LoRA does not match this at 1K but catches up with 30K–65K samples. The same pipeline shows that base models consistently improve when patches are refined over several rounds with execution feedback, whereas fine-tuned models peak with few iterations on simple bugs; on the harder Defects4J benchmark, iterative refinement helps even fine-tuned models, with Llama3.1 Base rising from 28 to 74 plausible patches between the least and most iterative strategy. The paper concludes that the best configuration is a balanced strategy that combines multi-output generation with iterative refinement, tuned per model and per task complexity.
Load-bearing premise
The headline numbers all count a patch as a success if it compiles and passes the benchmark's test suite; if the roughly 4% overfitting rate found in the 3,298 manually inspected plausible patches extends to the other ~6,000, the reported gains and strategy comparisons shrink, though they need not reverse.
Editorial extensions
If this is right
- APR practitioners can obtain large gains from full fine-tuning on about 1,000 curated examples, instead of collecting and training on tens of thousands of fixes.
- Given a budget of ten patches, generating all ten at once is rarely optimal: base models repair more bugs with balanced iterative strategies such as 6-2-2 or 4-3-3, and only extreme iteration (1×10) shows diminishing returns.
- Fine-tuned models concentrate their successes in the first few outputs, so stopping after the first five patches captures about 90% of their plausible patches on HumanEval-Java; later outputs add less than 10% to total repairs.
- The choice between full fine-tuning and LoRA changes the data requirement: FFT works at 1K samples while LoRA needs 30K–65K, reversing the usual PEFT-led advice for APR.
- On complex real-world bugs, iterative refinement remains valuable even for fine-tuned models, so the best strategy depends on task difficulty, not just the model.
Reading between the lines
- If the overfitting interpretation is right, then increasing the diversity or quality of the fine-tuning set, or adding regularization, should push the 30K/65K FFT curves above the 1K peak; this is a direct, testable way to separate overfitting from model-capacity limits.
- The 4% overfitting rate found in the manually-checked plausible patches (131 of 3,298) implies the absolute gains over base models are upper bounds on true fixes; ranking the strategies under a stricter correctness oracle (extra tests or human review) could change which strategy looks best, especially for iterative variants that produce many similar patches.
- Because base models with iteration solve a set of problems that fine-tuned models miss (about 12% unique on HumanEval-Java for Llama3.1, 19% on Defects4J), a hybrid that routes easy bugs to a fine-tuned single-shot model and hard bugs to an iterative base model could solve more total problems than either alone.
- The position analysis suggests an adaptive stopping rule: for fine-tuned models on simple benchmarks, the 10-patch budget is excessive, and compute could be reallocated to more bugs or to iterative refinement on complex ones.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates an APR pipeline that limits patch generation to 10 per bug and combines instruction-tuned LLMs with fine-tuning at three dataset sizes (1K, 30K, 65K) and two methods (FFT and LoRA). The evaluation on HumanEval-Java and Defects4J compares seven generation strategies that trade off outputs per iteration against number of iterations. The central claims are that FFT with only 1K samples yields large increases in the number of plausible patches (up to 78%), that larger fine-tuning datasets produce diminishing returns, and that base models benefit from iterative refinement more than fine-tuned models do.
Significance. If the results hold, the finding that FFT with around 1K samples is effective and underused is practically important for APR, and it challenges prior studies that reported limited gains from full fine-tuning. The paper's strengths include the release of the full pipeline, nine fine-tuned models, generated patches, and manual assessments, which enable replication and further analysis. The systematic comparison of iterative strategies and the use of two benchmarks are also valuable. However, the headline numbers rest on a parsing assumption that has not been validated across model conditions, and the RQ2 analysis selects variants using the evaluation benchmark itself; both issues need to be addressed before the central claims can be accepted.
major comments (3)
- [3.4, Table 1] Section 3.4 states that the generated output is parsed by looking for the triple backquote symbol, with no fallback described for responses that contain the fixed method without code fences. The fine-tuning data are rephrased to conform with an instruction format (Section 3.3), so fine-tuned models are trained to emit fenced code, while the base models are not. The paper does not report how often each model condition produced no parsable fenced block or an empty extraction. Because Table 1 compares base and fine-tuned models under this same parser, a systematic undercounting of base-model outputs due to format non-compliance would inflate the reported improvements (e.g., CodeLlama 60 to 107, DeepSeek-Coder 76 to 129). Please report parse-failure and uncompilable rates per model condition and recompute the main tables with a fallback parser that treats the whole response as code when no fence is found.
- [4.2] Section 4.2 states that for RQ2 the authors selected 'the base model as well as the best performing variants trained with FFT and LoRA on HumanEval-Java according to RQ1'. Because the RQ2 evaluation uses the same HumanEval-Java benchmark, the reported comparisons in Figure 4a are based on variants selected for their RQ1 performance on that benchmark, which is a form of selection bias. This is load-bearing for Finding 6 and the fine-tuned-versus-base conclusions in RQ2. The authors should either use a separate selection benchmark or a pre-registered criterion, or explicitly present the HumanEval-Java RQ2 numbers as post-selection estimates rather than unbiased evaluations.
- [4.1.1, Findings 2 and 6] The claims that diminishing returns and reduced iterative effectiveness are 'likely due to overfitting' are not supported by direct evidence in the manuscript. No training loss curves, held-out performance, output diversity measurements, or other overfitting indicators are reported. This is a plausible hypothesis, but as stated it is speculation, and the causal language is stronger than the evidence warrants. Please add supporting evidence or soften the claims.
minor comments (5)
- [3.1] In Section 3.1, 'placement of plausible patches withing the generated outputs' should be 'within the generated outputs'.
- [4.1.1] In Section 4.1.1, 'Both, FTT and LoRA methods' should read 'Both FFT and LoRA methods'.
- [Figure 3] The Venn diagram in Figure 3 lists counts without explaining how the regions correspond to the four strategy/condition combinations; a legend or a short description of the circle ordering would improve readability.
- [4.2.2, Figure 4] The heatmaps in Figure 4 show differences relative to Strategy A, but the body text does not restate this when discussing the numeric values; consider making the baseline explicit in the text for readers who do not inspect the caption.
- [References] Reference [27] is missing the publication venue and year; please complete the bibliographic details.
Circularity Check
No circular derivation found: all headline numbers are external-benchmark measurements and the only self-citation is non-load-bearing background.
full rationale
The paper's core claims — FFT with 1K samples increases plausible patches, and base models benefit more from iterative strategies than fine-tuned models — are measured on held-out external benchmarks (HumanEval-Java and Defects4J) using an independently defined success criterion: a patch compiles and passes all tests (Section 3.6). The fine-tuning corpus is a separate dataset, and no fitted parameter, selected hyperparameter, or training-set statistic is reused as an outcome. The only self-citation, reference [40], appears in related work supporting the general usefulness of feedback loops alongside independent citations, and it is not used to justify any reported number. The acknowledged test-overfitting caveat (Sections 3.6 and 5) and the unexamined triple-backtick parsing rule (Section 3.4) are measurement-validity threats rather than circular reductions: they could affect how many plausible patches are counted, but they do not make the predicted quantity equal to an input by construction. Selecting the best RQ1 variants for RQ2 is model selection, not circularity, because RQ2 evaluates new strategy comparisons rather than refitting the reported results. No circular step can be quoted from the paper.
Assumptions & free parameters
free parameters (3)
- Beam search width
- Fine-tuning hyperparameters (learning rate, epochs, LoRA rank/alpha)
- Fine-tuning dataset sizes (1K, 30K, 65K) =
1,000 / 30,000 / 65,000 samples
assumptions (3)
- domain assumption A patch that compiles and passes all tests in the benchmark test suite is treated as a successful fix (plausible patch, Section 3.6).
- domain assumption The 217-bug subset of Defects4J used for evaluation is a fair representation of Defects4J and comparable to prior work.
- domain assumption The fine-tuning data (GitHub Java single-hunk fixes) is disjoint from the evaluation benchmarks.
Cite this review
Pith. "Pith review of The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models." pith.science (2026). https://pith.science/paper/VNEM7UI7
@misc{pith2026250502931,
author = {Pith},
title = {Pith review of: The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/VNEM7UI7}},
note = {Machine review of arXiv:2505.02931}
}
read the original abstract
Automatic program repair (APR) aims to reduce the manual efforts required to identify and fix errors in source code. Before the rise of LLM-based agents, a common strategy was to increase the number of generated patches, sometimes to the thousands, to achieve better repair results on benchmarks. More recently, self-iterative capabilities enabled LLMs to refine patches over multiple rounds guided by feedback. However, literature often focuses on many iterations and disregards different numbers of outputs. We investigate an APR pipeline that balances these two approaches, the generation of multiple outputs and multiple rounds of iteration, while imposing a limit of 10 total patches per bug. We apply three SOTA instruction-tuned LLMs - DeepSeekCoder-Instruct, Codellama-Instruct, Llama3.1-Instruct - to the APR task. We further fine-tune each model on an APR dataset with three sizes (1K, 30K, 65K) and two techniques (Full Fine-Tuning and LoRA), allowing us to assess their repair capabilities on two APR benchmarks: HumanEval-Java and Defects4J. Our results show that by using only a fraction (<1%) of the fine-tuning dataset, we can achieve improvements of up to 78% in the number of plausible patches generated, challenging prior studies that reported limited gains using Full Fine-Tuning. However, we find that exceeding certain thresholds leads to diminishing outcomes, likely due to overfitting. Moreover, we show that base models greatly benefit from creating patches in an iterative fashion rather than generating them all at once. In addition, the benefit of iterative strategies becomes more pronounced in complex benchmarks. Even fine-tuned models, while benefiting less from iterations, still gain advantages, particularly on complex benchmarks. The research underscores the need for balanced APR strategies that combine multi-output generation and iterative refinement.
Figures
Reference graph
Works this paper leans on
-
[1]
K. Herb. Cost of Poor Software Quality in the U.S.: A 2020 Report . Tech. rep. Jan. 2021
work page 2020
-
[2]
D. H. O’Dell. “The Debugging Mindset: Understanding the Psychol- ogy of Learning Strategies Leads to Effective Problem-Solving Skills.” In: Queue 15.1 (Feb. 2017), pp. 71–90. doi: 10.1145/3055301.3068754
arXiv 2017
-
[3]
AVATAR: Fixing Semantic Bugs with Fix Patterns of Static Analysis Violations
K. Liu, A. Koyuncu, D. Kim, and T. F. Bissyandè. “AVATAR: Fixing Semantic Bugs with Fix Patterns of Static Analysis Violations. ” In: 2019 IEEE 26th International Conference on Software Analysis, Evolu- tion and Reengineering (SANER) . Feb. 2019, pp. 1–12. doi: 10.1109/ saner.2019.8667970
arXiv 2019
-
[4]
TBar: Revisiting Template-Based Automated Program Repair
K. Liu, A. Koyuncu, D. Kim, and T. F. Bissyandé. “TBar: Revisiting Template-Based Automated Program Repair. ” In:International Sym- posium on Software Testing and Analysis (ISSTA). Beijing China: ACM, July 2019, pp. 31–42. doi: 10.1145/3293882.3330577
arXiv 2019
-
[5]
FixMiner: Mining Relevant Fix Patterns for Auto- mated Program Repair
A. Koyuncu, K. Liu, T. F. Bissyandé, D. Kim, J. Klein, M. Monperrus, and Y. Le Traon. “FixMiner: Mining Relevant Fix Patterns for Auto- mated Program Repair. ” In:Empirical Software Engineering 25.3 (May 2020), pp. 1980–2024. doi: 10.1007/s10664-019-09780-z
-
[6]
ARJA: Automated Repair of Java Programs via Multi-Objective Genetic Programming
Y. Yuan and W. Banzhaf. “ARJA: Automated Repair of Java Programs via Multi-Objective Genetic Programming. ” In:IEEE Transactions on Software Engineering 46.10 (Oct. 2020), pp. 1040–1067. doi: 10.1109/ TSE.2018.2874648
arXiv 2020
-
[7]
Elixir: Effective Object-Oriented Program Repair
R. K. Saha, Y. Lyu, H. Yoshida, and M. R. Prasad. “Elixir: Effective Object-Oriented Program Repair. ” In:2017 32nd IEEE/ACM Interna- tional Conference on Automated Software Engineering (ASE). Oct. 2017, pp. 648–659. doi: 10.1109/ASE.2017.8115675
arXiv 2017
-
[8]
M. Motwani and Y. Brun. Better Automatic Program Repair by Using Bug Reports and Tests Together. Feb. 2023. arXiv: 2011.08340. 5https://doi.org/10.5281/zenodo.15294695 10 The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models Accepted at EASE 2025, 17–20 June 2025, Istanbul, Türkiye
arXiv 2023
Show all 67 references
-
[9]
Martinez and M
M. Martinez and M. Monperrus. Ultra-Large Repair Search Space with Automatically Mined Templates: The Cardumen Mode of Astor . July
-
[10]
DynaMoth: Dynamic Code Synthesis for Automatic Program Repair
T. Durieux and M. Monperrus. “DynaMoth: Dynamic Code Synthesis for Automatic Program Repair. ” In:2016 IEEE/ACM 11th International Workshop in Automation of Software Test (AST). May 2016, pp. 85–91. doi: 10.1145/2896921.2896931
2016
-
[11]
CoCoNuT: Combining Context-Aware Neural Translation Models Using Ensem- ble for Program Repair
T. Lutellier, H. V. Pham, L. Pang, Y. Li, M. Wei, and L. Tan. “CoCoNuT: Combining Context-Aware Neural Translation Models Using Ensem- ble for Program Repair. ” In:SIGSOFT International Symposium on Software Testing and Analysis. Virtual Event USA: ACM, July 2020, pp. 101–114....
2020
-
[12]
Automated Program Repair in the Era of Large Pre-Trained Language Models
C. S. Xia, Y. Wei, and L. Zhang. “Automated Program Repair in the Era of Large Pre-Trained Language Models. ” In:45th International Conference on Software Engineering . ICSE ’23. Melbourne, Victoria, Australia: IEEE, July 2023, pp. 1482–1494. doi: 10.1109/icse48619. 2023.00129
2023
-
[13]
Less Training, More Repairing Please: Revis- iting Automated Program Repair via Zero-Shot Learning
C. S. Xia and L. Zhang. “Less Training, More Repairing Please: Revis- iting Automated Program Repair via Zero-Shot Learning. ” In:30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering . Singapore Singapore: ACM, Nov. 2...
2022
-
[14]
Rozière et al
B. Rozière et al. Code Llama: Open Foundation Models for Code . Aug
- [15]
- [16]
-
[17]
A Survey of Learning- based Automated Program Repair
Q. Zhang, C. Fang, Y. Ma, W. Sun, and Z. Chen. “A Survey of Learning- based Automated Program Repair. ” In:ACM Transactions on Software Engineering and Methodology 33.2 (Feb. 2024), pp. 1–69. doi: 10.1145/ 3631974
2024
- [18]
-
[19]
Xiang, X
J. Xiang, X. Xu, F. Kong, M. Wu, H. Zhang, and Y. Zhang.How Far Can We Go with Practical Function-Level Program Repair? Apr. 2024. arXiv: 2404.12833 [cs]
2024 arXiv
-
[20]
Noller, R
Y. Noller, R. Shariffdeen, X. Gao, and A. Roychoudhury. Trust En- hancement Issues in Program Repair . Feb. 2022. arXiv: 2108.13064 [cs]
2022 arXiv
-
[21]
Zhang et al.Instruction Tuning for Large Language Models: A Survey
S. Zhang et al.Instruction Tuning for Large Language Models: A Survey. Mar. 2024. arXiv: 2308.10792[cs]
2024
-
[22]
Impact of Code Language Models on Automated Program Repair
N. Jiang, K. Liu, T. Lutellier, and L. Tan. “Impact of Code Language Models on Automated Program Repair. ” In: 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . May 2023, pp. 1430–1442. doi: 10.1109/ICSE48619.2023.00125
2023
-
[23]
Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs
R. Just, D. Jalali, and M. D. Ernst. “Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. ” In: International Symposium on Software Testing and Analysis (ISSTA) . San Jose, CA, USA: ACM, 2014, pp. 437–440. doi: 10.1145/2610384. 2628055
2014 doi
-
[24]
Dubey et al
A. Dubey et al. The Llama 3 Herd of Models . Aug. 2024. arXiv: 2407. 21783[cs]
2024
-
[25]
J. Lu, L. Yu, X. Li, L. Yang, and C. Zuo. LLaMA-Reviewer: Advanc- ing Code Review Automation with Large Language Models through Parameter-Efficient Fine-Tuning. Sept. 2023. arXiv: 2308.11148[cs]
2023 arXiv
- [26]
-
[27]
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
F. Cassano, L. Li, A. Sethi, N. Shinn, A. Brennan-Jones, A. Lozhkov, C. J. Anderson, and A. Guha. “Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions. ” In: ()
- [28]
-
[29]
X. Luo, Q. Zhu, Z. Zhang, X. Wang, Q. Yang, D. Xu, and W. Che. Semi-Instruct: Bridging Natural-Instruct and Self-Instruct for Code Large Language Models. Mar. 2024. arXiv: 2403.00338[cs]
2024 arXiv
-
[30]
J. He, M. Vero, G. Krasnopolska, and M. Vechev.Instruction Tuning for Secure Code Generation . Feb. 2024. arXiv: 2402.09497[cs]
2024 arXiv
- [31]
-
[32]
Muennighoff et al
N. Muennighoff et al. OctoPack: Instruction Tuning Code Large Lan- guage Models. Feb. 2024. arXiv: 2308.07124[cs]
2024 arXiv
- [33]
-
[34]
Magicoder: Empow- ering Code Generation with OSS-Instruct
Y. Wei, Z. Wang, J. Liu, Y. Ding, and L. Zhang. “Magicoder: Empow- ering Code Generation with OSS-Instruct. ” In: ()
-
[35]
Shen et al
B. Shen et al. PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback . July 2023. arXiv: 2307.14936 [cs]
2023 arXiv
-
[36]
Z. Yu, X. Zhang, N. Shang, Y. Huang, C. Xu, Y. Zhao, W. Hu, and Q. Yin. WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning . June 2024. arXiv: 2312.14187 [cs]
2024 arXiv
-
[37]
Z. Yuan, J. Liu, Q. Zi, M. Liu, X. Peng, and Y. Lou.Evaluating Instruction- Tuned Large Language Models on Code Comprehension and Generation. Aug. 2023. arXiv: 2308.01240[cs]
2023 arXiv
-
[38]
T. Y. Zhuo, A. Zebaze, N. Suppattarachai, L. von Werra, H. de Vries, Q. Liu, and N. Muennighoff. Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models . Jan. 2024. arXiv: 2401.00788 [cs]
2024 arXiv
- [39]
-
[40]
Fully Au- tonomous Programming with Large Language Models
V. Liventsev, A. Grishina, A. Härmä, and L. Moonen. “Fully Au- tonomous Programming with Large Language Models. ” In:Genetic and Evolutionary Computation Conference (GECCO) . ACM, 2023, pp. 1146–1155. doi: 10.1145/3583131.3590481
2023
-
[41]
Self-Refine: Iterative Refinement with Self-Feedback
A. Madaan et al. “Self-Refine: Iterative Refinement with Self-Feedback. ” In: Advances in Neural Information Processing Systems 36 (Dec. 2023), pp. 46534–46594
2023
-
[42]
H. Jin, Z. Sun, and H. Chen. RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance. Oct. 2024. arXiv: 2410.01242 [cs]
2024 arXiv
- [43]
- [44]
-
[45]
ITER: Iterative Neural Repair for Multi- Location Patches
H. Ye and M. Monperrus. “ITER: Iterative Neural Repair for Multi- Location Patches. ” In:IEEE/ACM 46th International Conference on Software Engineering . Feb. 2024, pp. 1–13. doi: 10.1145/3597503. 3623337. arXiv: 2304.12015[cs]
2024 arXiv
-
[46]
Gehring, K
J. Gehring, K. Zheng, J. Copet, V. Mella, T. Cohen, and G. Synnaeve. RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning. https://arxiv.org/abs/2410.02089v1. Oct. 2024. 11 Accepted at EASE 2025, 17–20 June 2025, Istanbul, Türkiye. Fernando Vallecillos R...
2024 arXiv
-
[47]
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
T. Zheng, G. Zhang, T. Shen, X. Liu, B. Y. Lin, J. Fu, W. Chen, and X. Yue. “OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement. ” In:Findings of the Association for Com- putational Linguistics: ACL 2024 . Ed. by L.-W. Ku, A. Martins, and V. Srikuma...
2024 doi
-
[48]
Silva, S
A. Silva, S. Fang, and M. Monperrus. RepairLLaMA: Efficient Repre- sentations and Fine-Tuned Adapters for Program Repair . June 2024. arXiv: 2312.15698 [cs]
2024 arXiv
- [49]
- [50]
- [51]
-
[52]
G. Li, C. Zhi, J. Chen, J. Han, and S. Deng.A Comprehensive Evaluation of Parameter-Efficient Fine-Tuning on Automated Program Repair. June
-
[53]
B. Yang, H. Tian, J. Ren, H. Zhang, J. Klein, T. F. Bissyandé, C. L. Goues, and S. Jin. Multi-Objective Fine-Tuning for Enhanced Program Repair with LLMs. Apr. 2024. arXiv: 2404.12636[cs]
2024
- [54]
- [55]
-
[56]
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek-AI et al. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism. Jan. 2024. arXiv: 2401.02954
2024 arXiv
-
[57]
Taori, I
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto.Stanford Alpaca: An Instruction-Following LLaMA Model. 2023
2023
-
[58]
A Syntax-Guided Edit Decoder for Neural Program Repair
Q. Zhu, Z. Sun, Y.-a. Xiao, W. Zhang, K. Yuan, Y. Xiong, and L. Zhang. “A Syntax-Guided Edit Decoder for Neural Program Repair. ” In:29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . Athens Greece: ACM...
2021
-
[59]
Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs
R. Just, D. Jalali, and M. D. Ernst. “Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. ” In: 2014 International Symposium on Software Testing and Analysis. ISSTA
2014
- [60]
- [61]
- [62]
-
[63]
Automated Patch Correctness Assessment: How Far Are We?
S. Wang, M. Wen, B. Lin, H. Wu, Y. Qin, D. Zou, X. Mao, and H. Jin. “Automated Patch Correctness Assessment: How Far Are We?” In: 35th IEEE/ACM International Conference on Automated Software Engineering. Virtual Event Australia: ACM, Dec. 2020, pp. 968–980. doi: 10.1145/332488...
2020
-
[2014]
437–440.doi: 10.1145/ 2610384.2628055
New York, NY, USA: ACM, July 2014, pp. 437–440.doi: 10.1145/ 2610384.2628055
2014
- [2018]
- [2023]
-
[2024]
arXiv: 2406.05639[cs]
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.