REVIEW 3 major objections 6 minor 1 cited by
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Two AI agents, a coder and a reviewer, synthesize code-training data that beats datasets several times larger.
desk verdict CodeEvo is a solid, data-efficient synthesis pipeline whose gains are real, though the filter's self-validation and the OSS comparison leave the mechanism story partly unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the hybrid feedback $f_{hybrid}$, formed when the Reviewer fuses deterministic compiler feedback $f_{comp}$ with its own natural-language evaluation $f_{NL}$. This fused signal does two jobs: it filters which instruction-code pairs enter the dataset, and it gives the Coder a concrete reason to refine. The other pillar is keyword-guided instruction generation: new instructions are created by conditioning on a subset of the seed's keyword tags, and the same keywords are removed to simplify a problem when the Coder fails. That bidirectional keyword control keeps evolved instructions semantically grounded while pushing difficulty upward.
What would settle it
Audit a random sample of instruction-code pairs that CodeEvo's pipeline retains by running them against independently written hidden tests that the Coder did not produce. If a nontrivial fraction fails those tests, or solves a different task than the instruction states, the hybrid filter is not doing what the data-quality claim requires.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a feedback-driven agent loop can replace rigid prompting heuristics in code-data synthesis. Starting from only a few thousand seed instructions and their keyword tags, the Reviewer proposes new instructions by keeping or removing keywords, the Coder attempts them, and a hybrid signal—compiler pass/fail plus the Reviewer's semantic read of the code—decides whether the pair is kept or sent back for refinement. The resulting trajectories yield instruction-code pairs that pass both execution and logical checks. Fine-tuning four base models on this data gives higher pass@1 on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench than training on Evol-Instruct or OSS-Instruct data, using 17K samples against 25K and 75K for those baselines.
Load-bearing premise
The pipeline assumes that the Reviewer's natural-language judgment, combined with compiler pass/fail on test cases the Coder wrote itself, reliably identifies correct, well-grounded solutions; the paper concedes that agent-generated test coverage can be incomplete.
Editorial extensions
If this is right
- Synthetic code data can be produced from natural-language seeds alone, with no gold code and no human-authored tests, which removes a major bottleneck in dataset construction.
- Data efficiency improves: comparable or better downstream pass@1 is reached with 17K CodeEvo samples than with 25K Evol-Instruct or 75K OSS-Instruct samples.
- Compiler-in-the-loop validation pays off most where test coverage is stricter, as suggested by gains on HumanEval+ and MBPP+ over the base versions.
- The approach transfers across general-purpose and code-specialized base models, and it works with smaller synthesis backbones, so frontier or proprietary models are not required.
- Because surviving instructions get harder as rounds progress, the retained pairs form a natural difficulty ladder, which could support curriculum-style fine-tuning.
Reading between the lines
- The same keyword-anchored, bidirectional evolution could be lifted to other symbolic or structured domains—math word problems, reasoning chains, or agent trajectories—where difficulty can be parameterized by a small vocabulary of techniques.
- The hybrid feedback signal, which already encodes a retain/refine verdict, could be reused as a reward signal for preference optimization or reinforcement-learning fine-tuning of code models.
- A stronger test of the data-efficiency claim would hold the training budget fixed at 17K samples for every method, rather than comparing against baselines that used 25K or 75K samples.
- Because the Reviewer is judging the same kind of output it helps produce, an independent human audit of a random retained subset would quantify how much of the gain comes from filtering versus from the seed and benchmark mix.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CodeEvo, a dual-agent framework for synthesizing instruction-code pairs to train code-generation LLMs. A Coder agent generates candidate solutions and self-written test cases, while a Reviewer agent produces keyword-conditioned new instructions and judges candidate solutions through a hybrid feedback mechanism that combines compiler pass/fail signals with natural-language assessment. The authors construct a dataset (CodeEvo-100K, with 17K examples used in experiments), fine-tune four base models, and report pass@1 on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench, claiming consistent improvements over Evol-Instruct and OSS-Instruct baselines as well as superior data efficiency. The paper also presents analyses of instruction diversity, difficulty, scaling behavior, and data survival rates.
Significance. If the results hold, CodeEvo provides a data-efficient and largely reference-free method for synthesizing code-centric training data, reducing reliance on human curation and on large or proprietary generator models. The keyword-guided instruction evolution and the hybrid verification loop are interesting and potentially transferable ideas, and the paper ships code and data (per the abstract). However, the current evidence is weakened by the absence of variance reporting, a confounded comparison against OSS-Instruct, and the lack of any independent validation of the self-assessed filtering criterion. The significance is therefore conditional on addressing these experimental-rigor issues.
major comments (3)
- [§4.4, Table 1] The central claim that CodeEvo-trained models "significantly outperform" baselines is not supported statistically. No variance, confidence intervals, or significance tests are reported for any pass@1 number, and several differences are within one point (e.g., DeepSeek-Coder-6.7B HumanEval: CodeEvo 77.4 vs OSS-Instruct 76.8; HE+ 71.3 vs 70.7; StarCoder2-7B HumanEval: 50.0 vs 50.6). I request at least three fine-tuning runs per condition with mean and standard deviation, or a paired bootstrap over benchmark problems, so that the claimed improvements can be distinguished from noise.
- [§4.2 / Appendix C.2, Table 1] The comparison to OSS-Instruct is confounded. CodeEvo is trained on 17K examples derived from roughly 5K LeetCode/Codeforces-style seeds with keyword tags and includes some seed reference solutions, whereas OSS-Instruct uses the released 75K Magicoder dataset generated from open-source code snippets with a different pipeline and generator model. The "4–5x fewer data" efficiency claim relative to OSS-Instruct may therefore reflect seed distribution or generator differences rather than the CodeEvo mechanisms themselves. Please either re-run OSS-Instruct on the same seed set with the same Qwen-based generator models (as was done for Evol-Instruct), or explicitly delimit the claim and discuss this confound.
- [§3.3 and Algorithm 1] The hybrid feedback acceptance criterion is self-referential and not independently verified. The same Reviewer agent that generates new instructions also judges whether the Coder's solution is correct, and the compiler signal only checks test cases that the Coder wrote for its own solution. The paper's own Limitations (Test Case Quality) concedes that "it remains difficult to ensure complete correctness, especially at scale." To substantiate the characterization of retained pairs as "well-grounded and executable" and to explain the data-efficiency advantage, the authors should validate a random sample of the 17K retained pairs against an independent oracle (e.g., hidden tests from the seed platforms, human annotation, or a separate strong model) and report the true positive rate. Without this, the survival-rate analysis in Section 5.5 (Figure 9) only demonstrates internal consistency of the filter, not its accuracy.
minor comments (6)
- [Abstract vs. §4.4] The abstract states that CodeEvo-100K is constructed, but all fine-tuning experiments and the scaling analysis (Figure 8) use at most 17K examples; clarify the relationship between CodeEvo-100K and the 17K training subset, and state whether any experiment uses the full 100K dataset.
- [§4.4] There is a duplicated or broken sentence in the main-results paragraph: "On StarCoder2-7B, included to align with the evaluation setups of WizardCoder and Magicoder, CodeEvo significantly surpasses their original synthesis designs, achieving striking gains on the more challenging benchmarks—nearly doubling the performance in some cases. CodeEvo delivers striking gains on challenging benchmarks." Please revise for clarity.
- [Appendix A] There is a typo in the DeepSeek-Coder model description: "we adopot this model" should be "we adopt this model."
- [§5.3 and Appendix G.2] The human evaluation of instruction difficulty uses five participants, and the solvability study does not report a sample size or inter-annotator agreement; please include these details and state whether raters were blinded to the synthesis method.
- [Algorithm 2 and Appendix E] The hyperparameters [rmin, rmax], tmax, and the maximum iterations N are never given concrete values or a sensitivity analysis; please report the values used in the experiments.
- [§3.3] The computation of fhybrid from fcomp and fNL is underspecified: it is not clear which prompt or decision rule the Reviewer uses to produce the final "valid" verdict, nor whether a separate LLM call is made; providing the exact prompt or decision rule would improve reproducibility.
Circularity Check
No significant circularity: CodeEvo's benchmark claims are evaluated on external test sets and are not fitted by the synthesis pipeline.
full rationale
The paper's central claim is that models fine-tuned on CodeEvo-synthesized data outperform baselines on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench (Table 1). These benchmarks are external to the Coder-Reviewer loop: the hybrid filter fhybrid in Section 3.3 decides retention, but no benchmark score is fed back as a fitting signal, and no evaluation metric is defined in terms of the Reviewer's judgment. The keyword-guided instruction generation (Eq. 1) and hybrid validation (Eq. 3) are internal design choices; their correctness is a robustness question, not a circular reduction. The paper's own Limitations section ('Test Case Quality') concedes that agent-written tests may have incomplete coverage, which is exactly the validity risk flagged by skeptics, but the risk does not make the downstream performance claim definitionally equivalent to the input. The self-citations (e.g., Sun et al., 2023; Sun et al., 2024b; Xu et al., 2024b; Xu et al., 2024c; Ma et al., 2025; Shao et al., 2025) appear in related-work context and are not used as load-bearing uniqueness theorems or as justifications that forbid alternative methods. The survival-rate analysis in Section 5.5 is definitionally tied to the hybrid filter, but it is a descriptive statistic about the pipeline, not a prediction validated against or used to fit any external benchmark. The ablations in Table 2 further show that removing seed code does not collapse performance, supporting the claim that the synthesized trajectories carry the training signal. Accordingly, no step reduces a reported prediction or first-principles result to its own inputs by construction.
Assumptions & free parameters
free parameters (3)
- maximum synthesis iterations N
- keyword sampling range [rmin, rmax]
- hybrid feedback validity threshold
assumptions (4)
- domain assumption The Reviewer's natural language evaluation is a reliable proxy for solution correctness and alignment with the instruction.
- domain assumption Compiler feedback combined with LLM review is sufficient to filter out functionally incorrect solutions.
- domain assumption Seed instructions from programming platforms are representative of the target evaluation distribution.
- domain assumption Fine-tuning on a filtered subset of synthetically generated data improves code generation without harmful distributional shift.
Cite this review
Pith. "Pith review of CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback." pith.science (2026). https://pith.science/paper/MBSXNRFQ
@misc{pith2026250722080,
author = {Pith},
title = {Pith review of: CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback},
year = {2026},
howpublished = {\url{https://pith.science/paper/MBSXNRFQ}},
note = {Machine review of arXiv:2507.22080}
}
read the original abstract
Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation. While automated synthesis has emerged as an alternative to expensive manual curation, current approaches often rely on rigid heuristics, yielding data that is ungrounded or lacks logical complexity. We propose CodeEvo, a dual-agent architecture comprising a Coder for iterative solution synthesis and a Reviewer to orchestrate the generation trajectory. To transcend the limitations of existing heuristics, the Reviewer formulates a Schema to systematically architect logic and complexity through an interleaved synthesis of instructions and code. This process is further reinforced by a hybrid verification protocol synergizing deterministic compiler feedback with semantic evaluation. Under this framework, we construct CodeEvo-100K, a large-scale dataset of instruction-code pairs with stepped difficulty levels. Extensive experiments demonstrate that models fine-tuned on CodeEvo data consistently outperform established baselines across code generation benchmarks. In-depth analyses further provide insights into effective code-centric data synthesis. Code and data are available at https://github.com/QiushiSun/CodeEvo.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
BenchEvolver evolves coding problem solutions to generate harder, valid tasks, producing LiveCodeBench-Plus where frontier models score 27.5-62.6% and enabling RL gains on held-out tests.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732
arXiv 2021
-
[4]
Zhangqian Bi, Yao Wan, Zheng Wang, Hongyu Zhang, Batu Guan, Fangxin Lu, Zili Zhang, Yulei Sui, Hai Jin, and Xuanhua Shi. 2024. https://doi.org/10.18653/v1/2024.findings-acl.138 Iterative refinement of project-level code context for precise code generation with compiler feedback . In Findings of the Association for Computational Linguistics: ACL 2024, page...
-
[5]
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang, Penglong Jiao, Zhenjiang Jin, Zhikai Lei, Jiaxing Li, Jingwen Li, Linyang Li, S...
arXiv 2024
-
[6]
Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca
2023
-
[7]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian...
arXiv 2021
-
[8]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
arXiv 2021
Show all 61 references
-
[9]
DeepSeek-AI, :, Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, Huazuo Gao, Kaige Gao, Wenjun Gao, Ruiqi Ge, Kang Guan, et al. 2024. Deepseek llm: Scaling open-source language models with longtermism. arXiv...
2024 arXiv
-
[10]
DeepSeek-AI. 2024. https://arxiv.org/abs/2412.19437 Deepseek-v3 technical report . Preprint, arXiv:2412.19437
2024 arXiv
-
[11]
Yu Feng, Ruben Martins, Osbert Bastani, and Isil Dillig. 2018. https://doi.org/10.1145/3296979.3192382 Program synthesis using conflict-driven learning . SIGPLAN Not., 53(4):420–435
2018
-
[12]
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al. 2024. Deepseek-coder: When the large language model meets programming--the rise of code intelligence. arXiv preprint arXiv:2401.14196
2024 arXiv
-
[13]
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and J \"u rgen Schmidhuber. 2024. https://openreview.net/forum?id=VtmBAGCN7o Meta GP...
2024
-
[14]
Mengkang Hu, Pu Zhao, Can Xu, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo, and Saravan Rajmohan. 2025. https://doi.org/10.1145/3690624.3709321 Agentgen: Enhancing planning abilities for large language model based agent via environment and task generation . In Proceedin...
2025
-
[15]
Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui
Dong Huang, Jie M. Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui. 2024 a . https://arxiv.org/abs/2312.13010 Agentcoder: Multi-agent-based code generation with iterative testing and optimisation . Preprint, arXiv:2312.13010
2024 arXiv
-
[16]
Siming Huang, Tianhao Cheng, Jason Klein Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J. Yang, J. H. Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Zhaoxiang Zhang, Jie Fu, Qian Liu, Ge Zhang, Zili Wang, Yuan Qi, Yinghui Xu, and Wei Chu. 2024 b . https://arxiv.org/pdf/2411.0490...
2024 arXiv
-
[17]
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, Kai Dang, Yang Fan, Yichang Zhang, An Yang, Rui Men, Fei Huang, Bo Zheng, Yibo Miao, Shanghaoran Quan, Yunlong Feng, Xingzhang Ren, Xuancheng Ren, Jingren Zhou...
2024 arXiv
-
[18]
Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez
Md. Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez. 2024. https://doi.org/10.18653/v1/2024.acl-long.269 M ap C oder: Multi-agent code generation for competitive problem solving . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguisti...
2024 doi
-
[19]
Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez
Md. Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez. 2025. https://aclanthology.org/2025.findings-naacl.285/ C ode S im: Multi-agent code generation and problem solving through simulation-driven planning and debugging . In Findings of the Association for Computational...
2025
-
[20]
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974
2024 arXiv
-
[21]
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. https://openreview.net/forum?id=chfJJYC3iL Livecodebench: Holistic and contamination free evaluation of large language models for code . I...
2025
-
[22]
Zaid Khan, Elias Stengel-Eskin, Jaemin Cho, and Mohit Bansal. 2025. https://openreview.net/forum?id=00SnKBGTsz Dataenvgym: Data generation agents in teacher environments with student feedback . In The Thirteenth International Conference on Learning Representations
2025
-
[23]
Shuvendu K Lahiri, Sarah Fakhoury, Aaditya Naik, Georgios Sakkas, Saikat Chakraborty, Madanlal Musuvathi, Piali Choudhury, Curtis von Veh, Jeevana Priya Inala, Chenglong Wang, et al. 2022. Interactive code generation via test-driven user-intent formalization. arXiv preprint ar...
2022 arXiv
-
[24]
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. https://openreview.net/forum?id=3IyL2XWDkG CAMEL : Communicative agents for ``mind'' exploration of large language model society . In Thirty-seventh Conference on Neural Informati...
2023
-
[25]
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and LINGMING ZHANG. 2023. https://openreview.net/forum?id=1qvx610Cu7 Is your code generated by chat GPT really correct? rigorous evaluation of large language models for code generation . In Thirty-seventh Conference on Neural Informa...
2023
-
[26]
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Z...
2024 arXiv
-
[27]
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Qingwei Lin, Jianguang Lou, Shifeng Chen, Yansong Tang, and Weizhu Chen. 2024 a . https://arxiv.org/abs/2407.10627 Arena learning: Build data flywheel for llms post-training via simulated chatbot arena . Preprint, arXiv:2407.10627
2024 arXiv
-
[28]
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2024 b . https://openreview.net/forum?id=UnUwSIgK5W Wizardcoder: Empowering code large language models with evol-instruct . In The Twelfth International Co...
2024
-
[29]
Wei Ma, Shangqing Liu, Zhihao Lin, Wenhan Wang, Qiang Hu, Ye Liu, Cen Zhang, Liming Nie, Li Li, and Yang Liu. 2024. https://arxiv.org/abs/2305.12138 Lms: Understanding code syntax and semantics for code analysis . Preprint, arXiv:2305.12138
2024 arXiv
-
[30]
Yichuan Ma, Yunfan Shao, Peiji Li, Demin Song, Qipeng Guo, Linyang Li, Xipeng Qiu, and Kai Chen. 2025. https://arxiv.org/abs/2502.11460 Unitcoder: Scalable iterative code synthesis with unit test guidance . Preprint, arXiv:2502.11460
2025 arXiv
-
[31]
Somshubra Majumdar, Vahid Noroozi, Mehrzad Samadi, Sean Narenthiran, Aleksander Ficek, Wasi Uddin Ahmad, Jocelyn Huang, Jagadeesh Balam, and Boris Ginsburg. 2025. https://arxiv.org/abs/2407.21077 Genetic instruct: Scaling up synthetic generation of coding instructions for larg...
2025 arXiv
-
[32]
Arindam Mitra, Luciano Del Corro, Guoqing Zheng, Shweti Mahajan, Dany Rouhana, Andres Codas, Yadong Lu, Wei ge Chen, Olga Vrousgos, Corby Rosset, Fillipe Silva, Hamed Khanpour, Yash Lara, and Ahmed Awadallah. 2024. https://arxiv.org/abs/2407.03502 Agentinstruct: Toward generat...
2024 arXiv
-
[33]
O'Brien, Carrie J
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. https://arxiv.org/abs/2304.03442 Generative agents: Interactive simulacra of human behavior . Preprint, arXiv:2304.03442
2023 arXiv
-
[34]
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. https://doi.org/10.18653/v1/2024.acl-long.810 C hat D ev: Communicative agents for software development . ...
2024 doi
-
[35]
Yunfan Shao, Linyang Li, Yichuan Ma, Peiji Li, Demin Song, Qinyuan Cheng, Shimin Li, Xiaonan Li, Pengyu Wang, Qipeng Guo, Hang Yan, Xipeng Qiu, Xuanjing Huang, and Dahua Lin. 2025. https://aclanthology.org/2025.coling-main.733/ C ase2 C ode: Scalable synthetic data for code ge...
2025
-
[36]
Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas Griffiths. 2024. https://openreview.net/forum?id=1i6ZCvflQJ Cognitive architectures for language agents . Transactions on Machine Learning Research. Survey Certification
2024
-
[37]
Qiushi Sun, Zhirui Chen, Fangzhi Xu, Kanzhi Cheng, Chang Ma, Zhangyue Yin, Jianing Wang, Chengcheng Han, Renyu Zhu, Shuai Yuan, et al. 2024 a . A survey of neural code intelligence: Paradigms, advances and beyond. arXiv preprint arXiv:2403.14734
2024 arXiv
-
[38]
Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, et al. 2024 b . Os-genesis: Automating gui agent trajectory construction via reverse task synthesis. arXiv preprint arXiv:2412.19723
2024 arXiv
-
[39]
Qiushi Sun, Zhangyue Yin, Xiang Li, Zhiyong Wu, Xipeng Qiu, and Lingpeng Kong. 2023. Corex: Pushing the boundaries of complex reasoning through multi-model collaboration. arXiv preprint arXiv:2310.00280
2023 arXiv
-
[40]
Xin Wang, Yasheng Wang, Yao Wan, Fei Mi, Yitong Li, Pingyi Zhou, Jin Liu, Hao Wu, Xin Jiang, and Qun Liu. 2022. https://doi.org/10.18653/v1/2022.findings-acl.2 Compilable neural code generation with compiler feedback . In Findings of the Association for Computational Linguisti...
2022 doi
-
[41]
Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brenna...
2025
-
[42]
Xinyu Wang, Isil Dillig, and Rishabh Singh. 2017. https://doi.org/10.1145/3158151 Program synthesis using abstraction refinement . Proc. ACM Program. Lang., 2(POPL)
2017 doi
-
[43]
Yaoxiang Wang, Haoling Li, Xin Zhang, Jie Wu, Xiao Liu, Wenxiang Hu, Zhongxin Guo, Yangyu Huang, Ying Xin, Yujiu Yang, Jinsong Su, Qi Chen, and Scarlett Li. 2025 b . https://arxiv.org/abs/2501.04694 Epicoder: Encompassing diversity and complexity in code generation . Preprint,...
2025
-
[44]
Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual Mee...
2023 doi
-
[45]
Yuxiang Wei, Federico Cassano, Jiawei Liu, Yifeng Ding, Naman Jain, Zachary Mueller, Harm de Vries, Leandro Von Werra, Arjun Guha, and LINGMING ZHANG. 2024 a . https://openreview.net/forum?id=xXRnUU7xTL Selfcodealign: Self-alignment for code generation . In The Thirty-eighth A...
2024
-
[46]
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2024 b . https://proceedings.mlr.press/v235/wei24h.html Magicoder: Empowering code generation with OSS -instruct . In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceed...
2024
-
[47]
XTuner Contributors . 2023. Xtuner: A toolkit for efficiently fine-tuning llm. https://github.com/InternLM/xtuner
2023
-
[48]
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024 a . https://openreview.net/forum?id=CfXh93NDgH Wizard LM : Empowering large pre-trained language models to follow complex instructions . In The Twelfth Interna...
2024
-
[49]
Fangzhi Xu, Qiushi Sun, Kanzhi Cheng, Jun Liu, Yu Qiao, and Zhiyong Wu. 2024 b . Interactive evolution: A neural-symbolic self-training framework for large language models. arXiv preprint arXiv:2406.11736
2024 arXiv
-
[50]
Fangzhi Xu, Zhiyong Wu, Qiushi Sun, Siyu Ren, Fei Yuan, Shuai Yuan, Qika Lin, Yu Qiao, and Jun Liu. 2024 c . https://doi.org/10.18653/v1/2024.acl-long.707 Symbol- LLM : Towards foundational symbol-centric interface for large language models . In Proceedings of the 62nd Annual ...
2024 doi
-
[51]
Fangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu, and Zhiyong Wu. 2025 a . Genius: A generalizable and purely unsupervised self-training framework for advanced reasoning. arXiv preprint arXiv:2504.08672
2025 arXiv
-
[52]
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. 2025 b . https://openreview.net/forum?id=Pnk7vMbznK Magpie: Alignment data synthesis from scratch by prompting aligned LLM s with nothing . In The Thirteenth International...
2025
-
[53]
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...
2024 arXiv
-
[54]
Zonghan Yang, Peng Li, Ming Yan, Ji Zhang, Fei Huang, and Yang Liu. 2024 b . https://openreview.net/forum?id=0VLBwQGWpA React meets actre: Autonomous annotation of agent trajectories for contrastive self-training . In First Conference on Language Modeling
2024
-
[55]
Zhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, and Qiufeng Yin. 2024. https://doi.org/10.18653/v1/2024.acl-long.280 W ave C oder: Widespread and versatile enhancement for code large language models by instruction tuning . In Proceedings o...
2024 doi
-
[56]
Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Wang Yongji, and Jian-Guang Lou. 2023. https://doi.org/10.18653/v1/2023.acl-long.411 Large language models meet NL 2 C ode: A survey . In Proceedings of the 61st Annual Meeting of the Association for Comp...
2023 doi
-
[57]
Huaye Zeng, Dongfu Jiang, Haozhe Wang, Ping Nie, Xiaotong Chen, and Wenhu Chen. 2025. Acecoder: Acing coder rl via automated test-case synthesis. ArXiv, 2502.01718
2025 arXiv
-
[58]
Kechi Zhang, Ge Li, Yihong Dong, Jingjing Xu, Jun Zhang, Jing Su, Yongfei Liu, and Zhi Jin. 2024. https://arxiv.org/abs/2410.05605 Codedpo: Aligning code models with self generated and verified source code . Preprint, arXiv:2410.05605
2024 arXiv
-
[59]
Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.762 O pen C ode I nterpreter: Integrating code generation with execution and refinement . In Findings of the Associatio...
2024 doi
-
[60]
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. 2024 b . http://arxiv.org/abs/2403.13372 Llamafactory: Unified efficient fine-tuning of 100+ language models . In Proceedings of the 62nd Annual Meeting of the Association for Co...
2024 arXiv
-
[61]
Terry Yue Zhuo, Vu Minh Chien, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen GONG, James Hoang, Armel Randy Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Ming Xu, Zhihan Zhang, Prateek Ya...
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.