REVIEW 2 major objections 4 minor 92 references
Self-Evolving Coding Agents
T0 review · 2 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper establishes self-evolving coding agents as a distinct research category and organizes them by what evolves: framework, memory, skills and tools, model, or workflow and topology, with timing and evidence as orthogonal dimensions.
desk verdict A useful, well-organized survey with a genuine taxonomy, but its own definition of self-evolution is at odds with the task-time category, and the anonymous citation is a real problem. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the object-centered taxonomy. It classifies self-evolving coding agents by the primary artifact being updated: agent framework self-evolution (the agent rewrites or searches over versions of its own scaffold), memory self-evolution (experience banks, repository memory, repair experience), skill and tool self-evolution (distilling trajectories into reusable procedures or creating tools during a task), model self-evolution (policy updates from self-play, coder-verifier co-evolution, adversarial tests), and workflow and topology self-evolution (changing multi-agent collaboration, prompts, or communication structures). Two orthogonal axes complete the framework: evolving tim
What would settle it
Run the same base coding agent in two modes on a set of public repository issues: a static baseline and a version that evolves by retaining memory and skills from its attempts. Then evaluate both on a fresh, contamination-free set of issues from the same repositories. If the evolved version does not outperform the static baseline on the fresh issues despite matching or beating it on the evolution-time issues, the premise that executable feedback grounds generalizable self-improvement is falsified.
Extended reading notes
Core claim
The paper's central claim is that 'self-evolving coding agent' names a distinct kind of system: an agentic software engineering system that updates its behavior or internal components based on previous coding attempts and software-specific feedback. The survey organizes this family by the object that evolves: agent framework, memory, skills and tools, model-side components, or workflow and topology, with timing (task-time, post-task, stage-wise) and evidence type (outcome, environmental, trajectory-derived) as complementary axes. It argues that repository-level context and executable feedback make software engineering a natural domain for agent self-evolution, while feedback unreliability, b
Load-bearing premise
The load-bearing premise is that executable software feedback, such as tests, compiler diagnostics, CI logs, generated tests, and reward models, is trustworthy enough that incorporating it into memory, skills, workflows, or model weights makes an agent better rather than more confidently wrong; the paper itself concedes these signals are imperfect and may carry biases and blind spots.
Editorial extensions
If this is right
- SWE-bench-style repository issue resolution becomes the canonical test bed, since repository context and executable validation are what make self-evolution observable and measurable.
- Evaluation must move beyond one-shot pass rates to include cost, runtime, token use, contamination resistance, and transfer to held-out repositories or tasks.
- Each evolution category needs its own safeguards: scaffold rewrites need validation and rollback, memory needs filtering and abstraction, skills need cross-repository generalization checks, and model updates need verifier calibration.
- The taxonomy gives system designers a checklist: decide what evolves, when it evolves, and what evidence authorizes the change; systems that evolve several objects should state which one is primary.
- The definition draws a line between model post-training on SWE data and genuine self-evolution: the training signal must be closed around the agent's own attempts and executable outcomes.
Reading between the lines
- The authors treat feedback unreliability as a challenge, but the taxonomy implies something stronger: when benchmark outcomes are the evolution evidence, benchmark overfitting is not an evaluation artifact, it is the objective being optimized. A design principle follows: evolution benchmarks need contamination-resistant validation by construction.
- The five categories are not mutually exclusive, and the compound case is left unexplored. A testable prediction is that combining memory (what to attend to) with skills (how to act) compounds gains on repository-level tasks more than either alone.
- Holding the base model fixed and varying only the evolving object would separate model capability from agent adaptation. The framework suggests task-time tool and skill evolution helps short tasks, while stage-wise workflow evolution matters more for long-horizon repository work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of "self-evolving coding agents," defined as agentic software-engineering systems that update their own behavior or internal components using previous coding attempts and software-specific feedback. It proposes an object-centered taxonomy with five categories (agent framework, memory, skills/tools, model, workflow/topology), complemented by temporal dimensions (task-time, post-task, stage-wise) and evidence dimensions (outcome, environmental, trajectory-derived). The survey maps roughly thirty systems onto this taxonomy, reviews benchmarks and evaluation metrics, and discusses open challenges such as feedback reliability, benchmark overfitting, safety, and generalization. The central contribution is the taxonomy itself, framed as a foundation for designing and comparing self-evolving coding agents.
Significance. If the taxonomy is accepted, it provides a useful organizing framework for a rapidly growing but fragmented literature. The paper is broad in coverage, identifies concrete software-specific feedback signals, and draws attention to important open problems, especially the risk that unreliable executable feedback may undermine self-evolution. The survey's value is primarily conceptual rather than empirical; it does not introduce new algorithms or results. Its main strength is the systematic grouping of heterogeneous systems, and its main weakness is that the boundary of the central concept is not yet operationally precise enough to make the taxonomy's inclusion/exclusion decisions reproducible.
major comments (2)
- [Section 2.3 vs Section 4.1] The core definition of self-evolving coding agents is internally inconsistent with the task-time evolution category. Section 2.3 distinguishes self-evolving agents from conventional ones by "turning these interactions into sources of persistent adaptation," but Section 4.1 defines task-time evolution as change occurring "while the agent is still solving the current coding task," including "an immediate change in the current patch, tool use, or workflow," and later states that such adaptations are "often local to the current task." If changes to the current patch count as self-evolution, then any coding agent that retries after a test failure is self-evolving, collapsing the boundary that the survey claims to draw. If, instead, persistence across tasks is required, then the task-time examples (Live-SWE-Agent, SEMAG, SEW, EvoMAC, AgentConductor) are either misclassified or need to show tha
- [Section 3.4 and Table 2] The boundary between model self-evolution and ordinary SWE-oriented model optimization is not applied consistently. Section 3.4 states that "ordinary post-training becomes self-evolution only when software-specific experience is fed back into the model-side components that govern later agent behavior," and later restricts the category further to signals "closed around the agent's own evolving attempts." However, Agent-RLVR is included in Table 2 as a model self-evolution system even though Section 3.4 describes it as receiving "guidance and environment rewards" and using "guided reattempts" to update the policy. If the guidance is not generated by the agent's own attempts, the "closed loop" criterion is not met. The criterion needs an operational definition of whose attempts generate the signal, otherwise the inclusion/exclusion of systems such as Agent-RLVR, SWE-RL, and SWE-Gym is subje
minor comments (4)
- [References / Table 2] The reference "Anonymous. Mendel Gödel machine..., 2026" and its use in Table 2 and Figure 2 as "Mendel GM [2026]" violate normal scholarly practice for a survey. An anonymous, unpublished manuscript cannot be independently verified by readers. The authors should either replace it with a citable, attributed source or remove it from the taxonomy.
- [Section 4.1] The phrase "immediate change in the current patch" in the definition of task-time evolution should be removed or qualified, since changing the current patch is exactly what a conventional coding agent does after a failed test. Clarifying this wording would help address the major concern above.
- [Section 5.2] The discussion of evaluation metrics mentions Pass@k, solve rate, and repair rate, but does not define Pass@k or explain how it is computed in this context. A brief definition or citation would improve clarity for readers outside the code-generation subfield.
- [Section 2.3 / Table 1] Table 1's row for "General self-evolving agents" lists "Prompts, memory, tools, policies, or architectures" as what changes, but the text in Section 2.2 also includes "workflows" and "multi-agent co-evolution." Making the table consistent with the text would avoid confusion.
Circularity Check
No significant circularity: the survey's taxonomy and definitions do not reduce to their inputs; self-citations are illustrative only.
full rationale
The paper is a literature survey and taxonomy, not a derivation with equations or fitted parameters. Its central claim—that self-evolving coding agents form a distinct research category organized by object of evolution (framework, memory, skills/tools, model, workflow/topology)—is supported by enumerating many independent systems and is applied broadly outside the authors' own work (e.g., SICA, SEW, SEMAG, AFlow, Live-SWE-Agent, Socratic-SWE). No prediction reduces to an input by construction. The only clear self-citation is EvoRepair [Hu et al., 2026a], used as an example of vulnerability-repair memory, and possibly the anonymous Mendel Gödel Machine [Anonymous, 2026]; both are representative instances rather than load-bearing evidence for the framework, so removing them would not change any category's support. Section 6's caveat that 'tests, compilers, CI logs, generated tests, and reward models are imperfect' and may carry 'biases and blind spots' is a genuine field-level limitation, but not a circular step. The definitional tension between Section 2.3's 'persistent adaptation' and Section 4.1's task-time changes that are 'often local to the current task' is a boundary weakness and a correctness risk, but it is an inconsistency in scope rather than a circular derivation: the survey does not use the taxonomy to prove the definition, nor vice versa. The 'Anonymous, 2026' citation is a transparency issue—an anonymous under-review manuscript cited in a named-author survey—but it is not load-bearing. Verdict: no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The set of surveyed papers is representative of the field and correctly characterized.
- domain assumption Self-evolving coding agents can be cleanly distinguished from conventional coding agents and general self-evolving agents.
- domain assumption Executable feedback is a sufficiently reliable signal for self-evolution.
Cite this review
Pith. "Pith review of Self-Evolving Coding Agents." pith.science (2026). https://pith.science/paper/4IMFZYTJ
@misc{pith2026260803392,
author = {Pith},
title = {Pith review of: Self-Evolving Coding Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/4IMFZYTJ}},
note = {Machine review of arXiv:2608.03392}
}
read the original abstract
Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet most existing agents remain largely static after deployment, even though software development is a dynamic, feedback-rich process in which repositories evolve, dependencies change, tests fail, and repair attempts leave reusable experience. This tension has motivated a growing body of work on self-evolving coding agents, where the agent improves its future behavior by updating its framework, memory, skills, tools, models, or collaboration structures from prior coding interactions. In this survey, we provide a systematic synthesis of this emerging area. We first define self-evolving coding agents and distinguish them from conventional coding agents and general self-evolving agents. We then develop an object-centered taxonomy that characterizes what evolves in these systems, and complement it with two orthogonal perspectives: when evolution occurs and what software-specific evidence drives it. Across the literature, we find that executable feedback, repository-level context, and coding trajectories give software engineering a distinctive role as a natural domain for agent self-evolution, but also introduce new challenges in feedback reliability, benchmark overfitting, safety, maintainability, cost, and generalization. By organizing existing work around these dimensions, this survey aims to clarify the conceptual boundaries of self-evolving coding agents and provide a foundation for designing more adaptive, reliable, and software-aware agentic systems. The papers we collect can be found at https://github.com/zhouhao1024/Awesome-Self-Evolving-Coding-Agents.
Figures
Reference graph
Works this paper leans on
-
[1]
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , year =. 2310.06770 , archivePrefix =
-
[2]
Yang, John and Jimenez, Carlos E. and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , year =. 2405.15793 , archivePrefix =
-
[3]
Wang, Xingyao and Li, Boxuan and Song, Yufan and Xu, Frank F. and Tang, Xiangru and Zhuge, Mingchen and Pan, Jiayi and Song, Yueqi and Li, Bowen and Singh, Jaskirat and Tran, Hoang H. and Li, Fuqiang and Ma, Ren and Zheng, Mingzhang and Qian, Bill and Shao, Yanjun and Muennighoff, Niklas and Zhang, Yizhe and Hui, Binyuan and Lin, Junyang and Brennan, Robe...
-
[4]
Qian, Chen and Liu, Wei and Liu, Hongzhang and Chen, Nuo and Dang, Yufan and Li, Jiahao and Yang, Cheng and Chen, Weize and Su, Yusheng and Cong, Xin and Xu, Juyuan and Li, Dahai and Liu, Zhiyuan and Sun, Maosong , year =. 2307.07924 , archivePrefix =
-
[5]
2023 , eprint =
Hong, Sirui and Zhuge, Mingchen and Chen, Jiaqi and Zheng, Xiawu and Cheng, Yuheng and Zhang, Ceyao and Wang, Jinlin and Wang, Zili and Yau, Steven Ka Shing and Lin, Zijuan and Zhou, Liyang and Ran, Chenyu and Xiao, Lingfeng and Wu, Chenglin and Schmidhuber, J. 2023 , eprint =
2023
-
[6]
and Luck, Michael and Bu, Qingwen and Qing, Yuhao and Cui, Heming , year =
Huang, Dong and Zhang, Jie M. and Luck, Michael and Bu, Qingwen and Qing, Yuhao and Cui, Heming , year =. 2312.13010 , archivePrefix =
-
[7]
Tufano, Michele and Agarwal, Anisha and Jang, Jinu and Moghaddam, Roshanak Zilouchian and Sundaresan, Neel , year =. 2403.08299 , archivePrefix =
-
[8]
Zhang, Yuntong and Ruan, Haifeng and Fan, Zhiyu and Roychoudhury, Abhik , year =. 2404.05427 , archivePrefix =
Show all 92 references
-
[9]
2403.17134 , archivePrefix =
Bouzenia, Islem and Devanbu, Premkumar and Pradel, Michael , year =. 2403.17134 , archivePrefix =
-
[10]
2406.11638 , archivePrefix =
Arora, Daman and Sonwane, Atharv and Wadhwa, Nalin and Mehrotra, Abhav and Utpala, Saiteja and Bairi, Ramakrishna and Kanade, Aditya and Natarajan, Nagarajan , year =. 2406.11638 , archivePrefix =
-
[11]
2406.01304 , archivePrefix =
Chen, Dong and Lin, Shaoxin and Zeng, Muhan and Zan, Daoguang and Wang, Jian-Gang and Cheshkov, Anton and Sun, Jun and Yu, Hao and Dong, Guoliang and Aliev, Artem and Wang, Jie and Cheng, Xiao and Liang, Guangtai and Ma, Yuchi and Bian, Pan and Xie, Tao and Wang, Qianxiang , y...
-
[12]
2408.02232 , archivePrefix =
Ruan, Haifeng and Zhang, Yuntong and Roychoudhury, Abhik , year =. 2408.02232 , archivePrefix =
-
[13]
2407.01489 , archivePrefix =
Xia, Chunqiu Steven and Deng, Yinlin and Dunn, Soren and Zhang, Lingming , year =. 2407.01489 , archivePrefix =
-
[14]
2024 , eprint =
Antoniades, Antonis and. 2024 , eprint =
2024
-
[15]
Bui, Nghi D. Q. , year =. Building. 2603.05344 , archivePrefix =
-
[16]
Executable Code Actions Elicit Better
Wang, Xingyao and Chen, Yangyi and Yuan, Lifan and Zhang, Yizhe and Li, Yunzhu and Peng, Hao and Ji, Heng , year =. Executable Code Actions Elicit Better. 2402.01030 , archivePrefix =
-
[17]
You Name It, I Run It: An
Bouzenia, Islem and Pradel, Michael , year =. You Name It, I Run It: An. 2412.10133 , archivePrefix =
-
[18]
2401.07339 , archivePrefix =
Zhang, Kechi and Li, Jia and Li, Ge and Shi, Xianjie and Jin, Zhi , year =. 2401.07339 , archivePrefix =
-
[19]
2602.18571 , archivePrefix =
Garg, Spandan and Huang, Yufan , year =. 2602.18571 , archivePrefix =
-
[20]
2606.14061 , archivePrefix =
Ma, Dongjian and Chen, Silin and Yang, Yufei and Shi, Yulin and Yan, Yanfu and Gu, Xiaodong , year =. 2606.14061 , archivePrefix =
-
[21]
2509.16941 , archivePrefix =
Deng, Xiang and Da, Jeff and Pan, Edwin and He, Yannis Yiming and Ide, Charles and Garg, Kanak and Lauffer, Niklas and Park, Andrew and Pasari, Nitin and Rane, Chetan and Sampath, Karmini and Krishnan, Maya and Kundurthy, Srivatsa and Hendryx, Sean and Wang, Zifan and Zhang, C...
-
[22]
Zhang, Jenny and Hu, Shengran and Lu, Cong and Lange, Robert and Clune, Jeff , year =. Darwin. 2505.22954 , archivePrefix =
-
[23]
2511.13646 , archivePrefix =
Xia, Chunqiu Steven and Wang, Zhe and Yang, Yan and Wei, Yuxiang and Zhang, Lingming , year =. 2511.13646 , archivePrefix =
-
[24]
2507.23361 , archivePrefix =
Chen, Silin and Lin, Shaoxin and Shi, Yuling and Lian, Heng and Gu, Xiaodong and Yun, Longfei and Chen, Dong and Cao, Lin and Liu, Jiyang and Xia, Nu and Wang, Qianxiang , year =. 2507.23361 , archivePrefix =
-
[25]
2026 , eprint =
Improving Code Localization with Repository Memory , author =. 2026 , eprint =
2026
-
[26]
2605.30105 , archivePrefix =
Hu, Haichuan and Xie, Guoqing and Zhang, Quanjun and Liu, Jiawei and Yu, Shengcheng and Fang, Chunrong and Chen, Zhenyu and Xiao, Liang , year =. 2605.30105 , archivePrefix =
-
[27]
2308.10144 , archivePrefix =
Zhao, Andrew and Huang, Daniel and Xu, Quentin and Lin, Matthieu and Liu, Yong-Jin and Huang, Gao , year =. 2308.10144 , archivePrefix =
-
[28]
2507.06229 , archivePrefix =
Tang, Xiangru and Qin, Tianrui and Peng, Tianhao and Zhou, Ziyang and Shao, Daniel and Du, Tingting and Wei, Xinming and Xia, Peng and Wu, Fang and Zhu, He and Zhang, Ge and Liu, Jiaheng and Wang, Xingyao and Hong, Sirui and Wu, Chenglin and Cheng, Hao and Wang, Chi and Zhou, ...
-
[29]
2026 , eprint =
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence , author =. 2026 , eprint =
2026
-
[30]
A Comprehensive Survey of Self-Evolving
Fang, Jinyuan and Peng, Yanwen and Zhang, Xi and Wang, Yingxu and Yi, Xinhao and Zhang, Guibin and Xu, Yi and Wu, Bin and Liu, Siwei and Li, Zihao and Ren, Zhaochun and Aletras, Nikos and Wang, Xi and Zhou, Han and Meng, Zaiqiao , year =. A Comprehensive Survey of Self-Evolvin...
-
[31]
2023 , eprint =
Reflexion: Language Agents with Verbal Reinforcement Learning , author =. 2023 , eprint =
2023
-
[32]
2023 , eprint =
Self-Refine: Iterative Refinement with Self-Feedback , author =. 2023 , eprint =
2023
-
[33]
2023 , eprint =
Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution , author =. 2023 , eprint =
2023
-
[34]
and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , year =
Khattab, Omar and Singhvi, Arnav and Maheshwari, Paridhi and Zhang, Zhiyuan and Santhanam, Keshav and Vardhamanan, Sri and Haq, Saiful and Sharma, Ashutosh and Joshi, Thomas T. and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , year =. 2310.03714...
-
[35]
2406.07496 , archivePrefix =
Yuksekgonul, Mert and Bianchi, Federico and Boen, Joseph and Liu, Sheng and Huang, Zhi and Guestrin, Carlos and Zou, James , year =. 2406.07496 , archivePrefix =
-
[36]
and Stoica, Ion and Gonzalez, Joseph E
Packer, Charles and Wooders, Sarah and Lin, Kevin and Fang, Vivian and Patil, Shishir G. and Stoica, Ion and Gonzalez, Joseph E. , year =. 2310.08560 , archivePrefix =
-
[37]
2504.19413 , archivePrefix =
Chhikara, Prateek and Khant, Dev and Aryan, Saket and Singh, Taranjeet and Yadav, Deshraj , year =. 2504.19413 , archivePrefix =
-
[38]
2512.18746 , archivePrefix =
Zhang, Guibin and Ren, Haotian and Zhan, Chong and Zhou, Zhenhong and Wang, Junhao and Zhu, He and Zhou, Wangchunshu and Yan, Shuicheng , year =. 2512.18746 , archivePrefix =
-
[39]
2023 , eprint =
Voyager: An Open-Ended Embodied Agent with Large Language Models , author =. 2023 , eprint =
2023
-
[40]
2025 , eprint =
Self-Challenging Language Model Agents , author =. 2025 , eprint =
2025
-
[41]
2402.17574 , archivePrefix =
Zhang, Wenqi and Tang, Ke and Wu, Hai and Wang, Mengna and Shen, Yongliang and Hou, Guiyang and Tan, Zeqi and Li, Peng and Zhuang, Yueting and Lu, Weiming , year =. 2402.17574 , archivePrefix =
-
[42]
2406.14228 , archivePrefix =
Yuan, Siyu and Song, Kaitao and Chen, Jiangjie and Tan, Xu and Li, Dongsheng and Yang, Deqing , year =. 2406.14228 , archivePrefix =
-
[43]
2026 , eprint =
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing , author =. 2026 , eprint =
2026
-
[44]
Richard , year =
Peng, Yulin and Hou, Haowen and Zhu, Xinxin and He, Ying Tiffany and Yu, F. Richard , year =. 2603.15707 , archivePrefix =
-
[45]
2024 , eprint =
Self-Evolving Multi-Agent Collaboration Networks for Software Development , author =. 2024 , eprint =
2024
-
[46]
2505.18646 , archivePrefix =
Liu, Siwei and Fang, Jinyuan and Zhou, Han and Wang, Yingxu and Meng, Zaiqiao , year =. 2505.18646 , archivePrefix =
-
[47]
2410.10762 , archivePrefix =
Zhang, Jiayi and Xiang, Jinyu and Yu, Zhaoyang and Teng, Fengwei and Chen, Xionghui and Chen, Jiaqi and Zhuge, Mingchen and Cheng, Xin and Hong, Sirui and Wang, Jinlin and Zheng, Bingnan and Liu, Bang and Luo, Yuyu and Wu, Chenglin , year =. 2410.10762 , archivePrefix =
-
[48]
2507.03616 , archivePrefix =
Wang, Yingxu and Liu, Siwei and Fang, Jinyuan and Meng, Zaiqiao , year =. 2507.03616 , archivePrefix =
-
[49]
2602.17100 , archivePrefix =
Wang, Siyu and Lu, Ruotian and Yang, Zhihao and Wang, Yuchao and Zhang, Yanzhou and Xu, Lei and Xu, Qimin and Yin, Guojun and Chen, Cailian and Guan, Xinping , year =. 2602.17100 , archivePrefix =
-
[50]
2024 , eprint =
Language Agents as Optimizable Graphs , author =. 2024 , eprint =
2024
-
[51]
2025 , eprint =
Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies , author =. 2025 , eprint =
2025
-
[52]
2410.11782 , archivePrefix =
Zhang, Guibin and Yue, Yanwei and Sun, Xiangguo and Wan, Guancheng and Yu, Miao and Fang, Junfeng and Wang, Kun and Chen, Tianlong and Cheng, Dawei , year =. 2410.11782 , archivePrefix =
-
[53]
2510.01617 , archivePrefix =
Leong, Hui Yi and Li, Yuheng and Wu, Yuqing and Ouyang, Wenwen and Zhu, Wei and Gao, Jiechao , year =. 2510.01617 , archivePrefix =
-
[54]
A Dynamic
Liu, Zijun and Zhang, Yanzhe and Li, Peng and Liu, Yang and Yang, Diyi , year =. A Dynamic. 2310.02170 , archivePrefix =
-
[55]
Cut the Crap: An Economical Communication Pipeline for
Zhang, Guibin and Yue, Yanwei and Li, Zhixun and Yun, Sukwon and Wan, Guancheng and Wang, Kun and Cheng, Dawei and Yu, Jeffrey Xu and Chen, Tianlong , year =. Cut the Crap: An Economical Communication Pipeline for. 2410.02506 , archivePrefix =
-
[56]
2605.25430 , archivePrefix =
Li, Yanzhou and Zhang, Yiran and Zhang, Xiaoyu and Liu, Xiaoxia and Liu, Yang , year =. 2605.25430 , archivePrefix =
-
[57]
Proceedings of the ACM Conference on AI and Agentic Systems , year =
Automatically Learning Skills for Coding Agents , author =. Proceedings of the ACM Conference on AI and Agentic Systems , year =. doi:10.1145/3786335.3813196 , url =
-
[58]
2606.07412 , archivePrefix =
Xiao, Chuan and Jiao, Zhengbo and Wang, Shaobo and Wang, Wei and Zhao, Bing and Wei, Hu and Zhang, Linfeng and Qu, Lin , year =. 2606.07412 , archivePrefix =
-
[59]
2026 , eprint =
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent , author =. 2026 , eprint =
2026
-
[60]
2605.27366 , archivePrefix =
Lin, Huawei and Li, Peng and Song, Jie and Jiang, Fuxin and Zhang, Tieying , year =. 2605.27366 , archivePrefix =
-
[61]
2025 , eprint =
A Self-Improving Coding Agent , author =. 2025 , eprint =
2025
-
[62]
2026 , note =
Mendel. 2026 , note =
2026
-
[63]
Wang, Wenyi and Pi. Huxley. 2025 , eprint =
2025
-
[64]
Novikov, Alexander and Vu, Ngan and Eisenberger, Marvin and Dupont, Emilien and Huang, Po-Sen and Wagner, Adam Zsolt and Shirobokov, Sergey and Kozlovskii, Borislav and Ruiz, Francisco J. R. and Mehrabian, Abbas and Kumar, M. Pawan and See, Abigail and Chaudhuri, Swarat and Ho...
-
[65]
2510.14150 , archivePrefix =
Assumpcao, Henrique and Ferreira, Diego and Campos, Leandro and Murai, Fabricio , year =. 2510.14150 , archivePrefix =
-
[66]
2026 , eprint =
Controlled Self-Evolution for Algorithmic Code Optimization , author =. 2026 , eprint =
2026
-
[67]
, year =
Wei, Yuxiang and Duchenne, Olivier and Copet, Jade and Carbonneaux, Quentin and Zhang, Lingming and Fried, Daniel and Synnaeve, Gabriel and Singh, Rishabh and Wang, Sida I. , year =. 2502.18449 , archivePrefix =
-
[68]
Training Software Engineering Agents and Verifiers with
Pan, Jiayi and Wang, Xingyao and Neubig, Graham and Jaitly, Navdeep and Ji, Heng and Suhr, Alane and Zhang, Yizhe , year =. Training Software Engineering Agents and Verifiers with. 2412.21139 , archivePrefix =
-
[69]
Da, Jeff and Wang, Clinton and Deng, Xiang and Ma, Yuntao and Barhate, Nikhil and Hendryx, Sean , year =. Agent-. 2506.11425 , archivePrefix =
-
[70]
2504.07164 , archivePrefix =
Jain, Naman and Singh, Jaskirat and Shetty, Manish and Zheng, Liang and Sen, Koushik and Stoica, Ion , year =. 2504.07164 , archivePrefix =
-
[71]
Shum, KaShun and Hui, Binyuan and Chen, Jiawei and Zhang, Lei and X. W. and Yang, Jiaxi and Huang, Yuzhen and Lin, Junyang and He, Junxian , year =. 2512.21919 , archivePrefix =
-
[72]
Toward Training Superintelligent Software Agents through Self-Play
Wei, Yuxiang and Sun, Zhiqing and McMilin, Emily and Gehring, Jonas and Zhang, David and Synnaeve, Gabriel and Fried, Daniel and Zhang, Lingming and Wang, Sida , year =. Toward Training Superintelligent Software Agents through Self-Play. 2512.18552 , archivePrefix =
-
[73]
2026 , eprint =
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain , author =. 2026 , eprint =
2026
-
[74]
2602.03411 , archivePrefix =
Song, Huatong and Huang, Lisheng and Sun, Shuang and Jiang, Jinhao and Le, Ran and Cheng, Daixuan and Chen, Guoxin and Hu, Yiwen and Chen, Zongchao and Jia, Yiming and Zhao, Wayne Xin and Song, Yang and Zhang, Tao and Wen, Ji-Rong , year =. 2602.03411 , archivePrefix =
-
[75]
and Wettig, Alexander and Khandpur, Kabir and Zhang, Yanzhe and Hui, Binyuan and Press, Ofir and Schmidt, Ludwig and Yang, Diyi , year =
Yang, John and Lieret, Kilian and Jimenez, Carlos E. and Wettig, Alexander and Khandpur, Kabir and Zhang, Yanzhe and Hui, Binyuan and Press, Ofir and Schmidt, Ludwig and Yang, Diyi , year =. 2504.21798 , archivePrefix =
-
[76]
2021 , eprint =
Evaluating Large Language Models Trained on Code , author =. 2021 , eprint =
2021
-
[77]
2021 , eprint =
Program Synthesis with Large Language Models , author =. 2021 , eprint =
2021
-
[78]
Measuring Coding Challenge Competence with
Hendrycks, Dan and Basart, Steven and Kadavath, Saurav and Mazeika, Mantas and Arora, Akul and Guo, Ethan and Burns, Collin and Puranik, Samir and He, Horace and Song, Dawn and Steinhardt, Jacob , year =. Measuring Coding Challenge Competence with. 2105.09938 , archivePrefix =
-
[79]
2403.07974 , archivePrefix =
Jain, Naman and Han, King and Gu, Alex and Li, Wen-Ding and Yan, Fanjia and Zhang, Tianjun and Wang, Sida and Solar-Lezama, Armando and Sen, Koushik and Stoica, Ion , year =. 2403.07974 , archivePrefix =
-
[80]
Competition-Level Code Generation with
Li, Yujia and Choi, David and Chung, Junyoung and Kushman, Nate and Schrittwieser, Julian and Leblond, R. Competition-Level Code Generation with. 2022 , eprint =
2022
-
[81]
2026 , eprint =
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces , author =. 2026 , eprint =
2026
-
[82]
International Conference on Learning Representations , year =
Self-Improvement via Fast Tree-Search , author =. International Conference on Learning Representations , year =
-
[83]
Self-Taught Optimizer (
Zelikman, Eric and Lorch, Eliana and Mackey, Lester and Kalai, Adam , booktitle =. Self-Taught Optimizer (. 2024 , eprint =
2024
-
[84]
2025 , eprint =
Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement , author =. 2025 , eprint =
2025
-
[85]
2411.13941 , archivePrefix =
Lin, Yalan and Ma, Yingwei and Cao, Rongyu and Li, Binhua and Huang, Fei and Gu, Xiaodong and Li, Yongbin , year =. 2411.13941 , archivePrefix =
-
[86]
and Wan, Chengcheng and Gu, Xiaodong , year =
Wang, Zimu and Shi, Yuling and Li, Mengfan and Liu, Zijun and Zhang, Jie M. and Wan, Chengcheng and Gu, Xiaodong , year =. 2603.27850 , archivePrefix =
-
[87]
2026 , eprint =
Structurally Aligned Subtask-Level Memory for Software Engineering Agents , author =. 2026 , eprint =
2026
-
[88]
2506.11442 , archivePrefix =
Jin, Yiyang and Xu, Kunzhao and Li, Hang and Han, Xueting and Zhou, Yanmin and Li, Cheng and Bai, Jing , year =. 2506.11442 , archivePrefix =
-
[89]
2506.03136 , archivePrefix =
Wang, Yinjie and Yang, Ling and Tian, Ye and Shen, Ke and Wang, Mengdi , year =. 2506.03136 , archivePrefix =
-
[90]
2604.07864 , archivePrefix =
Fan, Lishui and Chen, Mouxiang and Zhu, Tingwei and Liu, Kui and Xia, Xin and Li, Shanping and Liu, Zhongxin , year =. 2604.07864 , archivePrefix =
-
[91]
2025 , eprint =
Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation , author =. 2025 , eprint =
2025
-
[92]
2605.16299 , archivePrefix =
Huang, Yixu and Yu, Xinglei and Wei, Zhongyu , year =. 2605.16299 , archivePrefix =
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.