{"id":"ba408e55-25f3-4473-903b-b677838ea005","arxiv_id":"2608.10299","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper organizes research on co-evolving agentic AI into three progressive stages, from agent-to-agent to agent-environment to meta co-evolution.","lead":"This survey proposes a three-stage taxonomy of co-evolution in agentic systems, from agents adapting to each other, to agents and environments adapting together, to an evolvable evolution mechanism. It organizes a fast-growing literature and maps open problems in evaluation, scaling, and safety for self-improving AI agents.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stage 3 of the taxonomy rests on a single unverified classification (RQGM); if RQGM is a single-entity meta-evolution precursor, the three-stage taxonomy reduces to two stages plus a speculative direction.","rationale":"The reader's weakest assumption identifies exactly this: Stage 3 hinges on the RQGM classification. My analysis confirms this is the most load-bearing point: the central claim of a three-stage taxonomy is falsifiable at exactly this junction. I also considered Figure 4's Panel C selection bias; it weakens the empirical motivation but does not threaten the taxonomy's internal consistency or its organization of Stages 1–2, so I do not elevate it to the primary concern. The paper handles its own limitation honestly in the Limitations section, but that does not remove the need for verification. The proposed check is concrete and decisive: if RQGM fails either condition, Stage 3 loses its only exemplar, and the survey should be reframed. Because the reader's CONDITIONAL verdict already requires this verification, no verdict change is needed.","tokens_in":30353,"tokens_out":4661,"duration_ms":42802,"concrete_test":"Read Iacob et al. (2026) and, from its method section, extract the state update equations. Verify: (1) the lower-level state St contains at least two components — a task-agent and an evaluator — whose update rules depend on each other's current state and persist across rounds; and (2) there exists a distinct, self-generated process Γt that outputs a new mechanism Ωt+1, with Ωt+1 ≠ Ωt, and that Ωt+1 is applied to produce St+1. If (1) fails, RQGM reduces to Stage 2 or single-entity meta-evolution; if (2) fails, it is at most meta-evolution with a fixed controller. Then re-run the Figure 3/Table classification with RQGM removed; if Stage 3 contains no remaining work, the taxonomy should be presented as two realized stages plus an explicitly speculative Stage 3.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a three-stage taxonomy that organizes existing work. Stage 3 (Meta Co-Evolution) is defined in §2.3/§5.1 by two joint conditions: (i) a lower-level co-evolving system St with at least two persistent adaptive units, and (ii) a self-generated revision process Γt that changes the evolution mechanism Ωt and is itself driven by the joint trajectory of that lower-level system. The paper places exactly one work in Stage 3: RQGM (Iacob et al., 2026, §5.2), described in a single sentence as 'co-evolving task agents and evaluators, while its meta-agent uses joint feedback to guide later evolution.' This sentence does not establish that RQGM satisfies either condition under the survey's own Appendix A definitions. In particular, it does not show that the task agents and evaluators are distinct persistent units that continually reshape each other's further evolution (as opposed to a single agent training against a learned reward model, which would be Stage 2 feedback-space co-evolution), nor that the meta-agent's revision process is self-generated rather than a fixed higher-level controller of the kind that Appendix A explicitly permits in mere 'meta-evolution.' The paper's own Limitations concede that 'only limited work currently meets our definition of Stage 3' and that much of the discussion draws on single-entity precursors. If RQGM is reclassified as a Stage 2 system or as a single-entity meta-evolution precursor like PromptBreeder or Gödel Agent, then Stage 3 is currently empty, and the paper's headline contribution — a progressive three-stage taxonomy — becomes a two-stage taxonomy plus a forward-looking research direction. That would not destroy the survey's value, but it would substantially weaken the specific claim that the taxonomy organizes the existing literature into three realized stages.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of co-evolution in agentic systems, defined as a multi-component form of self-evolution in which multiple agents and their environment impose adaptive pressure on one another. The authors propose a progressive three-stage taxonomy: Agent–Agent Co-Evolution, Agent–Environment Co-Evolution, and Meta Co-Evolution, in which the evolution mechanism itself becomes evolvable. The survey reviews a large body of work under this taxonomy, provides formal definitions and comparisons with adjacent surveys, and includes a cross-paper evidence figure (Figure 4) intended to show that co-evolution improves performance and that gains plateau over time. The authors also discuss evaluation, scaling, and safety challenges, and they explicitly acknowledge in the Limitations that Stage 3 is supported by very little work and that much of the related material is single-entity precursors.","tokens_in":30609,"tokens_out":3369,"duration_ms":34523,"significance":"If the taxonomy holds, this would be a useful organizing contribution: it unifies scattered work on adversarial, collaborative, environment-driven, and meta-level adaptation under a single progressive framework, and it provides a vocabulary for an important emerging direction. The manuscript's strengths include its breadth of coverage, its formal distinction between co-evolution and adjacent concepts (interaction, self-play, evolutionary optimization, meta-evolution), its systematic comparison with other surveys, and its transparent description of the construction of Figure 4 in Appendix C. The Stage 3 category is the most fragile part of the contribution: it rests on a single paper (RQGM) whose fit to the paper's own definition is not established in the text. The empirical narrative built around Figure 4 is also weakened by the exclusion of non-improving trajectories in Panel C. These issues are significant because the progressive, three-stage structure and the plateau/convergence claim are central to the survey's thesis, but they are fixable by either verifying the classification or reframing Stage 3 as a speculative future direction.","major_comments":[{"comment":"The placement of RQGM in Stage 3 is not established against the paper's own formal definition. The definition in §2.3 requires a lower-level co-evolving system with at least two persistent adaptive units, plus a self-generated revision process Γt that changes the evolution mechanism Ωt and is driven by the joint trajectory of that lower-level system. The single sentence about RQGM in §5.2 — \"co-evolving task agents and evaluators, while its meta-agent uses joint feedback to guide later evolution\" — does not rule out the possibility that the meta-agent is a fixed higher-level controller, which Appendix A explicitly permits for mere meta-evolution. It also does not show that the task agents and evaluators are distinct persistent units that continually reshape each other's further evolution, as opposed to a single agent training against a learned reward model, which would be Stage 2 feedback-space co-evolution. Because Stage 3 is one of the three claimed contributions, this classification must be either verified in detail or downgraded to a precursor/future direction.","section":"Section 5.2 and Appendix A"},{"comment":"Panel C of Figure 4 is constructed from trajectories that are included only if they have an observed positive improvement, and the construction relies on approximate digitization of published plots. Section 5.1 uses this figure to claim that \"Stage 1 and Stage 2 improve performance across most settings, but the gains become smaller as evolution approaches a plateau.\" Excluding non-improving trajectories biases the normalized curves upward and makes the claim \"across most settings\" unsupported by the evidence as presented; the figure can at most describe the average shape of improving trajectories. The authors should either include non-improving trajectories (with an explicit statement of how they are handled) or substantially qualify the plateau claim. This is load-bearing because the convergence narrative motivates the need for Stage 3.","section":"Appendix C and Section 5.1"},{"comment":"The formal definition of co-evolution uses the condition \"x evolutionary pressure ←→ y,\" but the arrow is never defined. Since the entire contribution of the paper hinges on separating co-evolution from mere interaction or correlation, this is the most important primitive in the survey. Without an operational criterion for \"exerts evolutionary pressure on,\" it is hard to audit the inclusion decisions in the taxonomy, especially for cases where two components update simultaneously for external reasons. The authors should provide at least a checklist or a directional dependence criterion, such as whether a change in x changes the learning problem faced by y in the next step.","section":"Section 2.2"}],"minor_comments":[{"comment":"There is a typo in the sentence beginning \"where peer agents share a single task reward and co-adapt as the others chang\": \"chang\" should be \"change.\"","section":"Section 3.2.1"},{"comment":"The reference list contains two entries for what appears to be the same paper by Zhai et al. (2025a and 2025b), both titled \"AgentEvolver: Towards efficient self-evolving agent system.\" This should be consolidated.","section":"References"},{"comment":"In Table 1, several cells contain missing spaces after periods, e.g., \"No dedicated coverage.Section 3.4 lists\" and similar patterns. These should be fixed for readability.","section":"Table 1"},{"comment":"Figure 3 is extremely dense, with tiny font and overlapping annotations in the reproduced version. The authors should consider splitting the landscape into separate figures or providing a higher-resolution version, and should spell out the legend items that currently read as fragments such as \"Paper backgrounds.\"","section":"Figure 3"},{"comment":"The appendix states that some values were \"approximately digitized from the published plots\" but does not say which papers were digitized rather than read from explicit labels. Listing these papers would improve reproducibility and would let readers judge the precision of Panel C.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a legitimate survey contribution, and the self-citations appear to be natural for an active group working in this area. The main concern for the editor is the evidential weight of Stage 3: the paper is careful to flag it as early-stage, but the single-paper classification of RQGM is presented as a fact rather than an interpretive claim. If the authors can either substantiate that classification or reframe Stage 3 as a research direction, the survey would be solid. The Figure 4 issue is also worth watching, since the exclusion of non-improving trajectories is a clear selection bias that the authors acknowledge only implicitly in Appendix C."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This survey gives the agentic-AI literature a genuinely new organizing scheme: a three-stage progressive taxonomy from agent-agent co-evolution through agent-environment co-evolution to meta co-evolution, defined formally with an evolvable evolution mechanism. It is well written, honest about its own limitations, and better positioned against existing self-evolution surveys than the usual survey filler. The catch is that Stage 3 is thin: it rests on a single paper (RQGM), described in one sentence, and the paper's own limitations concede that only limited work meets its definition.\n\nThe taxonomy does real work. The formal machinery in Section 2 is simple but useful: it distinguishes co-evolution from mere interaction or self-play, clarifies what counts as an evolving unit, and gives a clean decomposition of the evolution mechanism into what/when/how/where/evaluate. The comparison table with adjacent surveys is a useful reference. Covering several hundred papers with sensible organization is no small feat, and the authors are transparent about the proxy nature of their evidence: Appendix C explicitly says non-improving trajectories are dropped from Panel C, and that some values are digitized from plots. Self-citation is present but not abusive.\n\nThe soft spots are in proportion. The big one is Stage 3. RQGM is described as \"co-evolving task agents and evaluators, while its meta-agent uses joint feedback to guide later evolution.\" That does not show that the meta-agent's revision process is self-generated rather than a fixed higher-level controller, nor that the lower-level units are persistent and mutually shaping. The stress-test concern is right: if RQGM does not satisfy the paper's own Appendix A conditions, Stage 3 is currently empty, and the progressive taxonomy is really two realized stages plus a speculative direction. This is addressable: either verify RQGM against the definition, or reframe Stage 3 as a forward-looking research program rather than a stage that organizes existing work. The Figure 4 convergence narrative is also weaker than it looks because Panel C selects only improving runs; the authors should temper the \"plateau\" claim or release the scripts and data.\n\nBottom line: this deserves a serious referee. The taxonomy is a useful contribution even if Stage 3 is scaled back. Anyone working on self-evolving or multi-agent systems should read this; it will likely become a common citation. Send it to review, and ask the authors to justify the RQGM classification and provide the Figure 4 data.","headline":"A useful, honest survey whose new three-stage taxonomy is worth taking seriously, provided Stage 3 is presented as a research direction rather than an established category.","tokens_in":31275,"tokens_out":2409,"would_cite":true,"duration_ms":23697,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that agentic improvement is fundamentally co-evolutionary: multiple agents, their environments, and eventually their own evolution mechanisms change in response to one another, and this survey organizes that literature…","keywords":["co-evolution","agentic systems","self-evolving agents","multi-agent systems","open-endedness","meta co-evolution","adaptive environments","LLM agents"],"falsifier":"Inspect RQGM's update loop: if its meta-agent revises the evolution mechanism using only the trajectory of a single entity, rather than joint feedback from at least two mutually adapting units, then the paper's Stage 3 definition is unmet and the third tier has no current instantiation.","tokens_in":30092,"feed_emoji":"🧬","tokens_out":6895,"duration_ms":57779,"temperature":0.7,"pith_summary":"Agentic systems improve not only by evolving as single entities but by co-evolving: multiple agents, and eventually their environments and even their own evolution mechanisms, change in response to one another. The paper proposes a progressive three-stage taxonomy—agent–agent co-evolution, agent–environment co-evolution, and meta co-evolution—that traces how a system sheds human-engineered constraints and expands what it is allowed to evolve. The aim is to unify scattered work on adversarial, collaborative, environment-driven, and meta-level adaptation under one definition: co-evolution occurs when at least two evolving units jointly adapt and continually reshape each other's further evolution, rather than merely exchanging information. If the taxonomy holds, it gives researchers a shared map for building self-directed, open-ended agentic systems that keep improving after deployment.","feed_headline":"AI agents improve by co-evolving, not just self-evolving","feed_subtitle":"A new survey maps adversarial, environmental, and meta-level adaptation onto one expanding boundary of evolutionary freedom.","key_machinery":"The machinery is the formal definition of co-evolution and the three-stage taxonomy built on it. A system is $S = (A, E)$ with agent collective $A = (\\{a_1,\\ldots,a_n\\}, \\Pi)$ and environment $E$; an agent $a_i = (m_i, h_i)$ has a model backbone and a harness (memory, tools, skills, prompts, workflows); an evolution mechanism $\\Omega$ drives state transitions $S_{t+1} = \\Omega(S_t, \\tau_t)$, specifying what can evolve, when, how, where, and how to evaluate. Co-evolution requires at least two units $x \\neq y$ in $S_t$ that both change and exert evolutionary pressure on each other. Stage 3 adds the meta recursion $\\Omega_{t+1} = \\Gamma_t(S_t, \\Omega_t, \\tau_t)$, where a self-generated revision process changes the mechanism that then drives lower-level evolution. The five-way decomposition of $\\Omega$ into what, when, how, where, and evaluate is what lets the survey classify methods and delimit the frontier of the field.","core_discovery":"The paper's central claim is that co-evolution is a distinct and increasingly central mode of agentic self-improvement, and that the field can be organized by an expanding boundary of evolutionary freedom. Stage 1 covers agents adapting through dynamic peers—adversarial, collaborative, and organizational—within a fixed environment. Stage 2 extends adaptation to the environment itself: tasks, feedback, and interaction spaces that change with the agents. Stage 3, meta co-evolution, is defined by a recursion in which the evolution mechanism $\\Omega$ is revised by a self-generated process $\\Gamma$, so the system decides what, when, how, and where to evolve, and how to evaluate those changes. The paper presents meta co-evolution as a pathway toward open-endedness, characterized by continuous novelty ($\\Omega_{t+1} \\neq \\Omega_t$) and unbounded adaptive capacity ($\\lim_{t\\to\\infty} H(S_t, \\Omega_t) = \\infty$), and identifies RQGM as the one cited work currently meeting that definition.","pith_inferences":["A direct test of Stage 3 is to ablate the lower-level co-evolution: if the meta-agent's revision signal is built from a single entity's trajectory, the system is a precursor, not meta co-evolution.","The Red Queen logic implies that static benchmarks will increasingly understate what co-evolved systems can do, because the counterpart's adaptation is itself the source of difficulty; evaluation may eventually need to co-evolve with the systems it measures.","The taxonomy suggests a scaling route from pairwise loops (attacker–defender, policy–reward, agent–task) to full multi-component systems, with the meta-mechanism deciding which components should change at each step."],"forward_implications":["Co-evolution should be treated as distinct from interaction, self-play, evolutionary optimization, and continual learning, so evaluation and benchmarks must track whether multiple components improve together.","Adversarial, collaborative, and organizational adaptation are one stage, not separate trends, because they differ only in what evolves within the agent collective.","Adaptive tasks, feedback, and interaction spaces belong to a single second stage, so progress on curriculum learning, reward-model updates, and world-model construction can be compared on the same axis.","Realizing meta co-evolution would let systems move from adaptation under designed rules toward self-expanding open-endedness, but requires governance that keeps the process auditable and interruptible.","Fixed benchmarks should be paired with process-level testing—historical cross-play, component ablations, and held-out evaluators—to catch evaluator exploitation, partner overfitting, and diversity collapse."],"supporting_citations":[{"why":"Supplies the Red Queen principle that motivates co-evolution as mutual adaptation rather than one-sided learning.","marker":"Van Valen, 2014"},{"why":"Origin of co-evolving parasites, the canonical example that co-evolution improves optimization.","marker":"Hillis, 1990"},{"why":"The one cited system the paper classifies as genuine meta co-evolution, making Stage 3 non-empty.","marker":"Iacob et al., 2026"},{"why":"Defines open-endedness, the long-term outcome the taxonomy is built toward.","marker":"Stanley et al., 2017"},{"why":"Position that open-endedness is essential, used to argue meta co-evolution matters.","marker":"Hughes et al., 2024"},{"why":"PromptBreeder as a single-entity precursor that evolves mutation prompts, delimiting what Stage 3 is not.","marker":"Fernando et al., 2024"},{"why":"Gödel Agent as a single-entity precursor that evolves self-modification machinery, another boundary case.","marker":"Yin et al., 2025"},{"why":"MemEvolve as a precursor evolving memory architecture, used to separate meta-evolution from meta co-evolution.","marker":"Zhang et al., 2025a"},{"why":"Closest existing survey on self-evolving agents, which the paper distinguishes by making co-evolution the central axis.","marker":"Gao et al., 2026"}],"fun_headline_variants":["Co-evolving agents: the next leap beyond self-improvement","AI co-evolution: from fixed tasks to self-directed change","Three stages to open-ended AI: co-evolution's expanding arc","Meta co-evolution: when AI evolves its own evolution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that RQGM (Iacob et al., 2026) genuinely meets the paper's definition of meta co-evolution—a lower-level co-evolving system whose joint feedback revises its own evolution mechanism—so that Stage 3 is occupied rather than speculative.","fun_headline_variants_meta":{"raw":{"variants":["Co-evolving agents: the next leap beyond self-improvement","AI co-evolution: from fixed tasks to self-directed change","Three stages to open-ended AI: co-evolution's expanding arc","Meta co-evolution: when AI evolves its own evolution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000327,"raw_usage":{"total_tokens":1823,"prompt_tokens":931,"completion_tokens":892,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":831}},"tokens_in":547,"tokens_out":892,"duration_ms":8150,"temperature":1.0,"reasoning_tokens":831,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:19.484749+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect RQGM's update loop: if its meta-agent revises the evolution mechanism using only the trajectory of a single entity, rather than joint feedback from at least two mutually adapting units, then the paper's Stage 3 definition is unmet and the third tier has no current instantiation.","supporting_citations":[],"review_version":1}