{"id":"08494fe7-9a75-4605-9f09-aa8509c1715a","arxiv_id":"2608.03764","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Rule hybridization makes agent self-evolution gains attributable; the best evolved configuration reaches 67.07% against a 91.6% fully-informed oracle ceiling.","lead":"GDPevo is a new benchmark that tests whether AI agents can learn reusable business rules from earlier tasks and apply them to related held-out tasks, using a fully automated pipeline that regenerates tasks to fight data contamination. It tracks accuracy and cost for four agent systems across CRM, ERP, finance, healthcare, legal, and data workflows.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibration thresholds in §3.4 select task groups by requiring a 10–30 pp fewshot lift, so the headline 'self-evolution consistently improves' is partly a selection artifact; sensitivity is never quantified.","rationale":"The paper is a serious, transparent benchmark-construction effort: rule hybridization is a principled way to create transferable skills, deterministic graders and container isolation are real controls, and the appendix case studies (F.1–F.5) provide concrete evidence of rule-level transfer and negative transfer. Those strengths mean the conditional verdict is appropriate, not rejection. The weakest point is exactly the one the reader flagged: Section 3.4's calibration criteria define the benchmark in terms of the outcome being measured. Since the calibration agent is Codex/GPT-5.5, the acceptance rule is effectively a selection filter for groups where that specific model's fewshot lift is 10–30 pp. Thus the aggregate gains in Table 3 are not an independent estimate of how much current agents benefit from self-evolution; they are partly guaranteed by construction. The fact that DS-V4's fewshot lift is only 5.21 pp and Appendix F shows a negative group (TG018) shows the effect is not fully forced across all agents/tasks, but the paper's unqualified 'consistently improves' overstates what the data support. The concrete test—regenerating the benchmark with altered thresholds—directly quantifies how much of the headline is selection. Until then, CONDITIONAL remains the right verdict, so no adjustment is needed.","tokens_in":21018,"tokens_out":7353,"duration_ms":90717,"concrete_test":"Run the released pipeline a second time with the §3.4 calibration thresholds changed to, e.g., base 0.2–0.8, fewshot lift >0.05 (or no fewshot-lift requirement), and fewshot accuracy <0.9, generating new task groups with the same rule-hybridization mechanism. Evaluate the four agents/supervision types on this unconstrained set and compare macro-average lifts. If the mean fewshot lift for GPT-5.5/Codex remains around 15 pp and DS-V4 still improves, calibration is not the main driver of the headline; if gains shrink toward zero or reverse, the headline is a selection artifact. A cheaper supplementary check: use the released per-group leaderboard data to recompute macro-average fewshot lift after dropping groups with the smallest and largest base accuracies, and test whether any of the four agents' 'consistent improvement' survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.4 accepts a task group only if a calibration agent (Codex+GPT-5.5 per Figure 1) scores 40–60% base, gains 10–30 pp under fewshot supervision, and stays below 80% after fewshot. The central empirical claim—that self-evolution consistently improves held-out accuracy by up to 16.44 pp—is therefore not an unbiased measurement of agent self-evolution. It is partly a consequence of keeping only groups on which the calibration model exhibits a strong fewshot gain. GPT-5.5's reported fewshot lift (15.14 pp) and Opus-4.8's (16.44 pp) sit exactly inside the calibrated 10–30 pp band, which is expected under the acceptance rule. Moreover, the task set is filtered for learnability by one specific model (GPT-5.5), so macro-average gains across all four agents are not independent evidence of general self-evolution ability. The paper does not report how the headline changes when calibration thresholds are varied, and it does not separate the calibrated fewshot-lift requirement from the models' actual behavior. The appendix itself shows group-level negative transfer (e.g., TG018, DS-V4 fewshot 39.83 vs base 48.36), so 'consistently improves' is also stronger than the per-group data support. This does not invalidate GDPevo as a calibrated evolution testbed, but it undermines the unqualified empirical conclusion that current agents reliably improve from self-evolution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GDPevo, a benchmark and automated pipeline for evaluating self-evolution of AI agents on enterprise workflows. The core proposal is rule hybridization: each workflow is decomposed into atomic business rules, different subsets are planted in five training tasks, and recombinations are used in five held-out test tasks, so that test-time gains can be attributed to training experience. The release contains 240 tasks in 24 groups across six domains (V1+V2). The evaluation covers four harness+model agents and four supervision types (base, self, reflect-3, fewshot), using deterministic rule-based graders and cost metrics. The authors report that every evolved configuration improves over its base by up to 16.44 pp, that fewshot is the most reliable supervision type, that reflect transfers across domains more robustly than fewshot, and that the evolution method matters less than the underlying model. They also report a fully informed oracle ceiling of 91.6%, which the evolved agents remain far below.","tokens_in":21403,"tokens_out":4262,"duration_ms":55881,"significance":"If the benchmark construction is accepted, GDPevo is a useful contribution to the evaluation methodology for agent self-evolution. Its strengths are real: rule hybridization is a concrete mechanism for making train-test transfer attributable; the pipeline is automated and can regenerate tasks to resist contamination; scoring is deterministic and rule-based; the experimental isolation controls are carefully described; and the appendix case studies (Appendix F) provide genuinely diagnostic qualitative evidence, including honest examples of negative transfer. The public release of the pipeline, benchmark, and per-task artifacts is a substantial asset. However, the empirical headline that 'self-evolution consistently improves held-out accuracy' is stronger than the evidence. The calibration procedure in Section 3.4 selects task groups precisely for a 0.1-0.3 fewshot lift, which makes the headline improvement partly a property of the acceptance rule rather than an unbiased measurement. Group-level results in Table 5 and Figure 5 also show negative transfer and many groups with no improvement. These issues are local to the interpretation of the empirical results rather than to the benchm","major_comments":[{"comment":"The calibration acceptance rule states that a task group is kept only if a calibration agent has base accuracy 0.4-0.6, fewshot lift 0.1-0.3, and fewshot accuracy below 0.8. The headline result that every agent improves under fewshot (Section 4.2, Table 3) is therefore partly engineered: the benchmark excludes task groups where the calibration model does not exhibit a 10-30 pp fewshot gain. The observed gains of GPT-5.5 (+15.14 pp) and Opus-4.8 (+16.44 pp) fall inside this band, as expected under the acceptance rule. The paper never reports how the headline changes if the bands are varied, nor how many candidate groups were rejected at each stage. Since the central empirical claim depends on this selection, the manuscript should provide a sensitivity analysis over the calibration thresholds and report rejection rates. Without this, 'self-evolution consistently improves' should be restate","section":"Section 3.4"},{"comment":"The claim that self-evolution 'consistently improves' is contradicted by the paper's own per-group data. Table 5 shows that for TG018 (DeepSeek-V4-Pro-Preview), every evolved supervision type is below base (base 48.36 vs fewshot 39.83, self 41.34, reflect-3 45.71); for TG016 (GLM-5.2), every evolved type is below base. Figure 5(a) states that self improves only 16 of 24 task groups and reflect 20 of 24 for Opus-4.8. The macroscopic statement 'every evolved combination improves over its base' is therefore a statement about macroaverages, not about consistent improvement at the group level. In addition, the paper reports no statistical significance tests. The smallest reported gain, DeepSeek self +2.59 pp, is within one standard deviation of the group-level distribution (STD 7.89 in Table 3) and is not shown to be significant. The authors should report per-group sign tests or confidence in","section":"Tables 3 and 5; Figure 5"},{"comment":"The RQ2 cross-domain transfer experiment uses only one task group per domain (tg02, tg06, tg10) and only two supervision types (fewshot and reflect-3). The claim that 'fewshot overfits its source group and can hurt on others, whereas reflect transfers more robustly' is based on a 3x3 matrix with six off-diagonal cells. With one group per domain, the observed deltas may reflect group-specific properties rather than domain-level or supervision-type-level regularities. The paper should either expand this experiment to more task groups or substantially qualify the conclusion. At minimum, the variance across groups should be reported so the reader can see whether the five negative fewshot cells are robust or driven by the choice of a single group.","section":"Section 4.3"},{"comment":"The calibration step is performed with a single calibration agent (Codex with GPT-5.5), and the same model is used as the calibration oracle and as one of the evaluated models. This makes the cross-model claims less independent than they appear. A task group is retained only if that particular model exhibits a fewshot lift; the fact that a different model (e.g., DeepSeek-V4-Pro-Preview) also improves on many groups is not independent evidence of general self-evolution ability, because the task set has been filtered for a property measured through GPT-5.5. The manuscript should disclose this dependency explicitly and, ideally, rerun calibration with at least one other model or report how much of the macro gain is attributable to the filtering. This is a load-bearing issue for the general claim that 'current agents' self-evolve reliably.","section":"Section 3.4 and Section 4.2"}],"minor_comments":[{"comment":"There are typographical and spacing issues such as 'GDPevoaddresses' (Introduction, first paragraph after the limitations) and 'GDPevois' in the same paragraph. These should be corrected in a final pass.","section":"Abstract / Introduction"},{"comment":"The four evaluated 'agents' are not factorial in harness and model: GPT-5.5 is paired only with Codex, while the other three models are paired with Claude Code. This confounds model and harness in the RQ1 comparisons. The paper should state this limitation clearly in the main text rather than only in the appendix, and the comparisons in Table 3 should be read as configurations rather than as isolated model or harness effects.","section":"Section 4.1 / Table 3"},{"comment":"Error bars are one standard deviation across task groups, not standard errors. For small gains such as DeepSeek self (+2.59 pp), the error bars suggest that the effect could be zero or negative. Please clarify in the caption what the error bars represent and, ideally, add standard errors or confidence intervals.","section":"Figure 2"},{"comment":"The main text (Section 4.2) describes the oracle as 'a fully informed model' that is given all hidden rules and gold answers, while Appendix D describes a human operator paired with Codex GPT-5.5 who receives the training packages but not the held-out test answer or evaluator. These are not the same protocol. The text should be reconciled so readers understand which procedure produced the 91.6% ceiling.","section":"Appendix D"},{"comment":"The heatmap uses only one randomly chosen group per domain, but the randomness is not described (seed, number of draws). Since the conclusion rests on six off-diagonal cells, the selection procedure and its stability should be documented.","section":"Section 4.3 / Figure 3"},{"comment":"The claim that V2 was produced 'within two days' would be more convincing with documentation of compute, number of agent runs, and failure/retry counts. This is not a central issue, but it supports the reproducibility of the pipeline.","section":"Section 3.1 / Table 2"}],"recommendation":"major_revision","confidential_remarks":"The benchmark and pipeline are potentially valuable community resources, and the authors are commendably candid in Appendix F about negative transfer and in Appendix G about methodological pitfalls. The barrier to acceptance is not the construction of the benchmark but the gap between the headline empirical claims and the selection procedure and statistics. I would encourage the editor to require a sensitivity analysis of the calibration bands and a statistically careful re-statement of the consistency claim. I do not see the need for a full rejection, since the core methodology is sound and the limitations are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: GDPevo is a genuinely useful benchmark with a real construction mechanism, but its headline empirical claim is partly a product of its own calibration filter. The benchmark is worth having; the 'consistently improves' language is not supported by the paper's own data.\n\nWhat is actually new: rule hybridization. Building the train-test relationship forward from atomic business rules instead of deriving it post hoc from an existing suite is a real departure from EvoAgentBench, SkillFlow, and SEA-Eval, and it gives the attribution story a firm footing. The automated pipeline, deterministic graders, isolated containers, cost as a first-class metric, and the 91.6% human-plus-Codex oracle ceiling are solid, useful engineering. The paper also deserves credit for honesty: Appendix G is a frank failure-mode writeup, and Appendix F includes a clean negative-transfer case (TG018) rather than only successes.\n\nThe main soft spot is the calibration circularity, and the stress-test is right about it. Section 3.4 accepts a task group only when a calibration agent shows a 10-30pp fewshot lift, so the suite is populated with groups where evolution is visible by design. The headline gains — and the ordering fewshot > reflect > self — are partly a selection effect, the more so because fewshot was itself the calibration condition. The paper discloses all of this, which mitigates the problem, but it never quantifies how the results move when the bands change. As a calibrated testbed for detecting learnability, the design choice is defensible; as an unbiased measurement of whether current agents evolve, it is not.\n\nTwo smaller issues. First, 'consistently improves' overstates the evidence: DeepSeek-V4-Pro-Preview's self gain of +2.59pp sits well within the cross-group noise (std around ±8 across 24 groups), no significance tests are reported, and Appendix F shows per-group negative transfer — so the claim holds at best at the macro level, not per group. Second, RQ2 rests on three hand-picked groups, so the SFT-overfits-vs-RL-transfers story is suggestive rather than established. Reproducibility depends on the public repo plus closed models I cannot inspect from the text, so the numbers need independent runs before I would lean on them.\n\nWho this is for: anyone building or evaluating agent self-evolution benchmarks, and anyone who wants a contamination-resistant, deterministic testbed with traceable rule-level scoring. It deserves a serious referee. The requested revisions are bounded: add a sensitivity analysis over calibration thresholds, report significance or soften the claims, and scope RQ2 accordingly. Send it to review.","headline":"Genuinely useful evolution benchmark with a real mechanism (rule hybridization), but the headline 'consistently improves' claim is partly a calibrated selection effect and the paper overstates what its own data show.","tokens_in":21870,"tokens_out":5696,"would_cite":true,"duration_ms":64397,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GDPevo: a benchmark that measures agent self-evolution on real business workflows and finds held-out accuracy gains up to 16.44 percentage points.","keywords":["agent self-evolution","rule hybridization","enterprise workflow benchmark","train-test attribution","LLM agents","data contamination","economically valuable tasks","skill-based evolution"],"falsifier":"Rerun the pipeline without the calibration thresholds (base accuracy 0.4-0.6, few-shot lift 0.1-0.3, few-shot accuracy below 0.8) and measure whether self-evolution gains persist; if the up-to-16.44-point advantage shrinks or reverses under other bands, the headline result is a selection artifact rather than a general property of self-evolution.","tokens_in":20921,"feed_emoji":"🧩","tokens_out":5955,"duration_ms":67835,"temperature":0.7,"pith_summary":"This paper introduces GDPevo, a benchmark and fully automated pipeline for measuring whether AI agents get better at real business work by reusing experience from earlier tasks. Its central design, rule hybridization, breaks each enterprise workflow into atomic business rules, scatters subsets of those rules across five training tasks, and recombines them in five held-out test tasks, so any test-time improvement has to come from training. The authors evaluate four agents under four supervision types and report that self-evolution consistently raises held-out accuracy, up to 16.44 percentage points, while sometimes also cutting test-time cost. They also estimate a fully informed oracle ceiling of 91.6%, and the best evolved agents stay well below it, suggesting current self-evolution is functional but far from complete.","feed_headline":"Self-evolution lifts held-out accuracy by up to 16.44 points","feed_subtitle":"Even the best evolved agents sit far below the 91.6% fully informed oracle ceiling.","key_machinery":"Rule hybridization is the mechanism that carries the argument. Each task group's business logic is decomposed into atomic rules, deliberate enterprise-specific conventions absent from the model's world knowledge. The rules are distributed as subsets across five training tasks and recombined in five held-out test tasks within the same shared environment. This design is what makes an accuracy gain after evolution attributable to training experience: a test task can require composing rules never seen together, so only an agent that inferred the rules during training and can combine them succeeds. The deterministic rule-based grader, which converts each rubric point into a code-based test case,","core_discovery":"The paper's central claim is that rule hybridization makes training-to-test generalization concrete and test-time gains attributable, and that, measured on the resulting benchmark, self-evolution works: every evaluated supervision type improves over the no-evolution base for all four agents, by 2.59 to 16.44 percentage points. At the same time, no evolved agent approaches the fully informed oracle ceiling of 91.6%, which the paper treats as evidence that self-evolution ability is still far from fully realized. The benchmark spans CRM, ERP, finance, healthcare, legal, and data-centric workflows, with 240 tasks in 24 task groups, and the pipeline regenerates fresh versions quickly to counter c","pith_inferences":["The headline 'evolution helps' result is partly a selection effect: calibration keeps only task groups with base accuracy of 40-60%, few-shot lift of 10-30 points, and post-evolution accuracy below 80%; changing those bands could change the measured gains.","The 91.6% oracle ceiling is a human-plus-model estimate given full training-side evidence, not an ideal mathematical upper bound, so improvements beyond it are not ruled out.","The finding that a minimal 'naive' skill creator matches or beats elaborate creators suggests that, for the skill-based evolution setting tested here, the model's intelligence matters more than the evolution method; testing parameter-updating evolution methods would be a natural extension.","Cross-domain transfer results imply that deployment choices about supervision type should weigh cost and negative-transfer risk: reflection-based feedback may be preferable when off-target transfer is expensive, while few-shot supervision is strongest within a domain."],"forward_implications":["Evolution-native benchmarks can be constructed from the outset rather than carved out of existing task suites, making the transferable ability explicit and testable.","Self-evolution is measurable and reproducible on economically valuable tasks: every tested supervision type improved every agent on held-out tasks.","Few-shot supervision behaves like supervised fine-tuning and can overfit its source domain, while reflection-based feedback transfers more robustly across domains.","Because all evolved agents remain below the 91.6% oracle ceiling, there is clear headroom; future self-evolution methods can be scored against this bound.","The fully automated pipeline can refresh the benchmark as public tasks become exposed, providing a practical response to data contamination."],"supporting_citations":[{"why":"Supplies a prior evolution-native benchmark organized into workflow families and motivates the skill-based evolution method used in the evaluation.","marker":"[8]"},{"why":"Provides a prior ability-graph train/test design that rule hybridization is positioned against and the skill-based evolution setup.","marker":"[10]"},{"why":"Provides GDP-related economically valuable seed scenarios from which task groups are constructed.","marker":"[14]"},{"why":"Provides rule-governed standard-operating-procedure seed scenarios used in scenario discovery.","marker":"[15]"},{"why":"Provides occupation-specific professional task seeds for scenario discovery.","marker":"[16]"},{"why":"Motivates the contamination-response requirement by showing how static benchmarks lose validity.","marker":"[17]"},{"why":"Supports the need for automatically constructed benchmarks that update with real-world knowledge, grounding the regeneration pipeline.","marker":"[18]"},{"why":"Defines reflective verbal-reinforcement learning, the analogue for the reflect supervision type.","marker":"[3]"}],"fun_headline_variants":["Self-evolution adds up to 16.44 points on held-out tasks","GDPevo: self-evolution gains up to 16.44 pts, yet below oracle","Rule hybridization makes self-evolution gains measurable","Evolved agents gain 16.44 pts, but lag oracle ceiling badly","Self-evolution works, but agents remain far from oracle"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The headline result depends on calibration filters: a task group is kept only when base accuracy is 40-60%, few-shot training lifts it by 10-30 percentage points, and evolved accuracy stays below 80%; those bands select for groups where evolution gains are visible, so the measured 'evolution helps' effect is partly a property of the filtered benchmark.","fun_headline_variants_meta":{"raw":{"variants":["Self-evolution adds up to 16.44 points on held-out tasks","GDPevo: self-evolution gains up to 16.44 pts, yet below oracle","Rule hybridization makes self-evolution gains measurable","Evolved agents gain 16.44 pts, but lag oracle ceiling badly","Self-evolution works, but agents remain far from oracle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000506,"raw_usage":{"total_tokens":2339,"prompt_tokens":815,"completion_tokens":1524,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":1445}},"tokens_in":559,"tokens_out":1524,"duration_ms":11284,"temperature":1.0,"reasoning_tokens":1445,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:02:12.999471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the pipeline without the calibration thresholds (base accuracy 0.4-0.6, few-shot lift 0.1-0.3, few-shot accuracy below 0.8) and measure whether self-evolution gains persist; if the up-to-16.44-point advantage shrinks or reverses under other bands, the headline result is a selection artifact rather than a general property of self-evolution.","supporting_citations":[{"cited_title":"EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer","cited_arxiv_id":"2607.05202","evidence_quote":"Provides a prior ability-graph train/test design that rule hybridization is positioned against and the skill-based evolution setup."},{"cited_title":"Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek","cited_arxiv_id":null,"evidence_quote":"Provides GDP-related economically valuable seed scenarios from which task groups are constructed."},{"cited_title":"JobBench: Aligning agent work with human will.CoRR, 2026","cited_arxiv_id":null,"evidence_quote":"Provides occupation-specific professional task seeds for scenario discovery."},{"cited_title":"LiveBench: A challenging, contamination-limited LLM benchmark","cited_arxiv_id":null,"evidence_quote":"Motivates the contamination-response requirement by showing how static benchmarks lose validity."},{"cited_title":"AntiLeakBench: Preventing data contamination by automatically constructing benchmarks with updated real-world knowledge","cited_arxiv_id":null,"evidence_quote":"Supports the need for automatically constructed benchmarks that update with real-world knowledge, grounding the regeneration pipeline."},{"cited_title":"Reflexion: Language 10 agents with verbal reinforcement learning","cited_arxiv_id":null,"evidence_quote":"Defines reflective verbal-reinforcement learning, the analogue for the reflect supervision type."}],"review_version":1}