{"id":"e2f4b67e-31c6-4b30-83a5-925bfc354bde","arxiv_id":"2608.09629","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A frontier language model can optimize an agent's skill without a prescribed improvement pipeline, matching structured baselines while using fewer target interactions, provided the optimizer is strong enough.","lead":"This paper tests whether self-improving AI agents still need hand-designed optimization pipelines, or whether a powerful language model can invent its own improvement process on the fly. Across benchmark comparisons, the open-ended approach matched or beat two prescribed pipelines while using about a third of the interaction budget, but weaker models still needed the structure.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 12-1-1 headline rests on single-run comparisons against only two prescribed pipelines, so the 'not necessary' conclusion may be an artifact of baseline selection and seed luck.","rationale":"The reader's weakest assumption identifies exactly the load-bearing point: the entire argument is comparative, and the comparators are two specific pipelines run once each. My reading confirms that the headline overgeneralizes from this narrow base. The paper is otherwise honest: it states the capability boundary explicitly, reports budget use, and includes a one-shot control. The concern is external validity rather than internal inconsistency, so it does not justify rejection; it justifies the conditional verdict until broader baseline coverage and variance reporting are provided.","tokens_in":10937,"tokens_out":14438,"duration_ms":144150,"concrete_test":"Run all 8 benchmark–target settings with 5 independent seeds using the official SkillOpt and GEPA repositories with their default/recommended hyperparameters, and repeat the protocol with a third prescribed pipeline (e.g., SkillOpt-Lite or AFlow). For each cell report the mean and 95% CI. If OEO's margin over a correctly configured prescribed baseline shrinks to within noise, or if a third prescribed pipeline beats OEO by more than the within-seed noise on multiple settings, the 'not necessary' conclusion should be narrowed to 'not necessary against these two implementations.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is categorical: prescribed optimization pipelines are not necessary for competitive self-evolution (Abstract, Section 5). The evidence is a 14-cell comparison against exactly two instantiations of 'prescribed'—SkillOpt and GEPA—with one stochastic run per cell and no released code, prompts, or seed control (Section 3.1, Appendix A.2). Two problems compound. First, the win/tie/loss count is not statistically grounded: two decisive margins are within 1–3 items out of 1,400 (OEO leads GEPA by 0.07 pp on SearchQA/Qwen and trails by 0.21 pp on SearchQA/GPT-5.5), so a single reseeding can flip the record. Second, the baselines are not demonstrated to be running at their official recommended configurations; without released artifacts an independent reviewer cannot tell whether SkillOpt or GEPA was under-tuned, mis-budgeted, or missing a component. The 14 head-to-head cells are also not 14 independent tasks—they are 8 benchmark–target settings with two comparators, so the category-level conclusion rests on a small, non-random sample. If SkillOpt is not representative of staged pipelines, or GEPA is not representative of evolutionary pipelines, the comparison tests only those two implementations, not 'prescription in general.' The capability-boundary experiments (Section 3.3) hedge the conclusion but do not repair the representativeness of the frontier comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Open-Ended Optimization (OEO), a protocol in which a frontier model acting as the optimizer composes the task-specific improvement process online, while the framework retains control over an external 'optimization contract' (objective, permitted interactions, budget, data boundaries, evaluation). The authors compare OEO against two prescribed pipelines, SkillOpt and GEPA, across 14 head-to-head cells spanning four benchmarks and two target models, reporting 12 wins, 1 tie, and 1 narrow loss (0.21 percentage points), along with a lower realized target-interaction token budget. Additional controls include a one-shot zero-interaction rewrite, a capability ladder showing that SkillOpt outperforms OEO at medium optimizer capability while a weak optimizer is blocked under OEO, and trajectory diagnostics showing that prescription changes the optimization path more consistently than the final evaluated behavior. The paper concludes that prescribed optimization pipelines are capability-dependent scaffolding rather than a prerequisite for frontier optimizers.","tokens_in":11188,"tokens_out":6155,"duration_ms":53776,"significance":"If the empirical pattern is robust, this is a useful conceptual reframing of self-evolving agent design. The separation between an externally governed optimization contract and a delegated task-specific meta-policy provides a clear vocabulary for a design question that is currently fragmented across many systems. The study also includes honest falsifiable controls: the one-shot rewrite shows that a static prior alone does not reproduce interactive gains, and the capability ladder demonstrates a crossover that restricts delegation's scope. The trajectory analysis is a thoughtful attempt to separate process from outcome. However, the headline claim that prescribed pipelines are 'not necessary' currently rests on a small, single-run, two-baseline comparison without released code or data. The qualitative pattern is plausible and worth reporting, but the strength of the categorical conclusion exceeds what the evidence can support without additional variance analysis and baseline validation.","major_comments":[{"comment":"The central '12 wins, 1 tie, 1 narrow loss' record is based on a single stochastic run per cell, with no seeds, confidence intervals, or significance testing reported in Appendix A.1. Two decisive margins are within a few items: OEO trails GEPA on SearchQA/GPT-5.5 by 0.21 percentage points (about 3 of 1,400 items) and leads GEPA on SearchQA/Qwen by 0.07 percentage points (about 1 item). A single reseeding could flip both cells, turning the claimed record into 10 wins and 3 losses (or similar). The paper should either provide multiple seeds with variance estimates, or explicitly downgrade the win/loss count to a descriptive point estimate and identify which margins are below a meaningful effect size.","section":"§3.1, Table 2"},{"comment":"The general conclusion that prescribed pipelines are 'not necessary' is supported by exactly two instantiations of 'prescribed'—SkillOpt and GEPA—and the paper does not demonstrate that these are run at their official recommended configurations. Appendix A.2 states that GEPA 'disables optional merging' and that the two SpreadsheetBench GEPA cells use a post-hoc matched checkpoint rather than GEPA's native full-budget selection. If optional merging is part of GEPA's default behavior, disabling it could handicap the baseline. The paper should either run both baselines under their documented defaults with released configurations, or narrow the claim explicitly to 'the two pipelines as configured here.' Without this, the categorical framing in Section 5 and the abstract is not supported.","section":"§2.2, Appendix A.2"},{"comment":"The abstract's 'not necessary' and Section 5's 'prescribed optimization pipeline is not a prerequisite' are categorical statements, but the evidence spans 4 benchmarks, 2 target models, 2 prescribed methods, and a single run per cell. This is a small, non-random sample, and the paper provides no argument that SkillOpt and GEPA represent the space of prescribed pipelines broadly. The conclusion in Section 7 is appropriately hedged ('Across the tested settings'), but the abstract and discussion go further. I recommend either rewording the abstract and Section 5 to 'not necessary in the tested settings' or 'among the tested pipelines,' or adding an explicit generalization argument (e.g., a survey of pipeline design axes and an argument that these two instantiations cover the extremes).","section":"§5"},{"comment":"The capability ladder is a valuable control that shows a crossover between medium and frontier optimizers, but it is also based on single runs. The medium-optimizer gap on LiveMath is substantial (9.68 percentage points), so this specific result is less sensitive to sampling noise than the head-to-head margins, but the weak-optimizer 'blocked' outcome and the medium-optimizer gap should still be accompanied by at least a seed or variance indicator (e.g., running the medium optimizer two or three times). Given that the paper's central claim is capability dependence, this table is load-bearing and would benefit from the same robustness treatment as Table 2.","section":"§3.3, Table 4"}],"minor_comments":[{"comment":"The method name is written inconsistently as 'SKILLOPT', 'SkillOpt', and 'SkillOPT'; please standardize to one capitalization scheme.","section":"Throughout"},{"comment":"The dagger symbol for SpreadsheetBench GEPA cells is explained in the text but not in the table caption or a table footnote; add a footnote so the table is self-contained.","section":"Table 2"},{"comment":"The right panel shows realized token fractions for OEO and GEPA, but the caption does not explain which bar color corresponds to which method; please add that information to the caption or a legend.","section":"Figure 2"},{"comment":"In the revision churn definition, if the sum of stepwise edits is zero, the denominator max(...,1) forces churn to be 1−0 = 1, which may be unintuitive; consider defining churn as 0 when no committed revisions occur.","section":"Appendix A.4, Eq. (2)"},{"comment":"The statement 'All GPT-5.5 calls use the GitHub Copilot Responses API' is unusual and may confuse readers; please clarify whether this is a standard API or a specific deployment detail, and if it is relevant to reproducibility, release the exact endpoint and version.","section":"Footnote 1"}],"recommendation":"major_revision","confidential_remarks":"The conceptual framework is a genuinely useful contribution to the self-evolving agent literature, and the capability-boundary control is a good idea. The paper's main weakness is that the empirical foundation is too thin for the strength of the claimed conclusion: two baselines, single runs, no released code or data, and margins that can flip within a few items. I think this can be fixed within the manuscript's scope by adding multiple seeds, releasing artifacts, verifying baseline configurations, and softening the categorical language. I would not reject it, but I would not accept it in its current form. Please weigh the single-run issue against the fact that the qualitative pattern (OEO improves all settings, uses less budget, and loses only at lower capability) is likely robust even if individual win/loss labels are not."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious systems question, clearly framed, and the paper does a real service by separating the externally fixed optimization contract from the task-specific meta-policy. That distinction is the main intellectual takeaway and it will survive even if every empirical result gets re-run.\n\nWhat's new: they run a frontier model (GPT-5.5) as an open-ended optimizer that picks its own evidence, edits, and stopping rule, and compare it against two stylized prescribed pipelines, SkillOpt (staged, bounded) and GEPA (evolutionary). They add a zero-interaction rewrite control, a capability ladder (weak/medium/frontier optimizers), and a trajectory analysis showing that different paths can converge to overlapping item-level behavior. Those are genuinely useful measurements, and the capability-boundary result—prescription helps when the optimizer is weaker—is the most credible part of the paper.\n\nWhere it's soft: the empirical basis for \"we don't need prescribed pipelines\" is a 14-cell comparison with a single run per cell, no standard errors, no released code, prompts, or seeds, and a proprietary API. Two cells swing on 0.07 and 0.21 percentage points, so the win/tie/loss record is not stable under reseeding. The two baselines are reasonable representatives of their families, but they are still just two systems, and there's no evidence they were tuned to their official recommended configurations. The paper is honest about the single-run protocol in Appendix A.2, but that honesty doesn't repair the fragility. The trajectory analysis is only for OEO vs SkillOpt, not GEPA, so the process-level claims are narrower than the title implies.\n\nBottom line: the question matters and the framework is solid. The empirical answer is plausible but not established. This deserves a serious referee because the field needs exactly this kind of clean conceptual decomposition, but the referee should push for multi-seed runs, released artifacts, and at least one or two additional prescribed baselines before the categorical claim is accepted.\n\nFor a reading group, I'd bring it: the contract/meta-policy split alone is worth an hour of discussion, and the path-vs-behavior analysis is a nice prompt for thinking about how to evaluate self-evolving agents. I'd cite it for the conceptual framework, not for the empirical numbers.","headline":"The contract/meta-policy split is the real contribution; the 12-1-1 headline is a single-run, two-baseline comparison that should be read as suggestive, not definitive.","tokens_in":11708,"tokens_out":2555,"would_cite":true,"duration_ms":21610,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"With a frontier model as optimizer, a self-evolving agent can compose its own improvement procedure online and stay competitive with—and often ahead of—hand-designed pipelines.","keywords":["self-evolving agents","open-ended optimization","optimization meta-policy","prescribed optimization pipelines","skill optimization","frontier language models","agentic self-improvement","capability-adaptive design"],"falsifier":"Hold the optimizer fixed and run OEO against a well-tuned prescribed pipeline on a new benchmark family with several settings; if the prescribed pipeline wins most comparisons within the same target-token budget, the claim that prescription is unnecessary for frontier optimizers fails.","tokens_in":10711,"feed_emoji":"🤖","tokens_out":7941,"duration_ms":63082,"temperature":0.7,"pith_summary":"This paper asks whether a frontier model, acting as the optimizer of a self-evolving agent, still needs a framework-prescribed procedure for how to improve. The authors introduce Open-Ended Optimization (OEO), which keeps the objective, allowed interactions, resource budget, data boundary, and evaluation fixed but lets the optimizer decide which evidence to inspect, what to revise, and when to stop. Across 14 head-to-head comparisons over 8 benchmark–target settings, OEO wins 12, ties 1, and loses 1 by 0.21 percentage points, while using a median of 34.3 percent of the SkillOpt reference target-token budget. A matched single-rewrite control does not explain the gains, and the pattern reverses at medium and weak optimizer capability. The paper's conclusion is that prescribed pipelines are capability-dependent scaffolding: the external optimization contract remains necessary, but a sufficiently capable optimizer can compose the route from feedback to persistent improvement itself.","feed_headline":"Frontier optimizers beat prescribed pipelines in self-evolving agents","feed_subtitle":"Open-ended optimization won 12, tied 1, lost 1 in 14 tests, using 34% of the reference target budget.","key_machinery":"The central object is the distinction between the optimization contract and the optimization meta-policy. The optimization contract fixes the objective, permitted interactions, resource budget, data boundary, and frozen evaluation; the optimization meta-policy is the task-specific sequence of evidence gathering, revision, selection, and stopping. OEO is the protocol that keeps the contract and a generic contract-enforcing action interface while leaving the meta-policy to the optimizer, so the only difference from the two baselines is who decides how improvement proceeds.","core_discovery":"The central claim is that, under the same external optimization contract, a sufficiently capable optimizer does not need a framework-supplied task-specific improvement procedure to achieve competitive self-evolution. OEO keeps the objective, allowed operations, budget, data boundary, and evaluator fixed, and delegates only the meta-policy—which evidence to inspect, which revision to try, when to stop. With GPT-5.5 as the optimizer, OEO improves every initial skill and beats SkillOpt in 7 of 8 settings and GEPA in 5 of 6, with the only loss by 0.21 percentage points. A one-shot, zero-interaction rewrite does not reproduce the gains, so interaction matters. At medium optimizer capability the result flips and SkillOpt wins, and a weak optimizer cannot act through the unchanged OEO interface; thus delegation is bounded by capability. The paper's positive claim is that prescription is an optional inductive bias or scaffold for weaker optimizers, not a prerequisite for frontier-model self-evolution.","pith_inferences":["Editorial inference: the efficiency result makes open-ended delegation especially attractive when target-model interaction is expensive; the paper does not perform a cost-benefit analysis, but the token savings point in that direction.","Editorial inference: a natural extension is to map capability to optimal scaffolding depth, since the crossover between OEO and SkillOpt between medium and frontier capability suggests a gradient that the paper does not measure.","Editorial inference: the path-versus-behavior finding suggests that self-evolution evaluations should routinely report trajectory-level metrics alongside item-level correctness, because final scores alone can miss large differences in how agents learn."],"forward_implications":["Hand-crafted task-specific pipelines become an optional inductive bias rather than a prerequisite once the optimizer is sufficiently capable.","Framework governance remains external even when the route to improvement is delegated: objectives, permissions, budgets, and evaluation boundaries stay fixed.","Designers can follow a capability-adaptive rule: delegate open-ended composition to strong optimizers, and reintroduce structured scaffolding for medium and weak optimizers.","Optimization studies should report both the committed trajectory and the selected skill's item-level behavior, because different routes can reach overlapping correct sets.","Delegation does not automatically raise resource use; in these runs OEO used about a third of the reference target-interaction token budget."],"supporting_citations":[{"why":"Defines SkillOpt, the staged bounded-edit prescribed pipeline used as the first baseline.","marker":"[16,17]"},{"why":"Defines GEPA, the reflective evolutionary search used as the second baseline.","marker":"[9]"},{"why":"Describes GPT-5.5, the frontier optimizer that drives all three methods in the main comparison.","marker":"[19]"},{"why":"Describes the Qwen3.5 target models that execute tasks in the benchmark settings.","marker":"[20]"},{"why":"Establishes that large language models can act as optimizers, grounding the delegation premise.","marker":"[4]"},{"why":"Supplies the SearchQA benchmark and evaluation split used in the comparison.","marker":"[21]"},{"why":"Supplies the SpreadsheetBench benchmark used in the comparison.","marker":"[22]"},{"why":"Supplies the OfficeQA benchmark used in the comparison.","marker":"[23]"},{"why":"Supplies the LiveMathematicianBench benchmark used in the comparison.","marker":"[25]"}],"fun_headline_variants":["Open-ended optimization beats prescribed pipelines in self-evolving agents","Self-evolving agents: open-ended optimizer wins 12 of 14 tests","Frontier agent writes its own optimization, beating prescribed pipelines","For frontier models, prescribed optimization pipelines are optional scaffolding","Capability matters: open-ended optimization only works with strong models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on SkillOpt and GEPA being faithful, representative examples of prescribed pipelines and on the 14 tested benchmark–target settings being enough to generalise about prescription.","fun_headline_variants_meta":{"raw":{"variants":["Open-ended optimization beats prescribed pipelines in self-evolving agents","Self-evolving agents: open-ended optimizer wins 12 of 14 tests","Frontier agent writes its own optimization, beating prescribed pipelines","For frontier models, prescribed optimization pipelines are optional scaffolding","Capability matters: open-ended optimization only works with strong models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001648,"raw_usage":{"total_tokens":6585,"prompt_tokens":1024,"completion_tokens":5561,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":5478}},"tokens_in":640,"tokens_out":5561,"duration_ms":32307,"temperature":1.0,"reasoning_tokens":5478,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:48:49.320928+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold the optimizer fixed and run OEO against a well-tuned prescribed pipeline on a new benchmark family with several settings; if the prescribed pipeline wins most comparisons within the same target-token budget, the claim that prescription is unnecessary for frontier optimizers fails.","supporting_citations":[{"cited_title":"Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl- Ong, Arnav Singhvi, Herumb Shandilya, Michael J","cited_arxiv_id":null,"evidence_quote":"Defines GEPA, the reflective evolutionary search used as the second baseline."},{"cited_title":"GPT-5.5 System Card","cited_arxiv_id":null,"evidence_quote":"Describes GPT-5.5, the frontier optimizer that drives all three methods in the main comparison."},{"cited_title":"Qwen3.5: Towards Native Multimodal Agents","cited_arxiv_id":null,"evidence_quote":"Describes the Qwen3.5 target models that execute tasks in the benchmark settings."},{"cited_title":"Le, Denny Zhou, and Xinyun Chen","cited_arxiv_id":null,"evidence_quote":"Establishes that large language models can act as optimizers, grounding the delegation premise."}],"review_version":1}