{"id":"9bf94f4f-23b4-4243-9394-81be470f6fe5","arxiv_id":"2608.10694","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Cheap-tier fitness evaluation plus a strong variation operator produces prompts that deploy upward across LLM tiers, matching or beating same-tier optimization at a fraction of the search cost.","lead":"This paper shows that evolutionary prompt optimization can run on the cheapest available LLM for nearly all of its workload, with a strong model used only for the rare edit step, and the resulting prompt then works as well or better on a more expensive model. This cuts search cost by 5.6 to 14x, and up to 25 to 54x on models that emit long chains of thought, while keeping deployment quality at parity.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cheap evaluator's rank correlation with the deployment tier is never measured, leaving the surrogate-validity assumption behind the central claim untested.","rationale":"The paper is methodologically strong, with honest limitations and extensive ablations. The reader's weakest assumption is the cheap evaluator's competence; I sharpen this to the unmeasured rank correlation between cheap and target validation scores. The end-to-end target regret bundles surrogate validity with everything else, so a favorable R_{s→t} does not by itself demonstrate that the cheap fitness signal is aligned with the deployment model. The role ablation locates the gain in the reflector, but it does not establish that the evaluator's ranking is useful; a useless evaluator would still allow the reflector to propose good prompts, making the method's success dependent on the reflector rather than on the cheap tier. Since all four tasks have moderate cheap-tier seed scores, the failure boundary is untested. This does not overturn the reported results, but it strengthens the case for conditional acceptance with a request for the surrogate-validity measurement.","tokens_in":30427,"tokens_out":8157,"duration_ms":84834,"concrete_test":"Compute, over the existing candidate pools of the 48 (task, search arm, deploy tier) cells, the Spearman rank correlation between the cheapest answerer's validation scores and the deployment tier's validation scores for all candidates scored during search. If the mean correlation is below roughly 0.3 on any task, the cheap fitness signal is not adequately steering selection toward target-optimal prompts, and the paper's central claim would need to be re-attributed to the reflector rather than to cheap-tier surrogate validity. This directly tests the load-bearing assumption of Eq. (2) and costs no new inference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that cheap-tier search substitutes for target-tier search, rests on Eq. (2)'s decomposition Jdep = Jtask + Δ(Π) being favorable for the prompts the search finds. This requires the cheap answerer's fitness signal to rank candidate prompts similarly to the deployment model. The paper measures the end-to-end target regret R_{s→t} (Eq. 3), but never measures the surrogate quality itself: no rank correlation, agreement, or calibration between cheap-tier and target-tier validation scores over the candidate pool is reported. The role ablation (Table 2) shows the reflector is the lever of transfer, but it does not show the evaluator's signal is aligned with the target; if the cheap model's ranking is largely uncorrelated with the target's, selection pressure is misdirected and observed success must be attributed to the reflector's independent proposal quality rather than to the cheap evaluator steering the search. Section 9 explicitly flags near-zero cheap-tier competence as a failure mode, yet no experiment probes this boundary: all four tasks have seed scores on the cheapest tier between 31% and 41% (Appendix F), so the method is only demonstrated in a moderate-competence regime. The claimed 'characterization of where it fails' is thus incomplete, and the substitution claim is conditional on an unmeasured surrogate-validity assumption that could break on tasks with lower cheap-tier competence or weaker cross-model ranking agreement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes decoupling the three roles an LLM plays inside evolutionary prompt optimization: a cheap answering model scores fitness, a strong model performs reflective variation, and the evolved prompt is deployed zero-shot on a stronger target tier. The authors instantiate this on GEPA, test it on four tasks and eleven models across four families, and report that the cheap search matches or beats full same-tier optimization in 36 of 48 deployments at 5.6–14x lower search cost (25–54x on Gemini), including a zero-API-cost self-hosted Qwen answerer. They also report a 2x2 role ablation locating the transfer effect in the variation operator, an MIPROv2 ablation, a break-even price analysis, and an explicitness analysis of the evolved prompts.","tokens_in":30669,"tokens_out":5897,"duration_ms":67556,"significance":"If the empirical claims hold, this is a practically important result: it separates search cost from deployment-tier cost and suggests that prompt optimization for large models can be done mostly on small, cheap models without losing deployment quality. The experimental design is unusually careful for this literature: fixed per-task budgets, validation-based model selection with held-out test scoring, n=3 seeds, seed-prompt controls, a weak-reflector arm, a 2x2 role ablation, an optimizer ablation, a neutral cross-family deploy target, and exact per-call token-log accounting. The central empirical claim is not circular: target regret is computed from measured held-out scores, and costs are measured rather than derived from the claim. The main gap is that the surrogate-validity assumption behind cheap-tier search is never directly measured, and the claimed 'characterization of where it fails' does not include the low-competence failure regime the authors themselves identify in Section 9.","major_comments":[{"comment":"The central substitution claim rests on the untested assumption that the cheap evaluator's fitness signal ranks candidate prompts similarly to the deployment model. Eq. (2) decomposes Jdep = Jtask + Delta(Pi), but the paper never measures rank agreement, correlation, or calibration between Jtask and Jdep over the candidate pool; it only measures the end-to-end target regret R_{s->t}. If the cheap tier's ranking is largely uncorrelated with the target tier's, selection pressure is misdirected and the observed deployment parity could be driven mainly by the strong reflector's independent proposal quality. I ask for a direct surrogate-validity analysis: for example, Kendall tau or rank overlap between cheap-tier and target-tier validation scores over the candidates generated in a sample of runs, reported per task and per tier pair. Without this, the claimed mechanism and the stated boundary conditions are not fully supported.","section":"Section 3(A), Section 2, Eq. (2)"},{"comment":"The paper claims a 'cost-controlled characterization of when cheap-tier search substitutes for target-tier search, and where it fails,' but the failure boundary is not actually characterized. Section 9 states that a cheap evaluator scoring near zero on a task would flatten the fitness landscape and stagnate the search, yet no experiment probes this regime: on all four tasks, the cheapest-tier seed scores are between 31% and 41% (Appendix F), so all demonstrations are in a moderate-competence regime. The only observed failure mode is the prompt-insensitive ceiling (LiveBench-Math), which is a limitation of prompt optimization generally rather than of cheap-tier transfer. The authors should either add a deliberately degraded or near-zero-competence cheap evaluator (e.g., a random-fitness control or a model with near-zero task competence) or soften the 'characterization of where it fails' claim to what the experiments actually support.","section":"Section 9 and Appendix F"},{"comment":"The pooled positive-transfer claim is more fragile than the headline 'mean residual +2.8%, 95% CI above zero' suggests. Appendix G notes that a variance-weighted mean sits slightly below zero, which indicates that the positive pooled mean is driven by large-margin points and is not a robust central tendency. In addition, the 48 residuals are not independent: they come from 12 setups, each contributing four dataset residuals, and the t-intervals do not appear to account for clustering by setup or by shared search arms. The paper should report a cluster-robust or setup-level analysis and should state clearly whether the strong claim (zero-shot positive transfer, R_{s->t}<0) is supported after this reanalysis or whether only the weak claim (R_{s->t} near zero) is.","section":"Appendix G, Figure 3"}],"minor_comments":[{"comment":"Figure 1 states that the cheap search is 'never worse than 3.8 points,' while Figure 4 and Appendix G report the largest shortfall as 6.12% of the full-cost score. The two statements are in different units and may be consistent, but the paper should clarify whether the 3.8 figure is in task-metric points and why it differs from the 6.12% figure.","section":"Figure 1 and Appendix G"},{"comment":"The sentence 'Raw totals are $16.39, $2.39, $2.10 and $30.74 in row order' does not match the row order of Table 5: the normalized search-cost column $1.95, $3.11, $22.17, $28.52 does not follow from the stated raw totals with the stated coverage fractions in the order given. The cost-basis paragraph should list the raw totals in the same order as the table rows or label each total explicitly.","section":"Appendix C, Table 5"},{"comment":"The abstract reports '5.6–14x lower search cost' as the headline range, but Figure 1 includes points up to 114x and the Qwen zero-API-cost rows in Table 11 reach 63–114x. The paper should state in the abstract or in the caption that the larger multipliers come from the zero-API-cost self-hosted answerer, so readers do not infer an inconsistency.","section":"Abstract and Section 5"},{"comment":"The paper reports that cheap arms end below full-cost arms in 29 of 32 validation-curve comparisons, by a median of 7.8 validation points, and then says these curves should not be used to rank configurations. This is a notable observation that is in tension with the idea that the cheap evaluator is a faithful surrogate; it deserves a fuller discussion in the main text or in the surrogate-validity analysis rather than a one-sentence dismissal.","section":"Appendix H"},{"comment":"The term 'variance-weighted mean' is introduced without defining the weights or the estimator variance; please specify the weighting scheme and why it is appropriate given that residuals are in percent-of-full-score units and vary across datasets with different score ranges.","section":"Appendix G"}],"recommendation":"major_revision","confidential_remarks":"The paper is well suited to a machine-learning venue and the empirical core is strong. My main concern is that the load-bearing surrogate-validity assumption and the claimed failure characterization are not directly tested, and the positive-transfer statistic is less robust than the headline suggests. These are fixable within the scope of the manuscript by adding a rank-agreement analysis and a low-competence or randomized-evaluator control. I do not see a reason to reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you the short version. The paper's central result is real and well-supported: on four tasks, eleven models, and four families, searching a prompt with the cheapest answerer and a strong reflector, then deploying one or two tiers up, matches or beats that tier's own full-price optimization at 5.6–54x lower search cost. That is a genuinely useful empirical result. The previous literature had only single-pair observations of weak-to-strong prompt transfer; this is the first systematic cost-aware characterization with a role ablation, a cross-optimizer check with MIPROv2, a seed-prompt control, and a break-even price analysis. The ablation is clean: the reflector carries the transfer, and cheap evaluation is not what produces the gain. The explicitness hypothesis is honestly labeled as a correlation, not an ablation.\n\nThe stress-test note worries that the cheap evaluator's rank correlation with the deployment tier is never measured. I don't think that breaks the central claim, because the target regret is measured end-to-end and that is the quantity that matters. The rank correlation is a mechanism-level assumption, and the paper never claims to have validated it. What is fair is that the 'where it fails' characterization is incomplete: all four tasks have seed scores 31–41% on the cheapest tier, so the method is only demonstrated in a moderate-competence regime. Section 9 flags the near-zero boundary but does not probe it. That is a scope limitation, not a fatal flaw.\n\nThe more practical soft spot is no code or data release; the appendix says nothing about artifacts, and reproducing this from tables alone would be painful. Also, as the paper itself honestly reports, the variance-weighted pooled residual sits near zero, so the aggregate claim is 'no worse, usually a bit better' rather than a robust positive edge. A few high-variance cells and the reflector-as-deployment-target confound in one family are real but partially controlled by the neutral deploy target.\n\nWho this is for: anyone doing budget-constrained prompt or agent optimization with tiered APIs. It deserves a serious referee and, with code and data attached, likely acceptance. The empirical design is unusually careful for this kind of work, and the paper is transparent about its own limits. I'd send it to review.","headline":"A careful, honest empirical paper that makes a strong case that cheap-tier search with a strong reflector can substitute for target-tier search at 5.6–54x lower cost; deserves peer review and should be held to artifact release.","tokens_in":31207,"tokens_out":2043,"would_cite":true,"duration_ms":22510,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A prompt evolved by searching on the cheapest answerer with a strong reflector, then deployed zero-shot on a stronger model, matches or beats the prompt that target tier optimized for itself at full price, at 5.6–14× lower search cost.","keywords":["evolutionary prompt optimization","cost-aware transfer","cross-tier deployment","fitness evaluation bottleneck","reflective mutation","zero-shot transfer","LLM tiers","surrogate optimization"],"falsifier":"Run the method on a task where the cheap model's score on the seed prompt is at floor and uncorrelated with the target model's scores across a pool of random prompts (Spearman $\\rho \\approx 0$). If the cheaply searched prompt still matches the target's own full-cost optimization, the central claim is wrong; the paper's own boundary condition predicts search would stagnate and target regret would grow large.","tokens_in":30226,"feed_emoji":"🧬","tokens_out":5108,"duration_ms":45886,"temperature":0.7,"pith_summary":"The paper claims that the dominant cost of evolutionary prompt optimization—running fitness evaluations on the target model—can be moved to the cheapest available model without losing quality at deployment. It decouples three roles an LLM plays inside the search loop: the high-volume answering/evaluation role runs on a cheap tier, a strong model handles the rare reflection/variation step, and the evolved prompt is deployed zero-shot on a stronger tier. Across four tasks and eleven models in four model families, the cheaply searched prompt matches or beats each tier's own full-cost optimization in 36 of 48 deployments, at 5.6–14× lower search cost (25–54× when reasoning tiers emit long chains of thought), and even a self-hosted 8B answerer at near-zero API cost matches paid tiers' own optimization. The paper's central claim is that cheap-tier search substitutes for target-tier search, and often benefits from it, because fitness evaluation only needs to rank candidates, not to measure deployed quality.","feed_headline":"Cheap search, strong deploy: prompts win at 5-14x less cost","feed_subtitle":"Evolving a prompt on the cheapest model, then deploying upward, matches or beats full-price same-tier optimization in 36 of 48 tests.","key_machinery":"The central mechanism is role decoupling with a cost asymmetry: three LLM roles—answering/evaluating, reflecting/varying, and deploying—are assigned to separate model tiers. Fitness evaluation runs on the cheapest model (high volume, needs only to rank), variation runs on a strong model (rare, needs precision), and deployment happens zero-shot on any stronger tier. The load-bearing cost identity is $C_{\\mathrm{opt}} \\approx (K N_{\\mathrm{val}} + 2Ab)\\, c(M_{\\mathrm{task}}) + A\\, c(M_{\\mathrm{refl}})$, which makes the answering tier the dominant term; the transfer quantities are the per-prompt residual $\\Delta(\\Pi)$ and the end-to-end target regret $R_{s\\to t}$, whose negated scale-free form $\\delta^\\%._{s\\to t}$ is pooled across runs. The method is a drop-in modification to reflective evolutionary optimizers such as GEPA.","core_discovery":"Searching for a prompt need not run on the model that will serve it. The paper restructures evolutionary prompt optimization so that fitness is evaluated by the cheapest available answering model, edits are proposed by a strong reflector, and the final prompt is deployed unchanged on a stronger tier. Formally, it optimizes the surrogate objective $J_{\\mathrm{task}}(\\Pi)$ while grading on $J_{\\mathrm{dep}}(\\Pi)$, and the cross-tier transfer residual $\\Delta(\\Pi) = J_{\\mathrm{dep}} - J_{\\mathrm{task}}$ is shown to be small or negative in practice: the target regret $R_{s\\to t}$ is near zero or favorable in 36 of 48 (task, search arm, deploy tier) cells, with pooled mean residual $+2.8\\%$ of the full-cost score. The saving follows from a cost identity: over 96% of search tokens go to answering, so replacing that single role with the cheapest tier moves almost the whole bill, while the strong reflector remains a bounded premium (under 5% of calls, median 27% of spend). The paper also locates the source of positive transfer in the variation operator, not cheap evaluation, and proposes that cheap evaluators force explicit, spelled-out prompts that stronger models then exploit.","pith_inferences":["The method suggests an adaptive fidelity schedule: instead of always using the cheapest tier, one could choose the evaluator tier per generation based on measured rank correlation with the target, spending more only when the cheap signal is weak.","The explicitness account is testable as a causal hypothesis: rewriting a same-tier-optimized prompt to the cheap prompt's explicitness level (or vice versa) and measuring transfer would separate the weak-evaluator-forcing effect from the strong-reflector-writing effect, which the paper notes it cannot.","A practical predictor of success would be the rank correlation between cheap and deployment model scores on a small fixed pool of candidate prompts before running the full search; near-zero correlation would predict failure and save the search cost.","The break-even volume $N^\\star$ means the method is a search-cost saving with a serving-cost caveat: for very high-volume deployments, the longer cheap-evolved prompt may erode the one-time saving, so the method is best for moderate-volume deployments or where the cheap prompt is not longer."],"forward_implications":["One cheap search produces a portable prompt deployable on any stronger tier with no mapping or re-optimization, amortizing search cost across deployment tiers.","Practitioners can cut prompt-search cost by an order of magnitude (5.6–14×, up to 25–54× on reasoning-heavy ladders) without sacrificing deployed accuracy.","The split is cost-robust: in 15 of 24 cells the cheap configuration stays cheaper even at price parity because it emits fewer output tokens, and break-even price ratios $\\lambda^\\star$ run 0.50–3.46.","A zero-API-cost local answerer (Qwen3-8B) with a paid reflector keeps matching paid tiers' own optimization for under $2 per run, so search can be run with almost no API spend.","The benefit appears precisely where prompt optimization has headroom; on near-saturated tiers the cheap recipe still matches full-cost optimization rather than beating it."],"supporting_citations":[{"why":"Supplies the GEPA formulation, reflective mutation operator, and the single-pair weak-to-strong observation that the paper generalizes into a cost strategy.","marker":"(Agrawal et al., 2026)"},{"why":"Provides prior evidence of optimizing on a weak model and deploying on a stronger one, plus the prompt-insensitive ceiling boundary condition.","marker":"(Gao et al., 2026)"},{"why":"Documents lateral 'model drift' degradation that the paper contrasts with reliable upward transfer.","marker":"(Wang et al., 2025b)"},{"why":"CAPO, a cost-aware prompt optimizer that cheapens scoring within the target tier; the paper's tier-switching approach is positioned against it.","marker":"(Zehle et al., 2025)"},{"why":"PMPO, which cheapens scoring while staying on the target model; used as a comparative baseline for cost reduction strategies.","marker":"(Zhao et al., 2025)"},{"why":"MIPROv2, the structurally different optimizer used in the ablation that shows the effect belongs to the cost-aware setup rather than one mutation operator.","marker":"(Opsahl-Ong et al., 2024)"}],"fun_headline_variants":["Cheap tier evaluates, strong tier deploys, 5-14x lower cost","Upward transfer saves 5-14x in prompt optimization","Evolve prompts on cheap models, deploy on strong ones","36 of 48 tests: cheap search beats same-tier optimization","Cost-aware cross-tier transfer: 5-14x cheaper search"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The cheap evaluator must be able to rank prompt candidates roughly as the deployment model does; if it scores near zero on a task, the fitness landscape is flat and search stagnates.","fun_headline_variants_meta":{"raw":{"variants":["Cheap tier evaluates, strong tier deploys, 5-14x lower cost","Upward transfer saves 5-14x in prompt optimization","Evolve prompts on cheap models, deploy on strong ones","36 of 48 tests: cheap search beats same-tier optimization","Cost-aware cross-tier transfer: 5-14x cheaper search"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000362,"raw_usage":{"total_tokens":1981,"prompt_tokens":1003,"completion_tokens":978,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":886}},"tokens_in":619,"tokens_out":978,"duration_ms":10426,"temperature":1.0,"reasoning_tokens":886,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:14:30.104504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the method on a task where the cheap model's score on the seed prompt is at floor and uncorrelated with the target model's scores across a pool of random prompts (Spearman $\\rho \\approx 0$). If the cheaply searched prompt still matches the target's own full-cost optimization, the central claim is wrong; the paper's own boundary condition predicts search would stagnate and target regret would grow large.","supporting_citations":[],"review_version":1}