{"id":"df44cab4-5a0a-4529-939d-0c182425ea81","arxiv_id":"2607.28568","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Execution-grounded SFT/RL on Draft/Improve/Debug/Crossover plus experience-guided search turns a 35B base model into a competitive open MLE agent with partial transfer to scientific AutoResearch tasks.","lead":"OpenMLE trains a 35B model on four reusable code-evolution operators and runs them in long-horizon search, lifting MLE-Bench Lite medal rate from 39% to 71% under a fixed 12-hour GPU budget. It is a concrete, open stack for studying AI that improves machine-learning pipelines, not a completed recursive self-improver.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The headline deltas rest on 22- and 10-task Lite sets with high run variance; that sample size is the softest load-bearing support for the claimed training and search gains.","rationale":"The reader correctly treats this as a strong open systems/MLE-agent paper whose soundness is capped by Lite-sample fragility, variance, and partial release, and scopes RSI as aspiration (Sec. 8). I agree that the weakest assumption is that these small held-out aggregates under a fixed sandbox budget suffice to underwrite the headline training and search gains (and, secondarily, that decontamination plus partial Gym release fully clear leakage/repro concerns). Stress-testing does not surface a deeper internal contradiction: operator-aligned SFT/RL, matched harness ablations, AIRA-Evo controls, and execution-based scoring are coherent; circularity is limited. The single most load-bearing concern is statistical fragility of the strongest_claim’s effect sizes and dual attribution on N=22/N=10 with SDs of ~8 pp—not a reason to reject, but the right gate for CONDITIONAL. Independent multi-seed confirmation of the 35B delta and leave-one-out stability on NatureBench Lite would settle it; failure would justify downgrading confidence or demanding full-bench numbers. No verdict change: CONDITIONAL remains appropriate.","tokens_in":50524,"tokens_out":792,"duration_ms":15135,"concrete_test":"Recompute Table 1 / D.1 Medal Average for Qwen3.6-35B-A3B vs Frontis-MA1-35B under identical OpenMLE-Evo with ≥5 independent seeds (or bootstrap over the 22 tasks×3 epochs). Require the mean delta to stay ≥15 pp with non-overlapping mean±1 SD bands (or a paired task-level test p<0.05). Separately, leave-one-out over NatureBench Lite’s 10 tasks: if model-swap Match-SOTA (70% vs 50%) or harness-swap (50% vs 20%) collapses when any 2 tasks are held out, treat dual-transfer attribution as unstable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that execution-grounded post-training and OpenMLE-Evo(-Max) produce large, attributable gains (39.39%→60.61%→71.21% Medal Average on MLE-Bench Lite; NatureBench Match-SOTA 50%→70% model swap and 20%→50% harness swap). Those numbers are means over only 22 and 10 tasks, three epochs where reported. Appendix D.1 shows Frontis-MA1-35B OpenMLE-Evo at 60.61%±7.73% and Evo-Max at 71.21%±8.57% Medal Average—SDs large enough that epoch-level intervals substantially overlap neighboring systems and weaken precise “exceeds GPT-5.5 + Codex” ordering. NatureBench Lite moves 10 pp per task, so 7/10 vs 5/10 is a two-task swing. Sec. 8 and 6.6 already flag small modality groups and the ten-task set; the load-bearing risk is not leakage rhetoric but that the strongest_claim’s magnitude and dual attribution are not yet stable under this N. Decontamination and partial data release (1,415/5,758 full packages) matter for reproduction but are secondary to whether the Lite aggregates pin the effect.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces OpenMLE, an open full-stack system for executable AI4AI research in machine learning engineering, comprising OpenMLE-Gym (5,758 quality-gated tasks with sandbox feedback), OpenMLE-ERL (execution-grounded SFT+RL on four atomic operators: Draft, Improve, Debug, Crossover), and OpenMLE-Evo (experience-guided long-horizon search). On this stack the authors post-train Frontis-MA1-35B (and a 30B companion) as a meta-evolution agent. Under a fixed 12h/RTX 4090 (12 GB) budget on MLE-Bench Lite, they report Medal Average rising from 39.39% (Qwen3.6-35B-A3B base) to 60.61% with OpenMLE-Evo and 71.21% with OpenMLE-Evo-Max, with matched harness comparisons against Claude Code/Codex and original AIRA-Evo, plus factorial transfer on NatureBench Lite (model swap 50%→70% Match-SOTA; harness swap 20%→50%). Weights and the full stack are released.","tokens_in":50927,"tokens_out":1502,"duration_ms":34693,"significance":"If the attributed gains hold under broader evaluation, this is a substantial contribution: a reproducible coupling of operator-level post-training with the same operators used in evolutionary test-time search, with open environments, training code, harness, and checkpoints—rare among public MLE-agent systems (Appendix Table 11). Strengths include matched model–harness factorials, dual-backbone replication (35B/30B), three-run means with standard deviations (App. D.1), mechanism case studies of Crossover/Improve and three-factor selection (§6.3–6.5), and explicit limitations on RSI (§8). The work is a concrete step toward studying meta-evolution in executable MLE rather than a demonstration of general RSI.","major_comments":[{"comment":"Table 1 and App. D.1: Headline Medal Average for Frontis-MA1-35B is 60.61%±7.73% (Evo) and 71.21%±8.57% (Evo-Max) over only 22 tasks × 3 runs. These SDs are large enough that epoch-level intervals substantially overlap neighboring systems (e.g., GPT-5.5+Codex at 68.18% point estimate). The claim that Evo-Max “exceeds GPT-5.5 + Codex” (§6.2, abstract, Fig. 1) is not statistically supported at this N. Please report pairwise differences with uncertainty (bootstrap or permutation over tasks), avoid strict ordering language against single-run external harnesses, and state the Lite sample-size limitation next to every headline comparison.","section":"§6.2, Table 1, Appendix D.1"},{"comment":"Table 2 / §6.6: NatureBench Lite has 10 tasks, so each task moves All M/All S by 10 pp. The model swap (5/10→7/10) and harness swap (2/10→5/10) are directionally informative factorial controls, but two- and three-task swings cannot pin “both components transfer” at the precision used in the abstract. Either expand the transfer set, report exact task-level outcomes with sensitivity (leave-one-task-out), or downgrade abstract/conclusion wording to “initial evidence on a 10-task subset,” consistent with the caveats already in §6.6 and §8.","section":"§6.6, Table 2"},{"comment":"§3.2–3.3 and §6.1: Training data are Kaggle-centric and “deduplicated against all evaluation benchmarks,” with MLE-Bench-overlapping competitions excluded from Gym construction. Given partial package release (1,415/5,758 full packages) and shared Kaggle distributional structure, the decontamination protocol should be specified operationally (slug lists, near-duplicate criteria, any embedding/n-gram checks, and whether NatureBench containers were in any teacher-rollout or prior pool). Without this, the mild train–eval distributional overlap risk remains a load-bearing reproducibility concern for the claimed generalization.","section":"§3.2–3.3, §6.1"},{"comment":"Abstract and §1 frame the work as progress “towards Recursive Self-Improvement,” while §8 correctly states OpenMLE does not realize RSI and evolution of the evolutionary system itself is future work. The gap between title/abstract RSI rhetoric and the actual contribution (meta-evolution of MLE program operators under fixed harness and objectives) should be tightened so central claims match the evaluated object: budgeted MLE search quality, not autonomous self-upgrade of the improver.","section":"Abstract, §1, §8"}],"minor_comments":[{"comment":"Abstract names the learning layer “OpenMLE-RL” while the body uses “OpenMLE-ERL”; unify the acronym.","section":"Abstract vs §4"},{"comment":"Figure 1 and leaderboard-style bars mix single-run general-agent results with three-run OpenMLE means; mark run counts and variance in the figure legend.","section":"Figure 1"},{"comment":"Eq. (2) omits max-centering used in the implementation (App. B.5); a forward reference would avoid mismatch for readers implementing from the main text.","section":"§4.3, Eq. (2)"},{"comment":"Parent-selection weights (λ_s, λ_Δ, λ_n, τ) and RL operator sampling probabilities are free parameters; a short sensitivity note or default table in the main text would help.","section":"§5.2, Eq. (4); App. B.3"},{"comment":"Typographical inconsistencies (e.g., “onthe-icml-2013-whale” spacing in §6.5; “OpenMLE-ERL” vs figure labels) should be cleaned in copy-edit.","section":"§6.5"}],"recommendation":"major_revision","confidential_remarks":"Solid systems paper with unusually complete release intentions; the main editorial risk is over-claiming from Lite aggregates and RSI framing rather than methodological error. If the authors add proper uncertainty on the 22-task comparisons, temper external model rankings, document decontamination, and align abstract language with §8, this could clear a top venue bar. Fit is appropriate for a methods/systems track in CL/AI; less so if the journal expects theoretical RSI results."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is not the RSI framing. It is a working open stack that trains the same Draft/Improve/Debug/Crossover operators used at search time, ships a 35B (and 30B) checkpoint, and shows matched gains under a fixed 12h/4090 budget.\n\nWhat is actually new is the alignment: execution-grounded SFT and RL on those four operators, then composition in OpenMLE-Evo with structured experience cards, three-factor parent selection, and operator-conditioned memory. Pieces exist (AIRA/AIDE, MLE-Dojo/Smith, various MLE RL papers). The joint loop plus an unusually complete release surface—Gym construction, sandbox, training, harness, weights—is the contribution. Controls are better than average: same-harness base vs post-trained on 35B and 30B, harness swaps on several frontier models, AIRA-Evo vs OpenMLE-Evo on the same checkpoint, and factorial NatureBench swaps holding model or harness fixed. Mechanism traces (token/context compression, targeted crossover, three-factor selection) are concrete. Sec. 8 scopes RSI honestly.\n\nSoft spots are real but proportional. MLE-Bench Lite is 22 tasks; NatureBench Lite is 10. App. D.1 gives Frontis-MA1-35B Evo 60.61%±7.73% and Evo-Max 71.21%±8.57% Medal Average—SDs large enough that fine orderings vs GPT-5.5+Codex are soft. NatureBench moves 10 pp per task. Partial data release (1,415/5,758 full packages) and free parameters (λs, entropic β, adaptive bounds) matter for reproduction more than for internal validity. Decontamination is claimed and eval is third-party execution; circularity is low. Math is standard adaptive-reward / entropic-advantage RL, not a proof burden.\n\nThis is for people building coding/MLE agents and open AI4AI infrastructure. Cite the stack and the controlled deltas; treat the leaderboard brag and RSI ladder as packaging. I would send it to peer review and bring it to reading group. Engage, try to reproduce the 35B delta, and keep the claims scoped as the authors mostly do in the limitations.","headline":"Solid open full-stack MLE meta-evolution paper: operator-aligned SFT/RL + search, real controls, rare release; headline medal numbers sit on small Lite sets with high variance.","tokens_in":51649,"tokens_out":576,"would_cite":true,"duration_ms":12546,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A full open stack trains one model on four program-evolution operators so learning and long-horizon evolutionary search share the same improver for machine learning engineering.","keywords":["recursive self-improvement","AI4AI","machine learning engineering","meta-evolution","evolutionary program search","execution-grounded RL","operator learning","test-time search"],"falsifier":"Rerun the matched harness comparisons with the same 12-hour single-GPU budget after stricter held-out task construction: if swapping the post-trained model or the experience-guided search no longer moves Medal Average on MLE-Bench Lite or Match-SOTA on NatureBench Lite, the claimed composition of training and search fails.","tokens_in":51399,"feed_emoji":"🔄","tokens_out":919,"duration_ms":26729,"temperature":0.7,"pith_summary":"This paper argues that recursive self-improvement needs systems that improve how AI is built, and treats machine learning engineering as a concrete, executable testbed for that goal. It introduces an open stack with quality-gated tasks and sandbox feedback, execution-grounded training of four reusable operators—Draft, Improve, Debug, and Crossover—and an experience-guided search harness that composes those same operators over long horizons. The trained 35B meta-evolution agent, under a fixed 12-hour budget on one mid-range GPU, raises medal rate on a standard MLE competition suite from about 39% for its base model to about 61%, and to about 71% with stronger search priors and parallel search. Controlled swaps on a held-out scientific AutoResearch suite attribute gains to both the trained model and the search framework. The authors release weights and the full stack so others can reproduce the learning–evolution loop.","feed_headline":"Trained evolution operators lift MLE medals from 39% to 71%","feed_subtitle":"One model learns Draft, Improve, Debug, and Crossover from sandbox feedback, then runs them in long-horizon search.","key_machinery":"The shared operator interface—Draft, Improve, Debug, Crossover—used both as execution-grounded SFT/RL targets and as the variation engine of long-horizon evolutionary search, closing a meta-evolutionary loop in which the improver itself is trained.","core_discovery":"Aligning post-training and inference on the same four atomic program-evolution operators, trained from sandbox execution scores and then composed in experience-guided search, produces large complementary gains: under identical search, the post-trained 35B model lifts MLE-Bench Lite Medal Average from 39.39% to 60.61% over its base, and with enhanced search reaches 71.21%, while NatureBench Lite controlled swaps raise Match-SOTA from 50% to 70% (model) and from 20% to 50% (framework).","pith_inferences":["If operator-level learning generalizes, the same Draft/Improve/Debug/Crossover interface could be pointed at improving training code or harnesses, not only task solutions.","The efficiency win from bounded, operator-conditioned memory suggests long-horizon agent search may be limited as much by context design as by raw model size.","Sustained late-horizon gains from Crossover and Improve imply evaluation protocols that stop at first valid submission will understate systems built for recombination."],"forward_implications":["Reusable program-edit skills can be learned from sandbox scores and dropped into different evolutionary controllers without retraining full trajectories.","Model post-training and test-time search supply additive gains when they share the same operator vocabulary.","Structured experience cards plus multi-factor parent selection can cut prompt bulk while raising useful new-best updates per token.","An open gym-plus-training-plus-search release makes meta-evolution experiments in executable MLE reproducible end to end."],"fun_headline_variants":["35B meta-agent lifts MLE medals 39% to 71% via four operators","Same Draft-Improve-Debug-Crossover loop trains and searches","Post-trained Frontis-MA1 hits 71% MLE medals, transfers on NatureBench","Execution-grounded operators couple learning and evolution in one loop","Model swap to 70% Match-SOTA; Evo framework swap to 50%"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"That a few dozen competition-style and scientific coding tasks under one fixed sandbox budget, after stated decontamination, are enough to stand in for general ability to improve how AI systems are built.","fun_headline_variants_meta":{"raw":{"variants":["35B meta-agent lifts MLE medals 39% to 71% via four operators","Same Draft-Improve-Debug-Crossover loop trains and searches","Post-trained Frontis-MA1 hits 71% MLE medals, transfers on NatureBench","Execution-grounded operators couple learning and evolution in one loop","Model swap to 70% Match-SOTA; Evo framework swap to 50%"]},"model":"grok-4.5","effort":"low","cost_usd":0.003196,"raw_usage":{"total_tokens":1223,"prompt_tokens":981,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":31964000,"prompt_tokens_details":{"text_tokens":981,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":150,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":981,"tokens_out":92,"duration_ms":4138,"temperature":1.0,"reasoning_tokens":150,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T03:29:25.049626+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the matched harness comparisons with the same 12-hour single-GPU budget after stricter held-out task construction: if swapping the post-trained model or the experience-guided search no longer moves Medal Average on MLE-Bench Lite or Match-SOTA on NatureBench Lite, the claimed composition of training and search fails.","supporting_citations":[],"review_version":1}