{"id":"59cd98f5-d661-4b3e-b93f-f4b098aa9d4a","arxiv_id":"2607.04108","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Parent-conditioned LLM evolution is sampling in disguise; set-level sparse selection over one-shot term dictionaries carries equation discovery.","lead":"LLM “evolution” loops for scientific equations do not beat matched independent sampling; they mainly produce term dictionaries. An external set-level sparse selector over those terms (PTB-Search) reaches 73% Acc0.1 on LLM-SRBench at one-tenth the usual call budget.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-scoped flat-component premise.","rationale":"The manuscript is unusually careful: pre-registered branches, leakage audits, identical-dictionary selector ablations, and an explicit program-domain stress test that delimits rather than overclaims. The strongest claim is empirical and holds under the stated protocol. The only genuine soft spot is the flat-component premise the reader already flagged; the paper itself treats non-additive structure as out of scope via the complexity guard and Principle 3. Manufacturing a further concern would violate the honest-non-finding rule. Therefore the CONDITIONAL / HIGH-confidence verdict stands without adjustment.","tokens_in":23042,"tokens_out":483,"duration_ms":4309,"concrete_test":"Re-run the official 239-problem Llama-3.1-8B PTB-Search pipeline with the generic-primitive dictionary alone (no LLM terms) under the identical frozen selector and report Acc0.1; if the hybrid-vs-generic gap collapses below ~5 points while still beating the 49.2% baseline, the 'LLM as material supplier' framing remains intact but the necessity of LLM proposals is weaker than the headline suggests.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim—that under matched budgets parent-conditioned LLM evolution is indistinguishable from fresh sampling, and that set-level selection over extracted term dictionaries carries discovery—is tightly supported by the budget-matched audit (Table 1), generation-flatness (Fig. 2 / D4), same-dictionary two-family selector separation (Table 3 / Fig. 3, p < 10^{-73}), and official Acc0.1 gains at 1/10 budget (Table 2). The reader's weakest assumption (flat reusable additive components with |S|≤6 and train-only joint scoring) is already the paper's own scope boundary (Principle 3, §8, complexity-guard experiment in §7). No additional internal inconsistency, leakage, or untested load-bearing step appears that would overturn the stated claims within that scope. Remaining caveats (single-seed DeepSeek anchor, strong generic-primitive baseline, caveated SA, weak LSR-Transform) are already reflected in the CONDITIONAL verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper audits parent-conditioned LLM evolution loops for scientific equation discovery under matched call budgets and finds them indistinguishable from fresh independent sampling (median OOD NMSE 0.045 vs 0.049), with instructed multi-parent crossover worse, final success predicted by initial proposals (AUC 0.959), and three iteration schemes flat or destructive. It reduces the loop operationally to a dictionary of candidate terms and proposes PTB-Search: one-shot independent LLM proposals, term extraction into a per-problem dictionary (pooled with generic primitives), and a single train-only set-level sparse subset search with least-squares coefficients (|S|≤6). On identical dictionaries, set-level selectors solve 165–169 of 717 cells versus 74–78 for single-term reductions. On the official 239-problem LLM-SRBench split, PTB-Search reaches 73.2% Acc0.1 (Llama-3.1-8B) and 77.0% (single-seed DeepSeek-V4 anchor) versus 49.2% for the best reported baseline at one-tenth the standardized LLM-call budget. A scoped program-domain stress test finds generation count still unreliable while retained external state can help on harder instances.","tokens_in":23325,"tokens_out":1765,"duration_ms":27586,"significance":"If the matched-budget audit and selector-family separation hold, the paper supplies a load-bearing correction to a widely used narrative in LLM scientific discovery: parent-conditioned “evolution” is often resampling, and discovery is carried by external set-level selection over reusable components. The work is unusually careful for this literature—budget-matched paired arms, train-only selection with leakage audits, pre-registered interpretation branches, same-dictionary selector contrasts with extreme paired significance, collapse ablations for every architectural piece, and an explicit scope boundary via the program stress test and complexity guard. The official Acc0.1 gains at 1/10 budget, together with the two-family selector result (165–169 vs 74–78), are of clear interest to symbolic regression, LLM-for-science, and sparse dictionary methods. Strengths that should be credited include the leakage contract, frozen blind splits, and the falsifiable Principles 1–3.","major_comments":[{"comment":"Appendix D / Table 10 (Llama 717-cell full benchmark): GenericPrimitiveSparse records more near and strict-solve cells (244 / 160) than Hybrid PTB-Search (239 / 135), while Hybrid wins only on median OOD NMSE (2.75e-2 vs 5.15e-2). The official headline in Table 2 reports Hybrid Acc0.1 = 73.2% but does not report the same Acc0.1 aggregation for a generic-primitives-only control under the identical frozen selector and protocol. Because the paper’s thesis is that “LLMs supply the material” while selection carries discovery, the official contribution of LLM-derived terms versus the selector-plus-generics stack is incompletely quantified on the metric used for the main claim. Please add generic-only (and, if feasible, LLM-only) Acc0.1 on the official 239-problem aggregation, or revise the material-supplier wording to match the Llama evidence that generics can dominate strict solves when the p","section":"Table 2; Appendix D / Table 10"},{"comment":"§5 / Table 2 and Appendix C: The DeepSeek-V4 official result is a single-seed (seed-0) anchor of 23,900 proposals, while the Llama result is a three-seed 717-cell grid. The manuscript labels this carefully, but the abstract and teaser still juxtapose 77.0% and 73.2% as parallel headline numbers against the multi-seed baseline composite. For the central cross-backbone stability claim (“within four points”), either restrict the abstract/teaser to the multi-seed Llama figure and treat DeepSeek strictly as a corroborating anchor in the body, or add at least a second DeepSeek seed on a declared subset so the gap is not seed-0-dependent.","section":"§5; Table 2; Abstract / Figure 1"},{"comment":"§4 hypothesis class and §7 complexity guard: The method is defined as sparse linear combinations with |S|≤6, and the nonlinear wrap/product-of-sums probe is excluded because it multiplies formula length (~8×) for only marginal near gains. That scope is honest, but the remaining-error autopsy (Fig. 6) attributes a non-zero composition gap on the Llama benchmark (5%) and a large availability gap (41%). The paper should state more explicitly in the main claims (not only Principle 3 / §8) that PTB-Search’s Acc0.1 gains are conditional on ground-truth structure being well approximated by a short additive support over extractable terms, and that LSR-Transform weakness is the designated failure regime under that assumption rather than a residual selector defect.","section":"§4; §7; Figure 6; Principle 3"}],"minor_comments":[{"comment":"Symbolic accuracy is correctly caveated, but Table 2 still leaves SA as ‘– †’ for PTB-Search while listing baseline SA percentages. A short footnote in the table itself (not only the caption) would prevent readers from treating the blank as missing data.","section":"Table 2"},{"comment":"Figure 5 (BPG10) is excellent for Principle 2; the main text could point more explicitly to Table 13’s train-tied / OOD-separated supports so readers can verify the one-term swap without hunting the appendix.","section":"Figure 5; Appendix E / Table 13"},{"comment":"Program-domain hard grid excludes seeds 3–4 after API/billing errors (Appendix F). The main text already scopes the result; one sentence noting that the sign-tests are on a 3×3 clean grid would reduce the chance of over-reading Table 4.","section":"§8; Table 4; Appendix F"},{"comment":"Typos / formatting: title line breaks (“NotDarwin”, “SET-LEVELSELECTION”) appear to be PDF extraction artifacts in places; ensure the camera-ready title spacing is clean. Also standardize “Acc 0.1” vs “Acc0.1” notation.","section":"Title; throughout"},{"comment":"Related work cites concurrent IGSR (Saveliev et al., 2026) appropriately for the per-term vs set-level distinction; a one-sentence clarification that the single-term reductions in Table 3 are re-implementations on identical dictionaries (not a re-run of their full system) would avoid a fairness quibble.","section":"§9; Table 3"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is stronger than most LLM-for-science submissions on experimental hygiene (leakage audits, pre-registration, same-dictionary selector contrasts). The generic-primitives result on Llama is the one place where the narrative could overclaim relative to the appendix tables; if the authors add official Acc0.1 for generic-only and slightly soften the abstract’s DeepSeek juxtaposition, I would expect a clean accept. Fit for a methods/ML venue is good; less so for a pure symbolic-regression journal that prioritizes compact closed forms, given the caveated SA."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is clean. Under matched call budgets on equation discovery, parent-conditioned “evolution” does not beat fresh independent sampling (median OOD NMSE 0.045 vs 0.049), instructed multi-parent crossover is worse, and final success is almost entirely predicted by the initial proposals (AUC 0.959). Three iteration schemes are flat or destructive. That null result is what the paper actually establishes, and it is well measured.\n\nWhat is new is not sparse dictionary regression—that is classical—but the controlled audit of the LLM lineage loop against its proper null, plus the same-dictionary two-family selector separation: set-level joint selectors solve 165–169 of 717 cells while single-term reductions (including marginal credit of the concurrent style) solve 74–78, with p-values that are not decorative. PTB-Search is the constructive reduction of that diagnosis: propose once, extract terms, one train-only set-level sparse recombination. On the official 239-problem split it hits 73.2% Acc0.1 with Llama-3.1-8B at 100 calls versus 49.2% for the best reported baseline at 1,000. The cross-backbone gap to a single-seed DeepSeek anchor is only four points, which supports the claim that selection carries a lot once the dictionary has material.\n\nSoft spots, in proportion. The method assumes flat reusable additive components with a support cap of six; the paper scopes this honestly (Principle 3, program stress test, complexity-guard experiment) and does not pretend to free-form program synthesis. Generic primitives alone are already strong, so LLM terms are complementary rather than sole drivers—stated, not hidden. DeepSeek official is a single-seed anchor; SA is caveated; LSR-Transform remains weak and is mostly availability-limited. None of these overturn the central audit or the selector-family result within the stated scope.\n\nMath and protocol look solid: train-only selection, leakage audits, pre-registered branches, paired cell-level tests, identical dictionaries across selector arms. Citations engage FunSearch/LLM-SR/LaSR and the sparse-regression lineage without padding.\n\nThis is for people building or auditing LLM scientific-search loops, and for anyone who still treats generation count as the active ingredient. Bring it to reading group. It deserves a serious referee; I would cite the audit and the set-level separation.","headline":"The matched-budget audit is the real contribution: parent-conditioned LLM evolution is just resampling, and set-level selection over term dictionaries does the work.","tokens_in":23917,"tokens_out":598,"would_cite":true,"duration_ms":8500,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"In scientific equation discovery, LLM evolution does not beat fresh sampling; discovery is carried by set-level selection over a term dictionary.","keywords":["scientific equation discovery","LLM evolution","symbolic regression","set-level selection","sparse dictionary regression","proposal banks","identifiability","PTB-Search"],"falsifier":"On a budget-matched equation-discovery grid with the same proposer and train-only selection, show that parent-conditioned multi-generation proposals systematically beat the same number of independent samples on held-out OOD accuracy, or that single-term credit selectors match set-level joint selectors on identical dictionaries.","tokens_in":23910,"feed_emoji":"📐","tokens_out":750,"duration_ms":5416,"temperature":0.7,"pith_summary":"The paper audits a common practice: using large language models as evolutionary engines that generate candidate equations, keep winners, and mutate them across generations. Under matched call budgets on scientific equation discovery, that parent-conditioned loop is no better than drawing the same number of independent proposals, multi-parent crossover is worse, and success is already fixed by the quality of the first proposals. What the loop actually produces is a dictionary of reusable terms. The authors turn that diagnosis into PTB-Search: sample proposals once, extract additive terms, and run one train-only sparse selection that scores candidate sets jointly with least-squares coefficients. On identical dictionaries, set-level selectors roughly double the solves of single-term credit rules. On the official 239-problem benchmark they report 73.2% numeric accuracy with a small open model and 77.0% with a stronger single-seed anchor, versus 49.2% for the best reported baseline, at one tenth of the standardized LLM budget. A program-domain stress test keeps generation count from becoming the active ingredient and points instead to retained external state. The upshot is a division of labor: LLMs supply material; discovery is done by external set-level selection over reusable components.","feed_headline":"LLM evolution fails; set-level term search wins equation discovery","feed_subtitle":"One-shot proposals plus joint sparse selection beat multi-generation loops at one tenth the budget.","key_machinery":"PTB-Search (proposal term bank search): one-shot independent LLM proposals, extraction of reusable additive terms into a per-problem dictionary, then a single global train-only best-subset sparse selection with least-squares coefficients. Its load-bearing principle is set-level identifiability—score candidate supports jointly rather than reduce to per-term statistics.","core_discovery":"Under matched LLM-call budgets in continuous equation discovery, parent-conditioned evolution is indistinguishable from fresh independent sampling, so the loop does not compound scientific structure. Operationally it reduces to a dictionary of candidate terms; discovery is then carried by train-only, set-level sparse recombination, because underdetermined data identifies the joint behavior of term sets rather than reliable per-term credit. PTB-Search implements that reduction and substantially outperforms the best reported baseline at one tenth of the standardized call budget.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Dictionaries beat Darwin: set-level search tops LLM evolution","LLM evolution fails to compound; set-level term selection wins","One-shot dictionaries plus joint sparse selection beat evo loops","Set-level term recombination, not multi-gen LLM loops, finds equations","Underdetermined data needs joint term sets, not parent-conditioned evolution"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The method assumes that useful scientific expressions can be broken into flat, reusable additive terms that can be extracted into a dictionary and scored jointly under a small support-size budget; if the true structure is non-additive or terms cannot be cleanly extracted, the reduction and the gains do not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Dictionaries beat Darwin: set-level search tops LLM evolution","LLM evolution fails to compound; set-level term selection wins","One-shot dictionaries plus joint sparse selection beat evo loops","Set-level term recombination, not multi-gen LLM loops, finds equations","Underdetermined data needs joint term sets, not parent-conditioned evolution"]},"model":"grok-4.5","effort":"low","cost_usd":0.006338,"raw_usage":{"total_tokens":1739,"prompt_tokens":928,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":63380000,"prompt_tokens_details":{"text_tokens":928,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":719,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":928,"tokens_out":92,"duration_ms":5696,"temperature":1.0,"reasoning_tokens":719,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T21:37:37.211830+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a budget-matched equation-discovery grid with the same proposer and train-only selection, show that parent-conditioned multi-generation proposals systematically beat the same number of independent samples on held-out OOD accuracy, or that single-term credit selectors match set-level joint selectors on identical dictionaries.","supporting_citations":[],"review_version":1}