{"id":"d71f4cd9-9c17-4a20-a7c7-ef397a6f7f2d","arxiv_id":"2608.00316","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An LLM agent that fully controls a reconfigurable Bayesian-optimization backend preserves standard BO reliability, outperforms LLM-only optimizers, and exploits natural-language priors and mid-run problem reformulation.","lead":"An LLM agent sits at the center of a Bayesian-optimization loop, querying and reconfiguring a probabilistic backend instead of replacing it. The paper shows this \"agentic BO\" matches standard BO with no domain hints, beats LLM-only optimizers, and turns natural-language chemistry and tuning priors into faster convergence.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Q2's 'natural-language priors improve beyond standard BO' rests only on priors that are correct by construction; no misleading or vague prior is tested, so the claimed improvement may not generalize.","rationale":"Read in good faith, the paper makes a three-part empirical claim: no-prior parity with state-of-the-art BO, superiority over LLM-only baselines, and improvement from natural-language priors. The first two are reasonably supported: the surrogate-ablation and GP-sample-path experiments provide clean evidence that the Bayesian backend adds value, and the LLAMBO/Centaur comparisons are consistent. The weakest load-bearing assumption is the prior-advantage claim. The reaction-yield benchmarks were constructed so that the natural-language context directly names the conditions that determine the optimum; this makes the prior informative by design, not by the agent's robust handling of imperfect knowledge. The paper never tests the regime where priors are misleading, vague, or partially wrong, even though its own system prompt encourages the agent to trust explicit context cues, which could amplify the harm of a bad prior. This is a missing control, not an internal contradiction. The reader's verdict of CONDITIONAL already hinges on this gap, so my assessment does not change the verdict; it reinforces it. The proposed test would settle whether the concern actually lands by measuring whether Sara's prior-informed gains survive adversarial or merely noisy priors.","tokens_in":36918,"tokens_out":12382,"duration_ms":117045,"concrete_test":"Run the same reaction-yield and LCBench protocols with deliberately mis-specified priors — e.g., describe Suzuki–Miyaura as 'hot, anhydrous, strong base' or swap reaction names/regimes — and also with a vague prior lacking trustworthy cues. Compare 10-seed final yield/regret against (a) Sara with the current correct prior, (b) Sara with no prior, and (c) Ax. If misleading-prior Sara falls below no-prior Sara or Ax, the Q2 claim is conditional on prior quality; if it does not, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's third clause — 'uses natural-language priors to improve beyond standard BO' — is supported by Section 6.3/Figure 7 and the LCBench results. In the reaction-yield family (Appendix A.2, Table 5), the same qualitative chemistry given in the prior string (Figure 15: 'Suzuki–Miyaura ... THF/water mixture with a mild base') is used to set the true optimum (μ_T=70 °C, μ_w=0.45). The prior is therefore not merely informative; it encodes the answer regime by construction. No experiment supplies a misleading, vague, or partially wrong prior — the regime where an agent explicitly instructed to 'trust explicit context cues and don't argue yourself out of them' (Section F) would confidently steer evaluations away from the optimum. The prior ablation (Figure 11) only contrasts an informative prior with no prior; it does not test robustness to plausible prior error. Because the paper claims a capability beyond standard BO, the absence of any wrong-prior condition is a missing control: if real-world priors are often approximate or wrong, the demonstrated gains could shrink or reverse, and the headline improvement is not established. The no-prior parity result (Q1) would survive, but the distinctive added value of agentic BO over standard BO would be conditional on prior quality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces agentic Bayesian optimization, in which an LLM agent is the central decision maker and delegates probabilistic modeling to a modular BoTorch-based backend (lenz). The framework is formalized as a metalevel decision process (Sec. 4) and instantiated as Sara, an LLM agent with a defined system prompt and CLI toolset. The experiments address four questions: no-prior parity with classical BO (Q1), use of natural-language priors (Q2), mid-run reconfiguration (Q3), and model/prompt ablations (Q4). The main empirical results are that Sara matches or slightly outperforms Ax on synthetic functions without a prior, outperforms LLAMBO and Centaur, improves faster with task descriptions on LCBench and synthetic reaction-yield functions, and can reformulate a constrained problem into a multi-objective one mid-run. The paper includes full system prompts and detailed appendices.","tokens_in":37105,"tokens_out":7460,"duration_ms":65692,"significance":"If the central claim holds, agentic BO is a meaningful advance: it combines an LLM's ability to consume unstructured priors and adapt strategy with calibrated GP-based search, and it demonstrates a dynamic-reconfiguration capability not available in standard BO. The paper is commendable for shipping full prompts, a detailed CLI reference, explicit anti-patterns, and honest discussion of benchmark recognition (Sec. D.1). The empirical support for the strongest headline—improvement beyond standard BO via natural-language priors—is not yet convincing, and the no-prior parity claim is partly confounded by pretraining memory. With additional controls, the contribution could be significant.","major_comments":[{"comment":"The Q2 claim that natural-language priors improve beyond standard BO is tested only with priors that are correct by construction. The reaction-yield contexts in Fig. 15 specify the reaction, catalyst, solvent, and base, and Table 5 places each function's optimum in exactly the regime those cues identify (e.g., Suzuki mu_T=70 C, mu_w=0.45). The ablation in Fig. 11 only contrasts this informative prior with no prior. No experiment supplies a misleading, vague, or partially wrong prior, so the demonstrated gain may reflect prior accuracy rather than agentic use of priors. Please add wrong/perturbed prior conditions and, if the claim is limited to correct priors, state that limitation.","section":"A.2, Table 5, Fig. 15; Sec. 6.3"},{"comment":"Sara's system prompt was explicitly iterated after observing failure modes in early experiments (Table 2 lists directives added because 'we observed' specific behaviors). The paper does not establish that the final benchmark tasks were held out from this prompt-development loop. Since the headline results compare Sara against baselines on these same tasks, prompt overfitting to the test suite is a live threat. Please clarify whether prompt tuning was performed on a separate development set; if not, add a validation split or a comparison with a generic untuned agent prompt.","section":"5.2, Table 2, Sec. 7"},{"comment":"The 'beyond standard BO' comparison in the prior-informed setting is asymmetric: Sara, LLAMBO, and Centaur receive the natural-language description, while Ax receives no prior at all. This conflates the availability of a domain prior with the agentic architecture. A standard BO method with the same prior injected through a conventional mechanism (e.g., piBO with user beliefs, or a narrow initial search region derived from the description) would separate prior encoding from agentic control. Without such a baseline, the distinctive added value of agentic BO over standard BO is not established.","section":"6.3, 6.2, 6.5"},{"comment":"The no-prior synthetic benchmarks are contaminated by pretraining memory. The paper itself reports that the bash-only agent explicitly identified the benchmark in 4/10 Hartmann, 3/10 constrained Hartmann, and 10/10 Ackley-10 runs despite renaming and shifting (Sec. D.1). This means the Q1 parity claim on these functions is not a clean measure of optimization ability. The GP-sample-path experiments are a good control and show a clear surrogate benefit, but they are multi-objective or high-dimensional and do not fully substitute for a single-objective no-prior parity test. Please either report no-prior parity on fresh single-objective surfaces or soften the claim.","section":"6.1, D.1"}],"minor_comments":[{"comment":"The formatting of Table 3 is ambiguous in the text: sub-/superscript q25/q75 values are run together with medians (e.g., '0.00060.0010 0.0005'). Please reformat so medians and quartiles are clearly distinguished and define the order explicitly.","section":"Table 3"},{"comment":"The summary statement 'Sara matches Ax' is slightly stronger than the table for Branin with Opus 4.8 (median 0.0020 vs Ax 0.0006). The difference is small, but please ensure the prose and table are consistent.","section":"Sec. 6.1, Table 3"},{"comment":"The finding that reasoning level 'off' outperforms higher reasoning levels on Mizoroki-Heck is interesting, but no statistical test is reported. A Mann-Whitney test or confidence intervals would clarify whether this is a robust effect.","section":"Sec. 6.5, C.1"},{"comment":"The paper would benefit from an explicit code/data availability statement. The full prompts and CLI reference are valuable, but releasing lenz and the evaluation harness would make the empirical claims much easier to reproduce and extend.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This paper is likely to interest the BO and LLM-agent communities. The authors are transparent about prompt design and benchmark recognition, which I appreciate. The largest risk is that the strongest claim (natural-language priors improve beyond standard BO) rests on priors that are correct by construction and on an asymmetric baseline comparison. Adding wrong-prior controls and a conventional prior-augmented BO baseline would substantially strengthen the contribution. I would be supportive if those experiments are added and the no-prior claim is qualified appropriately."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper is worth engaging with. The architecture is genuinely new — an LLM agent with full, lossless control of a BO backend — and the write-up is unusually honest about its own limitations. But you should not take the abstract's claims at face value. The headline \"preserves reliability of SOTA BO\" has a counterexample in Table 3 (Branin), and the \"natural-language priors improve beyond standard BO\" claim rests on priors that are correct by construction, with no misleading-prior control.\n\nWhat the paper does well: the framing of the design space is clear, and the related work is placed accurately. The experiments are more careful than the norm for LLM-agent papers. They sandbox the agent, rename and shift test-function optima, show the agent still recognizes benchmarks, and then add GP sample paths that cannot be memorized. Those GP-sample-path results are the cleanest evidence: the full Sara/lenz system beats the same agent with bash only by a wide margin in high dimensions. The ablations on model family, reasoning level, and prior conditions are useful. The limitation section is candid about prompt sensitivity, nondeterminism, and tool-use inertia.\n\nNow the soft spots. First, the \"matches Ax on synthetic benchmarks\" summary is false for Branin. In Table 3, Sara (Opus 4.8) ends at median regret ~0.23, Ax at ~0.0006. That is a large gap on a trivial 2-D problem. It does not sink the paper, but the claim needs to be qualified. Second, the prior-improvement result is narrower than advertised. The reaction-yield functions were constructed so the chemistry in the problem description lines up with the true optimum regime (Appendix A.2, Table 5). The agent is told to trust explicit context cues and not argue itself out of them, so the prior is effectively the answer key for the regime. There is no experiment with a misleading, vague, or partially wrong prior, which is exactly the regime where this capability could hurt. The stress-test note is right: this is a missing control. The no-prior parity result survives, but the distinctive added value of agentic BO over standard BO is conditional on prior quality. Third, the system prompt was iterated after observing agent failure modes on early runs, so the reported behavior is at least partly fitted to these tasks. The authors are transparent about this, but it still limits transferability. Fourth, the mid-run reformulation is a single trace, not a benchmark. Fifth, there is no code or data release, so the empirical claims cannot be independently checked.\n\nBottom line: this is a serious paper that deserves a serious referee. I would accept it for peer review, but I would send it back with three asks: qualify the SOTA-BO claim, add misleading-prior conditions, and release the code. Anyone working on LLM-BO hybrids or autonomous experimental design should read it; I'd bring it to a reading group.","headline":"A genuinely new agentic-BO architecture with careful, honest experiments, but the natural-language prior claim is only tested with priors that are correct by construction, and the Branin result contradicts the stated parity with SOTA BO.","tokens_in":37766,"tokens_out":3689,"would_cite":true,"duration_ms":36316,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Putting an LLM agent in charge of a Bayesian optimization loop preserves sample efficiency and lets natural-language descriptions act as priors that improve on standard BO.","keywords":["agentic Bayesian optimization","large language model agent","surrogate backend","Gaussian process","acquisition function","natural-language priors","run-time reconfiguration","autoresearch"],"falsifier":"Run the reaction-yield benchmarks again, but hand the agent a deliberately wrong context (for example, a cold anhydrous Grignard recipe for the Suzuki–Miyaura task). If the wrong-prior runs still beat uninformed BO, then the claimed prior advantage does not depend on prior correctness; if they fall below the no-prior runs, the 'natural-language priors improve beyond standard BO' claim requires accurate priors.","tokens_in":36667,"feed_emoji":"🤖","tokens_out":8457,"duration_ms":71889,"temperature":0.7,"pith_summary":"This paper introduces agentic Bayesian optimization: a paradigm in which a large language model (LLM) agent is the central decision maker in a Bayesian optimization (BO) campaign, while a Bayesian backend supplies uncertainty-aware candidate proposals. The central claim is that this division of labor preserves the sample efficiency of state-of-the-art BO when the agent has no domain knowledge, out-performs LLM-only optimizers that lack a calibrated surrogate, and converts natural-language problem descriptions into useful priors that improve on standard BO. The authors instantiate this in Sara, the agent, and lenz, a modular backend whose raw trial log survives reconfiguration, so the search strategy—bounds, acquisition function, even the objectives and constraints—can be revised mid-run without discarding data. Experiments span synthetic functions, hyperparameter tuning, and reaction-yield optimization; in a scaling-law study the agent re-frames a constrained single-objective problem as a multi-objective Pareto search after the user changes the requirements.","feed_headline":"Matches and, with priors, beats standard Bayesian optimization","feed_subtitle":"An LLM agent steering a surrogate backend preserves BO's sample efficiency and adds mid-run reconfiguration.","key_machinery":"The load-bearing object is the metalevel deliberation process defined in Section 4. The agent's policy A maps a state—trial data, the current configuration (surrogate, acquisition, bounds, objective/constraint partition), append-only context, and deliberation history—to either a computational action (probe, reconfigure, propose) or an evaluation action. Three design choices carry the argument: the separation of propose from commit, so the agent can accept, refine, or override the backend's candidate; the append-only context K_t, so later instructions can supersede earlier ones without invalidating data; and the backend (lenz), whose command-line interface exposes commands for creating proble","core_discovery":"Standard BO fixes its whole policy before the first evaluation: surrogate, acquisition function, search region, and the split of outcomes into objectives and constraints. The paper's central claim is that replacing this fixed policy with an LLM agent—which can inspect the surrogate, request and override proposals, and re-edit the configuration mid-run—costs nothing in sample efficiency and adds two capabilities. First, without any domain knowledge, the agent matches a well-tuned classical BO baseline on synthetic problems, while LLM-only baselines that discard the surrogate underperform, sometimes worse than random search. Second, with a natural-language problem description, the agent turns","pith_inferences":["This suggests a strict-generalization reading: any fixed BO policy is the special case where the agent always accepts the surrogate's proposal and never reconfigures; a formal regret comparison between the agent's policy and the fixed policy it could have run would clarify when deliberation pays.","The paper's own observation that the LLM can identify standard test functions from a few evaluations despite shifted optima implies synthetic-benchmark comparisons of LLM-based optimizers should be treated skeptically; random GP paths and fresh problem families are a more trustworthy evaluation surface.","A testable extension is fine-tuning the agent on the meta-MDP objective with token costs; if a small fine-tuned policy can replicate a frontier model's optimization decisions at a fraction of the token budget, the paradigm would become much cheaper to deploy.","The most direct practical consequence left implicit: agentic BO is best suited to expensive, evolving campaigns—such as experimental chemistry, hardware design, or ML system tuning—where requirements change mid-study and domain expertise exists mainly as text. A conversational interface that can re-target the objective is itself the product."],"forward_implications":["Without any natural-language context, the agentic system performs on par with a tuned classical BO policy across low- to high-dimensional synthetic problems, showing that the added agent layer does not sacrifice sample efficiency.","With a natural-language description, the same system achieves substantially better early convergence than BO that must discover productive regions from scratch, on both hyperparameter-tuning and reaction-yield benchmarks.","LLM-only optimizers that propose points from text summaries without a calibrated surrogate underperform, at times worse than random sampling, when the objective has no recognizable structure; the surrogate is what makes the search systematic.","Mid-run reconfiguration—promoting a constraint to an objective and switching to a hypervolume-based acquisition function—works without discarding any evaluations, a capability standard BO does not offer.","The agent's tool-use pattern is context-dependent: with a prior it moves quickly to local refinement around the incumbent, while without one it relies on generic surrogate proposals longer; persistent acquisitions and bound edits remain rare on the tested tasks."],"fun_headline_variants":["LLM agent steers Bayesian optimization, matches or beats it with priors","Agentic BO: LLM decides, Bayesian backend ensures sample efficiency","Mid-run reconfiguration: LLM agent revises Bayesian optimization on the fly","Surrogate-augmented autoresearch: LLM agent improves Bayesian optimization","LLM agent + Bayesian backend: sample-efficient optimization with adaptive strategy"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The demonstrated gains over standard BO assume the natural-language priors given to the agent are accurate and are correctly decoded; with misleading or vague priors the advantage would shrink or reverse, though the no-prior parity result would survive.","fun_headline_variants_meta":{"raw":{"variants":["LLM agent steers Bayesian optimization, matches or beats it with priors","Agentic BO: LLM decides, Bayesian backend ensures sample efficiency","Mid-run reconfiguration: LLM agent revises Bayesian optimization on the fly","Surrogate-augmented autoresearch: LLM agent improves Bayesian optimization","LLM agent + Bayesian backend: sample-efficient optimization with adaptive strategy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2647,"prompt_tokens":818,"completion_tokens":1829,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":1733}},"tokens_in":562,"tokens_out":1829,"duration_ms":11516,"temperature":1.0,"reasoning_tokens":1733,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:45:06.054027+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the reaction-yield benchmarks again, but hand the agent a deliberately wrong context (for example, a cold anhydrous Grignard recipe for the Suzuki–Miyaura task). If the wrong-prior runs still beat uninformed BO, then the claimed prior advantage does not depend on prior correctness; if they fall below the no-prior runs, the 'natural-language priors improve beyond standard BO' claim requires accurate priors.","supporting_citations":[],"review_version":1}