{"id":"d5b4c13c-c5d3-4558-8333-3fdde568e42e","arxiv_id":"2505.04169","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Multi-resolution Bayesian optimization with hierarchical coarse-graining finds molecules that promote lipid bilayer phase separation more strongly than single-resolution BO at the same evaluation budget.","lead":"This paper introduces a multi-level Bayesian optimization approach that searches chemical space at several coarse-grained resolutions at once, using cheap low-resolution simulations to guide expensive high-resolution ones. It demonstrates the funnel strategy on a lipid bilayer problem, finding molecules that promote phase separation with fewer evaluations than standard optimization.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-resolution prior in Eq. (5) is unvalidated: Section III D's neighborhood-size evidence is confounded by CS density, so the Fig. 9 advantage over standard BO lacks a demonstrated mechanism.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing point: the method's efficiency gain depends on lower-resolution landscapes being informative priors, but this is never directly measured. My stress-test sharpens that concern in two ways. First, the paper's Section III D evidence for smoothness is indirect and confounded: the neighborhood sizes in Fig. 10 and Table SI2.2 are computed from GP lengthscales fit to optimization-selected points, and the mapped neighborhood sizes scale almost exactly with the total number of molecules per resolution, so they mostly reflect the combinatorial multiplicity of the many-to-one mapping rather than an independent smoothness comparison. Second, the central claim is supported by only one real-system run of multi-level BO versus standard BO; the toy-model replication in SI2.7 uses a different representation and objective, so it cannot validate the bilayer result. The paper has real strengths—released code, an independent direct-simulation validation of the best candidate, and an honest discussion of limitations—so the appropriate disposition remains CONDITIONAL rather than rejection. The condition should be a direct cross-resolution predictivity check and, ideally, a repeated matched comparison on the real system or a more faithful toy model. Since my concern is the same one the reader flagged and the verdict already reflects it, no verdict adjustment is needed.","tokens_in":28635,"tokens_out":9024,"duration_ms":106009,"concrete_test":"Use the released repository and free-energy tables to identify the high-resolution evaluations whose mapped medium-resolution (or low-resolution) molecules were themselves evaluated during the optimization, and compute the paired residuals δ(x) = ΔΔG_high(x) − ΔΔG_medium(M(x)). Report correlation, mean, and standard deviation of δ, and also compare the high-resolution GP's prior predictive mean at x (from the lower-level model, as used in Eq. (8)) against the observed ΔΔG_high(x). If the lower-resolution values or prior means do not explain most of the variance in the high-resolution observations, the informative-prior assumption behind Eq. (5) is not supported and the Fig. 9 comparison cannot be attributed to the hierarchy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is Eq. (5): f_l(x) = f_{l−1}(M_l(x)) + δ_l(x), with δ_l modeled as a zero-mean GP. If lower-resolution free-energy landscapes are not informative about high-resolution values, then the funnel's guidance—especially the restriction of EI maximization to neighborhoods of low-level favorites (Section SI1.4)—can exclude the true optimum, and any advantage in Fig. 9 must be attributed to initialization or search restriction rather than to the hierarchy. The paper's only direct evidence for this premise is Section III D, where 'chemical neighborhood sizes' are obtained from RBF lengthscales fitted to the optimization trajectory's evaluated points. This does not test cross-resolution predictivity: the reported mapped neighborhood sizes (249 → 18,700 → 378,000) track the total CS sizes (9.0×10^4, 6.7×10^6, 1.37×10^8) almost exactly, meaning the quoted growth largely reflects molecule density and the many-to-one mapping, not a measured smoothness gradient. The paper never reports the distribution of δ_d(x) for any molecule, nor the correlation between lower-resolution priors and high-resolution observations. Because the headline comparison is a single run (Section III C), an unvalidated core assumption leaves the efficiency claim conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a multi-level Bayesian optimization framework for small-molecule discovery. Chemical space is represented at three coarse-grained resolutions with 15/45/96 Martini3-derived bead types; each space is embedded with a graph-neural-network regularized autoencoder. A delta-learning Gaussian process (Eq. 5) propagates lower-resolution free-energy information upward, with the lowest level initialized by a bead-additivity prior (Eq. SI3). The method is demonstrated by minimizing a free-energy proxy for phase separation in DPPC/DLiPC/cholesterol bilayers. The authors report that the multi-level funnel uses 327 evaluations to outperform standard high-resolution BO, finds candidates with ΔΔG below −1.3 kcal/mol, and validates the best candidate with direct contact simulations. They also argue that lower-resolution chemical neighborhoods are larger, implying smoother landscapes.","tokens_in":28988,"tokens_out":6277,"duration_ms":61886,"significance":"If the efficiency claim holds, the paper is a valuable contribution: it combines transferable CG models and active learning in a way that is broadly applicable to free-energy-based molecular optimization. Strengths include the release of code, models, and data; the independent validation of the top candidate with a distinct observable (DPPC–DLiPC contacts); a multi-run toy-model comparison that partially addresses initialization effects; and the identification of interpretable chemical design rules. However, the central efficiency claim rests on a single real-system run and on an unvalidated delta-learning premise, so the quantitative advantage over single-resolution BO should be regarded as conditional rather than established.","major_comments":[{"comment":"The evidence for the central premise that lower-resolution landscapes are informative priors is not direct. The neighborhood-size analysis in Section III D derives ξ_l by fitting RBF kernels to the evaluated molecules and then reports mapped neighborhood counts of 249 → 18,700 → 378,000. As shown in SI2.6, these counts are obtained by multiplying the low-resolution neighborhood size by the average many-to-one multiplicities (75 and 20), so the growth tracks the total number of molecules per resolution (9.0×10^4, 6.7×10^6, 1.37×10^8) rather than a measured smoothness gradient. The paper never reports the distribution of δ_l(x) (Eq. 5) or the correlation between lower-resolution predictions and high-resolution observations for the same molecule. Without such a test, the restriction of EI maximization to neighborhoods of low-resolution favorites (SI1.4) could be excluding the high-resolution optimum, and the Fig. 9 advantage cannot be attributed to the hierarchy rather than to search restriction.","section":"Section III D, Eq. (5)"},{"comment":"The headline comparison is a single run of each algorithm. The text acknowledges that averaging over multiple runs is 'computationally infeasible', but the resulting curves have no error bars, and standard BO is known to be initialization-sensitive. The toy-model average in SI2.7 uses a different synthetic score and cannot by itself establish the real-system comparison. I request either several independent standard-BO runs on the bilayer system at reduced cost, a statistical treatment of the single-run comparison, or a clear reframing of the claim as a case study rather than a general outperformance result.","section":"Section III C, Fig. 9"},{"comment":"The free-energy proxy ΔΔG is validated as a predictor of phase separation on only one molecule. The 1600 ns contact simulation shows the top candidate outperforms benzene and the control, but it does not test whether the ΔΔG ranking across the discovered candidates (Figs. 5 and 9) reflects demixing propensity. Given that the distribution-level claim in Section III C is based entirely on ΔΔG values, an additional validation on a non-top candidate, or on a small set spanning the observed range, would materially strengthen the claim that the method navigates chemical space toward the target property rather than toward an uncorrelated proxy.","section":"Section III B / Section II F"}],"minor_comments":[{"comment":"The main text says S is defined as a conditional weighted sum of ΔGwater−ΔGinterface and ΔGinterface−ΔGcenter, but Eq. SI2 uses conditions involving ΔGwater−ΔGcenter and ΔGinterface−ΔGcenter; please reconcile the description and the formula.","section":"Section II F and SI Eq. SI2"},{"comment":"The mapping M_l is defined via discrete molecular mapping between bead-type levels, but the GP is formulated on continuous latent spaces; please clarify how M_l(x) is evaluated for latent points that do not correspond exactly to an enumerated molecule during EI maximization.","section":"Eq. (8)"},{"comment":"The label 'DIPC' should be 'DLiPC' in Table SI2.1, and the sentence in SI2.6 explaining the multiplication by 75 and 20 should be more explicit that these are enumeration-count ratios rather than measurements of landscape correlation.","section":"Table SI2.1 and Section SI2.6"},{"comment":"The x-axis should state explicitly whether it is the total number of MD evaluations including initialization and prior evaluations for both methods; the current caption is ambiguous.","section":"Figure 9"},{"comment":"Two displayed equations in SI2.2 contain raw LaTeX markup; they should be typeset properly.","section":"SI Eq. SI3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the code/data release is a genuine asset. The main obstacles are the single-run benchmark and the indirect smoothness evidence; I do not see fatal internal contradictions. The authors should be encouraged to provide a direct delta-validation and a repeated-run or statistically bounded comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two headline points. First, this is a genuine piece of methodological integration: hierarchical bead-type coarse-graining used as a multi-fidelity ladder inside Bayesian optimization over learned latent spaces, with code, data, and a tutorial released. Second, the paper's central efficiency claim—that multi-level BO 'outperforms standard BO at a single resolution'—is thinner than the narrative suggests.\n\nWhat is new and right: combining the delta-learning multi-fidelity GP approach (Eq. 5) with a hierarchical CG model where higher-resolution bead types map deterministically to lower resolutions is a natural and useful integration, not a routine extension of Mohr et al.'s single-resolution BO. The demonstration is substantial: 327 molecules evaluated across three resolutions, a free-energy objective for membrane demixing, the top candidate independently validated with direct contact simulation and shown to beat benzene, and the design rules extracted by LASSO agree with prior physics. The released code and NOMAD data make this reproducible, which counts for a lot.\n\nSoft spots: the real-system comparison to standard BO is a single run. The authors acknowledge the computational cost and fall back on a toy-model comparison averaged over 50 runs, which is fine as supporting evidence but not the same as uncertainty on the main figure. The mechanism behind the advantage is the cross-resolution delta-learned prior, and the paper never directly measures how predictive low-resolution values are for high-resolution ones. Section III D's neighborhood-size evidence is partly circular—the lengthscales are fitted to the same evaluated data—and the quoted mapped sizes (249 to 18,700 to 378,000) largely reflect the density of the many-to-one mapping, not a measured smoothness gradient. The underlying fitted neighborhoods (249, 23, 37) actually show medium resolution rougher than high, which complicates the 'smoother at lower resolution' story. The free-energy proxy is validated on one molecule, which is a minor gap. These are addressable, not fatal; the method likely works.\n\nReadership: computational chemists doing inverse design with MD-based or free-energy objectives. It deserves a serious referee, and I'd recommend sending it out with a request for a direct delta-validation or repeated-run comparison rather than desk rejection.","headline":"A real integration of hierarchical CG with multi-level BO, reproducible and well-demonstrated on a membrane design problem, but the headline advantage over standard BO rests on a single run and an unmeasured delta-learning premise.","tokens_in":29432,"tokens_out":5087,"would_cite":true,"duration_ms":46569,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multi-level Bayesian optimization over hierarchically coarse-grained chemical spaces outperforms single-resolution Bayesian optimization for free-energy-based molecular discovery.","keywords":["Bayesian optimization","coarse-graining","chemical space exploration","free-energy differences","lipid bilayer phase separation","multi-fidelity optimization","latent space encoding","active learning"],"falsifier":"Compute $\\Delta\\Delta G$ at high resolution for roughly 50 molecules already evaluated at low resolution and measure the rank correlation between the two sets of values; if the correlation is near zero, the lower-resolution prior cannot carry the search.","tokens_in":28457,"feed_emoji":"🧪","tokens_out":6001,"duration_ms":59682,"temperature":0.7,"pith_summary":"The paper proposes a funnel-like strategy: enumerate chemical space at three coarse-grained resolutions (15, 45, and 96 bead types, spanning about 90,000 to 137 million molecules), embed each resolution in a smooth learned latent space, and run Bayesian optimization that starts at low resolution and uses lower-resolution Gaussian-process predictions as priors for higher-resolution searches. The claim is that this multi-level procedure beats standard Bayesian optimization performed only at the highest resolution, finding better candidates and a better distribution of candidates while evaluating fewer than 330 molecules. The demonstration optimizes small molecules that promote phase separation in a ternary lipid bilayer, scored by free-energy differences from molecular dynamics simulations. If true, the method offers a route to navigate very large chemical spaces when the objective is a free-energy difference, with interpretable design rules as a by-product.","feed_headline":"Coarse-grained funnel finds better molecules faster","feed_subtitle":"Layered chemical representations let Bayesian optimization skip unpromising regions and beat single-resolution search.","key_machinery":"The load-bearing mechanism is delta learning over a hierarchy of coarse-grained resolutions. Three levels share the same atom-to-bead mapping but use 15, 45, and 96 transferable bead types, so every high-resolution molecule maps to a unique lower-resolution molecule. Each resolution's chemical space is embedded separately by a regularized graph autoencoder into a five-dimensional latent space, and a many-to-one mapping $M_l$ transfers lower-resolution predictions into higher-resolution Gaussian processes. The model at level $l$ is $f_l(x) \\sim \\mathcal{GP}(f_{l-1}(M_l(x)), k_l(x,x'))$, with expected improvement maximized only inside neighborhoods of promising lower-resolution points. Resolution switching is triggered when the GP prediction error stays below 0.12 kcal/mol for three consecutive evaluations, and switching back occurs when the best candidate lies farther than $2\\xi_l$ from all evaluated points.","core_discovery":"On the paper's own terms, the central discovery is that coupling hierarchical coarse-graining to Bayesian optimization turns chemical-space resolution into a controllable exploration-exploitation dial. Each level's function $f_l$ is modeled as the next-lower function plus a Gaussian-process correction, $f_l(x) = f_{l-1}(M_l(x)) + \\delta_l(x)$, so low-resolution evaluations map out broad basins while higher-resolution evaluations refine them. In the bilayer application the algorithm evaluated 327 molecules, less than $3\\times10^{-4}\\%$ of the 137-million-molecule high-resolution space, shifted the distribution of $\\Delta\\Delta G$ values steadily downward, and produced top candidates with $\\Delta\\Delta G \\le -1.3$ kcal/mol, all composed of hydrophobic C4, C5, and C6 bead types. A direct 1200 ns validation simulation showed that the best candidate reduced DPPC-DLiPC contacts more than benzene, a known demixing agent. The paper takes this as evidence that multi-level BO outperforms single-resolution BO on both the best value and the spread of good candidates.","pith_inferences":["A natural next test, beyond the paper's analysis, would be an explicit check of the delta-correction assumption: computing high-resolution $\\Delta\\Delta G$ for a sample of molecules already evaluated at low resolution would show whether low-resolution values actually predict high-resolution values for identical molecules.","If the delta assumption holds, the same hierarchical funnel could be applied to other expensive molecular properties, such as binding free energies or solvation properties, with the cheap proxy prior replaced accordingly.","The bilayer comparison to standard BO rests on single runs; the paper's toy model suggests the advantage is reproducible on average, but repeated runs on a cheaper surrogate that preserves the three-resolution structure would make the real-system claim more decisive.","The neighborhood restriction uses the GP lengthscale as a fixed similarity radius; a testable extension would be to let that radius adapt per region, since the chemical space is likely not uniformly smooth."],"forward_implications":["A search that would require screening a 137-million-molecule high-resolution space can instead run a few hundred evaluations across three resolutions and still find multiple strong candidates.","Lower-resolution neighborhoods, each containing roughly 250 molecules, map to neighborhoods of about 378,000 high-resolution molecules, so a modest number of coarse evaluations can guide a large fraction of the fine-grained search.","The workflow returns interpretable chemical rules, here that hydrophobic C4, C5, and C6 beads in mixed sizes promote demixing, alongside the optimized molecules.","The same funnel applies to any molecular objective expressible as a free-energy difference, without requiring pretraining data for the target.","Because the switching thresholds are tied to the chemical space rather than to the specific application, the hyperparameters should transfer to other molecular optimization tasks."],"supporting_citations":[{"why":"Defines the coarse-grained force field whose 96 bead types form the high-resolution level of the hierarchy.","marker":"[24]"},{"why":"Establishes that reducing the number of bead types reduces the combinatorial complexity of chemical space.","marker":"[25]"},{"why":"Supplies the single-resolution Bayesian optimization over a coarse-grained chemical space that this work extends to multiple resolutions.","marker":"[26]"},{"why":"Provides the delta-learning assumption that each resolution's function is a correction to the next-lower one.","marker":"[27]"},{"why":"Gives the DPPC-DLiPC contact metric and benzene as a reference demixing agent for validation.","marker":"[32]"},{"why":"Establishes that demixing-relevant molecules localize near the bilayer center, justifying the free-energy-difference objective.","marker":"[33]"},{"why":"Provides the regularized autoencoder architecture used to embed each resolution's chemical space.","marker":"[43]"},{"why":"Defines the expected-improvement acquisition function used at every resolution.","marker":"[49]"}],"fun_headline_variants":["Funnel of coarse-grained layers finds molecules in 0.0003% of space","Funnel-shaped coarse-graining speeds Bayesian molecule search","327 molecules enough to beat 137M search","Funnel strategy: explore coarsely, exploit finely","Hierarchical Bayesian optimization skips 99.9997% of chemical space"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The funnel works only if a molecule's coarse-grained free-energy value is a sufficiently accurate guide to its finer-grained value, so the delta corrections stay small; the paper never directly measures that correspondence for identical molecules.","fun_headline_variants_meta":{"raw":{"variants":["Funnel of coarse-grained layers finds molecules in 0.0003% of space","Funnel-shaped coarse-graining speeds Bayesian molecule search","327 molecules enough to beat 137M search","Funnel strategy: explore coarsely, exploit finely","Hierarchical Bayesian optimization skips 99.9997% of chemical space"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001402,"raw_usage":{"total_tokens":5670,"prompt_tokens":951,"completion_tokens":4719,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":4632}},"tokens_in":567,"tokens_out":4719,"duration_ms":37059,"temperature":1.0,"reasoning_tokens":4632,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:35:41.328897+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute $\\Delta\\Delta G$ at high resolution for roughly 50 molecules already evaluated at low resolution and measure the rank correlation between the two sets of values; if the correlation is near zero, the lower-resolution prior cannot carry the search.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that reducing the number of bead types reduces the combinatorial complexity of chemical space."},{"cited_title":"Huang, T","cited_arxiv_id":null,"evidence_quote":"Provides the delta-learning assumption that each resolution's function is a correction to the next-lower one."},{"cited_title":"Centi, A","cited_arxiv_id":null,"evidence_quote":"Establishes that demixing-relevant molecules localize near the bilayer center, justifying the free-energy-difference objective."},{"cited_title":"Ghosh, M","cited_arxiv_id":null,"evidence_quote":"Provides the regularized autoencoder architecture used to embed each resolution's chemical space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the expected-improvement acquisition function used at every resolution."}],"review_version":1}