{"id":"bc07ced9-cb0e-4b8d-a439-f30d0f26e24d","arxiv_id":"2608.09696","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MDA combines LLM-proposed hypotheses with sequential Monte Carlo and value-of-information experiment design to discover mechanistic world models from very few experiments.","lead":"This paper presents MDA, a system that uses a large language model to propose candidate mechanistic models and Bayesian inference plus experiment design to identify the right one from few interventions. The authors report large data-efficiency gains over pure LLM agents on physics, chemistry, and a new neuroscience benchmark.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MDA's FORCEBENCH advantage may reflect prompt and design-menu pre-specification rather than Bayesian discovery; the M-open claim lacks a no-hint control.","rationale":"The reader's weakest_assumption identifies precisely the prompt vocabulary and design-menu pre-specification as load-bearing. My stress-test confirms this: the paper's own Section G.1.1 states the proposer is given the physical vocabulary, including the K1 Yukawa kernel, and Section C.5's control only varies the design menu, not the prompt. This is the single most consequential weakness because it directly undermines the M-open discovery claim for the flagship FORCEBENCH results, and it confounds the head-to-head comparison with the pure LLM agent. The reader's CONDITIONAL verdict already conditions on stronger baselines and code release; my analysis adds a specific missing control but does not change the verdict. I therefore recommend UNCHANGED, with the condition that the authors run the no-hint ablation to separate inference from pre-specification.","tokens_in":55409,"tokens_out":5664,"duration_ms":50534,"concrete_test":"Run the FORCEBENCH YUKAWA world with MDA's proposer prompt from Section G.1.1 edited to remove all screened/Helmholtz examples (delete 'K1(r/lam)/lam' and 'exp(-r/lam)/r'), keeping the same 13-launch menu and the same SMC/VoI pipeline. If MDA still recovers the K1(r/λ)/λ form within B=8 via the residual-triggered M-open expansion, the discovery claim holds; if it fails or degrades substantially, the headline result depends on the true family being named in the prompt.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that coupling an LLM proposer with SMC evidence and VoI design yields data-efficient mechanistic discovery, including in the M-open regime. The load-bearing weakness is that the empirical demonstration on FORCEBENCH appears to depend on the true mechanism family being explicitly named in the proposer prompt and on the design menu containing the decisive probes. Section G.1.1 shows the MDA proposer prompt names 'screened Poisson / Helmholtz -> ... K1(r/lam)/lam' among the allowed families, so the true Yukawa kernel is inside the initial hypothesis class, contradicting the claimed M-open status for that world. Table 5's 13-launch menu includes long-range probes (r0=5,6) that Section C.10 identifies as decisive for Yukawa, and the paper admits the menu was 'chosen by an LLM to cover the relevant dimensions.' The baseline LLM agent uses a different prompt and is not given this vocabulary, so the observed data-efficiency gap may be driven by prompt content rather than the SMC/VoI machinery. The paper's control in Section C.5 gives the LLM agent the same menu but not the same prompt, so it does not isolate the contribution of the Bayesian inference. Without a no-hint control, the 'discovery and design reinforce' claim is confounded by pre-specification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Model Discovery Agent (MDA), which couples an LLM-based proposer of mechanistic structures with sequential Monte Carlo (SMC) for parameter and structure posteriors, simulation-based likelihoods for intractable cases, and value-of-information (VoI) experiment design. MDA is designed to operate in the M-open regime: when the current hypothesis pool fits poorly, a predictive-check residual triggers the LLM to propose new mechanisms. The authors evaluate MDA on three benchmarks (FORCEBENCH physics, CHEMBENCH chemistry, and a new NEURONBENCH electrophysiology benchmark plus a stochastic extension), reporting substantially improved data efficiency over LLM-only baselines. The central claim is that MDAs coupling of Bayesian inference with LLM proposal yields faster held-out interventional prediction and more reliable mechanistic recovery than pure LLM agents, including in the M-open setting.","tokens_in":55807,"tokens_out":5645,"duration_ms":56263,"significance":"If the results hold, MDA is a valuable architecture for interactive scientific discovery, and NEURONBENCH and NEURONBENCHSTOCH are useful new benchmarks that add partial observability and stochastic dynamics to the existing evaluation landscape. The paper has several notable strengths: the VoI derivation in Eq. (5) is a standard, clearly stated Bayesian objective; the algorithms are given in sufficient pseudocode to be reproduced; the ablation study in Section D.3 and the base-model robustness study in Section C.8 are honest and informative; and the authors ship interactive apps and code for their new benchmark. However, the strongest empirical claims are currently supported by comparisons that do not fully isolate the contribution of the Bayesian machinery from prompt pre-specification, and the biology benchmark does not compare against a full LLM agent. The core idea is credible and the weaknesses are addressable, but the paper as written overstates the breadth of its M-open and SOTA claims.","major_comments":[{"comment":"The FORCEBENCH demonstration does not support the M-open discovery claim as cleanly as stated. The proposer prompt in Section G.1.1 explicitly names the screened-Poisson/Helmholtz family, including `K1(r/lam)/lam`, which is exactly the true Yukawa form, and Table 5 includes the long-range launches (r0=5,6) that Section C.10 identifies as decisive for Yukawa. Thus the correct mechanism is inside the initial hypothesis class and the decisive probe is inside the fixed design menu. The control in Section C.5 gives the LLM baseline the same menu but not the same physical vocabulary, so it does not isolate whether the data-efficiency gain comes from SMC/VoI or from prompt content. A no-hint control, in which the proposer prompt omits the true family and the menu omits the decisive probes, is needed before the paper can claim open-ended M-open discovery on FORCEBENCH.","section":"Section G.1.1, Table 5, Section C.10"},{"comment":"The NEURONBENCH evaluation does not compare MDA against a full LLM agent. Figure 4 compares the Bayes-forecaster with the in-context (ICL) forecaster; the LLM baselines in Section G.3.2 are acquisition and ICL-forecasting prompts, not an agent that proposes candidate mechanisms, simulates them, and forecasts from a fitted model. The paper's general claim that MDA is more data-efficient than pure LLM agents is therefore not established for the biology benchmark. In addition, Fig. 4 shows that within the Bayes-forecaster family, VoI and LLM acquisition perform similarly, and random acquisition is sometimes close; Section F.4 states that random designs are similar to, and arguably slightly better than, VoI designs in the stochastic setting. These results should be reported as a more nuanced boundary on the VoI contribution, and a full LLM-agent baseline should be added for NEURONBENCH.","section":"Section 4.3, Fig. 4, Section G.3.2"},{"comment":"The CHEMBENCH SOTA claim rests on thin evidence. The main comparison uses a 36-task subset with two seeds, and the 'reported' baseline row in Table 9 is based on a different, unreleased 36-task subset using gpt-4o-mini rather than the matched proposer model, so the headline comparison against the published number is not apples-to-apples. The head-to-head 'LLM-AUTOSCILAB(us)' row is fairer, but it is still only two seeds. Furthermore, key M-open and pool-size thresholds (tau_r=0.05, tau_p=0.9 in Table 2, and the Occam penalty lambda=2.5 in Section C.4) are set per benchmark with no sensitivity analysis, so it is unclear how robust the method is to these choices. Per-seed results, confidence intervals, and a sensitivity sweep over these thresholds are needed to support the data-efficiency and M-open claims.","section":"Section D.2, Table 9, Table 2"}],"minor_comments":[{"comment":"There is a typo in 'we summarize some our our experimental results' which should read 'some of our experimental results'.","section":"Section 4"},{"comment":"The text contains a typo: 'The summary statsitic we use' should be 'The summary statistic we use'.","section":"Section E.4"},{"comment":"The phrase 'candidate mechansims' should be 'candidate mechanisms'.","section":"Section E.3"},{"comment":"The caption and text refer to the 'LLM baseline' in NEURONBENCH, but the implemented baseline is an in-context forecaster; please use consistent terminology to avoid overstating the comparison.","section":"Section G.3.2"},{"comment":"The base model 'Fable 5' is used without a version or citation; please provide a reference or model card so that the robustness sweep can be reproduced.","section":"Section C.8"},{"comment":"The sentence 'Note only do we see that MDA is more sample efficient' has inverted word order; consider rewriting as 'Note that not only is MDA more sample efficient...'.","section":"Section F.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the core Bayesian-plus-LLM architecture is promising. The main risks are overclaiming the M-open regime and the data-efficiency advantage without fully matched baselines. If the authors add a no-hint proposer control on FORCEBENCH, a full LLM-agent baseline on NEURONBENCH, and more transparent CHEMBENCH statistics with threshold sensitivity, the paper could become acceptable. I do not see an unfixable internal error in the derivation; the concerns are about the strength of the empirical evidence rather than the method's soundness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Kevin, the thing to know: this is a serious, self-aware paper on LLM-assisted mechanistic discovery, and the central claim is plausible. But the headline FORCEBENCH advantage is partly an artifact of prompt and design-menu pre-specification. Worth reading and refereeing, with a control that is currently missing.\n\nWhat's genuinely new: MDA couples an LLM proposer with SMC evidence, VoI experiment design, and M-open residual-triggered expansion. The ablation in Section D.3 shows each piece helps. The NEURONBENCH stochastic extension—particle-filter likelihood, learned summaries, cost-aware auto-selection of the observation model with a spot-check—is substantial. The paper also does credit-earning things: it flags the unreliable explanation metric, discusses the fixed-sigma and Occam-paradox issues, and reports an honest boundary on the Fable-5 comparison.\n\nThe soft spot is real. On FORCEBENCH, the MDA proposer prompt (G.1.1) explicitly names screened Poisson/Helmholtz and the K1(r/lam)/lam form, and Table 5's 13-launch menu includes the long-range probes that C.10 identifies as decisive for Yukawa. The menu was chosen by an LLM to cover relevant dimensions. The C.5 control gives the LLM agent the same menu but not the same prompt vocabulary, so it does not isolate the Bayesian inference from the hint. A pure agent reaching 31% pass while MDA gets 93% may largely reflect who was told the family name. Without a no-hint control, the FORCEBENCH SOTA claim is confounded.\n\nThat said, the M-open case does not rest only on FORCEBENCH. NEURONBENCH is cleaner: the LLM sees only the phenotype, proposes candidate channels, and on Z-REBOUND initially omits the true current, with the residual check reopening the pool. The design menu still contains the decisive protocol, so the demonstration is partly pre-specified, but at least the mechanism is not named. CHEMBENCH states the baseline uses the same universal grammar; if that holds, it is the strongest evidence.\n\nOther soft spots: tau_r and tau_p are set per benchmark without sensitivity analysis (minor); the main MDA code is not released, only the NEURONBENCH benchmark (moderate); two seeds on CHEMBENCH (minor). The parsimony submission rule in C.8 is an honestly disclosed patch to the exact-form metric.\n\nThis is for anyone working on AI-for-science, automated discovery, or Bayesian experimental design. It deserves a serious referee; the confound needs a no-hint control and code release to settle. I'd send it to review.","headline":"Serious, unusually honest integration of LLM proposers with SMC/VoI, but the FORCEBENCH SOTA claim is confounded by prompt and design-menu pre-specification; still deserves refereeing.","tokens_in":56243,"tokens_out":4894,"would_cite":true,"duration_ms":43149,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian evidence layer turns an LLM proposer into a data-efficient discoverer of mechanistic laws, beating pure LLM agents even when the true mechanism is absent from the initial pool.","keywords":["mechanistic model discovery","Bayesian experiment design","value of information","sequential Monte Carlo","large language models","M-open model selection","simulation-based inference","interventional prediction"],"falsifier":"Run the Yukawa world with the design menu's longest launch kept inside the screening length ($r_0 \\le \\lambda$) and a proposer prompt that never mentions screened, Helmholtz, or Bessel kernels. The paper's own account predicts the true $K_1(r/\\lambda)/\\lambda$ law is never recovered — short-range probes leave the candidates indistinguishable and the residual check has no family to propose — so recovering the correct form under these restrictions would falsify the claimed necessity of discriminative design and supplied vocabulary.","tokens_in":2427,"feed_emoji":"🧪","tokens_out":3312,"duration_ms":146234,"temperature":0.7,"pith_summary":"Predicting the outcome of an intervention never performed requires a mechanistic causal model, and passive data cannot identify which mechanism generated what you saw — experiments are needed, and experiments are expensive. This paper claims that an LLM used only as a proposer of candidate mechanisms, paired with sequential Monte Carlo for posteriors and evidence, simulation-based inference for intractable likelihoods, and value-of-information maximization to choose the next intervention, identifies the true mechanism from only a handful of experiments, even in the M-open regime where the truth is not in the initial hypothesis pool. A residual-based predictive check detects when the current pool is inadequate and prompts the LLM to expand it, so discovery and experiment design reinforce each other. On physics, chemistry, and neuron-electrophysiology benchmarks, the resulting Model Discovery Agent beats pure LLM agents by a wide margin at matched budgets and matches the previous best agent's final accuracy with roughly a fifth of the experiments. If the claim holds, the bottleneck of automated scientific discovery shifts from proposing candidate forms to designing discriminative experiments and scoring evidence honestly.","feed_headline":"Bayesian design makes an LLM proposer 5x more data-efficient","feed_subtitle":"VoI design plus Bayesian evidence beats pure LLM agents on physics, chemistry, and neuron benchmarks.","key_machinery":"The load-bearing identity is the value-of-information reduction of Eq. (5): with Gaussian observation noise, the mutual information between the mechanism $M$ and the outcome of experiment $\\xi$ is $\\frac{1}{2}\\ln(1 + \\mathrm{Var}_{p(M|D)}[\\mu(\\xi)]/\\sigma^2)$, a monotone function of the between-class variance of the posterior predictive mean, so optimal design reduces to choosing the intervention on which the candidate mechanisms disagree most. The paper makes this concrete with the Yukawa world, where only a probe launched past the screening length $\\lambda$ splits the Bessel kernel from its power-law near-misses, and the Coulomb world, where only a charge intervention — a mechanism-level $do(a)$ — separates the true law from a curve fit that agrees with it on all unit-charge data. Around this identity sits the inference stack that feeds it: nested adaptive-tempered SMC computes each candidate's marginal likelihood (evidence) $Z_m = \\int p(D|m,\\theta)p(\\theta|m)\\,d\\theta$, whose automatic Occam factor trades complexity against fit; a model-level SMC maintains the posterior over structures, with LLM proposals conditioned on the whole pool's residuals; a residual-based predictive check triggers the M-open expansion of the hypothesis space; and an ESS-adaptive rule prunes near-duplicate pools. For the stochastic neuron benchmark the likelihood is intractable, so a bootstrap particle filter inside the tempered SMC estimates it, with a learned 1-D CNN summary statistic and synthetic Gaussian likelihood as a roughly $10^4\\times$ cheaper surrogate, guarded by a particle-filter spot-check that falls back to the exact filter when the cheap summary disagrees.","core_discovery":"Stated in the paper's own terms: discovery and experiment design reinforce each other. The value-of-information objective selects the intervention on which the candidate mechanisms' predictions disagree most; the outcome identifies the mechanism the LLM proposer introduced; and the identified mechanism improves interventional forecasts, shrinking residuals and exposing subtler predictive errors that trigger the next round of hypothesis expansion. The paper demonstrates this reinforcement on FORCEBENCH (force laws: MDA passes numerically in ~93% of 8-experiment runs versus ~31% for a budget-matched pure LLM agent, and reaches the previously reported state-of-the-art accuracy with ~5x fewer experiments), on CHEMBENCH (enzyme rate laws: ~56% symbolic accuracy within ~8 experiments versus ~42% for the previous state of the art at a budget of 60), and on NEURONBENCH (a new single-neuron electrophysiology benchmark with partial observability and, in its stochastic form, an intractable likelihood, where the model-based Bayes forecaster beats in-context LLM forecasting on every world). The recovered models are mechanistic — the screened Yukawa kernel $K_1(r/\\lambda)/\\lambda$, the enzyme-kinetic inhibition structure, the Hodgkin–Huxley channel composition — rather than numerically accurate but mechanistically meaningless expressions, and in the M-open miss cases (e.g., a neuron world whose novel current the LLM never proposed) the residual-triggered expansion still recovers the mechanism.","pith_inferences":["Editorial inference: the same VoI-as-disagreement identity should transfer to any domain where an LLM can propose mechanistic simulators rather than symbolic laws — dose-response models, policy simulators, engineering control laws — because the identity only requires simulating each candidate's predicted outcome, not closed-form expressions.","Editorial inference: the paper's base-model sweep, in which a much stronger LLM closes most of the gap and even leads on exact-form naming while scoring only ~11% numeric on Coulomb, suggests MDA's durable value is honest model selection — evidence over naming — and that 'can name the law' versus 'can compute with the law' will become the defining benchmark distinction as proposers improve.","Editorial inference: the paper's observation that random design matches VoI on the noisy stochastic benchmark hints that the marginal value of optimal design shrinks as observation noise dominates; repeat-aware VoI — spending budget re-running the discriminating experiment — is demonstrated on neurons and should generalize to any noisy experimental domain.","Editorial inference: the least automatic input left in the loop is the human-written prompt vocabulary (screened/Helmholtz kernels, enzyme-kinetics grammar, channel archetypes); a testable next step is to have the residual check propose that vocabulary itself, removing the last hand-engineered element from the discovery pipeline."],"forward_implications":["On FORCEBENCH, MDA's ~93% numeric pass rate with 8 one-per-round experiments versus ~31% for a budget-matched LLM agent implies that the bottleneck for pure LLM discovery is inference and experiment design, not capability.","In the M-open setting (e.g., a neuron world whose true current the LLM never proposed), residual-triggered expansion lets the agent recover mechanisms absent from the initial pool — a route to genuinely new knowledge rather than recall.","For interventional forecasting, a model-based Bayes forecaster built on the SMC posterior dominates in-context LLM forecasting by roughly two orders of magnitude on the noisy stochastic neuron benchmark.","Intractable-likelihood discovery is workable when the likelihood is replaced by a learned summary statistic plus synthetic likelihood, with cost-aware auto-selection of the observation model and a particle-filter spot-check as the safety anchor.","The benchmark's LLM-judged explanation metric is unreliable — flat and non-monotonic in the number of experiments — so held-out interventional forecast error, justified by the theorem that robust interventional prediction requires a causally correct model, is the more trustworthy evaluation."],"supporting_citations":[{"why":"Supplies FORCEBENCH/DiscoverPhysics, the pure LLM baseline agent, and the ~0.01 nMSE target that MDA matches with ~5x fewer experiments.","marker":"(Wiemann et al., 2026)"},{"why":"Supplies CHEMBENCH/ActiveSciBench-Chem and the LLM-AUTOSCILAB state-of-the-art baseline that MDA's symbolic-accuracy comparison must beat.","marker":"(Kabra et al., 2026)"},{"why":"Supplies SMC-S, the set-conditioned SMC over structures whose proposal kernel MDA adopts for its M-open expansion.","marker":"(Piriyakulkij et al., 2024)"},{"why":"Supplies ModelSMC, the vanilla SMC-over-LLM-structures framework whose 36% symbolic accuracy is MDA's ablation starting point.","marker":"(Wahl et al., 2026)"},{"why":"Supplies the mutual-information value-of-information objective that Eq. (5) reduces to between-class predictive variance.","marker":"(Lindley, 1956)"},{"why":"Supplies the theorem that robust interventional prediction implies a causal world model, justifying evaluation by held-out interventional forecast MSE.","marker":"(Richens & Everitt, 2024)"},{"why":"Supplies the synthetic-likelihood and simulation-based-inference methods used for the intractable neuron likelihood and learned summaries.","marker":"(Deistler et al., 2025)"},{"why":"Supplies the Langevin channel-noise SDE whose intractable transition density defines the stochastic neuron benchmark.","marker":"(Fox & Lu, 1994)"},{"why":"Supplies the Occam-factor argument that makes marginal likelihood the structure-selection score behind the complexity-fit tradeoff.","marker":"(MacKay, 1991)"}],"fun_headline_variants":["LLM proposes, Bayes designs: 5x fewer experiments","Discovery and design reinforce: MDA learns mechanisms fast","MDA: Bayesian design with LLM proposer beats pure LLM agents","Mechanistic models from few interventions: MDA sets SOTA on three benchmarks","LLM-assisted Bayesian experiment design: data-efficient world models"],"cache_read_input_tokens":58368,"weakest_assumption_plain":"MDA's gains rest on the pre-specified experiment menu and the proposer prompt containing, at least implicitly, the one intervention and the one mechanism family that can separate the true law from its near-misses — if no probe reaches past the screening length, or the prompt never names the screened-Bessel family, the value-of-information maximizer has no discriminating experiment to choose and the residual check has no vocabulary to invoke.","fun_headline_variants_meta":{"raw":{"variants":["LLM proposes, Bayes designs: 5x fewer experiments","Discovery and design reinforce: MDA learns mechanisms fast","MDA: Bayesian design with LLM proposer beats pure LLM agents","Mechanistic models from few interventions: MDA sets SOTA on three benchmarks","LLM-assisted Bayesian experiment design: data-efficient world models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000884,"raw_usage":{"total_tokens":3926,"prompt_tokens":1158,"completion_tokens":2768,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":774,"completion_tokens_details":{"reasoning_tokens":2680}},"tokens_in":774,"tokens_out":2768,"duration_ms":19454,"temperature":1.0,"reasoning_tokens":2680,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:31:09.066411+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Yukawa world with the design menu's longest launch kept inside the screening length ($r_0 \\le \\lambda$) and a proposer prompt that never mentions screened, Helmholtz, or Bessel kernels. The paper's own account predicts the true $K_1(r/\\lambda)/\\lambda$ law is never recovered — short-range probes leave the candidates indistinguishable and the residual check has no family to propose — so recovering the correct form under these restrictions would falsify the claimed necessity of discriminative design and supplied vocabulary.","supporting_citations":[],"review_version":1}