{"id":"cdd7e15b-49c8-4a5a-bf14-b29d697b0922","arxiv_id":"2505.21934","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A Bayesian calibration framework that combines empirical data with expert, theoretical, and qualitative constraints, demonstrated on coral, ecosystem, and biochemical models.","lead":"Researchers show how mathematical models can be calibrated using expert knowledge and theoretical expectations in addition to measured data. They combine Bayesian inference with approximate Bayesian computation and demonstrate it on coral recovery, ecosystem management, and biochemical adaptation examples.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation (14) treats non-empirical statements as exact constraints (I(ρ=0)); if expert knowledge is uncertain or misspecified, the constrained posterior is overconfident and may be worse than data-only inference.","rationale":"The reader's conditional verdict identifies exactly this weak point, and I agree with it. The formal contribution is Equation (14) plus the SMC algorithm, and these are coherent as far as they go: hard constraints are a legitimate limiting case, and the provided code is concrete support. The problem is the scope of the central claim. The paper says expert knowledge can be leveraged 'in a statistically rigorous manner,' but the rigor depends on treating elicited statements as exact. Expert knowledge is not exact; it is a source of information with uncertainty. Standard remedies exist—soft ABC tolerances, measurement-error models on summaries, or robust discrepancy functions—and none are implemented or tested here. The internal inconsistency in reported sample sizes (10,000 versus 100,000 parameter sets in Sections 2.3 and 3.1) weakens confidence, but it is secondary to the exactness issue. The proposed concrete test directly targets whether the central results and management implications would change under plausible departures from the exactness assumption. If they do not change, the paper's conclusions are substantially strengthened; if they do, the broad claims need qualification. This is consistent with the reader's CONDITIONAL verdict, so no verdict change is needed.","tokens_in":113,"tokens_out":2539,"duration_ms":42625,"concrete_test":"For Case studies 1 and 2, rerun the SMC algorithm with Equation (14) replaced by I(ρ ≤ ε) for ε corresponding to relaxing each elicited threshold by roughly 10–20% (e.g., early coral cover 10%→12%, recovery K−1%→K−2%, and allowing slightly positive eigenvalue real parts). Also run a version that places a prior distribution on the elicited bounds. Report changes in the predictive credible intervals in Figures 1, 3, and 4, and whether the fox-management conclusion in Figure 5 survives. If the posterior shifts materially or the management recommendation flips, the hard-constraint treatment is load-bearing; if the conclusions are stable, the exactness concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central target in Methods 4.3, Equation (14), combines a likelihood with the hard indicator constraint I(ρ(S_obs, S_sim(θ))=0). This treats each non-empirical statement as a logical certainty, not as an uncertain input. But the inputs are elicited summaries, which carry elicitation error, ambiguity, and potential bias. Under misspecification, Equation (14) is not a posterior under a reasonable prior over the non-empirical information; it is a point-mass constraint at the elicited value, producing overconfident predictions that can be confidently wrong. The paper acknowledges this in Discussion 3.3 ('treats information ... as generally irrefutable') but does not incorporate uncertainty into the discrepancy function or test sensitivity to the chosen thresholds. The demonstrations in Sections 2.1 and 2.2 use hand-selected thresholds (e.g., coral cover below 10% at 5 years, recovery within 1% of K, eigenvalue real parts negative) without showing that nearby, equally plausible expert statements give similar posteriors. Since the central claim is that the method guides models toward more realistic dynamics, the missing load-bearing evidence is robustness to plausible variation in the non-empirical input.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian calibration framework that combines empirical data, modeled through a likelihood, with non-empirical information (expert knowledge, qualitative observations, scientific theory) encoded as ABC-style discrepancy constraints. The target distribution, Equation (14), multiplies the usual posterior by an indicator enforcing zero discrepancy for the non-empirical constraints. The authors develop an SMC algorithm that simultaneously anneals the likelihood and reduces the discrepancy threshold, and they demonstrate the approach on three case studies: logistic coral growth with expert recovery-time constraints, a four-species ecosystem model with coexistence/stability constraints plus time-series data, and a biochemical adaptation network with sensitivity/precision constraints. The paper claims this framework yields more realistic dynamics and more informed predictions than data-only calibration, and it also positions the method as a simulation-based prior elicitation tool.","tokens_in":29402,"tokens_out":8840,"duration_ms":83999,"significance":"If the central claims hold, the paper offers a coherent statistical way to fold non-empirical knowledge into model calibration, with broad applicability in ecology, biology, and medicine. The strengths include a clearly specified target distribution, a complete SMC algorithm, public code, and case studies that generate testable findings (e.g., the K3/K4 relationship in Case study 3). The simulation-based prior elicitation viewpoint is a useful contribution. However, the case studies are mostly synthetic demonstrations, and the hard-constraint formulation, which the paper itself acknowledges as a limitation in Discussion 3.3, is not accompanied by any sensitivity analysis or uncertainty quantification for the elicited constraints. This gap, together with internal numerical inconsistencies in Case study 3, tempers the strength of the main claims.","major_comments":[{"comment":"The target distribution treats non-empirical statements as exact constraints via the indicator I(ρ=0). The paper acknowledges in Discussion 3.3 that elicited knowledge can contain uncertainties and biases, but it does not incorporate this uncertainty into the framework or test sensitivity to the chosen constraint thresholds. The case studies use hand-selected constraints (e.g., y(5)≤10%, y(50)≥K−1%, S>1, P>10) without demonstrating that nearby, equally plausible expert statements lead to similar posterior predictions. Since the central claim is that non-empirical information guides models toward more realistic dynamics, the missing load-bearing evidence is robustness to plausible variation in the non-empirical inputs.","section":"§4.3, Eq. (14) and Discussion 3.3"},{"comment":"The reported sample size for the biochemical case study is internally inconsistent. Section 2.3.1 and Supplementary S.1.3 state that the SMC-ABC ensemble consists of 10,000 parameter sets, while Discussion 3.1 claims \"we have obtained 10^5 parameter sets all of which are capable of adaptation.\" Additionally, Section 2.3 reports that Ma et al. (2009) found 1 in 10,000 (0.01%) adaptive parameterisations, whereas Discussion 3.1 says Ma et al. and Jeynes-Smith and Araujo (2023) tested 10^5 parameter sets with ~1% producing adaptation. These contradictions affect the paper's claims about computational efficiency and ensemble size and need to be resolved.","section":"Section 2.3.1 vs. Discussion 3.1"},{"comment":"The demonstrations partly measure success by construction. The posterior is constrained to satisfy the non-empirical conditions, and then those same conditions are presented as the outcome (e.g., Figure 1B shows the posterior meets the recovery constraints; Figure 6E shows S>1 and P>10). The non-circular evidence is the management-prediction change in Section 2.2.2 (Figure 5) and the K3/K4 bivariate relationship in Section 2.3.2 (Figure 7), which are not restatements of the target constraints. To support the claim of \"more realistic dynamics,\" the paper should either include an out-of-sample or additional-behavior validation or explicitly distinguish definitional success from independent predictive benefit.","section":"Sections 2.1 and 2.3"}],"minor_comments":[{"comment":"The caption says \"associated with the logistic growth model\" but the table describes the ecosystem population model; it should refer to Section 2.2.","section":"Table S3 caption"},{"comment":"The text states that 16 parameters require calibration, but Table S3 lists 20 parameters (4 initial populations, 4 growth rates, 4 intra-specific interactions, 4 negative interactions, and 4 positive interactions); the count should be corrected.","section":"Section S.1.2 and Table S3"},{"comment":"There is a typographical error in the exponent: the term should be (y_sim,i(θ)−y_obs,i)^2, and a closing parenthesis is missing in the display.","section":"Equation (13)"},{"comment":"The phrase \"we have obtained 105 parameter sets\" should read \"10^5 parameter sets\"; the superscript formatting appears to be lost.","section":"Discussion 3.1"},{"comment":"The caption contains a typo: \"Michaelis-Mentin\" should be \"Michaelis-Menten.\"","section":"Figure S6 caption"},{"comment":"The algorithm description should clarify that the target discrepancy threshold ε can be zero only for range-based constraints with a positive prior measure; for exact-value constraints, a tolerance would be needed, since a hard zero constraint on a continuous summary statistic generally has zero posterior mass.","section":"Algorithm S1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable methodology proposal, and the code release is commendable. The main weaknesses are (1) the hard-constraint treatment of non-empirical information is acknowledged but not addressed with sensitivity analysis or uncertainty modeling, and (2) the internal inconsistencies in Case study 3 sample sizes and acceptance rates undercut the computational-efficiency claims. The case studies are synthetic, so the \"more realistic dynamics\" language should be softened or supported by external validation. I would encourage the editor to request a revision that addresses these points, not a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'd say this paper is worth your time. The genuinely new piece is Equation (14): a single target distribution that multiplies a standard likelihood by an ABC-style indicator constraint for non-empirical summaries, plus an SMC sampler (Algorithm S1) that anneals both simultaneously. That's a practical, concrete contribution, not just a slogan. The three case studies are well chosen and illustrate different flavors of non-empirical information—expert bounds, ecological theory, and a qualitative systems-level property. The code is available, the sampler is described in enough detail to reproduce, and the management-relevant consequence in Case study 2 (fox culling has a much stronger effect when coexistence/stability constraints are imposed) is a genuinely striking demonstration. I also credit the K3/K4 finding in Case study 3 as an independent numerical check on an analytic condition from the literature. The soft spots are real but not fatal. The biggest one: the framework treats every non-empirical statement as an exact logical constraint via I(ρ=0), with no uncertainty or sensitivity analysis around the elicited thresholds. The paper acknowledges this in Section 3.3 but doesn't act on it. That matters because the central claim is that constraints make predictions more realistic; if an expert constraint is misspecified, the constrained posterior will be confidently wrong. A simple robustness check—perturb each threshold and see how much the posterior moves—would have addressed a large part of this. Relatedly, the demonstrations in Sections 2.1 and 2.2 are partly circular: the constraints are enforced, then displayed as posterior behavior. The independent content is the management outcome and the K3/K4 check, not the posterior satisfying the constraints. There is also a concrete internal inconsistency that needs fixing: Section 2.3 says an ensemble of 10,000 parameter sets was produced and that 1 in 10,000 (0.01%) of prior samples satisfy the adaptation criteria, but Discussion 3.1 says we have obtained 10^5 parameter sets and the supplementary says 0.18% of prior draws satisfy the criteria. Those numbers don't line up, and a careful reader will notice. Who is this for: applied statisticians and mechanistic modelers in ecology, biology, and related fields who calibrate models with scarce data and have access to domain knowledge. It's not a foundational statistical result, but it's a solid, reproducible extension of ABC. I would send it to a serious referee, with the expectation of major revision: add sensitivity analysis, soften the overclaims about more realistic dynamics without external validation, and fix the sample-size discrepancies. After that, it would be a genuinely useful methods paper.","headline":"A useful, moderately novel ABC framework for imposing non-empirical constraints in calibration, with real potential for data-limited modeling; the main weaknesses are the hard-constraint treatment of expert knowledge and a few internal inconsistencies.","tokens_in":732,"tokens_out":944,"would_cite":true,"duration_ms":29338,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","65C05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that expert knowledge, scientific theory, and qualitative observations can be encoded as hard constraints in an approximate Bayesian target, so that models can be calibrated from knowledge alone or from knowledge plus…","keywords":["approximate Bayesian computation","model calibration","expert knowledge elicitation","non-empirical information","prior elicitation","sequential Monte Carlo","discrepancy function","Bayesian inference"],"falsifier":"Generate data from a known simulation model, then calibrate with an intentionally false expert constraint (for example, a recovery-time bound the generating dynamics violates) together with a large dataset; if the Equation (14) posterior concentrates on parameter values the data-only likelihood rules out, the zero-discrepancy assumption has demonstrably propagated expert error into the predictions.","tokens_in":28964,"feed_emoji":"🧠","tokens_out":10365,"duration_ms":89166,"temperature":0.7,"pith_summary":"The paper sets out to show that model calibration does not have to be limited to empirical measurements. It builds an approximate Bayesian framework in which non-empirical information, including expert beliefs, ecological theory, and qualitative observations of system behavior, enters as directly specified summary statistics with hard constraints, and can be combined with any available data in a single target distribution. Three case studies (coral reef recovery, a four-species ecosystem under management, and biochemical adaptation) demonstrate that these constraints remove implausible predictions, keep long-term model behavior consistent with established theory, and can change management-relevant conclusions. If the framework is right, models in data-scarce settings gain information from knowledge scientists already hold, rather than waiting for measurements that may be costly or impossible to obtain.","feed_headline":"Calibrate models with expert knowledge when data is scarce","feed_subtitle":"A new Bayesian framework turns qualitative theory and expert opinion into exact constraints that sharpen predictions","key_machinery":"The load-bearing object is Equation (14), a target distribution that multiplies the prior by a likelihood for empirical data and by an indicator requiring the non-empirical discrepancy to vanish exactly: $$\\pi(\\$\\theta$ \\mid S_{\\text{obs}}, y_{\\text{obs}}) \\propto \\pi(\\$\\theta$)\\, f(y_{\\text{obs}} \\mid \\$\\theta$) \\int I\\big(\\rho(S_{\\text{obs}}, S(y_{\\text{sim}}(\\$\\theta$))) = 0\\big)\\, dy_{\\text{sim}}.$$ Non-empirical statements are treated as directly observed summary statistics with range constraints (for example $S_{\\text{obs}} \\le a$), enforced by one-sided discrepancy functions such as $\\rho = \\max(b - a, 0)$ rather than the two-sided distances used in ordinary ABC, and acceptance requires zero discrepancy instead of a tolerance. A bespoke sequential Monte Carlo algorithm simultaneously anneals the likelihood and shrinks the discrepancy threshold to zero, which makes the combined target computationally feasible, and the same construction doubles as a simulation-based prior elicitation method because it automatically yields joint priors that respect parameter interdependencies.","core_discovery":"The paper's central claim is that anything we know about how a system behaves, such as an expert's bound on coral recovery, the theoretical requirement that all species coexist at a stable equilibrium, or the observation that a biochemical network must both respond to and adapt to stimulation, can be converted into a summary statistic plus a discrepancy function and enforced inside an approximate Bayesian target. When empirical data exist, a likelihood term multiplies the same target, producing a posterior that satisfies the data and the knowledge at once. The case studies show this is more than a formalism: data alone still admits coral trajectories that never recover, whereas data plus expert constraints match both sources; coexistence and stability constraints reverse the predicted effect of fox control on small mammals; and the efficient sampler finds adaptation-capable biochemical parameter sets outside the region the literature claimed was necessary.","pith_inferences":["A natural extension the paper does not build: treat expert statements as noisy constraints with a tolerance or error model, so uncertainty in the knowledge itself propagates into the posterior; the paper flags this need but leaves it as future work.","A stress test the paper does not run: inject a deliberately wrong expert constraint alongside a rich dataset, and compare the constrained posterior with the data-only posterior to quantify how much overconfidence the exactness assumption can induce.","The $K_3$/$K_4$ result hints that the sampler's efficiency could turn calibration into an auditing tool for other literature-derived necessary conditions in biology and ecology.","Because empirical and non-empirical sources share one target, the framework could support sequential updating in which expert constraints are added, revised, or retracted as new data arrive."],"forward_implications":["Models can be calibrated in the complete absence of empirical data, using only expert statements or qualitative observations, as in the coral growth and biochemical adaptation case studies.","In sparse-data settings, adding non-empirical constraints eliminates predictions that conflict with established knowledge; the coral example shows that data alone still permits trajectories with virtually no recovery, which the constrained posterior rules out.","Theory can change decisions: enforcing coexistence and stability in the ecosystem model reverses the predicted benefit of fox population control for small mammals.","A more efficient search of the constrained parameter space can test claimed necessary conditions, since the biochemical case finds adaptation-capable parameter sets with $K_3 \\ge 1$ or $K_4 \\ge 1$, contrary to the conditions reported in the earlier literature.","The framework provides a general simulation-based prior elicitation method that builds informed joint priors with parameter interdependencies captured automatically."],"supporting_citations":[{"why":"supplies the ABC framework of summary statistics and discrepancy functions that the method adapts to non-empirical constraints.","marker":"Sisson et al. (2018)"},{"why":"the SMC-ABC algorithm the paper's combined sampler extends by simultaneously annealing likelihood and discrepancy.","marker":"Drovandi and Pettitt (2011)"},{"why":"defines the sensitivity and precision criteria used as the biochemical non-empirical constraints, plus the $K_3$, $K_4$ conditions the paper reassesses.","marker":"Ma et al. (2009)"},{"why":"provides the ABC methodological grounding for treating the discrepancy as an approximate likelihood.","marker":"Beaumont (2019)"},{"why":"an ad-hoc ecosystem calibration with equilibrium constraints and data that the formal combined target improves upon.","marker":"Baker et al. (2019)"},{"why":"supplies the feasibility and coexistence definition used as the equilibrium constraint in the ecosystem case study.","marker":"Grilli et al. (2017)"},{"why":"frames the prior elicitation challenge that the method addresses as a simulation-based elicitation procedure.","marker":"Mikkola et al. (2024)"},{"why":"underpins the discussion that expert knowledge carries uncertainty and bias, motivating the stated limitations.","marker":"O'Hagan (2019)"}],"fun_headline_variants":["Beyond data: Bayesian calibration with expert knowledge","Use expert insight when data is scarce for model fitting","Bayesian framework turns theory and opinion into constraints","Expert knowledge sharpens Bayesian model predictions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework treats every non-empirical statement as an exactly true constraint, requiring zero discrepancy, so if the expert knowledge is uncertain, biased, or poorly captured by the chosen summary statistics, the constrained posterior becomes overconfident and can be worse than ignoring the knowledge.","fun_headline_variants_meta":{"raw":{"variants":["Beyond data: Bayesian calibration with expert knowledge","Use expert insight when data is scarce for model fitting","Bayesian framework turns theory and opinion into constraints","Expert knowledge sharpens Bayesian model predictions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1194,"prompt_tokens":825,"completion_tokens":369,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":441,"tokens_out":369,"duration_ms":3983,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:19:55.188649+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate data from a known simulation model, then calibrate with an intentionally false expert constraint (for example, a recovery-time bound the generating dynamics violates) together with a large dataset; if the Equation (14) posterior concentrates on parameter values the data-only likelihood rules out, the zero-discrepancy assumption has demonstrably propagated expert error into the predictions.","supporting_citations":[],"review_version":1}