{"id":"148ca982-5c8e-46d3-823d-ac4d0a0d2d11","arxiv_id":"2608.06998","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A perspective contends that a model's scientific value comes from constrained parameters and analyzed dynamics, not from component count; whole-cell simulations currently reproduce data without delivering mechanistic insight.","lead":"A perspective argues that in cell biology modeling, understanding depends on how many free parameters a model has relative to available experimental data, not on how many components it simulates. It shows with a cell-cycle example that a detailed model fits nearly all parameter space while a minimal model constrained by a direct measurement fits only 2%, and calls for model hierarchies and dynamical analysis instead of bigger simulations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The leap from a 13-parameter oscillator module to whole-cell models rests on an unquantified scaling assumption; the paper's own Fig. 2 does not measure whether sloppiness worsens with model size.","rationale":"The reader's verdict of CONDITIONAL identifies the same load-bearing assumption: the paper generalizes from a small, well-characterized oscillator module to whole-cell models without a quantitative scaling argument. My stress-test agrees that this is the weakest point in the central claim. The paper's own disclaimers—that the claim is 'contingent, not principled'—partly mitigate the concern, and the philosophical core, that parameter-to-constraint ratio and interpretability matter more than component count, is defensible as a perspective. However, the strongest applied conclusion, that whole-cell models currently offer little mechanistic understanding, depends on sloppiness becoming more severe in larger models. That is an empirical claim about how the effective dimensionality of parameter space grows with model size, and it is not established by Figure 2. The figure's 85% versus 2% comparison is useful but is prior-dependent and does not by itself provide a scaling law. The reader's CONDITIONAL verdict remains appropriate: the paper is a legitimate contribution with a testable empirical gap. My proposed check, running a sloppiness analysis on an actual whole-cell model or on a nested family of models, would directly settle whether the scaling assumption holds. Until then, the concern is real but not fatal to the perspective's central argument, which is explicitly framed as contingent and as a call for further analysis rather than a proof that whole-cell models can never yield understanding.","tokens_in":10349,"tokens_out":6161,"duration_ms":68853,"concrete_test":"Take a published whole-cell-scale model (Karr et al. 2012 M. genitalium, or the Thornburg et al. 2026 syn3A model if code is available) and compute a local sensitivity or Fisher-information analysis for key outputs (division time, growth rate, a set of gene-expression levels) with respect to the full set of free parameters, using the model's existing simulation code or a Gaussian-process surrogate. Compute the participation ratio or the number of eigenvalues of the Fisher information matrix above a threshold (e.g., 1% of the largest eigenvalue). If this effective dimension is of order 10^1 to 10^2 while the nominal parameter count is 10^3 to 10^4, sloppiness does not monotonically scale with model size and the paper's central generalization fails. If, instead, the effective dimension grows proportionally to parameter count, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central, load-bearing step is the extrapolation from Figure 2 to whole-cell models. The figure compares a 5-variable/13-parameter mass-action model with a 2-variable/3-parameter minimal model of a single Xenopus embryonic oscillator and shows that, with only a period constraint, the mass-action model fits 85% of a scanned parameter plane while the minimal model fits about 2%. Even granting that result, it shows only that in this pair, replacing a fixed measured functional response with mass-action kinetics plus 10 free parameters makes one observable less informative. It does not show that sloppiness scales with model size. The paper's only scaling argument is the sentence: 'If parameter sloppiness is already severe in a two-variable approximation of a single embryonic oscillator module, it is difficult to see how it could be less severe in a model that is larger by several orders of magnitude.' This is an appeal to intuition. The systems-biology sloppiness literature, including the cited Gutenkunst et al., shows that many models have an effective parameter dimension much smaller than their nominal count; whether the effective dimension grows with model size is an empirical question. If whole-cell models are effectively low-dimensional in their stiff directions—if, say, a few hundred parameter combinations control the outputs—then the ratio of free parameters to constraints may not be the binding obstacle, and the conclusion that whole-cell models 'offer little by way of biological understanding' is unsupported. The 85% and 2% figures also depend on unspecified prior ranges and on 5000 random draws per point, so they are not a quantitative scaling law.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This perspective argues that model complexity in cell biology can obscure rather than enable understanding. The authors' central claim is that what matters for insight is not the number of components in a model but (1) the ratio of free parameters to available experimental constraints and (2) whether the relationship between parameters and behavior can be grasped. They illustrate this with a Xenopus cell-cycle oscillator: a 13-parameter mass-action model is consistent with the observed period across ~85% of a scanned parameter plane, while a minimal 3-parameter model built around a measured bistable functional response is consistent with only ~2% of the same plane. From this they conclude that whole-cell models, despite being impressive technical achievements and sources of higher-order evidence, have not yet delivered mechanistic understanding, and they propose explicit model hierarchies, dynamical analysis, and control-parameter identification as the route to such understanding. The argument is explicitly framed as contingent rather than principled, and the numerical code for Figure 2 is promised in a public repository.","tokens_in":10716,"tokens_out":4899,"duration_ms":51981,"significance":"If the argument holds, the paper would be a valuable corrective to the widespread assumption that adding biological detail automatically improves explanatory power. Its strengths are the clarity of the conceptual framework, the concrete and reproducible numerical illustration, the explicit acknowledgment that the whole-cell claim is contingent, and the meaningful engagement with the philosophical literature on minimal models and understanding. The paper also usefully distinguishes predictive success from causal understanding and connects this distinction to machine-learning models. As a perspective, its contribution is conceptual rather than technical, but the Figure 2 example and the proposed methodological agenda could influence modeling practice in systems and cell biology.","major_comments":[{"comment":"The 85% versus 2% acceptance fractions are point estimates from 5000 random draws, reported without confidence intervals or a sensitivity analysis over the scan ranges for the synthesis and degradation rates and over the 40±1 min acceptance window. As reported, the numbers depend on unspecified prior choices for the 11 internal parameters and on the random-draw protocol. The qualitative conclusion would survive, but the quantitative-looking claim should be accompanied by, at minimum, binomial confidence intervals and a brief statement of how the scan ranges were chosen.","section":"The identifiability problem: a concrete illustration (Figure 2)"},{"comment":"The extrapolation from a single 13-parameter embryonic oscillator module to whole-cell models is the load-bearing step for the paper's critique of whole-cell modeling, yet it rests on an appeal to intuition: 'If parameter sloppiness is already severe in a two-variable approximation of a single embryonic oscillator module, it is difficult to see how it could be less severe in a model that is larger by several orders of magnitude.' No quantitative scaling argument or identifiability measurement on an actual whole-cell model is provided. The explicit hedge that the claim is contingent softens this, but the conclusions still lean on the scaling sentence. The authors should either supply a quantitative analysis, draw more carefully on cited work such as Babtie and Stumpf (Ref. [20]) that addresses whole-cell parameter problems, or rephrase the whole-cell conclusion as an explicitly unverified conjecture about a class of models that has not yet been analyzed.","section":"Scale, parameters, and the conditions for insight; The identifiability problem: a concrete illustration"},{"comment":"The rebuttal to the objection that the minimal model was simply given more information is not demonstrated. The authors assert that fitting the mass-action model to the measured bistability curve would merely select another flat, degenerate region of parameter space, but no simulation or calculation supports this claim. Because the fairness of the Figure 2 comparison is central to the identifiability argument, the rebuttal should be backed by a numerical experiment: for example, fit the mass-action model to the measured functional response and show that the posterior over the 11 internal parameters remains flat or that the period constraint still admits a large fraction of the parameter space. Without this, the example remains open to the simpler reading that the minimal model is more identifiable because it incorporates an additional experimental observable by construction.","section":"The identifiability problem: a concrete illustration, paragraph beginning 'It is worth being explicit about why this…"}],"minor_comments":[{"comment":"There are missing spaces in the abstract, for example 'handfulofparameters' and 'towhole-cellsimulations'; please correct these typographical errors throughout the manuscript.","section":"Abstract"},{"comment":"The placement of 'Machine learning models' in the underconstrained region is not explained in the text. Since some machine-learning models are trained on very large datasets, the authors should add a sentence justifying this placement, for example by noting that the relevant constraint-to-parameter ratio refers to mechanistic parameters rather than training examples.","section":"Figure 1"},{"comment":"The central term 'grasped' in the two-factor account of understanding is not defined. The later connection to interpretability helps, but a brief definition or pointer to the interpretability literature (Refs. [10-12]) would make the criterion more precise.","section":"Scale, parameters, and the conditions for insight"},{"comment":"The models and measurements used in Figure 2 come from the authors' own prior work (Refs. [24,25]). This is transparent through the citations, but the text should state it explicitly in the figure description so that readers can assess potential confirmation bias.","section":"References [24,25] and Figure 2"}],"recommendation":"major_revision","confidential_remarks":"This is a perspective article, and its value depends on whether the journal sees conceptual argumentation as within scope. The main risk is overreach from one small module to whole-cell models; the authors have hedged appropriately, but they could strengthen the paper considerably by replacing the intuitive scaling sentence with either a cautious rephrasing or a quantitative analysis of a larger model. The code availability statement is a plus, and the Figure 2 illustration is effective even if not a rigorous identifiability study. I recommend major revision rather than rejection because the central conceptual claims are defensible and the missing support is additive rather than fatal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a well-argued perspective, not a research paper, and its central thesis is sound. The point that identifiability and interpretability matter more than component count is a useful corrective to the enthusiasm for whole-cell models. The novel contribution is the framing of sloppiness as both friend of prediction and obstacle to mechanistic understanding, plus a concrete comparison between a 13-parameter mass-action model and a 3-parameter minimal model of the same Xenopus embryonic oscillator. With only a period constraint, the mass-action model fits about 85% of the scanned parameter space while the minimal model fits about 2%. It's a memorable illustration of how adding mechanism without adding constraints can make a system harder to characterize.\n\nThe paper is honest. The authors explicitly say their critique of whole-cell models is contingent, not principled. They give whole-cell models credit as higher-order evidence and as tools for assembling knowledge. They also deal with the obvious objection that the minimal model is built to include the bistability measurement by explaining why the mass-action model cannot simply absorb that constraint. The proposed practices—model hierarchies and dynamical analysis—are concrete and sensible, not generic hand-waving.\n\nThe soft spot, which the reader and stress-test both flag, is the scaling step. The jump from a two-variable oscillator module to whole-cell models rests on one sentence: 'if sloppiness is already severe here, it's hard to see how it could be less severe in a model larger by orders of magnitude.' That's an appeal to intuition, not an argument. The systems-biology sloppiness literature shows many models are effectively low-dimensional; whether whole-cell models are effectively low-dimensional is an empirical question that they don't address. Also, the 85% and 2% figures are point estimates from 5000 random draws without confidence intervals or a clear statement of prior ranges. So the illustration is suggestive, not a quantitative scaling law.\n\nNone of this breaks the paper. The central argument doesn't depend on the scaling claim holding as a general law. It's a perspective, and it does its job: it reframes a practical problem in a clear way and points to what would be needed to make complex models more explanatory. I'd send it to peer review—it deserves referee time, even though the authors should soften the scaling claim or back it with an actual identifiability analysis on a larger model. For someone in modeling or philosophy of modeling, this is a good discussion paper to have on hand.","headline":"A well-argued perspective with a concrete illustration; the core message holds, but the extrapolation from a small oscillator module to whole-cell models is asserted more than shown.","tokens_in":11224,"tokens_out":3400,"would_cite":true,"duration_ms":31222,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A cell model generates understanding only when its free parameters are pinned down by experimental constraints and one can see why it behaves as it does; reproducing an observation alone is not enough.","keywords":["cellular modeling","parameter identifiability","parameter sloppiness","whole-cell models","model hierarchies","bifurcation analysis","cell cycle oscillator","scientific understanding"],"falsifier":"Run a systematic parameter scan or profile-likelihood analysis on a publicly available whole-cell model such as JCVI-syn3A, fixing an experimentally measured output like division time. If the model matches the output across most of its parameter space, the paper's claim is supported; if only a small fraction is consistent, its key assumption that sloppiness scales with model size fails.","tokens_in":10112,"feed_emoji":"🧬","tokens_out":8773,"duration_ms":84687,"temperature":0.7,"pith_summary":"A computational model earns scientific understanding only when it does something beyond reproducing the observations used to build it—predicting a new result, exposing an unsuspected coupling, or failing in a way that reveals a necessary component. The paper argues that the decisive feature is not model size or component count but identifiability: free parameters must be few enough relative to experimental constraints to be pinned down, and the relation between parameters and behavior must be graspable. A concrete comparison of two cell-cycle oscillator models shows a detailed 13-parameter version consistent with the observed period over about 85% of the scanned parameter space, while a minimal 3-parameter version built on a measured bistable response is consistent with only about 2%. On this view, whole-cell simulations are substantial technical achievements and valuable hypothesis generators, but they have not yet delivered mechanistic explanation because their parameter spaces have not been probed and reduced.","feed_headline":"More parameters can mean less understanding in cell models","feed_subtitle":"A detailed cell-cycle model fits 85% of parameter space; a minimal one fits only 2%, and that gap reveals the mechanism.","key_machinery":"The load-bearing object is parameter sloppiness, the near-universal phenomenon in systems-biology models in which a small number of stiff parameter combinations control predictions while most parameters can vary by orders of magnitude without affecting observable outputs. The paper's concrete demonstration is the comparison between the mass-action cell-cycle model and its minimal functional-response counterpart: in the detailed model the bistable switch is an emergent many-to-one function of 11 internal kinetic parameters, so fitting the oscillation period leaves the rates almost entirely unconstrained; in the minimal model the same switch enters as a fixed, directly measured input-output curve, so the remaining 3 parameters are informative. Sloppiness is what does the work of the argument: it explains why prediction can be robust while causal knowledge stays out of reach, and why adding mechanistic detail without adding experimental constraints makes a model harder, not easier, to learn from.","core_discovery":"The paper's central claim is that reproducing a known observation establishes consistency but not mechanism, and that a model becomes explanatory only when it makes a novel prediction, uncovers an unexpected coupling, or fails in a way that identifies a missing component. What determines whether a model can do any of these is not how many molecules or spatial dimensions it includes but the ratio of free parameters to independent experimental constraints and whether one can see why the model produces its behavior. The authors demonstrate the point with two models of the Xenopus embryonic cell cycle oscillator: a mass-action model with 5 variables and 13 free parameters reproduces the roughly 40-minute period across about 85% of the scanned synthesis-degradation parameter space because its 11 internal rates can be adjusted in flat, compensating directions, whereas a minimal model with 2 variables and 3 free parameters—built by fixing the measured bistable APC/C response—is consistent with only about 2% of the same space. From this they conclude that whole-cell models of Mycoplasma genitalium and JCVI-syn3A, however comprehensive, currently offer little mechanistic understanding; what they provide instead is higher-order evidence: an assembled, mutually consistent inventory of empirical knowledge that can guide hypothesis generation but not yet causal explanation.","pith_inferences":["A testable extension is that every mechanistic model should report an effective parameter dimension or identifiability measure alongside goodness of fit, allowing progress to be judged by how tightly parameters are constrained rather than by how many components are included.","If the argument generalizes, systematically reducing an existing whole-cell model should reveal that a small number of stiff parameter combinations suffice to reproduce any given phenotype, with most parameters free to float without changing the output.","The same logic predicts that 'virtual cell' foundation models will support causal intervention only if their training enforces identifiability constraints; interpolation accuracy alone will not ground mechanistic inference.","One could push the paper's logic further and propose a minimal whole-cell model in which each module's detailed kinetics is replaced by an experimentally measured input-output response function, exactly as the cell-cycle example replaces the PP2A-ENSA-GWL subnetwork."],"forward_implications":["Large agent-based models of cytoskeletal dynamics and vertex models of tissue mechanics can remain genuinely explanatory at scale, provided a small number of physically grounded parameters keep the parameter space explorable.","Whole-cell simulations should be treated as structured inventory-and-hypothesis tools rather than end-point explanations until they are subjected to systematic reduction and dynamical analysis.","A mechanistic claim from any model with many unconstrained parameters requires comparison with simpler representations; a good fit alone does not identify the mechanism.","Machine-learned and neural-differential-equation models inherit the same identifiability constraint: increasing expressive capacity without increasing experimental constraints worsens the parameter-data imbalance.","The two analytical practices—building explicit model hierarchies and mapping control parameters, bifurcations, and robustness boundaries—are the route that converts simulation into understanding."],"supporting_citations":[{"why":"The 2012 Mycoplasma genitalium whole-cell model is the principal example of a comprehensive simulation whose parameter space the paper argues is too large to probe systematically.","marker":"[6]"},{"why":"The spatially resolved JCVI-syn3A whole-cell model is the second major example of a large-scale simulation treated as a technical achievement without established mechanistic understanding.","marker":"[7]"},{"why":"Establishes that sloppy parameter sensitivities are near-universal in systems biology models, the phenomenon underlying the paper's identifiability argument.","marker":"[8]"},{"why":"Supplies the dynamical-analysis toolkit of bifurcation, sensitivity, and phase-plane methods that the paper argues must be applied to large-scale models.","marker":"[10]"},{"why":"Shows how agent-based cytoskeletal models self-organize from a few local rules, providing the template for parameter-lean large-scale models that remain exploratory.","marker":"[14]"},{"why":"Frames computer simulations as higher-order evidence, which the paper uses to credit whole-cell models with assembling and reconciling empirical knowledge without claiming explanation.","marker":"[19]"},{"why":"Supplies the adequacy-for-purpose criterion used to separate mechanistic explanation from other legitimate jobs a whole-cell model can do.","marker":"[21]"},{"why":"Provides the two cell-cycle oscillator models whose identifiability comparison is the paper's concrete, quantitative illustration.","marker":"[24]"},{"why":"Supplies the bistability measurements that fix the functional response in the minimal model, the source of its far tighter parameter constraints.","marker":"[25]"},{"why":"Supports the minimal-model explanatory account by showing that models can explain through universality classes, so omitted details can be irrelevant.","marker":"[29]"}],"fun_headline_variants":["Complex cell models often obscure, not clarify, biology","Fewer parameters can unlock real biological insight","Cell cycle model proves: complexity hides mechanism","Why whole-cell simulations may not advance understanding","Parameter count matters more than model size in biology"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The critique of whole-cell models assumes that parameter sloppiness becomes at least as severe in models thousands of times larger as it is in the small two-variable cell-cycle module analyzed, so that identifiability does not improve with scale.","fun_headline_variants_meta":{"raw":{"variants":["Complex cell models often obscure, not clarify, biology","Fewer parameters can unlock real biological insight","Cell cycle model proves: complexity hides mechanism","Why whole-cell simulations may not advance understanding","Parameter count matters more than model size in biology"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1579,"prompt_tokens":1044,"completion_tokens":535,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":660,"tokens_out":535,"duration_ms":5606,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:53:08.453286+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic parameter scan or profile-likelihood analysis on a publicly available whole-cell model such as JCVI-syn3A, fixing an experimentally measured output like division time. If the model matches the output across most of its parameter space, the paper's claim is supported; if only a small fraction is consistent, its key assumption that sloppiness scales with model size fails.","supporting_citations":[{"cited_title":"Karr, Jayodita C","cited_arxiv_id":null,"evidence_quote":"The 2012 Mycoplasma genitalium whole-cell model is the principal example of a comprehensive simulation whose parameter space the paper argues is too large to probe systematically."},{"cited_title":"Thornburg, Ansel Maytin, Jonghan Kwon, Troy A","cited_arxiv_id":null,"evidence_quote":"The spatially resolved JCVI-syn3A whole-cell model is the second major example of a large-scale simulation treated as a technical achievement without established mechanistic understanding."},{"cited_title":"Gutenkunst, Joshua J","cited_arxiv_id":null,"evidence_quote":"Establishes that sloppy parameter sensitivities are near-universal in systems biology models, the phenomenon underlying the paper's identifiability argument."},{"cited_title":"Adynamicalparadigmformolecularcellbiology.TrendsinCellBiology, 30:504–515, 2020","cited_arxiv_id":null,"evidence_quote":"Supplies the dynamical-analysis toolkit of bifurcation, sensitivity, and phase-plane methods that the paper argues must be applied to large-scale models."},{"cited_title":"Nédélec, Thomas Surrey, Anthony C","cited_arxiv_id":null,"evidence_quote":"Shows how agent-based cytoskeletal models self-organize from a few local rules, providing the template for parameter-lean large-scale models that remain exploratory."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames computer simulations as higher-order evidence, which the paper uses to credit whole-cell models with assembling and reconciling empirical knowledge without claiming explanation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the adequacy-for-purpose criterion used to separate mechanistic explanation from other legitimate jobs a whole-cell model can do."},{"cited_title":"A modular approach for modeling the cell cycle based on functional response curves.PLoS Computational Biology, 17:e1009008, 2021","cited_arxiv_id":null,"evidence_quote":"Provides the two cell-cycle oscillator models whose identifiability comparison is the paper's concrete, quantitative illustration."},{"cited_title":"Bistable,biphasicregulationofPP2A-B55accounts for the dynamics of mitotic substrate phosphorylation.Current Biology, 31(4):794–808.e6, 2021","cited_arxiv_id":null,"evidence_quote":"Supplies the bistability measurements that fix the functional response in the minimal model, the source of its far tighter parameter constraints."},{"cited_title":"Batterman and Collin C","cited_arxiv_id":null,"evidence_quote":"Supports the minimal-model explanatory account by showing that models can explain through universality classes, so omitted details can be irrelevant."}],"review_version":1}