{"id":"deb46f11-71f0-4606-8e9e-f0595deedccc","arxiv_id":"2411.17789","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":16,"one_line_summary":"A six-compartment ODE model of liver transplant rejection identifies cytotoxic T cell, IL-2, and liver-related parameters as the most influential drivers of simulated graft injury at day 30.","lead":"A new mathematical model simulates the immune response to a transplanted liver over 30 days, tracking liver cells, T cell types, and the cytokine IL-2. It finds that cytotoxic T cells, IL-2, and the liver itself dominate the simulated graft injury and thus may be the best targets for monitoring and therapy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Sobol' sensitivity ranking is computed under hand-estimated nominal parameters and narrow ±50% sampling ranges; the central claim that TC, IL-2, and liver parameters are key determinants may not survive alternative plausible parameter regimes.","rationale":"The reader's verdict is CONDITIONAL with high correctness risk, and the weakest assumption is precisely that sensitivity rankings are computed under largely estimated nominal parameter values and 50–150% sampling ranges. My analysis agrees: the single most load-bearing concern is that the Sobol' indices—and hence the central claim about TC, IL-2, and liver dynamics being key determinants—are internal to an unvalidated model with hand-picked nominal values. The paper's own text confirms that many top parameters are 'estimated' (Table 3, Estimated Parameters subsection) and that thresholds were chosen as 50% of initial conditions; carrying capacities are uniformly 6× initial values. This is not an internal inconsistency, but it is a correctness risk that a direct robustness check can settle. The paper includes code and extensive parameter derivation, which is credit-worthy, and its limitations are honestly stated. Yet the claim is conditional, as the reader concluded. I see no reason to move the verdict: CONDITIONAL remains appropriate, pending either validation against clinical data or an explicit multi-scenario sensitivity analysis that shows the top-7 ranking is stable across plausible parameter priors. My proposed concrete test—rerunning Sobol' with wide log-uniform ranges on the estimated parameters—would directly determine whether the concern lands. If the ranking persists, the central claim is substantially strengthened; if not, the conclusion would need to be weakened to 'within the nominal parameter box' rather than presented as a general finding about liver transplant rejection.","tokens_in":23600,"tokens_out":1632,"duration_ms":50997,"concrete_test":"Re-run the Sobol' sensitivity analysis with a 'parameter-uncertainty' protocol: replace each 'estimated' parameter in Table 3 (k2, K_L, parameters 8–11, 14–15, 21–22, 28–29, plus all carrying capacities) with a log-uniform prior spanning 0.1× to 10× its nominal value, while keeping measured parameters at their literature values; repeat the variance decomposition and ask whether the same top-7 parameters remain an order of magnitude above the rest. If the top-7 set changes or the total-variance share shifts materially, the paper's central conclusion is an artifact of the nominal box. As a complementary check, run a representative simulation and compare predicted ALT/AST/ALP trajectories or hepatocyte loss to a published liver transplant rejection cohort; even a qualitative match would strengthen the claim, and a mismatch would show the QOI is not clinically grounded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the global sensitivity analysis in §Sensitivity Analysis: the top-7 most-influential parameters are declared to drive hepatocyte loss at day 30. But these rankings are computed by sampling each parameter uniformly within 50–150% of its nominal value, and several of the most-influential nominals are explicitly 'estimated' in Table 3: the maximum TC killing effect (k2 = 10), the TC threshold for half-maximum killing (K_L = 200), the maximum IL-2 production rate by TC (28 = 0.36), and its threshold (352), plus all carrying capacities set to exactly six times initial values. The sensitivity index of a parameter measures the fraction of output variance attributable to that parameter over the prescribed box, not the parameter's biological importance. If the true values of these estimated parameters lie outside the ±50% box—which is plausible because some were anchored to 'values we felt were reasonable' and thresholds set to 50% of initial conditions—the variance decomposition could change qualitatively. More subtly, the model is not validated against any clinical data, so the QOI L(30) itself is a simulation output; if the nominal dynamics already drive L to near zero or to an exaggerated injury phenotype, the sensitivity ranking describes that model, not liver transplant rejection. The paper does acknowledge missing data and estimated parameters, but the headline claim is nevertheless conditional on these arbitrary nominal values and ranges. This is the load-bearing weak point: without demonstrating that the top-7 set is robust to alternative parameter constructions, the claim that IL-2 and TC dynamics are 'key determinants' is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a six-population ODE model of liver transplant rejection dynamics (hepatocytes, APCs, helper T cells, cytotoxic T cells, Tregs, and IL-2), parameterized from a literature search with a subset of parameters 'estimated' by the authors. The authors simulate the nominal model and then perform a global Sobol' sensitivity analysis over all 35 parameters, sampling each uniformly within 50% to 150% of its nominal value, with day-30 hepatocyte count L(30) as the quantity of interest. They report that seven parameters dominate the variance of L(30), that these parameters are related to cytotoxic T-cell dynamics, IL-2 dynamics, and hepatocyte loss, and they conclude that these are key determinants of liver graft injury with implications for monitoring and therapy. The paper acknowledges that data fitting and validation are future work.","tokens_in":24068,"tokens_out":6462,"duration_ms":61304,"significance":"The paper has clear strengths: transparent, literature-based parameter derivations; a large-scale global sensitivity analysis (3.7 million samples); the full code repository is provided; and the model schematic and pathway tables are useful. The PDF comparison in Fig. 4 is a sensible check that a small parameter subset reproduces most of the QOI variance within the sampled box. If the sensitivity ranking were robust to parameter uncertainty, the work would be a valuable mechanistic foundation for studying acute rejection in liver transplantation and for prioritizing future data collection and therapeutic targets. However, because the ranking is computed around an unvalidated, partly hand-estimated nominal parameter set, the abstract's clinical implications are currently stronger than the evidence supports. The paper is a good foundation, but the central claim needs to be either made conditional on the assumed parameter regime or hardened with additional sensitivity analyses.","major_comments":[{"comment":"The Sobol' indices are computed by varying each parameter uniformly within 50%–150% of its nominal value, but Table 3 marks several of the most influential parameters as 'estimated' (e.g., the maximum TC killing effect k2, the TC threshold k3, the maximum IL-2 production rates k28/k29, and the IL-2 thresholds), and the carrying capacities are all set to six times the initial values while the thresholds eta_A and eta_CD4 are set to 50% of initial conditions. The variance decomposition therefore describes an arbitrary box around hand-picked values, not the biological importance of the parameters. If the true values in transplant recipients fall outside this box, the top-7 list—and the abstract's claim that TC, IL-2, and liver dynamics are 'key determinants'—could change qualitatively. Please justify the sampling ranges with data, repeat the analysis under wider or multiple plausible parameter regimes, or explicitly reframe the results as conditional on the nominal parameterization.","section":"Parameter Values & Initial Values; Sensitivity Analysis"},{"comment":"The quantity of interest L(30) is a simulation output, and the model is not validated against any clinical or experimental data, as the paper itself states in the Discussion ('Future work to collect appropriate data and parametrize the model would be valuable'). The abstract nevertheless concludes that the findings 'have significant implications for the use of tests to monitor patients, and therapeutic strategies.' This overstates the current evidence. The results should be presented as model-generated hypotheses, and the paper should specify what clinical or experimental data would be needed to confirm or refute the predicted parameter ranking.","section":"Results; Abstract"},{"comment":"The concordance of the PDFs in Fig. 4 demonstrates that the seven selected parameters explain almost all QOI variance within the 50%–150% sampling box. It does not show that these parameters are the key determinants of graft injury in general: a parameter can have a large Sobol' index simply because its nominal value is uncertain and the model output is steep in that region. The text should explicitly separate 'explained variance within the assumed box' from 'key determinants of graft injury,' otherwise the phrasing in the Results and Discussion invites a stronger biological conclusion than the analysis supports.","section":"Sensitivity Analysis"},{"comment":"The derivation of k4 (the number of APCs primed by alloantigens per hepatocyte) is a chain of multiplicative assumptions: 2% necrosis, 8.7×10^9 protein molecules per hepatocyte, 10 potential antigens per protein, 10% alloantigen fraction, 10% graft antigens on an APC, and a factor of 1000 for lymphatic-vs-blood dilution. Although k4 is not labeled 'estimated' in Table 3, the final value of 4.52×10^-9 is effectively hand-built, and it directly scales the APC source term in Eq. (2). The sensitivity analysis should include wider ranges for k4 or a separate uncertainty analysis so that the ranking is not an artifact of this specific calculation.","section":"Parameter Values & Initial Values"}],"minor_comments":[{"comment":"The caption of Fig. 3 says the plot 'omits the other 25 parameters' (implying 10 parameters are shown), while the text later refers to the top 7 parameters and calls 'the other twenty-eight parameters' non-influential; please align these numbers.","section":"Figure 3; Results"},{"comment":"The caption of Fig. 4b says the histogram for the twenty-eight least-influential parameters is drawn with a red outline, but the red outline is also used for the all-parameters histogram; please correct the color/description so the two distributions are distinguishable.","section":"Figure 4b"},{"comment":"Several parameter symbols in the Results text, including the list of the seven most-influential parameters, appear to be missing in the manuscript version I reviewed; please ensure all mathematical symbols are typeset in the final version.","section":"Results"},{"comment":"The claim that this is 'the first mechanistic mathematical model of liver transplant and immune system dynamics' should be qualified, since the cited literature already contains mechanistic transplant models (Markovska, An, Gateno, Arciero, Lapp, Banks, Ciupe); if the novelty is liver-specific or component-specific, it should be stated that way.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its limitations and provides reproducible code, but the central claim is conditional on a sensitivity analysis whose parameter ranges and nominal values are largely unvalidated. The revision should either substantially broaden the sensitivity analysis or substantially weaken the conclusions in the abstract and discussion. This is a foundation-level modeling paper; it may be better positioned in a modeling-focused venue than in a clinical transplant journal as currently framed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a credible first-step modeling paper: a six-variable ODE model of acute rejection in liver transplant with an explicit IL-2 compartment, parameterized from the literature as thoroughly as the current data allow, plus a Sobol global sensitivity analysis. The code is on GitHub, the parameter derivations are transparent, and the authors are honest about what is estimated. The main finding—that cytotoxic T cell, IL-2, and hepatocyte-related parameters dominate simulated graft injury at day 30—is biologically plausible and consistent with what is known about TCMR.\n\nThe real novelty is fairly small: the model is liver-specific and includes IL-2 explicitly, but the framework is the same ODE structure used in prior heart, kidney, and general transplant models. The authors acknowledge this. What they do well is the painstaking derivation of parameter values from published studies and the careful Sobol implementation with PDF comparisons to show that seven parameters capture almost all the output variability within their sampled ranges.\n\nThe soft spot is exactly where the stress-test puts it: the sensitivity ranking is computed over ±50% uniform ranges around nominal values, and several of the most influential nominals are hand-estimated (\"values we felt were reasonable\"), with thresholds set at 50% of initial conditions and carrying capacities at six times initial values. If the true values lie outside those boxes, the ranking could change. More importantly, the model has not been validated against any clinical or experimental data, so L(30) is a simulation output under a particular parameter construction. The abstract's claim that TC, IL-2, and liver dynamics are \"key determinants\" of graft injury is therefore conditional, not established. The authors do flag this in the Limitations section and call for future data collection and fitting, so the paper is not deceptive—but the headline outruns the evidence.\n\nWho should read this: people working on mechanistic transplant models who want a liver-specific starting point and a clear template for sensitivity analysis. It does not deserve desk rejection; a serious referee could push for an explicit uncertainty analysis (e.g., sampling over alternative parameter constructions) and temper the clinical claims in the abstract. I'd send it to review.","headline":"A transparent, well-coded liver-transplant ODE model whose sensitivity ranking is real within its sampled box but conditional on hand-estimated nominal values; worth a serious referee, but the abstract overstates the result.","tokens_in":24591,"tokens_out":2043,"would_cite":false,"duration_ms":21246,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's central claim is that in the first mechanistic liver-transplant immune model, seven parameters governing cytotoxic T cell killing, IL-2-driven T cell proliferation, and IL-2 production dominate day-30 graft injury.","keywords":["liver transplantation","acute rejection","mechanistic mathematical model","ordinary differential equations","global sensitivity analysis","Sobol indices","interleukin-2","cytotoxic T cells"],"falsifier":"Measure serum IL-2, peripheral blood CD8+ T cell counts, and AST/ALT/ALP serially around rejection episodes in a cohort of liver transplant recipients, fit the model to each patient, and check whether variation in the seven influential parameters predicts the observed hepatocyte loss better than variation in the remaining 28; if the seven-parameter ranking does not survive refitting or reverses under bootstrap resampling of the data, the central claim would be falsified.","tokens_in":23403,"feed_emoji":"🩺","tokens_out":11009,"duration_ms":92955,"temperature":0.7,"pith_summary":"Liver transplant has a clinical puzzle: modern immunosuppression prevents rejection but its long-term toxicities are now a bigger cause of death than rejection itself. This paper tries to break that puzzle by building what it describes as the first mechanistic mathematical model of liver transplant immune dynamics—a system of six ordinary differential equations tracking hepatocytes, antigen-presenting cells, helper T cells, cytotoxic T cells, regulatory T cells, and IL-2. Simulating a 30-day acute rejection window starting roughly one year after transplant, the paper asks which of 35 parameters most control the number of healthy hepatocytes remaining at day 30. Its central claim is that seven parameters, all on the cytotoxic T cell–IL-2–liver axis, dominate graft injury; if correct, monitoring IL-2 and cytotoxic T cell activity and targeting their pathways should improve rejection management.","feed_headline":"Model pins liver graft injury on T cells and IL-2","feed_subtitle":"If right, monitoring IL-2 and CD8+ T cells—not just liver enzymes—could catch rejection earlier.","key_machinery":"The carrying object is a six-variable ODE system with one equation each for hepatocytes $L$, antigen-presenting cells $A$, helper T cells $T_H$, cytotoxic T cells $T_C$, regulatory T cells $T_R$, and IL-2 $I$, connected by Michaelis-Menten saturating interaction terms and logistic proliferation terms. The load-bearing coupling is that IL-2 internalization by effector T cells is what converts IL-2 loss into new T cell production (pathways p, q, r, s), and cytotoxic T cell attack is the only direct source of hepatocyte loss; this is why the sensitivity ranking lands on the IL-2–$T_C$–liver axis. The ranking itself is produced by Sobol' global sensitivity analysis, which decomposes the variance of the quantity of interest $L(30)$ across uniform 50–150% perturbations of each parameter.","core_discovery":"On the paper's own terms, the discovery is a quantitative ranking, not a measured biological fact: the day-30 healthy hepatocyte count $L(30)$ is almost entirely controlled by seven of the 35 model parameters. These seven parameters govern the maximum rate at which cytotoxic T cells kill hepatocytes, the T cell density at which that killing half-saturates, the strength of IL-2-driven proliferation of conventional T cells, and the rates at which IL-2 is produced by activated T cells. Varying just these seven across 50–150% of their nominal values reproduces nearly all of the variability in $L(30)$ that appears when all 35 parameters are varied, while varying the remaining 28 produces almost none. The paper concludes that cytotoxic T cell dynamics, IL-2 dynamics, and the liver's own loss kinetics are the key determinants of graft injury, and that this points toward IL-2 and T cell monitoring and therapies.","pith_inferences":["Editorial extension: the 50–150% sampling ranges are set around hand-estimated nominal values, so the seven-parameter dominance is a property of those ranges; a Bayesian calibration against longitudinal patient data could show whether the ranking persists under realistic uncertainty.","Editorial extension: because the model omits innate immune cells, B cells, and cytokines other than IL-2, the clinical payoff would be tested by asking whether IL-2/CD8 biomarkers outperform liver enzymes specifically in biopsy-confirmed T cell–mediated rejection episodes.","Editorial extension: an adaptive dosing rule could be derived directly from the model—raise or lower immunosuppression when simulated $L(30)$ drops below a threshold—which the paper motivates but does not implement.","Editorial extension: reparameterizing the hepatocyte-loss equation as a generic graft-mass equation would let the same sensitivity analysis rank drivers for other organs, giving a cross-organ comparison of whether IL-2/CD8 dominance is liver-specific or general."],"forward_implications":["Routine post-transplant monitoring should add IL-2 concentration and activated CD8+ T cell counts to liver enzyme tests, since those are the variables whose fluctuations most affect day-30 graft survival.","Therapies that suppress IL-2-driven cytotoxic T cell proliferation—such as IL-2 receptor–blocking antibodies—should be the most direct lever for preventing acute rejection, and could be combined rationally with calcineurin and mTOR inhibitors.","Future data collection and model fitting can focus on the seven influential parameters, because a seven-parameter version captures almost all of the model's output variability.","The same ODE skeleton, reparameterized for organ-specific cell turnover and immune kinetics, could be adapted to kidney, heart, or other transplanted organs.","Individualized simulations built on the seven parameters could serve as in silico trials for immunosuppression dosing before clinical trials."],"supporting_citations":[{"why":"Provides the maximum helper T cell activation rate and the Treg differentiation rate used in the model equations.","marker":"97"},{"why":"Supplies the peripheral blood helper and cytotoxic T cell counts used for initial values.","marker":"92"},{"why":"Supplies the circulating dendritic cell count used to set the initial antigen-presenting cell level.","marker":"89"},{"why":"Supplies doubling times and lifespans for helper and cytotoxic T cells used for proliferation and loss rates.","marker":"99"},{"why":"Supplies the IL-2 concentration used to estimate the Treg lifespan-extension and IL-2 consumption parameters.","marker":"105"},{"why":"Supplies the cell-concentration threshold at which IL-2 secretion saturates for helper T cells.","marker":"110"},{"why":"Supplies the IL-2 per cell needed to trigger helper T cell proliferation, used for an IL-2 proportionality constant.","marker":"103"},{"why":"Supplies the in vivo Treg disappearance rate used for the Treg natural loss constant.","marker":"109"},{"why":"Reports a clinical trial of low-dose IL-2 in liver transplant recipients and motivates modeling both pro- and anti-inflammatory IL-2 effects.","marker":"87"}],"fun_headline_variants":["Math model points to T cells and IL-2 in liver graft injury","Sensitivity analysis fingers T cells and IL-2 for graft injury","Model identifies T cells and IL-2 as key drivers of liver graft injury","T cells and IL-2 drive liver graft injury, model shows","Mathematical model: T cells and IL-2 control liver graft damage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hand-estimated nominal parameter values—including the seven flagged as most influential, several marked 'estimated' in Table 3—and the 50–150% uniform sampling ranges represent real transplant patients closely enough that the sensitivity ranking points at the true drivers of graft injury.","fun_headline_variants_meta":{"raw":{"variants":["Math model points to T cells and IL-2 in liver graft injury","Sensitivity analysis fingers T cells and IL-2 for graft injury","Model identifies T cells and IL-2 as key drivers of liver graft injury","T cells and IL-2 drive liver graft injury, model shows","Mathematical model: T cells and IL-2 control liver graft damage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000562,"raw_usage":{"total_tokens":2659,"prompt_tokens":929,"completion_tokens":1730,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":1635}},"tokens_in":545,"tokens_out":1730,"duration_ms":12837,"temperature":1.0,"reasoning_tokens":1635,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:00:42.838808+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure serum IL-2, peripheral blood CD8+ T cell counts, and AST/ALT/ALP serially around rejection episodes in a cohort of liver transplant recipients, fit the model to each patient, and check whether variation in the seven influential parameters predicts the observed hepatocyte loss better than variation in the remaining 28; if the seven-parameter ranking does not survive refitting or reverses under bootstrap resampling of the data, the central claim would be falsified.","supporting_citations":[],"review_version":1}