{"id":"781f693e-8126-45f2-8f5a-dc01e4e60ea6","arxiv_id":"2412.01844","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Shapley values from a Gaussian process regression rank income, workforce, HIV testing, unemployment, and population as the most influential socio-economic predictors of tuberculosis infection rates, but the regression only matches 10 of 87 Russian regions within 10 percent error.","lead":"This paper ranks which regional socio-economic statistics best predict tuberculosis transmission rates by combining an ODE co-infection model, inverse problem fitting, and Shapley value feature importance. The approach identifies five influential parameters but only reconstructs infection rates accurately for 10 of 87 Russian regions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The feature ranking is only as trustworthy as the unvalidated β_c targets from the simplified inverse problem, and with no error bars or robustness checks the Shapley feature list is not yet supported.","rationale":"The paper honestly reports a mostly negative result and uses Shapley values in a standard way, but the headline claim is a numerical ranking whose input is an inverse-problem output with no uncertainty quantification. The reader's weakest assumption identifies exactly this: the β_c values from the simplified ODE and manual data skipping are the load-bearing targets. My stress-test agrees and adds that the same concern is compounded by the fact that the regression used for Shapley values only achieves under 10% error for 10 of 87 regions, so the ranking describes a model that mostly fails. A synthetic-data recovery experiment is the cleanest check because it provides known ground truth, isolating bias in the inverse problem from the separate question of regression quality. I do not see an internal inconsistency that would require rejection, but the conditional verdict is appropriate until the robustness check is performed. No code or data are provided to verify the numerical claims independently, which further supports keeping the verdict conditional rather than accepting the ranking as established.","tokens_in":7118,"tokens_out":3927,"duration_ms":38771,"concrete_test":"Generate synthetic incidence data from the full latent-TB/HIV model of Kabanikhin et al. (2024) with known β_c values, subsample the outputs to reproduce the seasonal peaks, missing points, and manual skip pattern described in Section 2.1, and then run the paper's simplified inverse problem (Eq. 4) on the synthetic data. If the recovered β_c differs from the known values by more than the 10% threshold used in Section 4, the inverse-problem targets are corrupted and the Shapley ranking is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on treating the β_c values recovered by the inverse problem (Eq. 4, Section 3.2) as reliable regression targets. This is precarious for three reasons. First, Section 2.1 states that seasonal peaks and missing cumulative observations are skipped manually, with no anomaly detection or sensitivity analysis; if those choices alter the fitted β_c, the Shapley ranking changes. Second, the simplification in Section 3.1 absorbs latent TB and removes the treated compartment to restore identifiability, but no validation shows that the reduced 5-compartment model recovers the same β_c as a fuller model. The authors themselves cite prior work in which latent-TB assumptions produced unrealistic transition times, so the omission is not innocuous. Third, the Gaussian-process regression used for Shapley values reconstructs β_c to under 10% error for only 10 of 87 regions (Section 4). The feature ranking is thus computed on a model that fails for the large majority of regions; unless the ranking is stable when restricted to the 10 successful regions or weighted by predictive accuracy, it is not established that the selected features are meaningful for those regions. The paper reports no uncertainty intervals, bootstrap, or code, so these concerns cannot be dismissed from the text alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a workflow for modeling regional heterogeneity in tuberculosis-HIV co-infection in 87 regions of the Russian Federation. The authors simplify an SIR-type model to five compartments, solve per-region inverse problems to estimate transfer rates including the infection rate beta_c, train a Gaussian process regressor to predict beta_c from five socio-economic features, and compute Shapley values to rank feature importance. They report that the final regression achieves relative errors below 10% for only 10 of 87 regions and select mean income per capita, workforce, number of HIV tested, unemployment, and total population as the most influential features.","tokens_in":7400,"tokens_out":6090,"duration_ms":50515,"significance":"If the beta_c estimates are reliable, the paper offers a useful template for combining mechanistic epidemic models, inverse problems, and model-agnostic feature attribution to reduce data collection for regional calibration. The paper is honest about the low success rate and explicitly avoids causal claims. However, because the feature ranking is derived from a regression that fails for most regions and from an unvalidated inverse problem, the central claim is conditional. The paper also provides no code, data, or uncertainty quantification, limiting reproducibility and the confidence that can be placed in the numerical results.","major_comments":[{"comment":"The data are described as annual time series from 2009 to 2019 in Section 2.1, while Section 3.2 says the machine-learning tests use data from 2011 to 2019, and the abstract mentions 2009 to 2023. The misfit functional in Eq. (4) sums over 2007 to 2020. This inconsistency makes the inverse-problem setup irreproducible, and depending on which time window is actually used, the recovered beta_c values and the subsequent Shapley ranking could change. Please state the exact time window for each step and correct the equations and text accordingly.","section":"Section 2.1 and Eq. (4)"},{"comment":"The Shapley values are computed for a Gaussian process trained on all 87 recovered beta_c values, yet the text and Figure 7 state that relative regression errors are below 10% for only 10 of 87 regions. The paper does not report whether the feature ranking is stable when the regression is restricted to the 10 well-fitted regions, how per-region accuracy is weighted, or any uncertainty intervals or bootstrap replicates for the Shapley values. Additionally, Table 1 reports an averaged relative error of 0.030 for the Gaussian process, which is hard to reconcile with the 10/87 statement unless the average is dominated by regions with very small beta_c; this should be clarified. As it stands, the selected five features are not supported as the most influential for the full dataset.","section":"Section 4 and Figure 7"},{"comment":"The reduced five-compartment model is introduced primarily to avoid identifiability problems associated with latent and treated compartments, but no synthetic-data validation or comparison with the fuller model in Kabanikhin et al. (2024) demonstrates that the simplification preserves the value of beta_c. The authors cite the prior work's unrealistic latent-stage transition times as motivation, but that does not establish that the simplified model recovers the same beta_c as a fuller model. Without such a check, the beta_c targets used in the regression remain potentially biased, which would invalidate the Shapley-based feature selection.","section":"Section 3.1 and Section 3.2"},{"comment":"Outliers and missing observations are skipped manually, and the authors state that anomaly detection algorithms are not applicable in this scenario. No sensitivity analysis is reported for these skips, so it is unknown whether the fitted beta_c values—and the resulting Shapley ranking—are robust to the data-handling choices. A perturbation analysis or a comparison with alternative inclusion/exclusion rules is needed to show that the chosen five features are not artifacts of manual data cleaning.","section":"Section 2.1"}],"minor_comments":[{"comment":"The word 'Gussian' should be 'Gaussian', and the caption of Figure 3 says 'Flu diagram' rather than 'flow diagram'.","section":"Section 3.2"},{"comment":"The text says 'just 10 out of 87 parameters beta' but lists 11 regions: Kamchactka krai, Krasnoyarsk krai, Leningrad oblast, Republic of Dagestan, Republic of Mordovia, Republic of Northern Osetia, Rosvov oblast, Samara oblast, Smolensk oblast, Tver oblast, and Tomsk oblast. The count and the list need to be reconciled.","section":"Section 4"},{"comment":"The reference to 'Feng Z, 2000' and the associated author list are incomplete; the citation should follow the standard format used throughout the reference list.","section":"References"},{"comment":"The source of the epidemiological and socio-economic data, as well as the definitions of the variables, is not provided. A data availability statement or a link to the data sources would improve reproducibility.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for the journal and builds on the authors' prior work. The main risk is not circularity but the lack of validation of the inverse-problem targets and the absence of robustness checks. The authors deserve credit for reporting the 10/87 failure rate, but the inconsistencies in time windows and the region count should be corrected before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my take on the Neverov and Krivorotko preprint. The notable thing is the integration: nobody else seems to have paired Shapley values with a TB-HIV co-infection ODE inverse problem to rank socio-economic drivers of regional transmission rates. That is a genuinely useful idea for cutting data collection costs. The paper is also refreshingly honest: it tells you up front that regression only worked for 10 of 87 regions, and it explicitly avoids causal claims about the Shapley rankings.\n\nThe modeling choices are reasonable in context. The authors drop the latent-TB compartment because it creates identifiability problems, and they cite their own previous work showing the latent fraction had to be tuned to 1–5% to get realistic transition times. That is a legitimate practical justification, not a hand-wave. The Bayesian optimization approach is described well enough to follow.\n\nNow the soft spots, in proportion. The biggest issue is the one the stress-test flags: the β_c values that the Shapley ranking explains are themselves the output of a simplified inverse problem with manually skipped seasonal peaks and missing observations. The paper says anomaly detection wasn't feasible, so they skipped points by hand. That is fragile. No sensitivity analysis shows how the ranking changes if those skips are altered. No error bars or uncertainty intervals appear anywhere. The regression itself only works for 10 regions, yet the Shapley values are computed on all 87; the paper doesn't show whether the ranking is stable when you restrict to the regions where the GP actually predicts well. Without those checks, the feature list—mean income, workforce, HIV testing, unemployment, population—is not yet established as meaningful, even for the successful subset.\n\nThere are minor issues too: Eq. (4) sums over 2007–2020 while the text says data cover 2009–2019; Fig. 2 shows monthly data through 2022 for one region; and the list of 10 successful regions actually contains 11 names, with a typo (\"Rosvov\"). These are cosmetic but suggest the manuscript needs a careful pass.\n\nWho is this for? People working on regional epidemic modeling with limited data, inverse problem practitioners, and anyone interested in interpretable ML applied to public health. It deserves a serious referee because the idea is useful and the failures are reported rather than hidden. But it needs major revision: provide the code and data, add uncertainty quantification on the inverse problem, test sensitivity to the data-exclusion choices, and analyze the Shapley ranking on the successful subset. I'd send it to review, but I'd expect heavy revision before acceptance.\n\nRecommendation: engage with it, but require the robustness work. The core concept is worth the effort.\n\nBest,\n[Your name]","headline":"An honest proof-of-concept that applies Shapley-based feature ranking to a simplified TB-HIV ODE inverse problem across Russian regions, but the numerical ranking rests on unvalidated β_c targets and needs robustness work before it can be trusted.","tokens_in":7893,"tokens_out":1257,"would_cite":false,"duration_ms":13265,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["92D30","65J20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that five socio-economic indicators—mean income per capita, workforce, number of HIV tested, unemployment, and total population—are the most influential features for reconstructing regional tuberculosis infection rates…","keywords":["feature importance","Shapley values","tuberculosis","HIV co-infection","mathematical model","inverse problem","Gaussian process","socio-economic parameters"],"falsifier":"Compute Shapley values for the same regression task using the full latent-TB model from [Kabanikhin et al., 2024] on the same 87 regions; if the top-five feature set changes or the 10-region under-10% accuracy disappears, the paper's ranking depends on the model simplification rather than on a stable socio-economic signal.","tokens_in":6917,"feed_emoji":"📊","tokens_out":4169,"duration_ms":35290,"temperature":0.7,"pith_summary":"The paper asks whether a single epidemic model can be adapted to different regions by feeding it socio-economic data instead of re-fitting epidemiological parameters for each region separately. It solves an inverse problem for a simplified TB–HIV co-infection ODE model on 87 Russian regions to obtain per-region infection rates $\\beta_c$, then uses a Gaussian process regressor to predict $\\beta_c$ from socio-economic indicators. Applying Shapley values ranks these indicators, identifying mean income per capita, workforce, number of HIV tested, unemployment, and total population as the most influential. The authors stress that the regression was accurate to under 10% in only 10 regions, so the ranking is a promising step rather than a finished tool.","feed_headline":"Five socio-economic factors shape TB infection rate in model","feed_subtitle":"Shapley values rank income, workforce, HIV tests, unemployment, and population — but only 10 of 87 regions fit.","key_machinery":"The central mechanism is the combination of (a) a simplified SIR-like ODE model (Eq. 2) that collapses latent-TB transfer into the active-TB transition and removes the treated compartment, making the inverse problem for $q = \\{\\beta_c, \\lambda\\sigma, r_1, r^{*}, k\\}$ identifiable from measured compartments; and (b) Shapley values (Eq. 1), the unique payoff distribution satisfying linearity, symmetry, efficiency, and the null-player axiom, used to score each socio-economic feature's contribution to the Gaussian process regression of $\\beta_c$.","core_discovery":"The central claim is that the TB infection rate $\\beta_c$, recovered separately for each region from a five-compartment ODE model of TB–HIV co-infection, can serve as a regression target for regional socio-economic data, and that Shapley values computed on a Gaussian process regressor single out five features—mean income per capita, workforce, number of HIV tested, unemployment, and total population—as the most influential. This ranking, however, reproduces $\\beta_c$ within 10% relative error for only 10 of the 87 regions considered; for those regions the five parameters carry the reconstruction. The paper notes a negative correlation between infection rate and mean income or population below subsistence level, which the authors flag as unexpected and as a property of the chosen regression model rather than evidence of causation.","pith_inferences":["The 10-region success suggests the Shapley ranking may be an artifact of the Gaussian process kernel and the short 2009–2019 training window; a cross-validation with held-out years would test whether the five-feature ranking generalizes beyond the fitted regions.","Switching the regressor (e.g., to a tree ensemble or a linear model) would likely change Shapley rankings; the paper's choice of the dot-product-plus-white-noise Gaussian process is pragmatic but not theoretically defended, so the ranking is conditional on that choice.","The same pipeline could be applied to other endemic diseases whose regional transmission heterogeneity is suspected, provided the simplified compartment structure remains identifiable; the method is not TB-specific beyond the ODE structure.","If the manual exclusions of seasonal peaks and missing data were automated or replaced by an explicit missing-data model, the recovered $\\beta_c$ values could shift, potentially changing the feature ranking; this makes the manual preprocessing a hidden load-bearing step."],"forward_implications":["If the ranking is correct, public health agencies can prioritize collecting these five socio-economic indicators when building regional forecasts of TB incidence, reducing data collection burdens.","In the 10 regions where reconstruction stays under 10% error, the method supplies region-specific $\\beta_c$ without re-solving an inverse problem for each region, making multi-region modeling cheaper.","The simplification of omitting a latent-TB compartment makes the inverse problem identifiable; accepting this simplification justifies calibrating co-infection models with routinely measured surveillance data alone.","The approach reinterprets socio-economic covariates as drivers of the transmission rate, creating a concrete numerical link between economics and epidemic parameters.","The negative correlation with mean income, if it survives broader testing, would complicate the common assumption that TB concentrates in poorer areas, though the authors present it as model-specific."],"supporting_citations":[{"why":"Supplies the base compartment model of TB and HIV dynamics with latent and active stages that the paper simplifies into its five-equation ODE.","marker":"[Aparicio and Castillo-Chavez, 2009]"},{"why":"Proves the uniqueness of the Shapley value distribution, the theoretical foundation for the feature-importance scores used in the paper.","marker":"[Shapley, 1953]"},{"why":"Provides the inverse-problem methodology, parameter typical values, and Bayesian-type optimization approach used to recover $\\beta_c$ from incidence data.","marker":"[Kabanikhin et al., 2024]"},{"why":"Empirical Russian study linking socio-economic factors to tuberculosis indicators, motivating which parameters to include in the feature set.","marker":"[Podgayeva et al., 2011]"},{"why":"Adds Russian evidence on present-day socio-economic risk factors for tuberculosis, further informing the selection of candidate features.","marker":"[Aminev et al., 2013]"},{"why":"Documents how Shapley values are applied to machine-learning model interpretability, framing the paper's use of them on the Gaussian process regressor.","marker":"[Molnar, 2019]"}],"fun_headline_variants":["Shapley values rank 5 socio-economic keys to TB infection rate","TB-HIV model matches only 10 of 87 regions with five factors","Five socio-economic features drive TB in model, but fit is limited","Income and HIV tests top Shapley list for TB rate, but model fits few"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recovered infection rates $\\beta_c$ that the regression targets are reliable, even though the simplified model omits the latent-TB compartment and the data had seasonal peaks and missing values that were skipped manually.","fun_headline_variants_meta":{"raw":{"variants":["Shapley values rank 5 socio-economic keys to TB infection rate","TB-HIV model matches only 10 of 87 regions with five factors","Five socio-economic features drive TB in model, but fit is limited","Income and HIV tests top Shapley list for TB rate, but model fits few"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1470,"prompt_tokens":827,"completion_tokens":643,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":561}},"tokens_in":443,"tokens_out":643,"duration_ms":6633,"temperature":1.0,"reasoning_tokens":561,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:16:57.078069+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute Shapley values for the same regression task using the full latent-TB model from [Kabanikhin et al., 2024] on the same 87 regions; if the top-five feature set changes or the 10-region under-10% accuracy disappears, the paper's ranking depends on the model simplification rather than on a stable socio-economic signal.","supporting_citations":[{"cited_title":"Castillo-Chavez","cited_arxiv_id":null,"evidence_quote":"Supplies the base compartment model of TB and HIV dynamics with latent and active stages that the paper simplifies into its five-equation ODE."},{"cited_title":"A value for n-person games","cited_arxiv_id":null,"evidence_quote":"Proves the uniqueness of the Shapley value distribution, the theoretical foundation for the feature-importance scores used in the paper."},{"cited_title":"Identification of the mathematical model of tuberculosis and hiv co-infection dynamics","cited_arxiv_id":null,"evidence_quote":"Provides the inverse-problem methodology, parameter typical values, and Bayesian-type optimization approach used to recover $\\beta_c$ from incidence data."},{"cited_title":"Interpretable Machine Learning","cited_arxiv_id":null,"evidence_quote":"Documents how Shapley values are applied to machine-learning model interpretability, framing the paper's use of them on the Gaussian process regressor."}],"review_version":1}