{"id":"99b2eb5e-e554-4fdf-af8b-e35e02fdd31e","arxiv_id":"2506.07825","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"In SIR models with under-reporting and prior immunity, only the combinations beta*/p and beta*(1-pi) are identifiable from reported incidence; a single extra survey restores full identifiability.","lead":"This paper proves that in an SIR epidemic model with under-reporting and pre-existing immunity, the transmission rate, reporting fraction, and prior immunity fraction cannot all be estimated from reported case counts alone. It then shows that adding one population survey, measuring either prior immunity or the reporting fraction among the infected, restores identifiability.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 is proved for observed prevalence Ir(t), but the paper's data are cumulative reported incidence N1(t); no analytic proof shows identical N1(t) implies conditions (16)-(17), so Corollary 2 lacks direct support for the stated observation model.","rationale":"The central theorem is mathematically sound for the observed prevalence Ir(t). The proof in Appendix D correctly shows that identical Ir(t) forces the two invariant combinations, and conversely that these combinations produce identical Ir(t). Because N1(t) = p[S(0)-S(t)] and pS(0) is invariant under the conditions, the 'vice versa' direction also yields identical cumulative reported incidence, so the unidentifiability claim for reported incidence data is supported. The gap concerns the 'only if' direction for N1, which is needed for Corollary 2's assertion that knowing one parameter makes the others identifiable from incidence data. This is a secondary claim relative to the central unidentifiability result, and the simulation study in Section 3.2 provides practical evidence that estimation with one known parameter works. The Figure 3 example is a concrete error: the stated second parameter set does not satisfy condition (17), so the plotted trajectories are not identical under the theorem's own criteria. This does not invalidate the proof but indicates a need for correction. Overall, no fatal flaw was found in the central claim.","tokens_in":14732,"tokens_out":18617,"duration_ms":197234,"concrete_test":"Use the StructuralIdentifiability.jl package (or differential algebra) on the deterministic model (2)-(3) with observed output y(t)=dN1/dt (reported incidence) and known gamma, parameter vector theta=(beta*, p, pi). First verify unidentifiability; then fix pi to a known value and test for global structural identifiability of (beta*, p). If the fixed-pi model is globally identifiable, Corollary 2 extends to incidence data; if not, the paper's claim that knowing pi yields unique beta* and p from reported incidence is false.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central unidentifiability claim is established by Theorem 1's 'vice versa' direction: parameter sets satisfying beta*/p constant and beta*(1-pi) constant produce identical Ir(t), and hence identical cumulative reported incidence N1(t), so the three parameters are not identifiable from reported incidence curves. However, Corollary 2 (if one parameter is known, all three are identifiable) relies on the 'only if' direction: identical observed output must imply the two invariant combinations are equal. The Appendix D proof derives this only under the assumption that the observed output is the prevalence Ir(t), using the differential equation (12) and evaluation at t=0. The paper's data and simulations use the counting process N1(t) (cumulative reported infections), not Ir(t). For N1(t), the implication 'identical curves imply conditions (16)-(17)' is not proven analytically. If it failed, knowing pi would not fix unique (beta*, p), contradicting Corollary 2 as applied to incidence data. In addition, the illustrative parameter values in Figure 3 do not satisfy condition (17) (beta*1(1-pi1)=1.33 vs beta*2(1-pi2)=1.14), so the plotted trajectories cannot be identical as stated; this suggests the prevalence/incidence distinction was handled imprecisely.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyses structural identifiability in an SIR model with under-reporting (fraction p) and pre-existing immunity (fraction pi), using an effective transmission rate beta*. Theorem 1 (Appendix D) claims that the trajectory of reported infectious individuals Ir(t) identifies only the two combinations beta*/p and beta*(1-pi), so (beta*, p, pi) are not jointly identifiable; the converse is also proved. Corollary 2 states that knowledge of any one of the three parameters makes the triple identifiable. The paper then uses the initial-growth equation (5) and final-size equation (7), combined with sample-survey estimates of pi or p, to estimate the remaining parameters in 100 simulated stochastic epidemics (Table 2).","tokens_in":15009,"tokens_out":9559,"duration_ms":109437,"significance":"If the bridge to cumulative reported incidence is made explicit, the paper gives a clean and useful characterization of unidentifiability in a practically relevant model. The analytic proof is self-contained, includes the converse direction, is cross-checked with StructuralIdentifiability.jl, and the simulation code is publicly available. The paper will be a useful reference for modellers deciding what can and cannot be inferred from reported case curves. The main caveat is that Theorem 1 is stated for Ir(t) while the data are N1(t); this gap is easy to close but must be written down.","major_comments":[{"comment":"The theorem and corollary are stated for the trajectory of reported infectious individuals Ir(t), but the data considered in the paper (Section 2.4, Figure 2) are the cumulative reported incidence process N1(t). The proof never establishes that identical N1 trajectories imply identical Ir trajectories at the level of the deterministic model. This is true, because from (2), N1'(t) = beta* Ir(t) S(t)/n = dIr(t)/dt + gamma Ir(t); with gamma and Ir(0) known, N1(t) uniquely determines Ir(t). The paper should state this equivalence explicitly (as a lemma or remark) before applying Theorem 1 and Corollary 2 to reported incidence data, otherwise the central claim that 'only reported case data are available' is not directly supported by the proof.","section":"Appendix D, Theorem 1 and Corollary 2"},{"comment":"The parameter values used to illustrate Theorem 1 do not satisfy the theorem's conditions. For the first set {p1=0.4, pi1=0.3, beta*1=1.9} and the second set {p2=0.24, pi2=0.0, beta*2=1.14}, condition (16) holds (1.9/0.4 = 1.14/0.24 = 4.75), but condition (17) does not: beta*1(1-pi1)=1.33 whereas beta*2(1-pi2)=1.14. The two trajectories therefore cannot be identical; indeed their initial growth rates differ (0.33 vs 0.14 when gamma=1). Please correct the figure and the surrounding text, for example by taking beta*2=1.33 and p2=0.28, which satisfy both conditions.","section":"Section 3.1 and Figure 3"}],"minor_comments":[{"comment":"The text contains the duplicated phrase 'a linear relationship a linear relationship' in the description of the log-linear regression; please fix this typo.","section":"Section 2.3"},{"comment":"The phrase 'Taylor?series' appears to be a conversion artifact; please correct it to 'Taylor series'.","section":"Section 1"},{"comment":"The reported standard deviations reflect variation across the 100 simulated epidemics but do not propagate the uncertainty in the estimated growth rate rho-hat and final reported size z-hat-r from the regression and final-size estimation, nor the finite-sample survey uncertainty beyond what is captured by the simulation design; this limitation should be stated explicitly.","section":"Section 3.2 and Table 2"},{"comment":"The derivation leading to equation (16) assumes that the common coefficient in equation (14) is nonzero; the degenerate case of zero initial growth (where the coefficient vanishes) is not treated. Please either exclude this case in the theorem statement or add a short argument for it.","section":"Appendix D"},{"comment":"The proof evaluates the ODE at t=0, which requires observing the epidemic from its start with a known initial reported infectious count ni0 and known population size n; the Discussion acknowledges these assumptions, but it would be helpful to state this dependence explicitly next to Theorem 1.","section":"Appendix D, equation (15)"}],"recommendation":"major_revision","confidential_remarks":"The two major issues are technical presentation gaps rather than flaws in the core proof. The missing equivalence between observing N1(t) and observing Ir(t) is straightforward to add, and the Figure 3 values are simply wrong arithmetic that can be corrected. I found no circularity in the argument and no reason to doubt the theorem's validity. The paper fits the journal's scope and should be acceptable after a revision that addresses these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Core result is real and worth knowing: for a deterministic SIR with reporting probability p and pre-existing immunity π, observing only reported prevalence Ir(t) identifies the two combinations β*/p and β*(1−π), and no more. The proof via ODE rewriting is clean, the converse is shown, and the identifiability software check agrees. The extension showing that one extra survey measurement restores identifiability is sensible and well supported by simulation. Credit also for shipping the code.\n\nThe soft spots are modest. The theorem is stated for prevalence Ir(t), but the paper's data and simulations use cumulative reported incidence N1(t). The paper never explicitly proves that identical N1 curves imply identical Ir curves, so Corollary 2 as applied to the actual observation model is not directly proved. This is fixable: with γ known and the same initial condition, N1 determines Ir through dIr/dt = dN1/dt − γ Ir. The authors should state this one-liner. The stress-test note flags this as a possible fatal gap; I think it's a missing explanation rather than a flaw in the underlying claim.\n\nMore concretely: the parameter values in Figure 3 violate the second invariant. For the two listed triples, β*1(1−π1)=1.9·0.7=1.33 and β*2(1−π2)=1.14·1=1.14. The conditions required for identical trajectories are not satisfied, so the figure cannot be illustrating the theorem. The text claims they satisfy both invariants, but the arithmetic says otherwise. This is a localized error, likely a typo in one of the values, but it needs correcting before the figure makes sense.\n\nThe reader's soundness score of 7 is about right. The simulation table shows total SD across 100 epidemics but does not isolate the survey-sampling component; acceptable for a structural identifiability paper but worth a sentence. The assumptions (constant reporting fraction, known γ, known initial conditions) are stated and standard for this literature.\n\nWho is this for? Anyone doing inference from reported incidence curves with hidden immunity or reporting effects, and people who teach identifiability. I would send it to review.","headline":"Solid identifiability result for SIR with under-reporting and prior immunity, with a localized error in the illustrative figure and a small gap between the formal output and the data actually used.","tokens_in":15485,"tokens_out":4515,"would_cite":true,"duration_ms":51794,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["92D30","34A55"],"pacs":[],"model":"deepseek-v4-flash","headline":"Reported incidence alone cannot jointly identify the transmission rate, reporting fraction, and prior immunity; only two parameter combinations are identifiable.","keywords":["structural identifiability","SIR model","under-reporting","prior immunity","reported incidence","seroprevalence survey","basic reproduction number","parameter inference"],"falsifier":"Take two parameter triples that satisfy $\\beta^*_1/p_1=\\beta^*_2/p_2$ and $\\beta^*_1(1-\\pi_1)=\\beta^*_2(1-\\pi_2)$ but start with different reported initial counts $n i_0$, integrate the reduced ODE, and check whether the reported trajectories diverge; divergence would confirm that the theorem's claim depends on a known common starting value. Alternatively, simulate a stochastic outbreak and compute the profile likelihood for $(p,\\pi,\\beta^*)$ from reported incidence: the paper's structural result predicts the profile is flat along the invariant curve, so a clearly curved profile would indicate that the degeneracy does not survive finite-population stochasticity or unknown initial conditions.","tokens_in":14575,"feed_emoji":"🦠","tokens_out":11135,"duration_ms":120966,"temperature":0.7,"pith_summary":"An SIR model that splits infections into reported and unreported and starts with a fraction of the population already immune has three parameters that public-health decisions hinge on: the effective transmission rate $\\beta^*$, the reporting fraction $p$, and the initial immune fraction $\\pi$. The paper proves that a reported-incidence curve alone cannot identify all three. Two parameter triples generate exactly the same reported trajectory whenever they agree on the combinations $\\beta^*/p$ and $\\beta^*(1-\\pi)$, so the data fix only these two quantities; different triples in the same family imply different values of the basic reproduction number $R_0=\\beta^*/\\gamma$. The paper then shows analytically and by simulation that sampling the population to learn either $\\pi$ (prior-immunity serosurvey) or $p$ (prevalence at the epidemic's peak) unlocks all three parameters, with a time-zero immunity survey giving the more precise estimates. The message for practice is that adding case reports will not resolve this ambiguity; adding a different type of data will.","feed_headline":"Proof: reported cases alone cannot fix three epidemic parameters","feed_subtitle":"A proof shows two combinations of transmission, reporting, and immunity are identifiable; a serosurvey fixes all three.","key_machinery":"The carrying device is a one-dimensional reduction of the five-compartment ODE system. Using the identity $\\beta_r I_r+\\beta_u I_u=\\beta^* I_r/p$, the model collapses to a single integro-differential equation for reported infectious individuals, $$\\frac{d}{dt}I_r(t)= I_r(t)\\left[\\$\\beta$^*(1-\\pi)-\\frac{\\$\\beta$^*}{p}\\frac{I_r(0)}{n}\\right]\\exp\\left(-\\frac{\\$\\beta$^*}{p}\\frac1n\\int_0^t I_r(u)\\,du\\right)-\\gamma I_r(t),$$ with initial condition $I_r(0)=n i_0$. Comparing two supposedly identical reported trajectories, evaluating the equation at $t=0$ and then equating the exponential factors forces the two invariants $\\beta^*/p$ and $\\beta^*(1-\\pi)$ to match, and matching them is sufficient for identical reported curves. The same reduction makes it transparent that the other compartments ($S$, $I_u$, $R_r$, $R_u$) are not identical when only the reported curve coincides, so total epidemic size is not determined by the reported curve.","core_discovery":"The central claim is structural: in the deterministic SIR model with under-reporting and prior immunity, the triple $(\\beta^*, p, \\pi)$ is unidentifiable from the time series of reported infectious individuals. Theorem 1 gives the exact degeneracy: if $I_r(t,\\theta_1)=I_r(t,\\theta_2)$ for two parameter sets, then $\\beta^*_1/p_1=\\beta^*_2/p_2$ and $\\beta^*_1(1-\\pi_1)=\\beta^*_2(1-\\pi_2)$; conversely, any two triples satisfying these equalities and starting with the same number of reported infectious cases produce identical reported trajectories. Because the two invariants leave one degree of freedom, infinitely many triples match the same observed growth rate and final reported size, and those triples disagree about $R_0$. The paper also establishes (Corollary 2) that knowing any one of the three parameters removes the degeneracy, and it shows with a 100-epidemic simulation that a random sample of the population, either at time zero to obtain $\\hat{\\pi}$ or at the peak for $\\hat{p}$, lets the remaining parameters be estimated from the exponential growth rate and the final reported size.","pith_inferences":["Because the proof's decisive step evaluates the system at $t=0$, the clean two-invariant structure may break down when surveillance starts after the first cases; left-truncated epidemics would need separate identifiability analysis, and the practical problem could be even worse.","The same degeneracy should be expected in any compartment model in which reported and unreported infectives differ only by the reporting split and share a common recovery rate; testing the same remedies in SEIR and SIRS versions is a natural next step.","Because the unidentifiability is structural, no amount of additional case reporting can resolve it; the design implication is that surveillance budgets should include population surveys rather than only case-count expansion.","A directly testable extension is to compute the profile likelihood for stochastic simulated data along the curve $\\beta^*=p c_1$, $\\pi=1-c_2/(p c_1)$; the theory predicts a flat profile, and a strongly curved profile would signal that stochasticity or unknown initial conditions alter the result."],"forward_implications":["With only reported incidence, $R_0=\\beta^*/\\gamma$ is not estimable: the same growth rate and final reported size are compatible with a continuum of $(p,\\pi,\\beta^*)$ triples, so intervention decisions based on a single fitted $R_0$ are not uniquely supported by the data.","Knowing one parameter suffices: if $p$ or $\\pi$ is supplied by an external source, the remaining two parameters follow directly from the exponential growth rate and final reported size equations, without numerical optimization.","A time-zero serosurvey of prior immunity is a workable remedy: in 100 simulated outbreaks, $\\hat{\\pi}$ was close to the true value and produced accurate estimates of $p$ and $R_0$, with smaller standard deviations than the prevalence-at-the-peak survey.","Ignoring prior immunity can be seriously misleading even when the model fits: the paper's example shows the same reported trajectory is reproduced with $R_0=1.14$ and $\\pi=0$ instead of the true $R_0=1.9$ and $\\pi=0.3$.","Equality of reported curves does not imply equality of the whole epidemic: unreported incidence and total size can differ across parameter sets that generate the identical reported curve."],"supporting_citations":[{"why":"Supplies the structural-versus-practical identifiability distinction that frames the analysis.","marker":"Wieland et al. 2021"},{"why":"Provides the software used to verify structural unidentifiability before the analytical proof.","marker":"Dong et al. 2023"},{"why":"The Taylor-series identifiability approach to which the paper's explicit ODE reformulation is closely related.","marker":"Pohjanpalo 1978"},{"why":"The standard SIR system that the reduced model matches when reported and unreported compartments are pooled.","marker":"Kermack and McKendrick 1927"},{"why":"Supplies the exponential growth rate and final-size equations used to estimate parameters once one parameter is known.","marker":"Diekmann et al. 2013"},{"why":"Gives the algorithm used to simulate the stochastic epidemics that underpin the practical-identifiability study.","marker":"Gillespie 1976"},{"why":"Defines the general parametrized dynamical system and the structural identifiability condition stated in Equation (1).","marker":"Villaverde et al. 2016"}],"fun_headline_variants":["Reported cases alone can't pin three epidemic parameters","Proof: under-reporting and prior immunity defeat estimation","Three epidemic parameters need serosurvey data to be identifiable","Epidemic model shows case data insufficient for key estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof assumes the epidemic is observed from time zero with a known number of initially reported infectious cases $n i_0$ and a known population size $n$, because the decisive relation is obtained by evaluating the ODE at $t=0$; if the start of the outbreak is missing or the initial count is uncertain, the exact invariant relationships of Theorem 1 are not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Reported cases alone can't pin three epidemic parameters","Proof: under-reporting and prior immunity defeat estimation","Three epidemic parameters need serosurvey data to be identifiable","Epidemic model shows case data insufficient for key estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1307,"prompt_tokens":959,"completion_tokens":348,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":282}},"tokens_in":575,"tokens_out":348,"duration_ms":4761,"temperature":1.0,"reasoning_tokens":282,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:23:58.406272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take two parameter triples that satisfy $\\beta^*_1/p_1=\\beta^*_2/p_2$ and $\\beta^*_1(1-\\pi_1)=\\beta^*_2(1-\\pi_2)$ but start with different reported initial counts $n i_0$, integrate the reduced ODE, and check whether the reported trajectories diverge; divergence would confirm that the theorem's claim depends on a known common starting value. Alternatively, simulate a stochastic outbreak and compute the profile likelihood for $(p,\\pi,\\beta^*)$ from reported incidence: the paper's structural result predicts the profile is flat along the invariant curve, so a clearly curved profile would indicate that the degeneracy does not survive finite-population stochasticity or unknown initial conditions.","supporting_citations":[],"review_version":1}