REVIEW 2 major objections 5 minor 1 cited by
Identifiability in epidemic models with prior immunity and under-reporting
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Reported incidence alone cannot jointly identify the transmission rate, reporting fraction, and prior immunity; only two parameter combinations are identifiable.
desk verdict Solid identifiability result for SIR with under-reporting and prior immunity, with a localized error in the illustrative figure and a small gap between the formal output and the data actually used. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying device is a one-dimensional reduction of the five-compartment ODE system. Using the identity $\beta_r I_r+\beta_u I_u=\beta^* I_r/p$, the model collapses to a single integro-differential equation for reported infectious individuals, $$\frac{d}{dt}I_r(t)= I_r(t)\left[\$\beta$^*(1-\pi)-\frac{\$\beta$^*}{p}\frac{I_r(0)}{n}\right]\exp\left(-\frac{\$\beta$^*}{p}\frac1n\int_0^t I_r(u)\,du\right)-\gamma I_r(t),$$ with initial condition $I_r(0)=n i_0$. Comparing two supposedly identical reported trajectories, evaluating the equation at $t=0$ and then equating the exponential factors forces the two invariants $\beta^*/p$ and $\beta^*(1-\pi)$ to match, and matching them is sufficient for identical reported curves. The same reduction makes it transparent that the other compartments ($S$, $I_u$, $R_r$, $R_u$) are not identical when only the reported curve coincides, so total epidemic size is not determined by the reported curve.
What would settle it
Take two parameter triples that satisfy $\beta^*_1/p_1=\beta^*_2/p_2$ and $\beta^*_1(1-\pi_1)=\beta^*_2(1-\pi_2)$ but start with different reported initial counts $n i_0$, integrate the reduced ODE, and check whether the reported trajectories diverge; divergence would confirm that the theorem's claim depends on a known common starting value. Alternatively, simulate a stochastic outbreak and compute the profile likelihood for $(p,\pi,\beta^*)$ from reported incidence: the paper's structural result predicts the profile is flat along the invariant curve, so a clearly curved profile would indicate that the degeneracy does not survive finite-population stochasticity or unknown initial conditions.
Extended reading notes
Core claim
The central claim is structural: in the deterministic SIR model with under-reporting and prior immunity, the triple $(\beta^*, p, \pi)$ is unidentifiable from the time series of reported infectious individuals. Theorem 1 gives the exact degeneracy: if $I_r(t,\theta_1)=I_r(t,\theta_2)$ for two parameter sets, then $\beta^*_1/p_1=\beta^*_2/p_2$ and $\beta^*_1(1-\pi_1)=\beta^*_2(1-\pi_2)$; conversely, any two triples satisfying these equalities and starting with the same number of reported infectious cases produce identical reported trajectories. Because the two invariants leave one degree of freedom, infinitely many triples match the same observed growth rate and final reported size, and those triples disagree about $R_0$. The paper also establishes (Corollary 2) that knowing any one of the three parameters removes the degeneracy, and it shows with a 100-epidemic simulation that a random sample of the population, either at time zero to obtain $\hat{\pi}$ or at the peak for $\hat{p}$, lets the remaining parameters be estimated from the exponential growth rate and the final reported size.
Load-bearing premise
The proof assumes the epidemic is observed from time zero with a known number of initially reported infectious cases $n i_0$ and a known population size $n$, because the decisive relation is obtained by evaluating the ODE at $t=0$; if the start of the outbreak is missing or the initial count is uncertain, the exact invariant relationships of Theorem 1 are not guaranteed.
Editorial extensions
If this is right
- With only reported incidence, $R_0=\beta^*/\gamma$ is not estimable: the same growth rate and final reported size are compatible with a continuum of $(p,\pi,\beta^*)$ triples, so intervention decisions based on a single fitted $R_0$ are not uniquely supported by the data.
- Knowing one parameter suffices: if $p$ or $\pi$ is supplied by an external source, the remaining two parameters follow directly from the exponential growth rate and final reported size equations, without numerical optimization.
- A time-zero serosurvey of prior immunity is a workable remedy: in 100 simulated outbreaks, $\hat{\pi}$ was close to the true value and produced accurate estimates of $p$ and $R_0$, with smaller standard deviations than the prevalence-at-the-peak survey.
- Ignoring prior immunity can be seriously misleading even when the model fits: the paper's example shows the same reported trajectory is reproduced with $R_0=1.14$ and $\pi=0$ instead of the true $R_0=1.9$ and $\pi=0.3$.
- Equality of reported curves does not imply equality of the whole epidemic: unreported incidence and total size can differ across parameter sets that generate the identical reported curve.
Reading between the lines
- Because the proof's decisive step evaluates the system at $t=0$, the clean two-invariant structure may break down when surveillance starts after the first cases; left-truncated epidemics would need separate identifiability analysis, and the practical problem could be even worse.
- The same degeneracy should be expected in any compartment model in which reported and unreported infectives differ only by the reporting split and share a common recovery rate; testing the same remedies in SEIR and SIRS versions is a natural next step.
- Because the unidentifiability is structural, no amount of additional case reporting can resolve it; the design implication is that surveillance budgets should include population surveys rather than only case-count expansion.
- A directly testable extension is to compute the profile likelihood for stochastic simulated data along the curve $\beta^*=p c_1$, $\pi=1-c_2/(p c_1)$; the theory predicts a flat profile, and a strongly curved profile would signal that stochasticity or unknown initial conditions alter the result.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyses structural identifiability in an SIR model with under-reporting (fraction p) and pre-existing immunity (fraction pi), using an effective transmission rate beta*. Theorem 1 (Appendix D) claims that the trajectory of reported infectious individuals Ir(t) identifies only the two combinations beta*/p and beta*(1-pi), so (beta*, p, pi) are not jointly identifiable; the converse is also proved. Corollary 2 states that knowledge of any one of the three parameters makes the triple identifiable. The paper then uses the initial-growth equation (5) and final-size equation (7), combined with sample-survey estimates of pi or p, to estimate the remaining parameters in 100 simulated stochastic epidemics (Table 2).
Significance. If the bridge to cumulative reported incidence is made explicit, the paper gives a clean and useful characterization of unidentifiability in a practically relevant model. The analytic proof is self-contained, includes the converse direction, is cross-checked with StructuralIdentifiability.jl, and the simulation code is publicly available. The paper will be a useful reference for modellers deciding what can and cannot be inferred from reported case curves. The main caveat is that Theorem 1 is stated for Ir(t) while the data are N1(t); this gap is easy to close but must be written down.
major comments (2)
- [Appendix D, Theorem 1 and Corollary 2] The theorem and corollary are stated for the trajectory of reported infectious individuals Ir(t), but the data considered in the paper (Section 2.4, Figure 2) are the cumulative reported incidence process N1(t). The proof never establishes that identical N1 trajectories imply identical Ir trajectories at the level of the deterministic model. This is true, because from (2), N1'(t) = beta* Ir(t) S(t)/n = dIr(t)/dt + gamma Ir(t); with gamma and Ir(0) known, N1(t) uniquely determines Ir(t). The paper should state this equivalence explicitly (as a lemma or remark) before applying Theorem 1 and Corollary 2 to reported incidence data, otherwise the central claim that 'only reported case data are available' is not directly supported by the proof.
- [Section 3.1 and Figure 3] The parameter values used to illustrate Theorem 1 do not satisfy the theorem's conditions. For the first set {p1=0.4, pi1=0.3, beta*1=1.9} and the second set {p2=0.24, pi2=0.0, beta*2=1.14}, condition (16) holds (1.9/0.4 = 1.14/0.24 = 4.75), but condition (17) does not: beta*1(1-pi1)=1.33 whereas beta*2(1-pi2)=1.14. The two trajectories therefore cannot be identical; indeed their initial growth rates differ (0.33 vs 0.14 when gamma=1). Please correct the figure and the surrounding text, for example by taking beta*2=1.33 and p2=0.28, which satisfy both conditions.
minor comments (5)
- [Section 2.3] The text contains the duplicated phrase 'a linear relationship a linear relationship' in the description of the log-linear regression; please fix this typo.
- [Section 1] The phrase 'Taylor?series' appears to be a conversion artifact; please correct it to 'Taylor series'.
- [Section 3.2 and Table 2] The reported standard deviations reflect variation across the 100 simulated epidemics but do not propagate the uncertainty in the estimated growth rate rho-hat and final reported size z-hat-r from the regression and final-size estimation, nor the finite-sample survey uncertainty beyond what is captured by the simulation design; this limitation should be stated explicitly.
- [Appendix D] The derivation leading to equation (16) assumes that the common coefficient in equation (14) is nonzero; the degenerate case of zero initial growth (where the coefficient vanishes) is not treated. Please either exclude this case in the theorem statement or add a short argument for it.
- [Appendix D, equation (15)] The proof evaluates the ODE at t=0, which requires observing the epidemic from its start with a known initial reported infectious count ni0 and known population size n; the Discussion acknowledges these assumptions, but it would be helpful to state this dependence explicitly next to Theorem 1.
Circularity Check
No circular derivation: the unidentifiability proof flows from the ODEs themselves, and the simulation study re-estimates parameters rather than relabeling a fit as a prediction.
full rationale
The central unidentifiability result is not circular. The proof in Appendix D starts from the model ODEs (2)-(3), rewrites them as equation (12) using the identity beta_r I_r + beta_u I_u = (beta*/p) I_r, and then derives equations (15)-(17) by evaluating the identity at t=0 and comparing the remaining exponentials. The 'vice versa' direction is likewise proved directly: if (16)-(17) hold, the ODE for I_r(t) becomes identical, so the trajectories coincide. Corollary 2 is simple algebra from the two invariant combinations. No fitted parameter is renamed as a prediction: Section 3.2 estimates rho and z_r from simulated epidemics and then solves the closed-form equations (5) and (7) to recover the remaining parameters; this is a Monte Carlo consistency study, not a prediction forced by the model's own fitted output. The only citation involving an author of this paper (Diekmann et al. 2013) supplies standard exponential-growth and final-size relations, which are classical and not load-bearing. There are non-circular correctness concerns worth flagging: the theorem is stated for the prevalence I_r(t), while the data and simulation discussion use cumulative reported incidence N_1(t), and no direct proof shows that identical N_1(t) implies (16)-(17); also, Figure 3's printed second set {p_2=0.24, beta*_2=1.14} violates condition (17), since 1.9(1-0.3)=1.33 is not equal to 1.14(1-0)=1.14. These are internal-consistency and observation-model issues, not circularity, so they do not raise the circularity score above the minor-citation band.
Assumptions & free parameters
assumptions (5)
- domain assumption The population is closed and homogeneously mixing, with mass-action incidence.
- domain assumption Reporting fraction p is constant over time and applies identically to all infectious individuals.
- domain assumption The recovery rate gamma is known.
- domain assumption The initial number of reported infectious individuals ni0 and population size n are known.
- standard math Standard existence and uniqueness of solutions to the ODE system (2).
Cite this review
Pith. "Pith review of Identifiability in epidemic models with prior immunity and under-reporting." pith.science (2026). https://pith.science/paper/OTDKUTN4
@misc{pith2026250607825,
author = {Pith},
title = {Pith review of: Identifiability in epidemic models with prior immunity and under-reporting},
year = {2026},
howpublished = {\url{https://pith.science/paper/OTDKUTN4}},
note = {Machine review of arXiv:2506.07825}
}
read the original abstract
Identifiability is the property in mathematical modelling that determines if model parameters can be uniquely estimated from data. For infectious disease models, failure to ensure identifiability can lead to misleading parameter estimates and unreliable policy recommendations. We examine the identifiability of a modified SIR model that accounts for under-reporting and pre-existing immunity in the population. We provide a mathematical proof of the unidentifiability of jointly estimating three parameters: the fraction under-reporting, the proportion of the population with prior immunity, and the community transmission rate, when only reported case data are available. We then show, analytically and with a simulation study, that the identifiability of all three parameters is achieved if the reported incidence is complemented with sample survey data of prior immunity or prevalence during the outbreak. Our results show the limitations of parameter inference in partially observed epidemics and the importance of identifiability analysis when developing and applying models for public health decision making.
Forward citations
Cited by 1 Pith paper
-
Epidemic "momentum" and a conservation law for infectious disease dynamics
A conserved quantity, epidemic momentum, reveals that all renewal-equation outbreak models share a single phase-plane geometry and yields closed-form estimators for R0 and pre-outbreak immunity from incidence data.
Reference graph
Works this paper leans on
-
[1]
Bai Y , Y ao L, Wei T, Tian F, Jin DY , Chen L et al (2020) Presumed Asymptomatic Carrier Transmission of COVID-19. JAMA 14:1406–1407. https://doi.org/10.1001/jama.2020.2565 Bellman R, Åström KJ (1970) On structural identifiability. Math Biosci 7(3–4):329–339. https://doi.org/ 10.1016/0025-5564(70)90132-X Bellu G, Saccomani MP , Audoly S, D’Angiò L (2007) ...
-
[14]
https://doi.org/10.1186/1471-2334-14-480 Dankwa E, Brouwer A, Donnelly C (2022) Structural identifiability of compartmental models for infec- tious disease transmission is influenced by data type. Epidemics 41:100643. https://doi.org/10.1016/ j.epidem.2022.100643 Diekmann O (1977) Limiting behaviour in an epidemic model. Nonlinear Analysis Theory Methods & ...
arXiv 2022
-
[82]
https://doi.org/10.1007/s11538-020-00747-6 Kao YH, Eisenberg MC (2018) Practical unidentifiability of a simple vector-borne disease model: Impli- cations for parameter estimation and intervention assessment. Epidemics 25:89–100. https://doi.org/ 10.1016/j.epidem.2018.05.010 Kermack WO, McKendrick AG (1927) A contribution to the mathematical theory of epide...
arXiv 2018
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.