Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Identifiability in epidemic models with prior immunity and under-reporting

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Reported incidence alone cannot jointly identify the transmission rate, reporting fraction, and prior immunity; only two parameter combinations are identifiable.

desk verdict Solid identifiability result for SIR with under-reporting and prior immunity, with a localized error in the illustrative figure and a small gap between the formal output and the data actually used. read the letter →

arxiv 2506.07825 v2 pith:OTDKUTN4 submitted 2025-06-09 stat.ME

classification stat.ME MSC 92D3034A55
keywords structuralidentifiabilitySIRmodelunder-reportingpriorimmunityreportedincidenceseroprevalencesurveybasicreproductionnumberparameterinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

An SIR model that splits infections into reported and unreported and starts with a fraction of the population already immune has three parameters that public-health decisions hinge on: the effective transmission rate $\beta^*$, the reporting fraction $p$, and the initial immune fraction $\pi$. The paper proves that a reported-incidence curve alone cannot identify all three. Two parameter triples generate exactly the same reported trajectory whenever they agree on the combinations $\beta^*/p$ and $\beta^*(1-\pi)$, so the data fix only these two quantities; different triples in the same family imply different values of the basic reproduction number $R_0=\beta^*/\gamma$. The paper then shows analytically and by simulation that sampling the population to learn either $\pi$ (prior-immunity serosurvey) or $p$ (prevalence at the epidemic's peak) unlocks all three parameters, with a time-zero immunity survey giving the more precise estimates. The message for practice is that adding case reports will not resolve this ambiguity; adding a different type of data will.

What carries the argument

The carrying device is a one-dimensional reduction of the five-compartment ODE system. Using the identity $\beta_r I_r+\beta_u I_u=\beta^* I_r/p$, the model collapses to a single integro-differential equation for reported infectious individuals, $$\frac{d}{dt}I_r(t)= I_r(t)\left[\$\beta$^*(1-\pi)-\frac{\$\beta$^*}{p}\frac{I_r(0)}{n}\right]\exp\left(-\frac{\$\beta$^*}{p}\frac1n\int_0^t I_r(u)\,du\right)-\gamma I_r(t),$$ with initial condition $I_r(0)=n i_0$. Comparing two supposedly identical reported trajectories, evaluating the equation at $t=0$ and then equating the exponential factors forces the two invariants $\beta^*/p$ and $\beta^*(1-\pi)$ to match, and matching them is sufficient for identical reported curves. The same reduction makes it transparent that the other compartments ($S$, $I_u$, $R_r$, $R_u$) are not identical when only the reported curve coincides, so total epidemic size is not determined by the reported curve.

What would settle it

Take two parameter triples that satisfy $\beta^*_1/p_1=\beta^*_2/p_2$ and $\beta^*_1(1-\pi_1)=\beta^*_2(1-\pi_2)$ but start with different reported initial counts $n i_0$, integrate the reduced ODE, and check whether the reported trajectories diverge; divergence would confirm that the theorem's claim depends on a known common starting value. Alternatively, simulate a stochastic outbreak and compute the profile likelihood for $(p,\pi,\beta^*)$ from reported incidence: the paper's structural result predicts the profile is flat along the invariant curve, so a clearly curved profile would indicate that the degeneracy does not survive finite-population stochasticity or unknown initial conditions.

Watch

Extended reading notes

Core claim

The central claim is structural: in the deterministic SIR model with under-reporting and prior immunity, the triple $(\beta^*, p, \pi)$ is unidentifiable from the time series of reported infectious individuals. Theorem 1 gives the exact degeneracy: if $I_r(t,\theta_1)=I_r(t,\theta_2)$ for two parameter sets, then $\beta^*_1/p_1=\beta^*_2/p_2$ and $\beta^*_1(1-\pi_1)=\beta^*_2(1-\pi_2)$; conversely, any two triples satisfying these equalities and starting with the same number of reported infectious cases produce identical reported trajectories. Because the two invariants leave one degree of freedom, infinitely many triples match the same observed growth rate and final reported size, and those triples disagree about $R_0$. The paper also establishes (Corollary 2) that knowing any one of the three parameters removes the degeneracy, and it shows with a 100-epidemic simulation that a random sample of the population, either at time zero to obtain $\hat{\pi}$ or at the peak for $\hat{p}$, lets the remaining parameters be estimated from the exponential growth rate and the final reported size.

Load-bearing premise

The proof assumes the epidemic is observed from time zero with a known number of initially reported infectious cases $n i_0$ and a known population size $n$, because the decisive relation is obtained by evaluating the ODE at $t=0$; if the start of the outbreak is missing or the initial count is uncertain, the exact invariant relationships of Theorem 1 are not guaranteed.

Editorial extensions

If this is right

  • With only reported incidence, $R_0=\beta^*/\gamma$ is not estimable: the same growth rate and final reported size are compatible with a continuum of $(p,\pi,\beta^*)$ triples, so intervention decisions based on a single fitted $R_0$ are not uniquely supported by the data.
  • Knowing one parameter suffices: if $p$ or $\pi$ is supplied by an external source, the remaining two parameters follow directly from the exponential growth rate and final reported size equations, without numerical optimization.
  • A time-zero serosurvey of prior immunity is a workable remedy: in 100 simulated outbreaks, $\hat{\pi}$ was close to the true value and produced accurate estimates of $p$ and $R_0$, with smaller standard deviations than the prevalence-at-the-peak survey.
  • Ignoring prior immunity can be seriously misleading even when the model fits: the paper's example shows the same reported trajectory is reproduced with $R_0=1.14$ and $\pi=0$ instead of the true $R_0=1.9$ and $\pi=0.3$.
  • Equality of reported curves does not imply equality of the whole epidemic: unreported incidence and total size can differ across parameter sets that generate the identical reported curve.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the proof's decisive step evaluates the system at $t=0$, the clean two-invariant structure may break down when surveillance starts after the first cases; left-truncated epidemics would need separate identifiability analysis, and the practical problem could be even worse.
  • The same degeneracy should be expected in any compartment model in which reported and unreported infectives differ only by the reporting split and share a common recovery rate; testing the same remedies in SEIR and SIRS versions is a natural next step.
  • Because the unidentifiability is structural, no amount of additional case reporting can resolve it; the design implication is that surveillance budgets should include population surveys rather than only case-count expansion.
  • A directly testable extension is to compute the profile likelihood for stochastic simulated data along the curve $\beta^*=p c_1$, $\pi=1-c_2/(p c_1)$; the theory predicts a flat profile, and a strongly curved profile would signal that stochasticity or unknown initial conditions alter the result.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper analyses structural identifiability in an SIR model with under-reporting (fraction p) and pre-existing immunity (fraction pi), using an effective transmission rate beta*. Theorem 1 (Appendix D) claims that the trajectory of reported infectious individuals Ir(t) identifies only the two combinations beta*/p and beta*(1-pi), so (beta*, p, pi) are not jointly identifiable; the converse is also proved. Corollary 2 states that knowledge of any one of the three parameters makes the triple identifiable. The paper then uses the initial-growth equation (5) and final-size equation (7), combined with sample-survey estimates of pi or p, to estimate the remaining parameters in 100 simulated stochastic epidemics (Table 2).

Significance. If the bridge to cumulative reported incidence is made explicit, the paper gives a clean and useful characterization of unidentifiability in a practically relevant model. The analytic proof is self-contained, includes the converse direction, is cross-checked with StructuralIdentifiability.jl, and the simulation code is publicly available. The paper will be a useful reference for modellers deciding what can and cannot be inferred from reported case curves. The main caveat is that Theorem 1 is stated for Ir(t) while the data are N1(t); this gap is easy to close but must be written down.

major comments (2)
  1. [Appendix D, Theorem 1 and Corollary 2] The theorem and corollary are stated for the trajectory of reported infectious individuals Ir(t), but the data considered in the paper (Section 2.4, Figure 2) are the cumulative reported incidence process N1(t). The proof never establishes that identical N1 trajectories imply identical Ir trajectories at the level of the deterministic model. This is true, because from (2), N1'(t) = beta* Ir(t) S(t)/n = dIr(t)/dt + gamma Ir(t); with gamma and Ir(0) known, N1(t) uniquely determines Ir(t). The paper should state this equivalence explicitly (as a lemma or remark) before applying Theorem 1 and Corollary 2 to reported incidence data, otherwise the central claim that 'only reported case data are available' is not directly supported by the proof.
  2. [Section 3.1 and Figure 3] The parameter values used to illustrate Theorem 1 do not satisfy the theorem's conditions. For the first set {p1=0.4, pi1=0.3, beta*1=1.9} and the second set {p2=0.24, pi2=0.0, beta*2=1.14}, condition (16) holds (1.9/0.4 = 1.14/0.24 = 4.75), but condition (17) does not: beta*1(1-pi1)=1.33 whereas beta*2(1-pi2)=1.14. The two trajectories therefore cannot be identical; indeed their initial growth rates differ (0.33 vs 0.14 when gamma=1). Please correct the figure and the surrounding text, for example by taking beta*2=1.33 and p2=0.28, which satisfy both conditions.
minor comments (5)
  1. [Section 2.3] The text contains the duplicated phrase 'a linear relationship a linear relationship' in the description of the log-linear regression; please fix this typo.
  2. [Section 1] The phrase 'Taylor?series' appears to be a conversion artifact; please correct it to 'Taylor series'.
  3. [Section 3.2 and Table 2] The reported standard deviations reflect variation across the 100 simulated epidemics but do not propagate the uncertainty in the estimated growth rate rho-hat and final reported size z-hat-r from the regression and final-size estimation, nor the finite-sample survey uncertainty beyond what is captured by the simulation design; this limitation should be stated explicitly.
  4. [Appendix D] The derivation leading to equation (16) assumes that the common coefficient in equation (14) is nonzero; the degenerate case of zero initial growth (where the coefficient vanishes) is not treated. Please either exclude this case in the theorem statement or add a short argument for it.
  5. [Appendix D, equation (15)] The proof evaluates the ODE at t=0, which requires observing the epidemic from its start with a known initial reported infectious count ni0 and known population size n; the Discussion acknowledges these assumptions, but it would be helpful to state this dependence explicitly next to Theorem 1.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the unidentifiability proof flows from the ODEs themselves, and the simulation study re-estimates parameters rather than relabeling a fit as a prediction.

full rationale

The central unidentifiability result is not circular. The proof in Appendix D starts from the model ODEs (2)-(3), rewrites them as equation (12) using the identity beta_r I_r + beta_u I_u = (beta*/p) I_r, and then derives equations (15)-(17) by evaluating the identity at t=0 and comparing the remaining exponentials. The 'vice versa' direction is likewise proved directly: if (16)-(17) hold, the ODE for I_r(t) becomes identical, so the trajectories coincide. Corollary 2 is simple algebra from the two invariant combinations. No fitted parameter is renamed as a prediction: Section 3.2 estimates rho and z_r from simulated epidemics and then solves the closed-form equations (5) and (7) to recover the remaining parameters; this is a Monte Carlo consistency study, not a prediction forced by the model's own fitted output. The only citation involving an author of this paper (Diekmann et al. 2013) supplies standard exponential-growth and final-size relations, which are classical and not load-bearing. There are non-circular correctness concerns worth flagging: the theorem is stated for the prevalence I_r(t), while the data and simulation discussion use cumulative reported incidence N_1(t), and no direct proof shows that identical N_1(t) implies (16)-(17); also, Figure 3's printed second set {p_2=0.24, beta*_2=1.14} violates condition (17), since 1.9(1-0.3)=1.33 is not equal to 1.14(1-0)=1.14. These are internal-consistency and observation-model issues, not circularity, so they do not raise the circularity score above the minor-citation band.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central identifiability theorem rests on the model structure in (2)-(3): a closed, homogeneously mixing SIR model with constant reporting probability p, known gamma, known n, and known initial reported cases ni0. These are standard domain assumptions for structural identifiability analysis; none are fitted to data.

assumptions (5)
  • domain assumption The population is closed and homogeneously mixing, with mass-action incidence.
    Section 2.2 states homogeneous mixing and a closed population of size n; the ODEs in (2) rely on this.
  • domain assumption Reporting fraction p is constant over time and applies identically to all infectious individuals.
    Section 2.2 defines p as a fraction of infectious individuals reported; the proof and estimation assume p does not vary with time or disease state. The Discussion acknowledges this as a limitation.
  • domain assumption The recovery rate gamma is known.
    Section 2.4 states 'we consider the recovery rate gamma to be estimated separately and here assumed to be a known quantity'; the identifiability result holds conditional on gamma.
  • domain assumption The initial number of reported infectious individuals ni0 and population size n are known.
    Initial conditions (3) and the proof in Appendix D use Ir(0)=ni0 and known n; if ni0 is unknown, the parameter relationships could differ.
  • standard math Standard existence and uniqueness of solutions to the ODE system (2).
    The proof of Theorem 1 relies on comparing solutions of the rewritten ODE (12).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Identifiability in epidemic models with prior immunity and under-reporting." pith.science (2026). https://pith.science/paper/OTDKUTN4

@misc{pith2026250607825,
  author       = {Pith},
  title        = {Pith review of: Identifiability in epidemic models with prior immunity and under-reporting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OTDKUTN4}},
  note         = {Machine review of arXiv:2506.07825}
}
read the original abstract

Identifiability is the property in mathematical modelling that determines if model parameters can be uniquely estimated from data. For infectious disease models, failure to ensure identifiability can lead to misleading parameter estimates and unreliable policy recommendations. We examine the identifiability of a modified SIR model that accounts for under-reporting and pre-existing immunity in the population. We provide a mathematical proof of the unidentifiability of jointly estimating three parameters: the fraction under-reporting, the proportion of the population with prior immunity, and the community transmission rate, when only reported case data are available. We then show, analytically and with a simulation study, that the identifiability of all three parameters is achieved if the reported incidence is complemented with sample survey data of prior immunity or prevalence during the outbreak. Our results show the limitations of parameter inference in partially observed epidemics and the importance of identifiability analysis when developing and applying models for public health decision making.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Epidemic "momentum" and a conservation law for infectious disease dynamics

    q-bio.PE 2025-11 conditional novelty 8.0 of 10

    A conserved quantity, epidemic momentum, reveals that all renewal-equation outbreak models share a single phase-plane geometry and yields closed-form estimators for R0 and pre-outbreak immunity from incidence data.

Reference graph

Works this paper leans on

3 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    JAMA 14:1406–1407

    Bai Y , Y ao L, Wei T, Tian F, Jin DY , Chen L et al (2020) Presumed Asymptomatic Carrier Transmission of COVID-19. JAMA 14:1406–1407. https://doi.org/10.1001/jama.2020.2565 Bellman R, Åström KJ (1970) On structural identifiability. Math Biosci 7(3–4):329–339. https://doi.org/ 10.1016/0025-5564(70)90132-X Bellu G, Saccomani MP , Audoly S, D’Angiò L (2007) ...

  2. [14]

    Epidemics 41:100643

    https://doi.org/10.1186/1471-2334-14-480 Dankwa E, Brouwer A, Donnelly C (2022) Structural identifiability of compartmental models for infec- tious disease transmission is influenced by data type. Epidemics 41:100643. https://doi.org/10.1016/ j.epidem.2022.100643 Diekmann O (1977) Limiting behaviour in an epidemic model. Nonlinear Analysis Theory Methods & ...

  3. [82]

    Epidemics 25:89–100

    https://doi.org/10.1007/s11538-020-00747-6 Kao YH, Eisenberg MC (2018) Practical unidentifiability of a simple vector-borne disease model: Impli- cations for parameter estimation and intervention assessment. Epidemics 25:89–100. https://doi.org/ 10.1016/j.epidem.2018.05.010 Kermack WO, McKendrick AG (1927) A contribution to the mathematical theory of epide...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.