Pith. sign in

REVIEW 4 major objections 4 minor 33 references

The paper establishes that identifiability in epidemic transmission models is an emergent property of the interaction between latent dynamics, network structure, and the observation process, and it makes this precise through the decompositi

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 03:38 UTC pith:TELA3GV6

load-bearing objection Standard identifiability/missing-information machinery is mostly sound and usefully synthesized for network epidemics, but the paper's new beta/xi confounding result is assumed rather than proved, and the simulations don't back the strong claims. the 4 major comments →

arxiv 2607.23079 v1 pith:TELA3GV6 submitted 2026-07-25 math.ST stat.TH

Identifiability and Information-Based Inference for Epidemic Transmission Models Under Partial Observation

classification math.ST stat.TH MSC 62M0562F1262B10
keywords identifiabilityFisher informationdynamic contact networkspartially observed stochastic processesmissing informationSEIR modelexternal infection confoundingobservation design
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that whether per-contact transmission rate β and external infection rate ξ can be estimated is not fixed by the epidemic mechanism alone; it is co-determined by how much of the latent epidemic–network trajectory the observation process captures. When surveillance data carry only an aggregate exposure summary, the observed-data law depends on (β, ξ) through the single combination λ_agg = β·C̄ + ξ, so the two parameters are not separately identifiable unless exposure itself is identifiable from the data. To quantify this, the paper derives the missing-information decomposition I_C = I_O + I_M, separating complete-data information from information lost to latent infection times, missing contact histories, and measurement error, and uses it to define relative information and identifiability phase boundaries. Simulations in high, moderate, and sparse observation regimes show the theoretical information measures track finite-sample bias, RMSE, coverage, and interval width, which matters because the framework offers a way to choose surveillance frequency, network coverage, and reporting accuracy before data collection.

Core claim

The central claim is that identifiability of epidemic transmission parameters is not a property of the SEIR mechanism alone but of the joint system formed by the latent epidemic–network process and the observation process. The paper formalizes this by defining structural and local identifiability through the observed-data law, and by proving that if infection times and contact histories are unobserved and the observation process reduces the latent process to an aggregate exposure summary C̄(t), then the observed-data law depends on (β, ξ) only through λ_agg(t) = β C̄(t) + ξ, making β and ξ observationally equivalent unless C̄(t) varies in a way recoverable from the data. The complementary mi

What carries the argument

The load-bearing object is the observed-data law P^O_θ, the distribution over symptom and contact reports obtained by integrating out latent infection times and unobserved network trajectories; identifiability is defined as injectivity of the map θ ↦ P^O_θ. On top of this sits the missing-information identity I_C(θ) = I_O(θ) + I_M(θ), where I_M is the expected conditional variance of the complete-data score given the observations. The identity makes observation-driven information loss quantitative: local identifiability holds only if I_O(θ0) has full rank. The second mechanism is the aggregate exposure summary λ_agg(t) = β C̄(t) + ξ, which the paper shows is the only channel through which (β

Load-bearing premise

The results separating internal transmission from external infection rest on the assumption that the observed-data law depends on (β, ξ) only through the aggregate exposure summary λ_agg(t) = β C̄(t) + ξ, an aggregate-only condition the paper assumes rather than derives; a second unproved premise is that missing information I_M tends to zero as observation spacing Δ → 0.

What would settle it

Simulate a known SEIR dynamic-network process with fixed β and ξ, then estimate under two observation designs: one recording only aggregate prevalence plus mean exposure C̄(t), and one recording individual-level contact exposures. If the individual-level design separates β and ξ while the aggregate design cannot, Theorem 2's mechanism is supported; if either both succeed or both fail, the aggregate-only summary is not the operative source of confounding. Separately, compute I_M(θ; Δ) numerically for shrinking Δ and check whether it actually converges to zero; a single sequence showing it does

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the confounding theorem is correct, surveillance systems that record only aggregate epidemic curves cannot separate internal transmission from external importation; the estimable object is the total infection pressure β·C̄ + ξ.
  • The identity I_C = I_O + I_M assigns a quantitative price to missing infection times and contact histories: the information lost is exactly the expected conditional variance of the complete-data score.
  • The phase-diagram criterion — λ_min(I_O) above, within, or at zero — provides a pre-data rule for classifying observation regimes as identifiable, weakly identifiable, or non-identifiable.
  • Under increasing population size, consistency requires the normalized information matrix N^{-1}I_O to converge to a positive-definite limit; then information grows in every direction, with dense networks accumulating transmission information faster (O(N²) susceptible–infectious contacts) than sparse ones (O(N)).
  • Relative information R = I_C^{-1/2} I_O I_C^{-1/2}, whose eigenvalues lie in [0,1], is a single summary that predicts estimation quality; simulations show retention dropping from 0.90 (high) to 0.25 (sparse) with matching deteriorations in RMSE and interval width.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The aggregate-exposure assumption suggests a direct design corollary the paper leaves implicit: collecting contact-level rather than aggregate exposure data should break the β–ξ equivalence class whenever individual-level variation in exposure is retained in the observations.
  • Read as a design tool, the framework implies an information-optimal surveillance criterion — maximize λ_min(I_O) or relative information subject to a budget on observation frequency, network coverage, and reporting accuracy — which could be tested on real outbreak data.
  • The same missing-information decomposition should transfer to other partially observed interacting systems, such as information diffusion or behavioral contagion on networks, where latent event times and unobserved interaction structures create analogous confounds between spontaneous and contact-driven adoption.
  • A useful testable extension is to verify numerically whether I_M(θ; Δ) actually vanishes as Δ→0 for the specific contact-observation model; the paper asserts this convergence for 'regular observation schemes' without proof, so a counterexample or a formal proof would settle whether continuous monitoring truly restores complete-data identifiability.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper develops a framework for identifiability and Fisher information in SEIR-type epidemic models on dynamic contact networks under partial observation. It defines structural and local identifiability through the observed-data law, introduces a Louis-type decomposition I_C = I_O + I_M to quantify information loss, and uses this to discuss observation regimes. The central claimed result is Theorem 2, which states that the transmission rate β and external infection rate ξ are not separately structurally identifiable when the observed-data law depends on (β, ξ) only through the aggregate exposure λ_agg(t) = β Cbar(t) + ξ. The paper also states asymptotic identifiability results for increasing population size and observation frequency, and presents simulations in high, moderate, and sparse observation regimes.

Significance. If the main results were established, the paper would provide a useful framework for linking identifiability and information loss to surveillance design in epidemic-network models. The Louis identity is correctly invoked and the decomposition is a helpful organizational tool. However, the paper's headline theoretical result, Theorem 2, is not proved from the Section 2 model: the aggregate-only dependence premise is asserted, not derived, and the proof treats the latent exposure process as a fixed covariate. The asymptotic results in Section 6 are largely definitional or rely on unproved convergence assumptions. The simulations are suggestive but do not validate the key premise of Theorem 2 and lack Monte Carlo uncertainty quantification. The paper's contribution is therefore conditional; with substantial revision it could become a useful reference, but in its current form the central claim is unestablished.

major comments (4)
  1. [§5.4, Theorem 2] Theorem 2's conclusion is essentially assumed. The aggregate-only condition that the observed-data law depends on (β, ξ) through λ_agg(t) = β Cbar(t) + ξ is not derived from the Section 2 observation model, which records individual symptom reports Y_i(t_r) and dyad contacts B_ij(t_r). After integrating over latent states, the observed likelihood contains factors of the form E[exp(-∫(β C_i(t)+ξ) dt)] for each susceptible i, which depend on individual exposure, not only the susceptible-average Cbar(t). Moreover, Cbar(t) is latent and its distribution changes with (β, ξ). The proof compares β Cbar(t)+ξ and β' Cbar(t)+ξ' as though the same Cbar path were available under both parameter values; under θ' the latent process has a different law. Thus equality of realized aggregate hazard paths does not imply equality of the marginal observed-data laws. The paper neither proves the aggregate-only
  2. [§6.2, Proposition 1] The assertion that I_M(θ; Δ) → 0 as Δ → 0 'under regular observation schemes' is stated without proof or a precise definition of 'regular'. This is load-bearing because Proposition 1 and the conclusion that sufficiently frequent observation recovers complete-data identifiability depend on it. A regularity condition is needed, and a proof or a counterexample should be supplied. As written, the proposition is conditional on an unexplained assumption.
  3. [§6.1, Theorem 3] Theorem 3 is definitional rather than substantive. It assumes N^{-1} I_O^{(N)}(θ) → J(θ) with J positive definite, and concludes λ_min(I_O^{(N)}) → ∞. This is exactly the definition of positive definiteness scaled by N. The paper does not provide conditions on the epidemic-network process or observation scheme that ensure the required convergence and positive definiteness. Without such conditions, the theorem does not establish asymptotic identifiability for the model introduced in Section 2. Please rework the asymptotic analysis to give verifiable sufficient conditions or replace these claims with a discussion of what would be needed.
  4. [§7, Tables 1-2] The simulation section does not report Monte Carlo standard errors for the information and estimation quantities. For example, in Table 2 the difference in bias for β between High (0.0021) and Moderate (0.0156) may be within Monte Carlo noise with R=500 replicates; the same applies to the information traces in Table 1. The claim that 'theoretical predictions closely matching finite-sample performance' is not supported without uncertainty quantification. In addition, the empirical identifiability boundaries in §7.4 are based on posterior dependence between β and ξ, but they do not operationalize or test the aggregate-only premise of Theorem 2. Please add Monte Carlo SEs or confidence intervals and include a simulation scenario that actually satisfies (or deliberately violates) the aggregate-only condition.
minor comments (4)
  1. [§2] Typographical error: 'consier' should be 'consider'.
  2. [§7, Table 2] In the κ High row, the bias is written as '-0,0003' with a comma; use a standard decimal point.
  3. [§4.2] The matrix I_O(θ) is defined as E_θ[S_O(θ) S_O(θ)^T], i.e., the expected Fisher information. This is usually called the expected (or Fisher) information, not the 'observed' information, which is typically the curvature of the log-likelihood at a realized dataset. Consider clarifying the terminology to avoid confusion.
  4. [General] The phrase 'Figures 1' should be 'Figure 1' when referring to the single figure. Also, the caption 'Recovery and coverage of information and true parameter values' is vague; please expand the caption to explain panels (a)–(f).

Circularity Check

1 steps flagged

Theorem 2's external-infection non-identifiability restates its own aggregate-only assumption rather than deriving it from the observation model.

specific steps
  1. self definitional [Section 5.4, Theorem 2 (and its proof); invoked again in Section 7.4 and Section 8]
    "If the observed-data law depends on (β, ξ) only through λagg(t) = β ¯C(t) + ξ, then β and ξ are not separately structurally identifiable unless ¯C(t) varies over time in a way that is identifiable from O. ... Since, by assumption, the observed-data law depends on (β, ξ) only through λagg(t), it follows that P^O_{β,ξ} = P^O_{β′,ξ′}."

    The theorem's antecedent says exactly that the observed-data law is a function of (β, ξ) only through λagg = β ¯C + ξ. The conclusion that pairs (β, ξ) and (β′, ξ′) giving the same λagg are observationally equivalent is simply that antecedent restated; the proof adds no model-based argument. The aggregate-only dependence is not derived from the individual-level observation process in Section 2.3 (Y_i and B_ij), nor tested in Section 7. Thus the paper's central 'external infection confounding' result is assumed, not established: it reduces by construction to the claim that the observed law depends on (β, ξ) only through λagg.

full rationale

The paper's Theorem 1 (Louis identity) is a standard external result and is proved directly; the simulation studies are self-consistent checks of the information calculations, not circular derivations. The self-citation to Asaduzzaman (2026) for the model in Section 2 is not treated as load-bearing because the model equations are stated in the paper. The serious circularity is localized to Theorem 2: its non-identifiability conclusion is the aggregate-only dependence assumption in different words, and that assumption is never shown to follow from the Section 2 model. Because this theorem underpins the paper's headline conclusion about transmission/external-infection confounding and the surveillance-design implications in Section 8, the central claim is partially circular, though the information-decomposition framework itself remains independent. Score 6 reflects this partial but load-bearing by-construction element.

Axiom & Free-Parameter Ledger

2 free parameters · 6 axioms · 0 invented entities

The framework depends on standard stochastic-process assumptions plus several ad hoc assumptions. The most important is the aggregate-exposure condition in Theorem 2, which effectively assumes away individual-level exposure information and makes the confounding result true by construction. The model itself is imported from a self-cited companion preprint, and the simulation validation relies on hand-chosen parameters.

free parameters (2)
  • Simulation baseline parameter vector theta0 = beta=0.30, xi=0.03, kappa=0.40, gamma=0.25, pE=0.35, pI=0.80, s=0.85, c=0.90; network rates eta and tau are not specifie
    Chosen by hand in Section 7.1 'to produce realistic epidemic growth and network dynamics'; the empirical validation depends on these values, and the network rates are not reported at all.
  • Identifiability threshold k = small positive threshold, not quantified
    Section 6.4 defines strong/weak/absent identifiability by comparing Phi(theta;D) to an unspecified threshold k; the phase diagrams depend on this arbitrary choice.
axioms (6)
  • domain assumption The latent epidemic-network process {Z(t)} is a continuous-time Markov process with the specified transition intensities.
    Assumed in Sections 2.1-2.2 without justification; all subsequent likelihood and information calculations rest on this.
  • domain assumption The observation process is conditionally independent of the latent process given the latent state at observation times.
    Section 2.3 defines p_theta(O|Z) as a product of per-time conditional distributions; this is a structural modeling assumption.
  • standard math Standard regularity conditions hold: the observed log-likelihood is twice differentiable, differentiation and integration interchange, and relevant expectations exist.
    Invoked in Sections 3.2 and 4.1 to define score functions and Fisher information and to prove Theorem 1.
  • ad hoc to paper For Theorem 2, the observed-data law depends on (beta, xi) only through lambda_agg(t) = beta * Cbar(t) + xi.
    This assumption is introduced in Section 5.4 and is the entire source of the non-identifiability conclusion; it is not derived from the observation model.
  • ad hoc to paper As observation interval Delta -> 0, the missing-information matrix I_M(theta; Delta) -> 0 under 'regular observation schemes.'
    Stated in Section 6.2 without proof; it is needed for Proposition 1 on local identifiability under frequent observation.
  • domain assumption In Theorem 3, N^{-1} I_O^{(N)}(theta) converges in probability to a positive-definite matrix J(theta).
    The asymptotic identifiability result is conditional on this standard but nontrivial convergence assumption.

pith-pipeline@v1.3.0-alltime-deepseek · 13736 in / 11475 out tokens · 113236 ms · 2026-08-01T03:38:29.962646+00:00 · methodology

0 comments
read the original abstract

Inference for epidemic transmission on dynamic networks is fundamentally limited by latent infection times, incomplete contact histories, imperfect observation, and external sources of infection. Although coherent likelihood formulations are available for partially observed epidemic processes, considerably less is known about the theoretical limits of statistical inference under such observation mechanisms. This paper develops a unified framework for studying identifiability and Fisher information in epidemic transmission models observed on dynamic contact networks. We establish conditions for structural and local identifiability, derive observed and complete-data information matrices, and quantify information loss arising from unobserved transmission events and missing network information through a missing-information decomposition. We further investigate how observation frequency, network coverage, and measurement accuracy influence parameter estimability and statistical efficiency, providing a principled basis for evaluating surveillance strategies. Simulation studies demonstrate that the proposed framework accurately characterises the relationship between observation design, statistical information, and parameter estimation, with theoretical predictions closely matching finite-sample performance. The proposed framework clarifies the relationship between observation design, identifiability, and inferential precision, and provides a theoretical foundation for statistical inference in partially observed epidemic transmission models.

Figures

Figures reproduced from arXiv: 2607.23079 by Md Asaduzzaman.

Figure 1
Figure 1. Figure 1: Recovery and coverage of information and true parameter values [PITH_FULL_IMAGE:figures/full_fig_p019_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 4 linked inside Pith

  1. [1]

    Abed, A., Torabi, M., and Mashreghi, Z. (2026). Spatial individual-level models for transmission dynamics of seasonal infectious diseases. Statistics in Medicine , 45(3-5):e70384

  2. [2]

    and Deardon, R

    Almutiry, W. and Deardon, R. (2021). Contact network uncertainty in individual level models of infectious disease transmission. Statistical Communications in Infectious Diseases , 13(1):20190012

  3. [3]

    Asaduzzaman, M. (2026). A complete-data likelihood for epidemic processes on partially observed dynamic networks. arXiv preprint arXiv:2607.15179

  4. [4]

    Bansal, S., Read, J., Pourbohloul, B., and Meyers, L. A. (2010). The dynamic nature of contact networks in infectious disease epidemiology. Journal of Biological Dynamics , 4(5):478--489

  5. [5]

    Becker, N. G. and Britton, T. (1999). Statistical studies of infectious disease incidence. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 61(2):287--307

  6. [6]

    Black, A. J. (2019). Importance sampling for partially observed temporal epidemic models. Statistics and Computing , 29(4):617--630

  7. [7]

    Bret \'o , C. (2018). Modeling and inference for infectious disease dynamics: a likelihood-based approach. Statistical Science: A Review Journal of the Institute of Mathematical Statistics , 33(1):57

  8. [8]

    P., Drovandi, C., Turner, I

    Browning, A. P., Drovandi, C., Turner, I. W., Jenner, A. L., and Simpson, M. J. (2022). Efficient inference and identifiability analysis for differential equation models with random parameters. PLOS Computational Biology , 18(11):e1010734

  9. [9]

    E., Volfovsky, A., and Xu, J

    Bu, F., Aiello, A. E., Volfovsky, A., and Xu, J. (2025). Stochastic em algorithm for partially observed stochastic epidemics with individual heterogeneity. Biostatistics , 26(1):kxae018

  10. [10]

    E., Xu, J., and Volfovsky, A

    Bu, F., Aiello, A. E., Xu, J., and Volfovsky, A. (2022). Likelihood-based inference for partially observed epidemics on dynamic networks. Journal of the American Statistical Association , 117(537):510--526

  11. [11]

    Daley, D. J. and Gani, J. M. (1999). Epidemic modelling: an introduction . Number 15. Cambridge University Press

  12. [12]

    P., House, T., Jewell, C

    Danon, L., Ford, A. P., House, T., Jewell, C. P., Keeling, M. J., Roberts, G. O., Ross, J. V., and Vernon, M. C. (2011). Networks and the epidemiology of infectious disease. Interdisciplinary Perspectives on Infectious Diseases , 2011(1):284909

  13. [13]

    Eames, K., Bansal, S., Frost, S., and Riley, S. (2015). Six challenges in measuring contact networks for use in modelling. Epidemics , 10:72--77

  14. [14]

    Fintzi, J., Cui, X., Wakefield, J., and Minin, V. N. (2017). Efficient data augmentation for fitting stochastic epidemic models to prevalence data. Journal of Computational and Graphical Statistics , 26(4):918--929

  15. [15]

    Groendyke, C., Welch, D., and Hunter, D. R. (2011). Bayesian inference for contact networks given epidemic data. Scandinavian Journal of Statistics , 38(3):600--616

  16. [16]

    M., Andreasen, V., Bansal, S., De Angelis, D., Dye, C., Eames, K

    Heesterbeek, H., Anderson, R. M., Andreasen, V., Bansal, S., De Angelis, D., Dye, C., Eames, K. T., Edmunds, W. J., Frost, S. D., Funk, S., et al. (2015). Modeling infectious disease dynamics in the complex landscape of global health. Science , 347(6227):aaa4339

  17. [17]

    Hethcote, H. W. (2000). The mathematics of infectious diseases. SIAM Review , 42(4):599--653

  18. [18]

    Huang, J., Morsomme, R., Dunson, D., and Xu, J. (2024). Detecting changes in the transmission rate of a stochastic epidemic model. Statistics in Medicine , 43(10):1867--1882

  19. [19]

    S., and Peter, L

    Istvan, Z., MILLER, K., JOEL, C. S., and Peter, L. (2019). Mathematics of Epidemics on Networks: From Exact to Approximate Models . Springer

  20. [20]

    O., Njiasse, I

    Kamkumo, F. O., Njiasse, I. M., and Wunderlich, R. (2025). Estimating unobservable states in stochastic epidemic models with partial information. arXiv preprint arXiv:2506.00906

  21. [21]

    N., Docherty, P

    Lam, N. N., Docherty, P. D., and Murray, R. (2022). Practical identifiability of parametrised models: A review of benefits and limitations of various approaches. Mathematics and Computers in Simulation , 199:202--216

  22. [22]

    Louis, T. A. (1982). Finding the observed information matrix when using the em algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology , 44(2):226--233

  23. [23]

    and Xu, J

    Morsomme, R. and Xu, J. (2025). Exact bayesian inference for fitting stochastic epidemic models to partially observed incidence data. The Annals of Applied Statistics , 19(3):2279--2293

  24. [24]

    O’Neill, P. D. and Roberts, G. O. (1999). Bayesian inference for partially observed stochastic epidemics. Journal of the Royal Statistical Society Series A: Statistics in Society , 162(1):121--129

  25. [25]

    Pastor-Satorras, R., Castellano, C., Van Mieghem, P., and Vespignani, A. (2015). Epidemic processes in complex networks. Reviews of Modern Physics , 87(3):925--979

  26. [26]

    Pellis, L., Ball, F., Bansal, S., Eames, K., House, T., Isham, V., and Trapman, P. (2015). Eight challenges for network epidemic models. Epidemics , 10:58--62

  27. [27]

    P., Wilkinson, R

    Preston, S. P., Wilkinson, R. D., Clayton, R. H., Chappell, M. J., and Mirams, G. R. (2025). Think before you fit: parameter identifiability, sensitivity and uncertainty in systems biology models. Current Opinion in Systems Biology , page 100563

  28. [28]

    W., Lubold, S., Chandrasekhar, A

    Reeves, S. W., Lubold, S., Chandrasekhar, A. G., and McCormick, T. H. (2024). Model-based inference and experimental design for interference using partial network data. arXiv preprint arXiv:2406.11940

  29. [29]

    Saucedo, O., Laubmeier, A., Tang, T., Levy, B., Asik, L., Pollington, T., and Prosper, O. (2024). Comparative analysis of practical identifiability methods for an seir model. arXiv preprint arXiv:2401.15076

  30. [30]

    T \"o nsing, C., Timmer, J., and Kreutz, C. (2018). Profile likelihood-based analyses of infectious disease models. Statistical Methods in Medical Research , 27(7):1979--1998

  31. [31]

    N., Meyer, A

    Vecherin, S. N., Meyer, A. C., Cummings, C. L., Trump, B. D., Ehlschlaeger, C. R., and Linkov, I. (2026). Infection risk assessment for socially structured population using stochastic microexposure model. Journal of Exposure Science & Environmental Epidemiology , 36(2):386--397

  32. [32]

    and Meyers, L

    Volz, E. and Meyers, L. A. (2007). Susceptible--infected--recovered epidemics in dynamic contact networks. Proceedings of the Royal Society B: Biological Sciences , 274(1628):2925

  33. [33]

    M., Halloran, M

    Yang, Y., Longini Jr, I. M., Halloran, M. E., and Obenchain, V. (2012). A hybrid EM and M onte C arlo EM algorithm and its application to analysis of transmission of infectious diseases. Biometrics , 68(4):1238--1249