Pith. sign in

REVIEW 2 major objections 5 minor 22 references

Maximum likelihood abundance estimation from capture-recapture data when covariates are missing at random

T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper introduces a maximum empirical likelihood estimator for closed-population abundance when covariates are missing at random, proves it is asymptotically normal, and derives a chi-square limiting distribution for the likelihood…

desk verdict The method is promising and the simulations are convincing, but the central asymptotic theorem has a real linear-dependence flaw: the U functions sum to zero, so the claimed degrees of freedom are off by one and W is singular. read the letter →

arxiv 2507.14486 v1 pith:V4W4T3ZY submitted 2025-07-19 stat.ME

classification stat.ME MSC 62D0562F1262G05
keywords capture-recaptureabundanceestimationmissingatrandomempiricallikelihoodHuggins-Alhologisticmodelratioconfidenceintervalclosedpopulationcoverageprobability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper addresses a standard complication in capture-recapture studies: some captured individuals have missing covariate values, such as sex or tail length, especially when they were caught only a few times. It argues that the usual inverse-probability-weighting and multiple-imputation estimators, while consistent, are not efficient and produce Wald-type confidence intervals whose coverage can be far from nominal and whose lower limit can fall below the number of individuals actually captured. The paper's proposal is a maximum empirical likelihood estimator that combines the logistic capture probability model with a nonparametric treatment of the covariate distribution, yielding a full likelihood for the abundance. The paper proves that this estimator is asymptotically normal and that the empirical likelihood ratio statistic for the abundance converges to a chi-square distribution with one degree of freedom, and it reports simulations in which the estimator has smaller mean squared error and better interval coverage than the existing methods.

What carries the argument

The engine is the empirical likelihood constraint $\int U(z;\alpha,\beta)\,dF_{Z^*}(z)=0$, where $U_k(z;\alpha,\beta) = \binom{K}{k} g(z;\beta)^k (1-g(z;\beta))^{K-k} - \gamma_k$ and $g$ is the logistic capture probability. Profiling the empirical weights $p_i$ through a Lagrange multiplier $\lambda$ gives the profile empirical log-likelihood $\ell(\nu,\alpha,\beta)$; the likelihood ratio $R'(\nu) = 2\{\max_{\nu,\alpha,\beta}\ell - \max_{\alpha,\beta}\ell(\nu,\alpha,\beta)\}$ is the object whose $\chi_1^2$ limit under the null carries the confidence interval. The asymptotic covariance $W$ in Equation (8) is assembled from the missingness probabilities $h_k$, the capture-count probabilities $\gamma_k$, and the expectation of the estimating functions.

What would settle it

One concrete check is a large-scale simulation under the paper's Scenario A or B with $\nu_0 = 10{,}000$, comparing the empirical distribution of $R'(\nu_0)$ against $\chi_1^2$; the theorem predicts exact agreement in the limit, so any systematic divergence at large sample size would refute the central claim.

Watch

Extended reading notes

Core claim

Under the Huggins-Alho logistic model for capture probability (Equation 1) and the missing-at-random assumption (Equation 2), the paper constructs a partial likelihood in which the marginal distribution of the covariate $Z^*$ is profiled out by empirical likelihood. The maximum empirical likelihood estimator $(\hat{\nu}, \hat{\alpha}, \hat{\beta})$ is asymptotically normal with covariance matrix $W^{-1}$, and the profile likelihood ratio statistic for abundance $R'(\nu_0)$ converges in distribution to $\chi_1^2$ (Theorem 1; Theorem 2 covers the extension with an always-observed binary covariate). This yields the confidence interval $I = \{\nu : R'(\nu) \le \chi_{1,1-a}^2\}$, which has asymptotically correct coverage and, because $R'(\nu)$ is defined only on $[n,\infty)$, whose lower limit is never below the captured count $n$. In simulations the proposed estimator has uniformly smaller bias and relative mean squared error than the inverse-probability-weighting and multiple-imputation estimators, and in the yellow-bellied prinia example the proposed 95% interval is $[449, 1980]$ with point estimate 770, whereas the Wald intervals have lower limits below 163.

Load-bearing premise

The load-bearing premise is that capture probabilities follow the specified logistic model and that missingness of a covariate depends only on the number of captures, not on the value of the missing covariate itself.

Editorial extensions

If this is right

  • The proposed confidence interval requires no variance estimation and automatically respects the logical constraint that abundance cannot be less than the number of captured individuals.
  • When one covariate is fully observed and another is missing, the same construction works and retains the $\chi_1^2$ limit for the abundance statistic (Theorem 2).
  • The simulations imply that discarding incomplete capture records is not a benign fix under missing at random: complete-case estimates can be severely biased, so a full-likelihood treatment is needed.
  • The method can in principle be extended to capture models with time effects, behavioral effects, or both, as the paper notes in its discussion.
  • The likelihood ratio test for a covariate coefficient, used on the prinia data, provides a model-selection tool within the same framework.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An immediate testable extension is to check the Huggins-Alho model and the MAR condition before trusting the interval; a misspecified capture probability would break the moment constraint and the chi-square calibration, so a diagnostic based on the estimating function could be developed.
  • The same profile-likelihood logic could be transferred to continuous-time capture-recapture designs, where Wald lower limits below the observed count are also a known nuisance.
  • Because the covariate distribution is treated nonparametrically, the estimator should be insensitive to the shape of the covariate distribution; a targeted simulation with skewed or multimodal covariates would confirm whether that robustness actually holds at finite sample sizes.
  • The paper's finite-sample coverage results suggest the chi-square approximation is reliable at $\nu_0$ around 200; whether this extends to much smaller populations is an open question.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes a maximum empirical likelihood estimator for the abundance of a closed population in discrete-time capture-recapture experiments when individual covariates are missing at random. The method combines a conditional capture-history likelihood under the Huggins-Alho logistic model with empirical likelihood for the covariate distribution, using moment equations implied by the model. The authors claim the resulting estimator is asymptotically normal, that the joint empirical likelihood ratio statistic converges to a chi-square distribution with K+p+2 degrees of freedom, and that the profile likelihood ratio statistic for abundance converges to chi-square with one degree of freedom. Simulations compare the estimator and a likelihood-ratio confidence interval with inverse-probability-weighting and multiple-imputation competitors, and a real-data example is provided.

Significance. If the asymptotic results were correct, the paper would offer a practically attractive method: the empirical likelihood ratio interval is one-step, free of variance estimation, and its lower limit is automatically at least the number of captured individuals. The simulation study is fairly extensive and consistently shows better finite-sample coverage for the proposed interval than for Wald intervals. However, the central theoretical claims are not only unproved but, as stated, internally inconsistent due to a linear dependence in the estimating functions. The significance of the contribution therefore hinges on a repair of the asymptotics and a proof; the current version cannot be accepted as a rigorous contribution.

major comments (2)
  1. [Section 2.3, Eq. (6) and Eq. (8)] The estimating vector U(z;α,β) is linearly dependent: since Σ_{k=0}^K C(K,k)g(z;β)^k(1-g(z;β))^{K-k}=1 and Σ_{k=0}^K γ_k=1, we have 1^T U(z;α,β)=0 for every z. Consequently V44 = -λ00^{-2} E[U(Z^*)U(Z^*)^T/π(Z^*;β0)] satisfies V44 1 = 0 and is singular. The matrix W in Eq. (8) contains V44^{-1}, so the hypothesis 'W is nonsingular' cannot be satisfied under the model. Theorem 1(a)'s covariance matrix W^{-1} is therefore not defined, and the parameter α lies on a simplex, making the effective dimension K+p+1, not K+p+2. The claimed chi-square limit for R(ν0,α0,β0) in Theorem 1(b) is off by one degree of freedom. The same linear dependence appears in the extension of Section 2.4, where the sum of U0 and all Ujk equals 1-γ0-Σ_j Σ_k γjk = 0, so Theorem 2(b)'s df of 2K+p+2 is also incorrect. Remark 1 does not fix this: the redundancy persists even when every m_k > 0. The profile statistic R'(ν0) may still have a chi-square_1 limit because the redundant direction is profiled out, but this needs to be established explicitly after removing the redundant constraint.
  2. [Theorems 1 and 2] Theorems 1 and 2 are the central theoretical claims of the paper, but they are stated without proof and no supplementary material is provided. The asymptotic normality of the maximum empirical likelihood estimator and the chi-square limits are used to justify the confidence interval in Eq. (9) and to interpret the simulation results, yet the reader cannot verify the regularity conditions, the derivation of the covariance matrix W, or the handling of the linearly dependent moment functions. The finite-sample simulation study in Section 3 cannot substitute for a proof. This is a load-bearing gap that must be closed before the results can be considered established.
minor comments (5)
  1. [Title and Abstract] The title and abstract say 'maximum likelihood abundance estimation,' but the method is maximum empirical likelihood; please adjust the terminology for accuracy.
  2. [Section 3.2, Table 1 discussion] The text in Section 3.2 refers to comparisons with '˜ν1 and ˜ν3' when describing the bias of the complete-case estimator; the context indicates the comparison should be with '˜ν1 and ˜ν2'.
  3. [Section 4, Table 3] The text states the Wald-type lower limits are 'no greater than 133,' but for Model 1 the IPW and MI lower limits are 80 and 95; the sentence appears to refer only to Model 2 and should be clarified.
  4. [Section 4] The text reports the proposed 95% confidence interval for Model 1 as [449, 1981], while Table 3 lists [449, 1980]; please make the values consistent.
  5. [Throughout] There are several typographical errors, e.g., 'unstatisfactory' in Section 1 and 'covarites' in Section 2.2; a careful proofreading pass is needed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the empirical likelihood estimator and likelihood-ratio confidence interval are obtained from the stated model, not from fitted inputs or self-citation chains.

full rationale

The proposed construction is self-contained: the profile empirical log-likelihood ℓ(ν,α,β) is obtained by profiling the empirical distribution of complete-case covariates subject to the moment constraints (6), and the abundance ν enters only through the binomial count term log(ν/n)+(ν−n)log γ0 plus the relation γ0=pr(D*=0). Equation (6) is not a fitted input; it restates the definition γ_k=pr(D*=k) as E[Bin(k;K,g(Z))−γ_k]=0, and no component is solved for ν or for the likelihood-ratio cutoff. The comparators (Lee et al. 2016; Liu et al. 2017) are external benchmarks used in simulations and real data, not load-bearing inputs to the theorems; the few self-citations (Liu et al. 2017, 2018) supply background and comparator methods rather than an unverified uniqueness premise. The χ²₁ limit for R′(ν0) is an assertion about the profile likelihood under the stated nonsingularity condition; whether that condition can hold is a mathematical correctness question (indeed, the estimating functions U are linearly dependent since 1ᵀU=0, so V44 is singular and the displayed W may not be nonsingular), but that is an internal technical flaw, not a definitional or self-citation circularity. No prediction is equivalent to its input by construction, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical or conceptual entities. The model parameters β and α are standard estimable parameters, not ad hoc fitted constants, so the free parameter list is empty. The main burden is the assumed capture and missingness models and the unproved regularity conditions.

assumptions (4)
  • domain assumption Capture probability follows the Huggins-Alho logistic model: pr(D*_k=1|Z*=z) = exp(β'z)/(1+exp(β'z)).
    Statement in Eq (1); the estimating equations in Eq (6) and the entire likelihood are built on this model.
  • domain assumption Missingness is at random: pr(R*=1|Z*=z, D*=k) = pr(R*=1|D*=k).
    Statement in Eq (2); it justifies discarding the missingness model L0 and factoring the likelihood.
  • domain assumption The population is closed with no birth, death, or migration, so abundance ν is constant.
    Stated in Section 1; without it, abundance is not a well-defined parameter.
  • standard math Regularity conditions: ∫ π(z;β)^{-1} dF(z) < ∞ and the matrix W is nonsingular.
    Conditions for Theorem 1; they are assumed but not verified in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Maximum likelihood abundance estimation from capture-recapture data when covariates are missing at random." pith.science (2026). https://pith.science/paper/V4W4T3ZY

@misc{pith2026250714486,
  author       = {Pith},
  title        = {Pith review of: Maximum likelihood abundance estimation from capture-recapture data when covariates are missing at random},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V4W4T3ZY}},
  note         = {Machine review of arXiv:2507.14486}
}
read the original abstract

In capture-recapture experiments, individual covariates may be subject to missing, especially when the number of times of being captured is small. When the covariate information is missing at random, the inverse probability weighting method and multiple imputation method are widely used to obtain the point estimators of the abundance. These point estimators are then used to construct the Wald-type confidence intervals for the abundance. However, such intervals may have severely inaccurate coverage probabilities and their lower limits can be even less than the number of individuals ever captured. In this paper, we proposed a maximum empirical likelihood estimation approach for the abundance in presence of missing covariates. We show that the maximum empirical likelihood estimator is asymptotically normal, and that the empirical likelihood ratio statistic for abundance has a chisquare limiting distribution with one degree of freedom. Simulations indicate that the proposed estimator has smaller mean square error than the existing estimators, and the proposed empirical likelihood ratio confidence interval usually has more accurate coverage probabilities than the existing Wald-type confidence intervals. We illustrate the proposed method by analyzing the bird species yellow-bellied prinia collected in Hong Kong.

Figures

Figures reproduced from arXiv: 2507.14486 by the authors.

Figure 1
Figure 1. QQ-plots of R′ (ν0) and R′ e (ν0) (row 1), and two pivotal statistics (row 2) in Scenarios A-D with ν0 = 200 (columns 1-4). RMSE than Lee et al. (2016)’s two estimators ˜ν1 and ˜ν2. The proposed interval estimator always exhibits better performance than the two Wald-type interval estimators I1 and I2 in terms of one-sided and two-sided coverage accuracies. 4 Real data We illustrate the proposed estimation method by … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 22 canonical work pages

  1. [1]

    Alho, J. M. (1990). Logistic regression in capture-recapture models. Biometrics\/ 43\/ (3), 623--635

  2. [2]

    El Emam, and D

    Barnard, J., K. El Emam, and D. Zubrow (2003). Using capture-recapture models for the reinspection decision. Software Quality Professional\/ 5\/ (2), 11--20

  3. [3]

    Boden, L. I. and A. Ozonoff (2008). Capture--recapture estimates of nonfatal workplace injuries and illnesses. Annals of epidemiology\/ 18\/ (6), 500--506

  4. [4]

    Borchers, D. L., S. T. Buckland, and W. Zucchini (2002). Estimating animal abundance: closed populations . Springer Science & Business Media

  5. [5]

    Chao, A. (2001). An overview of closed capture-recapture models. Journal of Agricultural, Biological, and Environmental Statistics\/ 6\/ (2), 158--175

  6. [6]

    Huggins, R. (1989). On the statistical analysis of capture experiments. Biometrika\/ 76\/ (1), 133--140

  7. [7]

    and W.-H

    Huggins, R. and W.-H. Hwang (2011). A review of the use of conditional likelihood in capture-recapture experiments. International Statistical Review\/ 79\/ (3), 385--400

  8. [8]

    Hwang, and J

    Lee, S.-M., W.-H. Hwang, and J. de Dieu Tapsoba (2016). Estimation in closed capture--recapture models when covariates are missing at random. Biometrics\/ 72\/ (4), 1294--1304

Show all 22 references
  1. [9]

    Lin, D. and P. S. Yip (1999). Parametric regression models for continuous time removal and recapture studies. Journal of the Royal Statistical Society: Series B (Statistical Methodology)\/ 61\/ (2), 401--411

  2. [10]

    Little, R. J. and D. B. Rubin (2014). Statistical analysis with missing data , Volume 333. John Wiley &amp; Sons

  3. [11]

    Li, and J

    Liu, Y., P. Li, and J. Qin (2017). Maximum empirical likelihood estimation for abundance in a closed population from capture-recapture data. Biometrika\/ 104\/ (3), 527--543

  4. [12]

    Liu, Y., Y. Liu, P. Li, and J. Qin (2018). Full likelihood inference for abundance from continuous time capture--recapture data. Journal of the Royal Statistical Society: Series B (Statistical Methodology)\/ 80\/ (5), 995--1014

  5. [13]

    Otis, D. L., K. P. Burnham, G. C. White, and D. R. Anderson (1978). Statistical inference from capture data on closed animal populations. Wildlife monographs\/ (62), 3--135

  6. [14]

    Owen, A. (1990). Empirical likelihood ratio confidence regions. The Annals of Statistics\/ 18\/ (1), 90--120

  7. [15]

    Owen, A. B. (1988). Empirical likelihood ratio confidence intervals for a single functional. Biometrika\/ 75\/ (2), 237--249

  8. [16]

    Rubin, D. B. (1976). Inference and missing data. Biometrika\/ 63\/ (3), 581--592

  9. [17]

    Hwang, S.-H

    Stoklosa, J., W.-H. Hwang, S.-H. Wu, and R. Huggins (2011). Heterogeneous capture--recapture models with covariates: A partial likelihood approach for closed populations. Biometrics\/ 67\/ (4), 1659--1665

  10. [18]

    Tilling, K., J. A. Sterne, and C. D. Wolfe (2001). Multilevel growth curve models with covariate effects: application to recovery after stroke. Statistics in medicine\/ 20\/ (5), 685--704

  11. [19]

    Wang, Y. (2005). A semiparametric regression model with missing covariates in continuous-time capture-recapture studies. Australian &amp; New Zealand Journal of Statistics\/ 47\/ (3), 287--297

  12. [20]

    Wang, Y. and P. S. Yip (2003). A semiparametric model for recapture experiments. Scandinavian journal of statistics\/ 30\/ (4), 667--676

  13. [21]

    Watson, J.-P

    Xi, L., R. Watson, J.-P. Wang, and P. S. Yip (2009). Estimation in capture--recapture models when covariates are subject to measurement errors and missing data. Canadian Journal of Statistics\/ 37\/ (4), 645--658

  14. [22]

    Yip, P. S. and Y. Wang (2002). A unified parametric regression model for recapture studies with random removals in continuous time. Biometrics\/ 58\/ (1), 192--199

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.