{"id":"0d7e4277-0c98-4384-a7b9-a7befd5abc07","arxiv_id":"2507.14486","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A maximum empirical likelihood estimator for closed-population abundance with covariates missing at random is asymptotically normal, and its likelihood ratio interval has correct coverage with a lower bound at least the number of captured individuals.","lead":"This paper proposes a statistical method for estimating the size of a closed animal population when some individual covariates are missing, using empirical likelihood. The method gives confidence intervals for abundance that respect the logical lower bound of the number of captured animals and often have more accurate coverage than existing approaches.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The asymptotic theorems are internally inconsistent: the moment functions U are linearly dependent, making V44 singular and the stated chi-square degrees of freedom K+p+2 one too large.","rationale":"The reader located the weakest point in the Huggins-Alho/MAR specification. That is a legitimate robustness concern but it is a model assumption shared by the comparison methods and does not make the mathematics wrong. The linear dependence of U is more serious because it contradicts the paper's own theorem: V44 is singular, so W in Eq. (8) is not invertible and the chi-square df K+p+2 is overstated by one. This is an internal inconsistency, not a disagreement with consensus. I am not claiming the abundance estimator is inconsistent; the df used for the confidence interval (R'(ν0) vs chi-square_1) likely remains correct, and the simulations are extensive. For this reason I would not reject outright, but acceptance requires replacing Theorem 1/2 statements with a rank-reduced formulation (e.g., dropping one U component or imposing Σγ_k=1), verifying the resulting df, and supplying proofs. The reader's CONDITIONAL verdict is therefore upheld, but for a different and more concrete reason than model misspecification.","tokens_in":13697,"tokens_out":18130,"duration_ms":264203,"concrete_test":"Test 1 (analytic/numerical): for K=2, p=1 under Scenario A, compute the sample analogue V̂44 = m^{-1} Σ U(Z_i;α0,β0)U(Z_i;α0,β0)^T / π(Z_i;β0); confirm its smallest eigenvalue is 0 to numerical precision, so V44^{-1} in Eq. (8) is undefined. Test 2 (simulation): generate data under Scenario A with ν0=400, compute R(ν0,α0,β0) in 5000 replicates, and compare its empirical quantiles with chi-square_5 and chi-square_4. If the distribution matches chi-square_4, Theorem 1(b)'s degrees of freedom are wrong. Also replicate R'(ν0) versus chi-square_1; if this still calibrates, the abundance confidence interval is likely salvageable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central asymptotic claim is compromised by a linear dependence in the estimating functions. In Eq. (6), U(z;α,β) = (u0,...,uK)^T with u_k = {K choose k} g(z;β)^k(1-g(z;β))^{K-k} - γ_k. Since the first term sums to 1 over k and Σ γ_k = 1, we have 1^T U(z) = 0 for every z. Thus U has rank at most K, and the matrix V44 = -λ00^{-2} E[U U^T/π(Z)] appearing in Eq. (8) satisfies V44 1 = 0, so V44 is singular. The displayed W therefore contains V44^{-1} and cannot be nonsingular, contradicting the hypothesis of Theorem 1. Equivalently, α lies on the simplex and the parameter (ν,α,β) has effective dimension K+p+1, not K+p+2; any covariance matrix for (α̂-α0) has a zero eigenvector in the 1 direction. Theorem 1(b)'s claim R(ν0,α0,β0) -> chi-square_{K+p+2} is off by one degree of freedom. Remark 1 does not fix this: the redundancy persists even when every m_k > 0. The abundance statistic R'(ν0) -> chi-square_1 may survive because the redundant direction is profiled out, so the confidence interval in Eq. (9) is not necessarily invalidated, but the theorem as stated is not correct.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a maximum empirical likelihood estimator for the abundance of a closed population in discrete-time capture-recapture experiments when individual covariates are missing at random. The method combines a conditional capture-history likelihood under the Huggins-Alho logistic model with empirical likelihood for the covariate distribution, using moment equations implied by the model. The authors claim the resulting estimator is asymptotically normal, that the joint empirical likelihood ratio statistic converges to a chi-square distribution with K+p+2 degrees of freedom, and that the profile likelihood ratio statistic for abundance converges to chi-square with one degree of freedom. Simulations compare the estimator and a likelihood-ratio confidence interval with inverse-probability-weighting and multiple-imputation competitors, and a real-data example is provided.","tokens_in":13994,"tokens_out":4274,"duration_ms":46272,"significance":"If the asymptotic results were correct, the paper would offer a practically attractive method: the empirical likelihood ratio interval is one-step, free of variance estimation, and its lower limit is automatically at least the number of captured individuals. The simulation study is fairly extensive and consistently shows better finite-sample coverage for the proposed interval than for Wald intervals. However, the central theoretical claims are not only unproved but, as stated, internally inconsistent due to a linear dependence in the estimating functions. The significance of the contribution therefore hinges on a repair of the asymptotics and a proof; the current version cannot be accepted as a rigorous contribution.","major_comments":[{"comment":"The estimating vector U(z;α,β) is linearly dependent: since Σ_{k=0}^K C(K,k)g(z;β)^k(1-g(z;β))^{K-k}=1 and Σ_{k=0}^K γ_k=1, we have 1^T U(z;α,β)=0 for every z. Consequently V44 = -λ00^{-2} E[U(Z^*)U(Z^*)^T/π(Z^*;β0)] satisfies V44 1 = 0 and is singular. The matrix W in Eq. (8) contains V44^{-1}, so the hypothesis 'W is nonsingular' cannot be satisfied under the model. Theorem 1(a)'s covariance matrix W^{-1} is therefore not defined, and the parameter α lies on a simplex, making the effective dimension K+p+1, not K+p+2. The claimed chi-square limit for R(ν0,α0,β0) in Theorem 1(b) is off by one degree of freedom. The same linear dependence appears in the extension of Section 2.4, where the sum of U0 and all Ujk equals 1-γ0-Σ_j Σ_k γjk = 0, so Theorem 2(b)'s df of 2K+p+2 is also incorrect. Remark 1 does not fix this: the redundancy persists even when every m_k > 0. The profile statistic R'(ν0) may still have a chi-square_1 limit because the redundant direction is profiled out, but this needs to be established explicitly after removing the redundant constraint.","section":"Section 2.3, Eq. (6) and Eq. (8)"},{"comment":"Theorems 1 and 2 are the central theoretical claims of the paper, but they are stated without proof and no supplementary material is provided. The asymptotic normality of the maximum empirical likelihood estimator and the chi-square limits are used to justify the confidence interval in Eq. (9) and to interpret the simulation results, yet the reader cannot verify the regularity conditions, the derivation of the covariance matrix W, or the handling of the linearly dependent moment functions. The finite-sample simulation study in Section 3 cannot substitute for a proof. This is a load-bearing gap that must be closed before the results can be considered established.","section":"Theorems 1 and 2"}],"minor_comments":[{"comment":"The title and abstract say 'maximum likelihood abundance estimation,' but the method is maximum empirical likelihood; please adjust the terminology for accuracy.","section":"Title and Abstract"},{"comment":"The text in Section 3.2 refers to comparisons with '˜ν1 and ˜ν3' when describing the bias of the complete-case estimator; the context indicates the comparison should be with '˜ν1 and ˜ν2'.","section":"Section 3.2, Table 1 discussion"},{"comment":"The text states the Wald-type lower limits are 'no greater than 133,' but for Model 1 the IPW and MI lower limits are 80 and 95; the sentence appears to refer only to Model 2 and should be clarified.","section":"Section 4, Table 3"},{"comment":"The text reports the proposed 95% confidence interval for Model 1 as [449, 1981], while Table 3 lists [449, 1980]; please make the values consistent.","section":"Section 4"},{"comment":"There are several typographical errors, e.g., 'unstatisfactory' in Section 1 and 'covarites' in Section 2.2; a careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The linear-dependence issue is a genuine mathematical error in the stated theorems, not merely a matter of presentation. However, it is likely fixable by dropping one redundant component of U (or reparameterizing α to the simplex) and adjusting the chi-square degrees of freedom accordingly, so I do not recommend rejection if the authors can supply a corrected asymptotic theory with proofs. The absence of proofs for the main theorems is a serious concern for a statistics journal, and the revision should include a complete proof or a clearly identified supplement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The empirical likelihood approach to missing covariates in capture-recapture is a natural step forward, and the paper shows it can beat IPW and MI in simulations. The likelihood construction is clean, the simulation study is thorough, and the real-data analysis is a nice illustration. The lower bound of the proposed interval at n is a genuine practical advantage.\n\nBut there is a load-bearing problem in Theorem 1 (and Theorem 2 by inheritance). The estimating functions U have components u_k = C(K,k) g^k(1-g)^{K-k} - gamma_k, so summing over k gives 1 - sum gamma_k = 0 identically. That makes V44 singular, so the W matrix in Eq. (8) contains V44^{-1} and cannot be nonsingular. Equivalently, alpha lies on the simplex and has only K free components, not K+1. The covariance matrix of (alpha-hat - alpha0) is singular, and the correct degrees of freedom for the joint likelihood ratio statistic is K+p+1, not K+p+2. The paper states the theorems without proofs, so this issue is not caught internally.\n\nThe chi-square_1 claim for R'(nu0) may still survive because the redundant direction is profiled out, and the confidence interval in Eq. (9) is probably fine in practice. The simulations support this. But Theorem 1(a) as written is not correct, and the assumption 'if W is nonsingular' is vacuous.\n\nOther soft spots: no proofs appear in the paper or a supplement, and no code or data are provided. The MAR and logistic-model assumptions are standard, but there is no robustness check, which matters because misspecification would break the moment condition.\n\nOverall, this is a serious paper with a fixable flaw. The authors should be asked to correct the degrees of freedom, drop one redundant estimating equation or reparameterize alpha, and provide proofs or at least a detailed supplement. The empirical method itself deserves attention; the theory needs to be made right.\n\nRecommendation: send to peer review. The applied message is useful and the flaw is not in the core estimator but in the stated asymptotics for the joint statistic. A careful referee could help the authors fix this and produce a solid contribution.","headline":"The method is promising and the simulations are convincing, but the central asymptotic theorem has a real linear-dependence flaw: the U functions sum to zero, so the claimed degrees of freedom are off by one and W is singular.","tokens_in":14482,"tokens_out":5872,"would_cite":false,"duration_ms":76721,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62F12","62G05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a maximum empirical likelihood estimator for closed-population abundance when covariates are missing at random, proves it is asymptotically normal, and derives a chi-square limiting distribution for the likelihood…","keywords":["capture-recapture","abundance estimation","missing at random","empirical likelihood","Huggins-Alho logistic model","likelihood ratio confidence interval","closed population","coverage probability"],"falsifier":"One concrete check is a large-scale simulation under the paper's Scenario A or B with $\\nu_0 = 10{,}000$, comparing the empirical distribution of $R'(\\nu_0)$ against $\\chi_1^2$; the theorem predicts exact agreement in the limit, so any systematic divergence at large sample size would refute the central claim.","tokens_in":13527,"feed_emoji":"🐦","tokens_out":7625,"duration_ms":72952,"temperature":0.7,"pith_summary":"The paper addresses a standard complication in capture-recapture studies: some captured individuals have missing covariate values, such as sex or tail length, especially when they were caught only a few times. It argues that the usual inverse-probability-weighting and multiple-imputation estimators, while consistent, are not efficient and produce Wald-type confidence intervals whose coverage can be far from nominal and whose lower limit can fall below the number of individuals actually captured. The paper's proposal is a maximum empirical likelihood estimator that combines the logistic capture probability model with a nonparametric treatment of the covariate distribution, yielding a full likelihood for the abundance. The paper proves that this estimator is asymptotically normal and that the empirical likelihood ratio statistic for the abundance converges to a chi-square distribution with one degree of freedom, and it reports simulations in which the estimator has smaller mean squared error and better interval coverage than the existing methods.","feed_headline":"Full-likelihood abundance estimator beats missing-data rivals","feed_subtitle":"Empirical likelihood ratio intervals respect the captured-count floor and improve coverage.","key_machinery":"The engine is the empirical likelihood constraint $\\int U(z;\\alpha,\\beta)\\,dF_{Z^*}(z)=0$, where $U_k(z;\\alpha,\\beta) = \\binom{K}{k} g(z;\\beta)^k (1-g(z;\\beta))^{K-k} - \\gamma_k$ and $g$ is the logistic capture probability. Profiling the empirical weights $p_i$ through a Lagrange multiplier $\\lambda$ gives the profile empirical log-likelihood $\\ell(\\nu,\\alpha,\\beta)$; the likelihood ratio $R'(\\nu) = 2\\{\\max_{\\nu,\\alpha,\\beta}\\ell - \\max_{\\alpha,\\beta}\\ell(\\nu,\\alpha,\\beta)\\}$ is the object whose $\\chi_1^2$ limit under the null carries the confidence interval. The asymptotic covariance $W$ in Equation (8) is assembled from the missingness probabilities $h_k$, the capture-count probabilities $\\gamma_k$, and the expectation of the estimating functions.","core_discovery":"Under the Huggins-Alho logistic model for capture probability (Equation 1) and the missing-at-random assumption (Equation 2), the paper constructs a partial likelihood in which the marginal distribution of the covariate $Z^*$ is profiled out by empirical likelihood. The maximum empirical likelihood estimator $(\\hat{\\nu}, \\hat{\\alpha}, \\hat{\\beta})$ is asymptotically normal with covariance matrix $W^{-1}$, and the profile likelihood ratio statistic for abundance $R'(\\nu_0)$ converges in distribution to $\\chi_1^2$ (Theorem 1; Theorem 2 covers the extension with an always-observed binary covariate). This yields the confidence interval $I = \\{\\nu : R'(\\nu) \\le \\chi_{1,1-a}^2\\}$, which has asymptotically correct coverage and, because $R'(\\nu)$ is defined only on $[n,\\infty)$, whose lower limit is never below the captured count $n$. In simulations the proposed estimator has uniformly smaller bias and relative mean squared error than the inverse-probability-weighting and multiple-imputation estimators, and in the yellow-bellied prinia example the proposed 95% interval is $[449, 1980]$ with point estimate 770, whereas the Wald intervals have lower limits below 163.","pith_inferences":["An immediate testable extension is to check the Huggins-Alho model and the MAR condition before trusting the interval; a misspecified capture probability would break the moment constraint and the chi-square calibration, so a diagnostic based on the estimating function could be developed.","The same profile-likelihood logic could be transferred to continuous-time capture-recapture designs, where Wald lower limits below the observed count are also a known nuisance.","Because the covariate distribution is treated nonparametrically, the estimator should be insensitive to the shape of the covariate distribution; a targeted simulation with skewed or multimodal covariates would confirm whether that robustness actually holds at finite sample sizes.","The paper's finite-sample coverage results suggest the chi-square approximation is reliable at $\\nu_0$ around 200; whether this extends to much smaller populations is an open question."],"forward_implications":["The proposed confidence interval requires no variance estimation and automatically respects the logical constraint that abundance cannot be less than the number of captured individuals.","When one covariate is fully observed and another is missing, the same construction works and retains the $\\chi_1^2$ limit for the abundance statistic (Theorem 2).","The simulations imply that discarding incomplete capture records is not a benign fix under missing at random: complete-case estimates can be severely biased, so a full-likelihood treatment is needed.","The method can in principle be extended to capture models with time effects, behavioral effects, or both, as the paper notes in its discussion.","The likelihood ratio test for a covariate coefficient, used on the prinia data, provides a model-selection tool within the same framework."],"supporting_citations":[{"why":"Supplies the logistic capture probability model (Equation 1) that links capture counts to covariates and underlies the full likelihood.","marker":"Huggins (1989)"},{"why":"Provides the logistic regression formulation for capture probabilities, the other half of the Huggins-Alho model.","marker":"Alho (1990)"},{"why":"Provides the empirical likelihood method used to profile the covariate distribution under the moment constraints.","marker":"Owen (1990)"},{"why":"Defines the inverse probability weighting and multiple imputation estimators that serve as the baselines for comparison in simulations and real data.","marker":"Lee et al. (2016)"},{"why":"Establishes the maximum empirical likelihood estimator for fully observed covariates that this paper extends to the missing-covariate setting.","marker":"Liu et al. (2017)"},{"why":"Formalizes the missing-at-random assumption (Equation 2) on which the likelihood factorization and the moment condition depend.","marker":"Rubin (1976)"}],"fun_headline_variants":["Empirical likelihood trims abundance errors","Abundance intervals never fall below captured count","Capture-recapture: empirical likelihood beats rivals","Maximum empirical likelihood boosts abundance accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that capture probabilities follow the specified logistic model and that missingness of a covariate depends only on the number of captures, not on the value of the missing covariate itself.","fun_headline_variants_meta":{"raw":{"variants":["Empirical likelihood trims abundance errors","Abundance intervals never fall below captured count","Capture-recapture: empirical likelihood beats rivals","Maximum empirical likelihood boosts abundance accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1390,"prompt_tokens":992,"completion_tokens":398,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":345}},"tokens_in":608,"tokens_out":398,"duration_ms":5345,"temperature":1.0,"reasoning_tokens":345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:53:56.588745+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete check is a large-scale simulation under the paper's Scenario A or B with $\\nu_0 = 10{,}000$, comparing the empirical distribution of $R'(\\nu_0)$ against $\\chi_1^2$; the theorem predicts exact agreement in the limit, so any systematic divergence at large sample size would refute the central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the logistic capture probability model (Equation 1) that links capture counts to covariates and underlies the full likelihood."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the logistic regression formulation for capture probabilities, the other half of the Huggins-Alho model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the empirical likelihood method used to profile the covariate distribution under the moment constraints."},{"cited_title":"Hwang, and J","cited_arxiv_id":null,"evidence_quote":"Defines the inverse probability weighting and multiple imputation estimators that serve as the baselines for comparison in simulations and real data."},{"cited_title":"Li, and J","cited_arxiv_id":null,"evidence_quote":"Establishes the maximum empirical likelihood estimator for fully observed covariates that this paper extends to the missing-covariate setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Formalizes the missing-at-random assumption (Equation 2) on which the likelihood factorization and the moment condition depend."}],"review_version":1}