Pith. sign in

REVIEW 3 major objections 4 minor 14 references

Moment convergence of the generalized maximum composite likelihood estimators for determinantal point processes

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read For stationary DPPs, the two-step composite likelihood estimator converges in every polynomial moment to a normal limit.

desk verdict A genuinely new idea for DPP composite likelihood estimation, but the main theorem's proof has an invalid moment calculation and the key assumption fails for the paper's own simulation kernels. read the letter →

arxiv 1909.01211 v1 pith:2NOHM33J submitted 2019-09-03 math.ST math.PRstat.TH

classification math.STmath.PRstat.TH MSC 62M8660G5562F12
keywords determinantalpointprocessescompositelikelihoodtwo-stepestimationmomentconvergenceinformationcriteriaasymptoticnormalitystationarylargedeviationinequality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Determinantal point processes (DPPs) are spatial point-process models in which points repel, and their joint intensities are determinants of a positive definite kernel, so the intensity of every order is available in closed form. This paper exploits that fact to define a two-step generalized maximum composite likelihood estimator: first estimate the intensity parameter with a quasi-likelihood, then estimate the interaction parameter by maximizing a composite likelihood built from p-th order joint intensities, for any integer $p\ge2$. The central claim is that, for stationary DPPs satisfying the paper's regularity assumptions, the scaled estimation error $\sqrt{|D_n|}(\hat{\theta}_n-\theta_0)$ converges in distribution to a centered normal and, more strongly, every polynomial moment of that error converges to the corresponding normal moment. That moment convergence carries the practical payoff: it justifies a bias-corrected information criterion for selecting among competing DPP models. A sympathetic reader should care because full likelihood inference for DPPs is difficult, and this provides a theoretically grounded composite-likelihood route with model selection.

What carries the argument

The machinery is the polynomial-type large deviation inequality, applied through four verifiable conditions (M1)-(M4) on the composite likelihood score, its derivatives, and the centered contrast. The central object that makes the whole construction possible is the p-th order joint intensity $\rho_\theta^{(p)}(x_1,\ldots,x_p)=\det[K_\theta](x_i-x_j)_{i,j\le p}$, which is explicit for every p because the process is determinantal; this lets the second step be run at any order $p\ge2$ rather than only $p=2$. A uniform lower bound on these determinants keeps the score integrands bounded, and the embedding inequality assumed on the compact parameter space converts pointwise moment bounds into uniform sup-norm bounds, which is what the large-deviation argument needs.

What would settle it

Evaluate the two-point determinant for the squared-exponential kernel $K_\theta(x,y)=\lambda\exp(-|x-y|^2/\alpha^2)$: $\det[K](x,y)=\lambda^2(1-\exp(-2|x-y|^2/\alpha^2))$, whose infimum over $x\neq y$ is 0, so Assumption 2.2(ii) cannot hold for that model. Where the assumptions do hold, simulate a stationary DPP satisfying Assumption 2.2(ii) and check whether $E[(\sqrt{|D_n|}(\hat{\alpha}_n-\alpha_0))^4]$ converges to the fourth moment of its limiting normal distribution; failure of that convergence would falsify Theorem 4.1.

Watch

Extended reading notes

Core claim

The core discovery is Theorem 4.1: under Assumptions 2.1 and 2.2, for every polynomial-growth function $f$, $\lim_{n\to\infty} E[f(\sqrt{|D_n|}(\hat{\theta}_n-\theta_0))] = E[f(u)]$, where $u \sim N(0, I^{(p)}(\theta_0)^{-1}\Sigma^{(p)}(\theta_0) I^{(p)}(\theta_0)^{-1})$. Here $\hat{\theta}_n=(\hat{\lambda}_n,\hat{\alpha}_n^{(p)})$ is the two-step generalized maximum composite likelihood estimator: $\hat{\lambda}_n$ is the quasi-likelihood intensity estimator and $\hat{\alpha}_n^{(p)}$ maximizes the p-th order composite likelihood built from the p-th order joint intensity $\det[K_\theta]$. The result packages consistency, asymptotic normality, uniform boundedness of every moment of the scaled error, and convergence of every polynomial moment into a single statement. A corollary is the information criterion $IC^{(2)}=-2CL_n^{(2)}(\hat{\alpha})+2\operatorname{tr}(\Sigma_{22}^{(2)} I_{22}^{(2)-1})$, whose bias correction is shown to make the estimated composite likelihood asymptotically unbiased.

Load-bearing premise

The load-bearing premise is that the joint intensity $\det[K_\theta](x_1,\ldots,x_p)$ stays uniformly bounded away from zero; for the three stationary kernels simulated in the paper, the two-point determinant $\lambda^2(1-C_\alpha(x-y)^2)$ tends to zero as two points coalesce, so the premise fails exactly where the numerics are run.

Editorial extensions

If this is right

  • Under the theorem's assumptions, the two-step generalized maximum composite likelihood estimator at any order $p\ge2$ is consistent and asymptotically normal, with asymptotic covariance $I^{(p)}(\theta_0)^{-1}\Sigma^{(p)}(\theta_0)I^{(p)}(\theta_0)^{-1}$.
  • The moment convergence implies the scaled estimation error has uniformly bounded moments of every order, so higher-order bias and risk expansions can be carried out without adding moment assumptions.
  • The information criterion $IC^{(2)}$ gives a concrete model-selection rule: choose the DPP model with the smaller value, and the criterion is an asymptotically unbiased estimator of the composite likelihood at the true parameter.
  • Because the matrices $\Sigma^{(p)}$ and $I^{(p)}$ are explicit for every p, analogous information criteria $IC^{(p)}$ are available for higher-order composite likelihoods as well.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's simulations use squared-exponential, exponential, and heavy-tailed kernels for which Assumption 2.2(ii) fails as points coalesce, yet the estimates remain well behaved; this suggests the moment convergence may survive under a weaker, local version of the uniform lower bound, and checking that extension is a natural next step.
  • The moment convergence for all polynomial f opens the door to higher-order bias corrections for composite-likelihood information criteria, refining the trace term by expanding the expected composite likelihood to the next order.
  • For parameters such as the shape index of the heavy-tailed kernel that are poorly identified by second-order composite likelihood, higher-order composite likelihoods are identifiable in principle; the obstacle is numerical, so approximating the multiple integrals inside the K-function is a concrete algorithmic target.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper studies two-step generalized maximum composite likelihood estimation for stationary determinantal point processes (DPPs). The first step estimates the intensity parameter by a Poisson-type quasi-likelihood, and the second step uses a p-th order composite likelihood built from the determinant form of DPP joint intensities. The main result, Theorem 4.1, asserts moment convergence of the scaled estimation error to a Gaussian limit for every polynomial growth function, under Assumptions 2.1 and 2.2. The paper also derives an AIC-type information criterion based on the second-order composite likelihood and reports a small simulation study using Gaussian, Laplace, and Cauchy kernels.

Significance. The idea of exploiting the closed-form higher-order intensities of DPPs to construct p-th order composite likelihood estimators is attractive, and a genuine moment convergence result would be more informative than asymptotic normality alone, especially for deriving information criteria. The bias-corrected composite likelihood criterion in Section 4.2 is a useful byproduct. However, the main theorem is not supported as stated: Assumption 2.2(ii) excludes the very models used in the numerical section, and the proof of Theorem 4.1 contains an invalid moment computation. The paper's central claim therefore does not currently hold for its demonstrated setting.

major comments (3)
  1. [Section 2, Assumption 2.2(ii)] Assumption 2.2(ii) is violated by every stationary kernel of the form K_theta(x,y) = lambda C_alpha(x-y) with C_alpha(0)=1 and continuous C_alpha, including the Gaussian, Laplace, and Cauchy kernels used in Section 5. For p=2, det[K_theta](x,y) = lambda^2(1 - C_alpha(x-y)^2), which tends to 0 as |x-y| tends to 0, so the required infimum over all configurations is 0. This uniform lower bound is used repeatedly in the proofs of Theorem 3.2, Lemma 3.4, and the verification of conditions (M1) and (M2) to bound the composite likelihood score integrands. Consequently, Theorem 4.1 does not apply to the models in Section 5, and the paper's headline claim is not supported for the demonstrated setting.
  2. [Section 4.1, proof of Theorem 4.1] The identity used to verify condition (M1), namely E[|U_{n2}^{(p)}(alpha_0)|^L] = |D_n|^{-L/2} times the single integral over D_n^p of |U_2^{(p)}(x_1,...,x_p)|^L lambda_0^p rho_tilde_{alpha_0}(x_1,...,x_p) dx_1...dx_p, is false for L>1. The L-th power of the sum over all p-tuples of the point process must be expanded using factorial moment measures of orders p through Lp; a single p-th order intensity integral is only valid for L=1. Since this step is the entire verification of the first inequality in (M1), the required moment bound is not established.
  3. [Section 4.1, verification of condition (M4)] The Taylor expansion in the proof of (M4) gives a pointwise relation involving I_22^{(p)}(tilde alpha), but condition (M4) requires a uniform quadratic inequality Y_2(alpha, alpha_0) lesssim -|alpha-alpha_0|^2 for all alpha in the compact parameter space. The proof does not justify the required uniformity, for instance by a positive lower bound on I_22^{(p)}(alpha) uniformly in alpha, so the verification of (M4) is incomplete.
minor comments (4)
  1. [Section 4.1, proof of Theorem 4.1] The phrase 'by the Brillinger mixing condition for X' is used without explicitly stating or verifying that condition; since it is essential for the moment bound in (M1), the condition should be stated or a precise citation given.
  2. [Section 3.1, Eq. (2)] The counting measure N^{(p)} and the integrals over D_n^p are used informally; a more precise definition of the product measure and the domain of integration would improve readability.
  3. [Section 3.2, Eq. (6)] The derivative expression involving rho_tilde_alpha^{(p)} and its logarithm should be stated with the explicit positivity condition on the determinant, since the logarithm and the inverse matrix are used.
  4. [Throughout] There are several typographical errors, including 'matirx' in Proposition 3.3, 'competes the proof' in Theorem 4.5, and the garbled '~CL^{(p)}~CL^{(p)}' passage at the beginning of Section 4.1.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the main theorem is verified against Yoshida's external polynomial large-deviation conditions; self-citations are not load-bearing.

full rationale

The derivation chain is self-contained: Theorem 4.1 is obtained by verifying Yoshida's conditions (M1)-(M4), and each verification uses the paper's regularity Assumptions 2.1-2.2 rather than the conclusion. The sandwich covariance and the moment convergence are not fitted from data or defined in terms of the target result. The only self-citation is to Shimizu and Zhang (2017) as a place where Yoshida's conditions are described; the actual theorem invoked is Yoshida (2011), an external source, so this citation is not load-bearing. The simulation section estimates model parameters and reports finite-sample bias, not predictions derived from fitted values, so no fitted input is relabeled as a prediction. A separate concern, outside circularity, is that Assumption 2.2(ii) appears incompatible with the Gaussian/Laplace/Cauchy kernels used in Section 5, since their determinants vanish as points coalesce; this is a correctness/assumption gap, not a circular reduction.

Assumptions & free parameters 3 free parameters · 8 assumptions · 0 invented entities

The central claim rests on a list of model assumptions, one of which (Assumption 2.2(ii)) is violated by the paper's own numerical examples, and one unstated mixing assumption that enters the main proof. These are load-bearing premises rather than mere technical conveniences.

free parameters (3)
  • r = r = n/8 in simulations
    Tuning parameter in the weight function w_r; chosen by hand, affects the composite likelihood and the estimator, with no data-driven selection rule.
  • nu (Cauchy shape parameter) = nu = 1 in text, nu = 0.5 in Table 3 caption (inconsistent)
    Fixed by hand before estimation; the manuscript is internally inconsistent about its value.
  • p (order of composite likelihood) = p = 2 in simulations
    The theory allows any p >= 2, but simulations only use p = 2; the choice is not data-driven.
assumptions (8)
  • domain assumption Assumption 2.1: existence of a stationary DPP with kernel K_theta, with spectral density bounded by 1/||rho||_infinity.
    This is stated as the model condition ensuring the DPP exists and is well-defined.
  • domain assumption Assumption 2.2(i): local identifiability of the kernel K_theta on a neighborhood of radius r.
    Used to ensure the true parameter is identifiable from the kernel on the support of the weight function.
  • domain assumption Assumption 2.2(ii): inf det[K_theta](x_1,...,x_p) > 0 uniformly over x_i.
    Load-bearing for boundedness of the score integrand and Hessian convergence; false for the Gaussian, Laplace and Cauchy kernels used in Section 5.
  • domain assumption Assumption 2.2(iii): fourth continuous differentiability of the kernel with bounded derivatives.
    Standard smoothness condition used for Taylor expansions and moment bounds.
  • domain assumption Assumption 2.2(iv): compact convex parameter space admitting a Sobolev inequality.
    Used to control suprema over the parameter space in the proof of condition (M1).
  • domain assumption Brillinger mixing condition for the stationary DPP, invoked in the proof of Theorem 4.1 but not listed among the assumptions.
    The proof of (M1) relies on the Brillinger mixing condition from Biscio and Lavancier (2016), yet the assumptions of the paper do not state it explicitly.
  • domain assumption Ergodicity of stationary DPPs, from Soshnikov (2000).
    Used to justify almost sure convergence of normalized composite likelihoods.
  • domain assumption Identifiability of the composite likelihood score: E[U_{n2}(alpha)] = 0 iff alpha = alpha0.
    Asserted in the proof of Theorem 3.2 without proof; a non-unique root would break consistency.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Moment convergence of the generalized maximum composite likelihood estimators for determinantal point processes." pith.science (2026). https://pith.science/paper/2NOHM33J

@misc{pith2026190901211,
  author       = {Pith},
  title        = {Pith review of: Moment convergence of the generalized maximum composite likelihood estimators for determinantal point processes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2NOHM33J}},
  note         = {Machine review of arXiv:1909.01211}
}
read the original abstract

The maximum composite likelihood estimator for parametric models of determinantal point processes (DPPs) is discussed. Since the joint intensities of these point processes are given by determinant of positive definite kernels, we have the explicit form of the joint intensities for every order. This fact enables us to consider the generalized maximum composite likelihood estimator for any order. This paper introduces the two step generalized composite likelihood estimator and shows the moment convergence of the estimator under a stationarity. Moreover, our results can yield information criteria for statistical model selection within DPPs.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

14 extracted references · 14 canonical work pages

  1. [1]

    Pure and Applied Mathematics (Amsterdam), Elsevier/Academic Press, Amsterdam

    Adams, R.\ A.\ and Fournier, J.\ J.\ F.\ Sobolev spaces . Pure and Applied Mathematics (Amsterdam), Elsevier/Academic Press, Amsterdam. 140 , second edition, Elsevier/Academic Press, Amsterdam. (2003)

  2. [2]

    IEEE Trans

    Akaike, H.\ A new look at the statistical model identification. IEEE Trans. Automatic Control . AC-19 , p.716-723. (1974)

  3. [3]

    and Lavancier, F

    Biscio, C.\ A.\ N. and Lavancier, F. Brillinger mixing of determinantal point processes and statistical applications. Electron. J. Stat. 10 , no.1, p.582-607. (2016)

  4. [4]

    and Lavancier, F

    Biscio, C.\ A.\ N. and Lavancier, F. Contrast estimation for parametric stationary determinantal point processes. Scand. J. Stat. 44 , no.1, p.204-229. (2017)

  5. [5]

    Statistical inference for ergodic point processes and application to limit order book

    Clinet, S.\ and Yoshida, N. Statistical inference for ergodic point processes and application to limit order book. Stochastic Process. Appl. 127, no.6, p.1800-1839. (2017)

  6. [6]

    Guan, Y.\ A composite likelihood approach in fitting spatial point process models. J. Amer. Statist. Assoc. 101 , no.476, p.1502-1512. (2006)

  7. [7]

    Zeros of G aussian analytic functions and determinantal point processes

    Hough, J.\ B., Krishnapur, M., Peres, Y.\ and Vir\' a g, B. Zeros of G aussian analytic functions and determinantal point processes. University Lecture Series.\ 51, American Mathematical Society, Providence, RI.\ (2009)

  8. [8]

    Lavancier, F., M ller, J.\ and Rubak, E.\ Determinantal point process models and statistical inference. J. R. Stat. Soc. Ser. B. Stat. Methodol. 77 , no. 4, p.853-877. (2016)

Show all 14 references
  1. [9]

    arXiv:1806.06231 [math.ST] (2018)

    Lavancier, F., Poinas, A., and Waagepetersen, R.\ Adaptive estimating function inference for non-stationary determinantal point processes. arXiv:1806.06231 [math.ST] (2018)

  2. [10]

    Advances in Appl

    Macchi, O.\ The coincidence approach to stochastic point processes. Advances in Appl. Probability. 7 p.83-122. (1975)

  3. [11]

    Insurance Math

    Shimizu, Y.\ and Zhang, Z.\ Estimating G erber- S hiu functions from discretely observed L \' e vy driven surplus. Insurance Math. Econom.\ 74, p.84-98. (2017)

  4. [12]

    Uspekhi Mat

    Soshnikov, A.\ Determinantal random point fields. Uspekhi Mat. Nauk. 55 , no.5 (335), p.107-160. (2000)

  5. [13]

    Waagepetersen, R.\ and Guan, Y.\ Two-step estimation for inhomogeneous spatial point processes. J. R. Stat. Soc. Ser. B Stat. Methodol. 71.\ no.3, 685-702.\ (2009)

  6. [14]

    Yoshida, N.\ Polynomial type large deviation inequalities and quasi-likelihood analysis for stochastic differential equations. Ann. Inst. Statist. Math. 63, no.3, 431-479. (2011)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.