Pith. sign in

REVIEW 3 major objections 4 minor 14 references

Sharp Convergence Rates of Empirical Unbalanced Optimal Transport for Spatio-Temporal Point Processes

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Under a sub-quadratic variance-growth condition, the empirical Kantorovich–Rubinstein distance between intensity measures of weakly stationary spatio-temporal point processes converges at sharp dimensional rates, with matching lower bounds.

desk verdict Solid new result: sharp upper and matching (Poisson-case) lower rates for empirical Kantorovich–Rubinstein distance under dependent spatio-temporal sampling; the central theorem holds up, soft spots are minor. read the letter →

arxiv 2509.04225 v1 pith:E7J4UJSU submitted 2025-09-04 math.ST stat.MLstat.TH

classification math.STstat.MLstat.TH MSC 62G0562G0762R2060D0560G60
keywords Kantorovich–Rubinsteindistanceunbalancedoptimaltransportspatio-temporalpointprocessesintensitymeasureestimationconvergenceratesminimaxoptimalityHawkesprocessWasserstein
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Unbalanced optimal transport generalizes classical optimal transport to measures with different total mass, making it natural for comparing event-stream intensities—rainfall records, seismic activity, neural spikes—where the amount of 'material' varies. This paper asks how quickly the empirical plug-in comparison based on observing a spatio-temporal point process up to time t approaches the true transport distance. It establishes that under a sub-quadratic variance-growth condition (PVar β), the expected error decays at three distinct rates determined by the intrinsic dimension of the spatial intensity, with a parametric mass-estimation term always present. The rates match classical i.i.d. Wasserstein rates when the temporal dependence is weak (β = 0), degrade as β approaches 1, and are complemented by matching lower bounds showing near minimax optimality.

What carries the argument

The machinery is the (p, C)-Kantorovich–Rubinstein distance—an unbalanced transport metric with a mass penalty of C^p/2 per unit of unmatched mass, so that C is the maximal distance across which mass can move—combined with a dyadic discretization on an ultrametric tree built from ε-coverings of the domain. The proof bounds the empirical error by flattening the continuous problem onto the tree and summing over resolution levels; at each level the stochastic error is controlled by (PVar β), a condition that the sum of variances over any partition of space grows at most like t^{1+β}. Dimensionality enters through (Dim α), the polynomial covering-number bound. This pairing separates the geometri

What would settle it

Observe a point process that violates (PVar β) by placing one copy of a random spatial point at every integer time: the count in any cell has variance proportional to t^2. The paper predicts the expected KR error cannot improve with t; measuring it directly at increasing t and finding decay would falsify the theorem's necessity claim. Conversely, for a Poisson process on [0,1]^d with β = 0, the predicted exponent t^{-1/(2p)} (d < 2p) or t^{-1/d} (d > 2p) can be checked by Monte Carlo; a systematic deviation in the exponent would refute Theorem 3.3.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central result is Theorem 3.3: for a weakly time-stationary L2 spatio-temporal point process whose partitional variance growth is subquadratic (PVar β, β ∈ [0, 1)) and whose spatial domain has covering-number dimension α, the expected (p, C)-Kantorovich–Rubinstein distance between the per-unit-time empirical measure μ̂_t and its intensity measure μ is bounded by three terms: a mass-estimation term C κ^{1/(2p)} t^{-(1-β)/(2p)}, a scale term that is C^{1-α/(2p)} t^{-(1-β)/(2p)} when α < 2p, a logarithmic term when α = 2p, and a dimension term t^{-(1-β)/α} when α > 2p. On the unit cube [0,1]^d this reads t^{-(1-β)/(2p)}, t^{-(1-β)/(2p)} log^{1/p} t, or t^{-(1-β)/d}

Load-bearing premise

Everything rests on (PVar β): over every measurable partition of space, the summed variances of the counts must grow no faster than t^{1+β} for some β < 1—if that fails (quadratic growth), the empirical measure need not concentrate and the stated rates collapse.

Editorial extensions

If this is right

  • For β = 0—covering the Poisson, binomial, Neyman–Scott, and subcritical Hawkes processes treated in Section 4—the empirical KRD converges at the same rates as the classical empirical Wasserstein distance in the i.i.d. case, with n replaced by m_μ t.
  • For d > 2p the mass-penalty parameter C drops out of the dominant rate in t, while for d < 2p smaller C improves the rate; in the critical case d = 2p there is a logarithmic correction.
  • The triangle inequality transfers the bounds to one- and two-sample plug-in estimators of KR_{p,C}(μ, ν), so that the statistical error for comparing two point processes decays at the rate set by the shorter observation time.
  • Lower bounds in Section 5 show that for Poisson observations the plug-in empirical measure is near minimax rate-optimal in both t and C, up to logarithmic factors at α = 2p.
  • The same variance-growth assumption yields comparable rates for other unbalanced divergences, including Hellinger–Kantorovich, Gaussian–Hellinger, and collider-event EMD, via a p-dominated divergence framework.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The critical dimension 2p marks a genuine qualitative boundary: two-dimensional images with p = 1 enjoy essentially parametric rates, while higher-dimensional intensity fields require substantially longer observation windows or dimension reduction to reach useful precision.
  • For dependent models with β > 0, the universal mass-estimation term C t^{-(1-β)/(2p)} appears in every regime, suggesting that tests or confidence intervals built on the KRD will inherit a total-mass bottleneck no matter the spatial dimension.
  • The paper leaves distributional limits open; the sharp rates provide the natural scaling for conjecturing central limit theorems, though their proof would need stronger mixing assumptions than (PVar β).
  • The long-range dependent construction used for β ∈ (0, 1) lower bounds could likely be adapted to other heavy-tailed renewal mechanisms, giving a family of models whose empirical KRD error decays exactly at t^{-(1-β)/α}.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies empirical plug-in estimation of the (p,C)-Kantorovich–Rubinstein distance (KRD) for spatial intensity measures of weakly time-stationary spatio-temporal point processes observed over an expanding time window (0,t]. The main result, Theorem 3.3, gives upper bounds for E[KR_{p,C}(μ̂_t, μ)] under a covering-dimension assumption (Dim α) and a partition-variance-growth assumption (PVar β), yielding rates t^{-(1-β)/(2p)} for α<2p, t^{-(1-β)/(2p)} log^{1/p}t for α=2p, and t^{-(1-β)/α} for α>2p, with explicit dependence on the mass-penalization parameter C. The paper also extends these bounds to other unbalanced OT divergences, connects (PVar β) to time-reduced factorial covariance measures via (RCov β), verifies the conditions for Poisson, binomial, Poisson-cluster, Hawkes, and log-Gaussian Cox processes, and proves pointwise and minimax lower bounds that match the upper rates up to logarithmic factors. The proofs use a dyadic ultrametric-tree discretization and detailed moment-measure calculations.

Significance. If the main results are correct, this is a substantial contribution. It moves statistical unbalanced OT beyond the i.i.d. finite-support setting, provides sharp dimension-dependent phase transitions, and gives a clean stochastic condition, (PVar β), that is verifiable for a broad class of dependent spatio-temporal point processes. The paper is also notable for its explicit treatment of the dependence on the mass-penalization parameter C, its matching lower bounds, and its careful spelling out of moment-measure computations for Hawkes, Neyman–Scott, and log-Gaussian Cox processes. There are no fitted parameters: the rates are genuine consequences of the stated assumptions. These strengths justify careful review despite the issues below.

major comments (3)
  1. [§3.1, Theorem 3.3, Eq. (3.2)]
  2. [§4.2, Proposition 4.5(i)]
  3. [§3.2, Theorem 3.12]
minor comments (4)
  1. [§3.1, Corollary 3.4]
  2. [§5.1, Example 5.3]
  3. [Appendix D.2–D.3]
  4. [§3.1, Corollary 3.5]

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The central claim (Theorem 3.3) is a genuine consequence of the stated assumptions (Dim α) and (PVar β). The proof uses a dyadic discretization / ultrametric-tree argument, where each step is an inequality (Jensen, Cauchy–Schwarz, covering-number bounds, variance control) and the final rate is obtained by minimizing in ε. (PVar β) is an assumption on the point process, not a fitted parameter; the paper verifies it for Poisson, binomial, Poisson-cluster, Hawkes, and log-Gaussian Cox processes by direct covariance calculations, and the lower bounds are obtained from independent constructions (mass-deviation lower bounds, long-range dependence examples, and minimax arguments for Poisson processes). The self-citations (Heinemann et al. 2023 for finite-support KRD structure; Hundrieser et al. 2025 for finite-space statistical analysis) are published mathematical theorems used as ingredients, not as conclusions, and do not presuppose the present results. A minor constant typo in Proposition 4.5(i) (κ should be m_μ + 2κ′, not 2κ′) affects only the constant, not the rate, and is not circular. Overall, the derivation chain is self-contained and no reduction of a prediction to an input by definition or by self-citation is present.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central result rests on two clearly stated assumptions: the variance-growth condition (PVar beta) and the intrinsic-dimension condition (Dim alpha). The paper verifies (PVar beta) for a range of standard point-process models, so the axioms are not ad hoc. No new entities are postulated; the Kantorovich-Rubinstein distance is adopted from prior literature.

assumptions (6)
  • domain assumption Weak time-stationarity and L2 integrability of the spatio-temporal point process (Definition 1.2).
    Ensures the spatial intensity measure mu exists via E[Exi(A x (0,t])] = t mu(A) (Proposition 4.1) and that the empirical measure (1.5) targets mu.
  • domain assumption (PVar beta): for any partition (A_i) of X and t >= t0, sum_i Var(Exi(A_i x (0,t])) <= kappa t^{1+beta}, beta in [0,1).
    Load-bearing stochastic control used in the proof of Theorem 3.3. If beta approaches 1 the convergence rates vanish, and the paper shows quadratic growth can break consistency.
  • domain assumption (Dim alpha): covering numbers N(epsilon, X, d) <= A epsilon^{-alpha}.
    Gives the intrinsic dimension control that drives the dyadic tree discretization and the bound on the number of nodes at each level in the proof of Theorem 3.3.
  • standard math Discrete KRD bounds on ultrametric trees from Heinemann et al. (2023, Theorem 2.3).
    Used as Lemma 3.15 to bound the discretized KRD; this is an established theorem cited from earlier work.
  • standard math Existence of optimal plans for the entropy-transport formulation (Liero et al., 2018, Theorem 3.1) and metrization of weak convergence by Wasserstein-type distances (Villani, 2008).
    Used in Section 2 to establish existence of unbalanced OT plans, metric properties, and consistency (Propositions 2.1 and 2.6).
  • standard math Packing-covering duality and minimax lower-bound results from Singh and Poczos (2018), Weed and Bach (2019), and Niles-Weed and Rigollet (2022).
    These external results supply the lower-bound machinery for the Poisson case in Section 5.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sharp Convergence Rates of Empirical Unbalanced Optimal Transport for Spatio-Temporal Point Processes." pith.science (2026). https://pith.science/paper/E7J4UJSU

@misc{pith2026250904225,
  author       = {Pith},
  title        = {Pith review of: Sharp Convergence Rates of Empirical Unbalanced Optimal Transport for Spatio-Temporal Point Processes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E7J4UJSU}},
  note         = {Machine review of arXiv:2509.04225}
}
abstract

We statistically analyze empirical plug-in estimators for unbalanced optimal transport (UOT) formalisms, focusing on the Kantorovich-Rubinstein distance, between general intensity measures based on observations from spatio-temporal point processes. Specifically, we model the observations by two weakly time-stationary point processes with spatial intensity measures $\mu$ and $\nu$ over the expanding window $(0,t]$ as $t$ increases to infinity, and establish sharp convergence rates of the empirical UOT in terms of the intrinsic dimensions of the measures. We assume a sub-quadratic temporal growth condition of the variance of the process, which allows for a wide range of temporal dependencies. As the growth approaches quadratic, the convergence rate becomes slower. This variance assumption is related to the time-reduced factorial covariance measure, and we exemplify its validity for various point processes, including the Poisson cluster, Hawkes, Neyman-Scott, and log-Gaussian Cox processes. Complementary to our upper bounds, we also derive matching lower bounds for various spatio-temporal point processes of interest and establish near minimax rate optimality of the empirical Kantorovich-Rubinstein distance.

Figures

Figures reproduced from arXiv: 2509.04225 by the authors.

Figure 1
Figure 1. Optimal plans for balanced and unbalanced optimal transport between two point clouds (32 red and 26 blue points). The presence of gray lines connecting two points indicates that mass is transported between them. (a): Optimal transport plan according to 2-Wasserstein distance between red and blue points, each weighted with mass 1/32 and 1/26, respectively. (b – d): Unbalanced optimal transport plan according to (2, C… view at source ↗
Figure 2
Figure 2. Estimation of KR distances, exemplified for two ST Poisson processes. (a): Left: a realization of a Poisson point process Ξ on R 2 × (0,∞) with spatial intensity measure 2N ((−0.5, 0) ⊺ , Σ) + N ((−3.5, 3.5) ⊺ , Σ), Σ = diag(0.5, 0.5). Right: Ξ(⋅ × (0, 50]) for the same real￾ization. The color shades indicate (from light to dark red/blue) the observation times of points for better visuals. (b): A realization of a Po… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

14 extracted references · 12 canonical work pages

  1. [1]

    Ajtai, M., Koml´ os, J., & Tusn´ ady, G. (1984). On optimal matchings.Combinatorica, 4(4), 259–264. Altschuler, J., Niles-Weed, J., & Rigollet, P. (2017). Near-linear time approximation algorithms for optimal transport via Sinkhorn iteration. In I. Guyon, U. von Luxburg, & others (Eds.), Advances in Neural Information Processing Systems , volume 30: Curra...

  2. [2]

    Baddeley, A

    CRC press Boca Raton. Baddeley, A. & Turner, R. (2000). Practical maximum pseudolikelihood for spatial point patterns. Australian & New Zealand Journal of Statistics , 42(3), 283–322. Bai, S. & Taqqu, M. S. (2018). How the instability of ranks under long memory affects large-sample inference. Statistical Science, 33(1), 96–116. Barthe, F. & Bordenave, C. ...

  3. [5]

    Bertsekas, D

    Athena Scientific. Bertsekas, D. P. & Castanon, D. A. (1989). The auction algorithm for the transportation problem. Annals of Operations Research, 20(1), 67–96. B laszczyszyn, B., Yogeshwaran, D., & Yukich, J. E. (2019). Limit theory for geometric statistics of point processes having fast decay of correlations. The Annals of Probability , 47(2), 835–895. ...

  4. [6]

    Diggle, P. J. (2013). Statistical analysis of spatial and spatio-temporal point patterns . CRC press. Divol, V. (2022). Measure estimation on manifolds: an optimal transport approach. Probability Theory and Related Fields , 183(1), 581–647. Dudley, R. M. (1969). The speed of mean Glivenko-Cantelli convergence.The Annals of Mathematical Statistics, 40(1), ...

  5. [7]

    Figalli, A

    Wiley, 2nd edition. Figalli, A. (2010). The optimal partial transport problem. Archive for Rational Mechanics and Analysis, 195(2), 533–560. Figalli, A. & Gigli, N. (2010). A new transportation distance between non-negative measures, with applications to gradients flows with Dirichlet boundary conditions. Journal de math´ ematiques pures et appliqu´ ees, ...

  6. [9]

    Parallel Unbalanced Optimal Transport Regularization for Large Scale Imaging Problems

    Cambridge University Press. Lavancier, F., Møller, J., & Rubak, E. (2015). Determinantal point process models and statistical inference. Journal of the Royal Statistical Society Series B: Statistical Methodology , 77(4), 853–877. Ledoux, M. (2019). On optimal matching of Gaussian samples. Journal of Mathematical Sciences , 238, 495–522. Lee, J., Bertrand,...

  7. [13]

    Santambrogio, F

    McGraw-Hill. Santambrogio, F. (2015). Optimal transport for applied mathematicians: Calculus of variations, PDEs, and modeling . Progress in Nonlinear Differential Equations and Their Applications. Springer. Savar´ e, G., Sodini, G. E., et al. (2022). A simple relaxation approach to duality for optimal transport problems in completely regular spaces. Jour...

  8. [30]

    Meyer, S., Elias, J., & H¨ ohle, M. (2012). A space-time conditional intensity model for invasive meningococcal disease occurrence. Biometrics, 68(2), 607–616. Mohler, G. O., Short, M. B., Brantingham, P. J., Schoenberg, F. P., & Tita, G. E. (2011). Self- exciting point process modeling of crime. Journal of the American Statistical Association , 106(493),...

Show all 14 references
  1. [47]

    Victor, J

    Cambridge university press. Victor, J. D. & Purpura, K. P. (1997). Metric-space analysis of spike trains: Theory, algorithms and application. Network: computation in neural systems , 8(2), 127–164. Villani, C. (2003). Topics in optimal transportation, volume 58 of Graduate Stu...

  2. [261]

    Boissard, E

    American Mathematical Society. Boissard, E. & Le Gouic, T. (2014). On the mean speed of convergence of empirical and occupation measures in Wasserstein distance. Annales de l’Institut Henri Poincar´ e, Probabilit´ es et Statistiques, 50(2), 539–563. Bonneel, N., van de Panne, ...

  3. [668]

    Rudin, W. (1991). Functional Analysis, volume

  4. [1139]

    & Scott, E

    Neyman, J. & Scott, E. L. (1958). Statistical approach to problems of cosmology. Journal of the Royal Statistical Society Series B: Statistical Methodology , 20(1), 1–29. Niles-Weed, J. & Berthet, Q. (2022). Minimax estimation of smooth densities in Wasserstein distance. The A...

  5. [2011]

    Bourne, D

    , 30(6). Bourne, D. P., Schmitzer, B., & Wirth, B. (2018). Semi-discrete unbalanced optimal transport and quantization. Preprint arXiv:1808.01962. Brillinger, D. R. (1975). Statistical inference for stationary point processes. Stochastic processes and related topics , 1, 55–99...

  6. [2420]

    Gonz´ alez, J

    Addison-Wesley. Gonz´ alez, J. A., Rodr ´ ıguez-Cort´ es, F. J., Cronie, O., & Mateu, J. (2016). Spatio-temporal point process statistics: A review. Spatial Statistics, 18, 505–544. Gramfort, A., Peyr´ e, G., & Cuturi, M. (2015). Fast optimal transport averaging of neuroimagin...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.