Pith. sign in

REVIEW 4 major objections 5 minor 27 references

An Independence Test Based on Recurrence Rates

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A test that compares joint and marginal recurrence rates detects dependence between random elements whenever dependence leaves a trace in the pairwise distances.

desk verdict A promising new independence test with strong simulations, but the proof of the null distribution has a concrete algebraic error that needs fixing before the results can be trusted. read the letter →

arxiv 1908.03305 v1 pith:XI7IDPXH submitted 2019-08-09 math.ST stat.TH

classification math.STstat.TH MSC 62H1562H20
keywords independencetestrecurrenceratesU-processCramér-vonMisesstatisticmetricspacesdistance-basedcontiguousalternativestimeseries
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes an independence test for random elements X and Y that uses only the distances between sample points, so it can be applied to variables, vectors, functions, and time series in any metric space. The test compares the joint recurrence rate, the fraction of pairs that are close in both coordinates, with the product of the two marginal recurrence rates, which are equal under independence. The statistic is a Cramér-von Mises functional that integrates the squared difference over all thresholds r and s with respect to a weight distribution, so no single radius has to be chosen. Under the null the statistic converges to the squared norm of a centered Gaussian process, and under alternatives for which the pair-distances $d(X_1,X_2)$ and $d(Y_1,Y_2)$ are not independent it diverges to infinity, making the test consistent. The paper also derives the limiting drift under contiguous alternatives and reports simulations where the test matches or outperforms several distance-based tests.

What carries the argument

The engine is the recurrence rate: for a threshold r, the fraction of pairs of observations whose distance is below r. The test looks at these rates for X, for Y, and for the pair (X,Y) simultaneously, and the key process is the normalized difference $E_n(r,s)$. To get asymptotic theory, the paper approximates $E_n$ by a four-index U-process $E'_n(r,s)$, a U-process being a normalized average of a kernel over all ordered 4-tuples of distinct sample indices, with the approximation error bounded by $4/\sqrt{n}$. This approximation puts the process under a U-process weak-convergence theorem that yields the Gaussian null limit. The consistency argument is carried by the deterministic discrepancy function $\mu(r,s)=P(d(X_1,X_2)<r,d(Y_1,Y_2)<s)-P(d(X_1,X_2)<r)P(d(Y_1,Y_2)<s)$: positivity of the weight density plus any point with $\mu^2>0$ forces $T_n$ to diverge. The weight G, typically a product of Gaussian densities centered at the average distances, is what lets the statistic average over all thresholds instead of picking one.

What would settle it

Fix a dependent distribution with dependent pair-distances, for example bivariate normal with nonzero correlation, and simulate $T_n$ for n = 50, 100, 200, 400, 800 using a weight with positive density. The consistency theorem predicts $T_n \to \infty$ in probability; if the statistic stays stochastically bounded for such a distribution, Theorem 4 is wrong. A complementary check is to estimate the surface $\mu(r,s)$ for any candidate alternative: if it is identically zero while X and Y are dependent, the theorem does not apply and the test should show no power.

Watch

Extended reading notes

Core claim

The central claim is that independence can be tested by contrasting the joint recurrence rate with the product of the marginal recurrence rates. With $E_n(r,s)=\sqrt{n}(\mathrm{RR}^{X,Y}_n(r,s)-\mathrm{RR}^X_n(r)\mathrm{RR}^Y_n(s))$, the proposed statistic is $T_n = n\int_0^\infty\int_0^\infty E_n(r,s)^2\,dG(r,s)$. The paper proves two complementary limit statements. Theorem 3 shows that when X and Y are independent and the distributions of the distances are continuous, the process $\{E_n(r,s)\}$ converges weakly to a centered Gaussian process $\{\mathcal{E}(r,s)\}$, so $T_n$ converges in distribution to $\|\mathcal{E}\|^2_{L^2(dG)}$. Theorem 4 shows that if $d(X_1,X_2)$ and $d(Y_1,Y_2)$ are continuous and not independent, and the weight function has density positive everywhere, then $T_n \to \infty$ in probability; this is the consistency result, with multivariate normal pairs as a corollary. Under contiguous alternatives the limit process becomes $\{\mathcal{E}(r,s)+\delta\mu(r,s)\}$, adding a deterministic drift that gives nontrivial local power. The paper notes that some dependent distributions produce independent pair-distances, and for those the consistency argument does not apply.

Load-bearing premise

All the consistency guarantees rest on the assumption that dependence between X and Y shows up as dependence between the pairwise distances $d(X_1,X_2)$ and $d(Y_1,Y_2)$; the paper itself exhibits dependent variables whose distances are independent, and for those distributions the test has no asymptotic power.

Editorial extensions

If this is right

  • If the paper is correct, any dependence between X and Y that makes the pair-distance variables dependent is eventually detected: $T_n$ grows without bound as $n$ increases.
  • Under $H_0$ the asymptotic null distribution is available, and for finite samples the paper's permutation scheme estimates p-values by breaking the X-Y pairing, so the test can be calibrated without simulating the limiting Gaussian process.
  • Because the statistic uses only distances, the same test applies to random vectors, function-valued data, and discrete or continuous time series; the multivariate normal case is covered by Corollary 1.
  • A sup-version $T'_n=\sqrt{n}\sup_{r,s>0}|E_n(r,s)|$ is allowed by the same arguments, giving a user an alternative that avoids choosing a weight function.
  • Under contiguous alternatives the limiting power is governed by the drift $\delta\mu(r,s)$; in the simulation study the test shows higher power than several distance-based competitors in most of the settings considered.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The theorem's condition is about the pair-distance variables, not about (X,Y) itself; effectively the test's null domain is 'the pair distances are independent,' so dependence that lives only in angles or absolute locations is invisible to the statistic.
  • An adaptive weight function could in principle concentrate mass where $\mu(r,s)$ is large for a suspected alternative; the paper fixes G for the theory, so data-dependent G would need new arguments but is a natural practical extension.
  • For time series, the permutation scheme removes the contemporaneous X-Y pairing while preserving each series' internal order, suggesting the test can target cross-dependence even when the marginal series are serially dependent; the paper's asymptotic proofs are for i.i.d. pairs, so the time-series case rests on this permutation calibration plus simulation.
  • A quick diagnostic for a new application is to estimate $\mu(r,s)$ before relying on the test: if the estimate is flat at zero while X and Y are believed dependent, the consistency theorem offers no power guarantee and another test should be considered.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a new test of independence between two random elements X and Y taking values in metric spaces. The test statistic T_n is a Cramér-von Mises-type functional of the U-process E_n(r,s) = √n(RR^{X,Y}_n(r,s) − RR^X_n(r) RR^Y_n(s)), where the RR terms are empirical recurrence rates based on indicators 1{d(X_i,X_j)<r} and 1{d(Y_i,Y_j)<s}. By integrating the squared process against a distribution G on (0,∞)^2, the test avoids choosing a threshold radius. The paper claims: under H0, T_n converges weakly to the squared norm of a centered Gaussian process (Theorem 3); under contiguous alternatives the process acquires a deterministic drift (Proposition 1 and the subsequent assertion); and under alternatives for which the pair-distance variables d(X_1,X_2) and d(Y_1,Y_2) are dependent, T_n diverges to infinity in probability (Theorem 4). Simulation studies compare the test with HHG, distance covariance, and HSIC for scalar, vector, and time-series data.

Significance. If the theoretical claims are correct, the test is an attractive addition to the independence-testing literature: it applies to general metric spaces, avoids threshold selection by integrating over all radii, has a computationally explicit statistic (Section 3.4), and the reported power is competitive with or better than several popular tests in many of the studied alternatives. The paper also ships a permutation procedure for calibration, which is practically useful. However, the proof of the null-distribution theorem has a concrete algebraic error (Lemma 2) that as written invalidates the transfer from the approximating U-process to the actual process; the consistency proof omits a required uniform-integration step; and the contiguity result is only sketched. These are load-bearing issues, although all appear repairable without changing the overall approach. The structural limitation identified in Remark 3, that the test only detects dependence between the distance variables, is real and should be prominently acknowledged.

major comments (4)
  1. [Section 2.1, Lemma 2] Lemma 2's assertion that 0 ≤ H_n(r,s) ≤ 4/√n is false. In the decomposition displayed after equation (23), the coefficient of the I_4 term is 1/N^2 − 1/(N(n−2)(n−3)), which is negative for every n>2; the manuscript's equation (24) treats this coefficient as positive. A concrete configuration with n=4, d(X_1,X_2)<r, d(Y_3,Y_4)<s, and no other recurrences gives RR^X_n=RR^Y_n=1/6, RR^{X,Y}_n=0, hence E_n = −1/18; the corresponding E'_n equals −1/3, so H_n = E'_n − E_n = −5/18 < 0. Thus the stated nonnegativity is false. Since Theorem 3 uses Lemma 2 to replace E_n by E'_n, the proof of Theorem 3 is invalid as written. The likely repair is to replace the nonnegativity claim with an absolute bound |H_n| ≤ C/√n, which follows from the same decomposition because all terms are bounded by constants times n^{−1/2}; but this correction is absent from the manuscript.
  2. [Section 2.1, Theorem 4 (proof)] The consistency proof uses the inequality T_n ≥ (n/2)∫µ² dG − n∫(RR^{X,Y}_n − RR^X_n RR^Y_n − µ)² dG and then asserts that the second term is negligible as n→∞. No argument is given for this negligibility. The difference process involves products of U-statistics, and while pointwise convergence of RR^{X,Y}_n, RR^X_n, and RR^Y_n suggests the integrated squared difference should be o_P(1), proving this requires either a Donsker-type tightness result in L²(G) or a dominated-convergence argument for n∫(·)²dG. Without this step, the claim T_n →_P ∞ does not follow from the displayed inequality. The gap is load-bearing for the consistency result, though it appears fixable with a standard empirical-process argument.
  3. [Section 2.2, contiguous alternatives] The weak convergence under H_n of {E_n(r,s)} to {E(r,s) + δµ(r,s)} is asserted after the sentence "With a little more work, using the Le Cam third lemma," but no proof is provided. This is one of the three main theoretical claims of the paper, alongside Theorem 3 and Theorem 4. The assertion is not a trivial corollary: one must verify the contiguity conditions, handle the U-statistic remainder H_n under the triangular array of densities f^{(n)}_{X,Y}, and justify the weak convergence of the process in the space where the functional T_n is continuous. The omission of these details makes the contiguous-alternative result unverifiable as written. At minimum, a full proof or a precise reference with the required verification should be supplied.
  4. [Section 4.1, Tables 1 and 2] In the "Four independent clouds" row, which is a null case (X and Y are independent), the column labeled N(1,1) reports power 0.512 for n=30 and n=50, while every other column in that row reports 0.046–0.057. At nominal level 0.05, a size of 0.512 is inconsistent with the paper's claim that the test has correct level under H0. Either there is a simulation error, a mislabeling, or a calibration problem specific to this weight function and sample size. The authors should explain or correct this entry; as printed, it undermines confidence in the numerical comparisons.
minor comments (5)
  1. [Section 2.1, Theorem 3 (statement)] Theorem 3 as stated asserts weak convergence of {E_n(r,s) − E(E_n(r,s))} without assuming H0, but the proof explicitly begins with "If H0 is true". The theorem statement should include the independence assumption, since the claimed limit process and covariance structure are derived under H0.
  2. [Section 2.1, Remark 1] Remark 1 says the test statistic T_n is the norm of the Gaussian process, but T_n is defined with an un-squared distance from the definition: T_n = n∫(RR^{X,Y}_n − RR^X_n RR^Y_n)² dG, which converges to the squared norm ||E||². The remark should say "squared norm" or adjust the notation consistently.
  3. [Section 2.1, Lemma 1 (proof)] The intermediate algebra in the proof of Lemma 1 contains several apparent typos, such as "p^{(3)}_X(r∧r′)p_Y(s)" where the context suggests "p^{(3)}_X(r∧r′)p_Y(s)p_Y(s′)" and analogous expressions for p^{(3)}_Y. The final covariance formula (4) is plausible, but the intermediate display is difficult to follow and should be corrected.
  4. [Section 3.2, permutation p-value] The unbiasedness argument for the permutation estimator is sketchy: the claim that each B_i is Binomial(m,1/n!) assumes that the values T_n under different permutations are distinct with probability one and that the permutation distribution is uniform over all n! permutations, but ties can occur and the double limit in n and m is not addressed. A more careful statement of the permutation test's exactness or asymptotic validity would improve the presentation.
  5. [General notation] The same letter φ is used both for the standard normal density and for the normal CDF in Section 3.1 (equations involving φ^{-1}(F_X(X))), which is confusing; the text should distinguish the density from the cumulative distribution function.

Circularity Check

0 steps flagged · score 1.0 of 10

Self-contained test derivation; the only self-citation (Kalemkerian 2017) defines a simulation model and is not load-bearing.

full rationale

The central derivation is self-contained. The statistic T_n in (3) is a direct functional of the empirical recurrence rates; no term in it is fitted to the quantity being predicted. Theorem 3 is obtained from a standard U-process weak convergence theorem (Arcones & Gine 1993), with the entropy integral verified for the class of rectangle indicators; Lemma 2 (even if its H_n bound has an algebraic issue) is an approximation device and not an input that already contains the Gaussian limit. Theorem 4 is a consequence of the SLLN for U-statistics: the integrand converges to mu^2, and the assumption that the distance variables are dependent is exactly what guarantees some positive-mu rectangle. This is a characterization of the population analogue, not a renaming of the conclusion: Remark 3 explicitly exhibits dependent (X,Y) for which the distance variables are independent and the test has no power, so the consistency claim is honestly scoped. The contiguous-alternative analysis uses Le Cam's third lemma and Cabana's contiguity criterion, both external. The only self-citation is reference [17] (Kalemkerian 2017), used in Section 4.3 only to define the FOU(2) simulation model; it is not used in any proof or in the definition of the test. The data-dependent choice of G in Section 3.3 is a practical weighting heuristic, not a fitted parameter that is later reported as a prediction. The algebra concern about Lemma 2's H_n bound is a correctness risk in a proof step, not a circularity: a wrong bound does not make a result identical to its assumptions. Overall, no circular step is identifiable.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The central claim rests on standard U-process theory and two domain assumptions: continuity of distance distributions and strict positivity of the weight density. The data-dependent choice of G is a heuristic free parameter not covered by the theory. No new entities are postulated.

free parameters (2)
  • Weight function G (Gaussian mean and variance) = Data-dependent: mu = E[d(X1,X2)], sigma^2 = V[d(X1,X2)] for g1; similarly for g2
    The test statistic T_n depends on a user-chosen distribution G; in Section 3.3 the authors propose estimating the Gaussian parameters from pairwise distances. This choice is heuristic and not accounted for in the asymptotic theory.
  • Distance metric = Euclidean norm in simulations
    The theory allows any metric, but simulations use Euclidean distance; the choice affects power.
assumptions (6)
  • standard math Strong law of large numbers for U-statistics (Hoeffding 1961) ensures convergence of recurrence rates to their expectations.
    Used in equation (1) to justify that RR values converge almost surely.
  • standard math Arcones-Giné Theorem 4.10 (1993) on weak convergence of U-processes under bracketing entropy conditions.
    Used in the proof of Theorem 3 to obtain weak convergence of the process E_n.
  • domain assumption Continuity of the distribution functions of d(X1,X2) and d(Y1,Y2).
    Assumed in Theorem 3 and Theorem 4 to ensure the bracketing arguments and separability.
  • domain assumption The weight function G has density g that is strictly positive everywhere on (0,infty)^2.
    Required in Theorem 4 for consistency; the data-driven Gaussian choice satisfies this.
  • standard math Le Cam third lemma (Le Cam & Yang 1990; Oosterhoff & Van Zwet 1979) is invoked for the contiguous alternative weak convergence.
    The paper states that with 'a little more work' the third lemma proves the result, but the proof is not written out.
  • standard math For Corollary 1, the identity COV(Z^2,T^2) = 2 COV(Z,T)^2 for centered normal bivariate vectors.
    Used to show that for multivariate normal X,Y, the squared norms are dependent if X,Y are not independent.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Independence Test Based on Recurrence Rates." pith.science (2026). https://pith.science/paper/XI7IDPXH

@misc{pith2026190803305,
  author       = {Pith},
  title        = {Pith review of: An Independence Test Based on Recurrence Rates},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XI7IDPXH}},
  note         = {Machine review of arXiv:1908.03305}
}
abstract

A new test of independence between random elements is presented in this article. The test is based on a functional of the Cram\'{e}r-von Mises type, which is applied to a $U$-process that is defined from the recurrence rates. Theorems of asymptotic distribution under $H_{0},$ and consistency under a wide class of alternatives are obtained. The results under contiguous alternatives are also shown. The test has a very good behaviour under several alternatives, which shows that in many cases there is clearly larger power when compared to other tests that are widely used in literature. In addition, the new test could be used for discrete or continuous time series.

Figures

Figures reproduced from arXiv: 1908.03305 by the authors.

Figure 1
Figure 1. σ 2 X,Y (r, r) in function of r, in the case X and Y are independent and N(0, 1). Tn, but given the observed value from our sample that we call tobs, we could generate, by a permutation procedure, a large sample of Tn with which we can estimate P (Tn ≥ tobs). Given (X1, Y1),(X2, Y2), ...,(Xn, Yn) i.i.d. sample of (X, Y ). Observe that the distri￾bution of Tn depends of the joint distribution of (X1, Y1),(X2, Y2), ..… view at source ↗
Figure 2
Figure 2. Parabola, Two parabolas, circle, diamond, wshape and four independent [PITH_FULL_IMAGE:figures/full_fig_p027_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 26 canonical work pages

  1. [1]

    Arcones, M. A. and Gin´ e, E. (1993). Limit Theorems for U-Processes. The Annals of Probability Vol 21-3, 1494-1542

  2. [2]

    & Caba˜ na, E., (2016)

    Arratia, A., Caba˜ na, A. & Caba˜ na, E., (2016). A construction of Continuous time ARMA models by iterations of Ornstein-Uhlenbeck process,SORT Vol 40 (2) 267-302

  3. [3]

    K., Rizzo, M

    Bakirov, N. K., Rizzo, M. L., and Sz´ ekely, G. J. (2006). A multivariate non- para- metric test of independence. Journal of Multivariate Analysis, 97(8):1742-1756

  4. [4]

    Beran, R., Bilodeau, M., and Lafaye de Micheaux, P. (2007). Nonparametric tests of independence between random vectors. Journal of Multivariate Analysis. 98(9):1805–1824

  5. [5]

    and Lafaye de Micheaux, P

    Bilodeau, M. and Lafaye de Micheaux, P. (2005). A multivariate empirical carac- teristic function test of independence with normal marginals. Journal of ultivariate Analysis, 95(2):345–369

  6. [6]

    Blomqvist, N. (1950). On a measure of dependence between two random variables. The Annals of Mathematical Statistics 593-600. 24

  7. [7]

    Boglioni, G. (2016). A consistent test of independence between random vectors. https://papyrus.bib.umontreal.ca/xmlui/bitstream/handle/1866/18773/Bo glioni Beaulieu Guillaume 2016 memoire.pdf?sequence=2

  8. [8]

    Caba˜ na, E. M. (1997). Contiguidad, pruebas de ajuste y Procesos Emp´ ı ricos Trans- formados. D´ ecima escuela venezolana de Matem´ aticas

Show all 27 references
  1. [9]

    P, Oliffson Kamphorst, S, Ruelle, D

    Eckmann, J. P, Oliffson Kamphorst, S, Ruelle, D. (1987). Recurrence plots of dy- namical systems. Europhys. Lett. 4, 973-977

  2. [10]

    Galton, F. (1888). Co-relations and their measurement, chiefly from anthropometric data. Proceedings of the Royal Society of London. 45(273-279):135–145

  3. [11]

    and Zinn, J

    Gin´ e, E. and Zinn, J. (1986). Lectures on the central limit theorem for empirical processes. Lecture Note in Math. 1221. 50-113. Springer, New York

  4. [12]

    and Sch¨ olkopf, B

    Gretton, A., Bousquet, O., Smola, A. and Sch¨ olkopf, B. (2005). Measuring statisti- cal dependence with hilbert-scmidts norms. International Conference on algorithmic learning theory. 63-77. Springer

  5. [13]

    Heller, R., Heller, Y., and Gorfine, M. (2012). A consistent multivariate test of association based on ranks of distances. Biometrika. 100(2):503–510

  6. [14]

    Hoeffding, W. (1961). The strong law of large numbers for U-statistics. http://www.lib.ncsu.edu/resolver/1840.4/2128

  7. [15]

    Hoefdding, W. (1948). A non-parametric test of independence. The Annals of Math- ematical Statistics. 546-557

  8. [16]

    and R´ emillard, B

    Genest, G. and R´ emillard, B. (2004). Test of independence and randomness based on the empirical copula process. Test, 13(2): 335-369

  9. [17]

    Kalemkerian, J. (2017). Fractional Iterated Ornstein-Uhlenbeck Processes. arXiv preprint arXiv:1709.07143v1[math.ST]

  10. [18]

    and Holmes, M

    Kojadinovic, I. and Holmes, M. (2009). Tests of independence among continuous random vectors based on cram´ er-von mises functional of the empirical copula process. Journal of Multivariate Analysis, 100(6): 1137-1154

  11. [19]

    Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1/2):81–93

  12. [20]

    & Yang, G

    Le Cam, L. & Yang, G. L. (1990). Asymptotics in Statistics. Some Basic Concepts. Springer, New York

  13. [21]

    Marwan, N. (2008). A Historical Review of Recurrence Plots. European Physical Journal—Special Topics, 164, 3-12. 25

  14. [22]

    & Van Zwet, W

    Oosterhoof, J. & Van Zwet, W. R. (1979). A note on contiguity and Hellinger distance. Contributions to Statistics. Jaroslav H´ ajek Memorial Volume (J. Jorecov´ a, ed) Reidel Dordrecht. 157-166

  15. [23]

    Pearson, K. (1898). Mathematical contributions to the theory of evolution on a form of spurious correlation which may arise when indices are used in the measurement of organs. Proceedings of the Royal Society of London. 60(359-367):489–498

  16. [24]

    Spearman, C. (1904). The proof and measurement of association between two things. The American journal of psychology. 15(1):72–101

  17. [25]

    J., Rizzo, M

    Sz´ ekely, G. J., Rizzo, M. L., Bakirov, N. K., et al. (2007). Measuring and testing dependence by correlation of distances. The Annals of Statistics. 35(6):2769– 2794

  18. [26]

    J., Rizzo, M

    Sz´ ekely, G. J., Rizzo, M. L., et al. (2009). Brownian distance covariance. The annals of applied statistics, 3(4):1236–1265

  19. [27]

    Wilks, S. (1935). On the independence of k sets of normality distributed statistical variables. Econometrica, Journal of the Econometric Society. 309-326. 26 Figure 2: Parabola, Two parabolas, circle, diamond, wshape and four independent clouds. 27

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.