REVIEW 4 major objections 5 minor 27 references
An Independence Test Based on Recurrence Rates
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A test that compares joint and marginal recurrence rates detects dependence between random elements whenever dependence leaves a trace in the pairwise distances.
desk verdict A promising new independence test with strong simulations, but the proof of the null distribution has a concrete algebraic error that needs fixing before the results can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the recurrence rate: for a threshold r, the fraction of pairs of observations whose distance is below r. The test looks at these rates for X, for Y, and for the pair (X,Y) simultaneously, and the key process is the normalized difference $E_n(r,s)$. To get asymptotic theory, the paper approximates $E_n$ by a four-index U-process $E'_n(r,s)$, a U-process being a normalized average of a kernel over all ordered 4-tuples of distinct sample indices, with the approximation error bounded by $4/\sqrt{n}$. This approximation puts the process under a U-process weak-convergence theorem that yields the Gaussian null limit. The consistency argument is carried by the deterministic discrepancy function $\mu(r,s)=P(d(X_1,X_2)<r,d(Y_1,Y_2)<s)-P(d(X_1,X_2)<r)P(d(Y_1,Y_2)<s)$: positivity of the weight density plus any point with $\mu^2>0$ forces $T_n$ to diverge. The weight G, typically a product of Gaussian densities centered at the average distances, is what lets the statistic average over all thresholds instead of picking one.
What would settle it
Fix a dependent distribution with dependent pair-distances, for example bivariate normal with nonzero correlation, and simulate $T_n$ for n = 50, 100, 200, 400, 800 using a weight with positive density. The consistency theorem predicts $T_n \to \infty$ in probability; if the statistic stays stochastically bounded for such a distribution, Theorem 4 is wrong. A complementary check is to estimate the surface $\mu(r,s)$ for any candidate alternative: if it is identically zero while X and Y are dependent, the theorem does not apply and the test should show no power.
Extended reading notes
Core claim
The central claim is that independence can be tested by contrasting the joint recurrence rate with the product of the marginal recurrence rates. With $E_n(r,s)=\sqrt{n}(\mathrm{RR}^{X,Y}_n(r,s)-\mathrm{RR}^X_n(r)\mathrm{RR}^Y_n(s))$, the proposed statistic is $T_n = n\int_0^\infty\int_0^\infty E_n(r,s)^2\,dG(r,s)$. The paper proves two complementary limit statements. Theorem 3 shows that when X and Y are independent and the distributions of the distances are continuous, the process $\{E_n(r,s)\}$ converges weakly to a centered Gaussian process $\{\mathcal{E}(r,s)\}$, so $T_n$ converges in distribution to $\|\mathcal{E}\|^2_{L^2(dG)}$. Theorem 4 shows that if $d(X_1,X_2)$ and $d(Y_1,Y_2)$ are continuous and not independent, and the weight function has density positive everywhere, then $T_n \to \infty$ in probability; this is the consistency result, with multivariate normal pairs as a corollary. Under contiguous alternatives the limit process becomes $\{\mathcal{E}(r,s)+\delta\mu(r,s)\}$, adding a deterministic drift that gives nontrivial local power. The paper notes that some dependent distributions produce independent pair-distances, and for those the consistency argument does not apply.
Load-bearing premise
All the consistency guarantees rest on the assumption that dependence between X and Y shows up as dependence between the pairwise distances $d(X_1,X_2)$ and $d(Y_1,Y_2)$; the paper itself exhibits dependent variables whose distances are independent, and for those distributions the test has no asymptotic power.
Editorial extensions
If this is right
- If the paper is correct, any dependence between X and Y that makes the pair-distance variables dependent is eventually detected: $T_n$ grows without bound as $n$ increases.
- Under $H_0$ the asymptotic null distribution is available, and for finite samples the paper's permutation scheme estimates p-values by breaking the X-Y pairing, so the test can be calibrated without simulating the limiting Gaussian process.
- Because the statistic uses only distances, the same test applies to random vectors, function-valued data, and discrete or continuous time series; the multivariate normal case is covered by Corollary 1.
- A sup-version $T'_n=\sqrt{n}\sup_{r,s>0}|E_n(r,s)|$ is allowed by the same arguments, giving a user an alternative that avoids choosing a weight function.
- Under contiguous alternatives the limiting power is governed by the drift $\delta\mu(r,s)$; in the simulation study the test shows higher power than several distance-based competitors in most of the settings considered.
Reading between the lines
- The theorem's condition is about the pair-distance variables, not about (X,Y) itself; effectively the test's null domain is 'the pair distances are independent,' so dependence that lives only in angles or absolute locations is invisible to the statistic.
- An adaptive weight function could in principle concentrate mass where $\mu(r,s)$ is large for a suspected alternative; the paper fixes G for the theory, so data-dependent G would need new arguments but is a natural practical extension.
- For time series, the permutation scheme removes the contemporaneous X-Y pairing while preserving each series' internal order, suggesting the test can target cross-dependence even when the marginal series are serially dependent; the paper's asymptotic proofs are for i.i.d. pairs, so the time-series case rests on this permutation calibration plus simulation.
- A quick diagnostic for a new application is to estimate $\mu(r,s)$ before relying on the test: if the estimate is flat at zero while X and Y are believed dependent, the consistency theorem offers no power guarantee and another test should be considered.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a new test of independence between two random elements X and Y taking values in metric spaces. The test statistic T_n is a Cramér-von Mises-type functional of the U-process E_n(r,s) = √n(RR^{X,Y}_n(r,s) − RR^X_n(r) RR^Y_n(s)), where the RR terms are empirical recurrence rates based on indicators 1{d(X_i,X_j)<r} and 1{d(Y_i,Y_j)<s}. By integrating the squared process against a distribution G on (0,∞)^2, the test avoids choosing a threshold radius. The paper claims: under H0, T_n converges weakly to the squared norm of a centered Gaussian process (Theorem 3); under contiguous alternatives the process acquires a deterministic drift (Proposition 1 and the subsequent assertion); and under alternatives for which the pair-distance variables d(X_1,X_2) and d(Y_1,Y_2) are dependent, T_n diverges to infinity in probability (Theorem 4). Simulation studies compare the test with HHG, distance covariance, and HSIC for scalar, vector, and time-series data.
Significance. If the theoretical claims are correct, the test is an attractive addition to the independence-testing literature: it applies to general metric spaces, avoids threshold selection by integrating over all radii, has a computationally explicit statistic (Section 3.4), and the reported power is competitive with or better than several popular tests in many of the studied alternatives. The paper also ships a permutation procedure for calibration, which is practically useful. However, the proof of the null-distribution theorem has a concrete algebraic error (Lemma 2) that as written invalidates the transfer from the approximating U-process to the actual process; the consistency proof omits a required uniform-integration step; and the contiguity result is only sketched. These are load-bearing issues, although all appear repairable without changing the overall approach. The structural limitation identified in Remark 3, that the test only detects dependence between the distance variables, is real and should be prominently acknowledged.
major comments (4)
- [Section 2.1, Lemma 2] Lemma 2's assertion that 0 ≤ H_n(r,s) ≤ 4/√n is false. In the decomposition displayed after equation (23), the coefficient of the I_4 term is 1/N^2 − 1/(N(n−2)(n−3)), which is negative for every n>2; the manuscript's equation (24) treats this coefficient as positive. A concrete configuration with n=4, d(X_1,X_2)<r, d(Y_3,Y_4)<s, and no other recurrences gives RR^X_n=RR^Y_n=1/6, RR^{X,Y}_n=0, hence E_n = −1/18; the corresponding E'_n equals −1/3, so H_n = E'_n − E_n = −5/18 < 0. Thus the stated nonnegativity is false. Since Theorem 3 uses Lemma 2 to replace E_n by E'_n, the proof of Theorem 3 is invalid as written. The likely repair is to replace the nonnegativity claim with an absolute bound |H_n| ≤ C/√n, which follows from the same decomposition because all terms are bounded by constants times n^{−1/2}; but this correction is absent from the manuscript.
- [Section 2.1, Theorem 4 (proof)] The consistency proof uses the inequality T_n ≥ (n/2)∫µ² dG − n∫(RR^{X,Y}_n − RR^X_n RR^Y_n − µ)² dG and then asserts that the second term is negligible as n→∞. No argument is given for this negligibility. The difference process involves products of U-statistics, and while pointwise convergence of RR^{X,Y}_n, RR^X_n, and RR^Y_n suggests the integrated squared difference should be o_P(1), proving this requires either a Donsker-type tightness result in L²(G) or a dominated-convergence argument for n∫(·)²dG. Without this step, the claim T_n →_P ∞ does not follow from the displayed inequality. The gap is load-bearing for the consistency result, though it appears fixable with a standard empirical-process argument.
- [Section 2.2, contiguous alternatives] The weak convergence under H_n of {E_n(r,s)} to {E(r,s) + δµ(r,s)} is asserted after the sentence "With a little more work, using the Le Cam third lemma," but no proof is provided. This is one of the three main theoretical claims of the paper, alongside Theorem 3 and Theorem 4. The assertion is not a trivial corollary: one must verify the contiguity conditions, handle the U-statistic remainder H_n under the triangular array of densities f^{(n)}_{X,Y}, and justify the weak convergence of the process in the space where the functional T_n is continuous. The omission of these details makes the contiguous-alternative result unverifiable as written. At minimum, a full proof or a precise reference with the required verification should be supplied.
- [Section 4.1, Tables 1 and 2] In the "Four independent clouds" row, which is a null case (X and Y are independent), the column labeled N(1,1) reports power 0.512 for n=30 and n=50, while every other column in that row reports 0.046–0.057. At nominal level 0.05, a size of 0.512 is inconsistent with the paper's claim that the test has correct level under H0. Either there is a simulation error, a mislabeling, or a calibration problem specific to this weight function and sample size. The authors should explain or correct this entry; as printed, it undermines confidence in the numerical comparisons.
minor comments (5)
- [Section 2.1, Theorem 3 (statement)] Theorem 3 as stated asserts weak convergence of {E_n(r,s) − E(E_n(r,s))} without assuming H0, but the proof explicitly begins with "If H0 is true". The theorem statement should include the independence assumption, since the claimed limit process and covariance structure are derived under H0.
- [Section 2.1, Remark 1] Remark 1 says the test statistic T_n is the norm of the Gaussian process, but T_n is defined with an un-squared distance from the definition: T_n = n∫(RR^{X,Y}_n − RR^X_n RR^Y_n)² dG, which converges to the squared norm ||E||². The remark should say "squared norm" or adjust the notation consistently.
- [Section 2.1, Lemma 1 (proof)] The intermediate algebra in the proof of Lemma 1 contains several apparent typos, such as "p^{(3)}_X(r∧r′)p_Y(s)" where the context suggests "p^{(3)}_X(r∧r′)p_Y(s)p_Y(s′)" and analogous expressions for p^{(3)}_Y. The final covariance formula (4) is plausible, but the intermediate display is difficult to follow and should be corrected.
- [Section 3.2, permutation p-value] The unbiasedness argument for the permutation estimator is sketchy: the claim that each B_i is Binomial(m,1/n!) assumes that the values T_n under different permutations are distinct with probability one and that the permutation distribution is uniform over all n! permutations, but ties can occur and the double limit in n and m is not addressed. A more careful statement of the permutation test's exactness or asymptotic validity would improve the presentation.
- [General notation] The same letter φ is used both for the standard normal density and for the normal CDF in Section 3.1 (equations involving φ^{-1}(F_X(X))), which is confusing; the text should distinguish the density from the cumulative distribution function.
Circularity Check
Self-contained test derivation; the only self-citation (Kalemkerian 2017) defines a simulation model and is not load-bearing.
full rationale
The central derivation is self-contained. The statistic T_n in (3) is a direct functional of the empirical recurrence rates; no term in it is fitted to the quantity being predicted. Theorem 3 is obtained from a standard U-process weak convergence theorem (Arcones & Gine 1993), with the entropy integral verified for the class of rectangle indicators; Lemma 2 (even if its H_n bound has an algebraic issue) is an approximation device and not an input that already contains the Gaussian limit. Theorem 4 is a consequence of the SLLN for U-statistics: the integrand converges to mu^2, and the assumption that the distance variables are dependent is exactly what guarantees some positive-mu rectangle. This is a characterization of the population analogue, not a renaming of the conclusion: Remark 3 explicitly exhibits dependent (X,Y) for which the distance variables are independent and the test has no power, so the consistency claim is honestly scoped. The contiguous-alternative analysis uses Le Cam's third lemma and Cabana's contiguity criterion, both external. The only self-citation is reference [17] (Kalemkerian 2017), used in Section 4.3 only to define the FOU(2) simulation model; it is not used in any proof or in the definition of the test. The data-dependent choice of G in Section 3.3 is a practical weighting heuristic, not a fitted parameter that is later reported as a prediction. The algebra concern about Lemma 2's H_n bound is a correctness risk in a proof step, not a circularity: a wrong bound does not make a result identical to its assumptions. Overall, no circular step is identifiable.
Assumptions & free parameters
free parameters (2)
- Weight function G (Gaussian mean and variance) =
Data-dependent: mu = E[d(X1,X2)], sigma^2 = V[d(X1,X2)] for g1; similarly for g2
- Distance metric =
Euclidean norm in simulations
assumptions (6)
- standard math Strong law of large numbers for U-statistics (Hoeffding 1961) ensures convergence of recurrence rates to their expectations.
- standard math Arcones-Giné Theorem 4.10 (1993) on weak convergence of U-processes under bracketing entropy conditions.
- domain assumption Continuity of the distribution functions of d(X1,X2) and d(Y1,Y2).
- domain assumption The weight function G has density g that is strictly positive everywhere on (0,infty)^2.
- standard math Le Cam third lemma (Le Cam & Yang 1990; Oosterhoff & Van Zwet 1979) is invoked for the contiguous alternative weak convergence.
- standard math For Corollary 1, the identity COV(Z^2,T^2) = 2 COV(Z,T)^2 for centered normal bivariate vectors.
Cite this review
Pith. "Pith review of An Independence Test Based on Recurrence Rates." pith.science (2026). https://pith.science/paper/XI7IDPXH
@misc{pith2026190803305,
author = {Pith},
title = {Pith review of: An Independence Test Based on Recurrence Rates},
year = {2026},
howpublished = {\url{https://pith.science/paper/XI7IDPXH}},
note = {Machine review of arXiv:1908.03305}
}
abstract
A new test of independence between random elements is presented in this article. The test is based on a functional of the Cram\'{e}r-von Mises type, which is applied to a $U$-process that is defined from the recurrence rates. Theorems of asymptotic distribution under $H_{0},$ and consistency under a wide class of alternatives are obtained. The results under contiguous alternatives are also shown. The test has a very good behaviour under several alternatives, which shows that in many cases there is clearly larger power when compared to other tests that are widely used in literature. In addition, the new test could be used for discrete or continuous time series.
Figures
Reference graph
Works this paper leans on
-
[1]
Arcones, M. A. and Gin´ e, E. (1993). Limit Theorems for U-Processes. The Annals of Probability Vol 21-3, 1494-1542
work page 1993
-
[2]
Arratia, A., Caba˜ na, A. & Caba˜ na, E., (2016). A construction of Continuous time ARMA models by iterations of Ornstein-Uhlenbeck process,SORT Vol 40 (2) 267-302
work page 2016
-
[3]
Bakirov, N. K., Rizzo, M. L., and Sz´ ekely, G. J. (2006). A multivariate non- para- metric test of independence. Journal of Multivariate Analysis, 97(8):1742-1756
work page 2006
-
[4]
Beran, R., Bilodeau, M., and Lafaye de Micheaux, P. (2007). Nonparametric tests of independence between random vectors. Journal of Multivariate Analysis. 98(9):1805–1824
work page 2007
-
[5]
Bilodeau, M. and Lafaye de Micheaux, P. (2005). A multivariate empirical carac- teristic function test of independence with normal marginals. Journal of ultivariate Analysis, 95(2):345–369
work page 2005
-
[6]
Blomqvist, N. (1950). On a measure of dependence between two random variables. The Annals of Mathematical Statistics 593-600. 24
work page 1950
-
[7]
Boglioni, G. (2016). A consistent test of independence between random vectors. https://papyrus.bib.umontreal.ca/xmlui/bitstream/handle/1866/18773/Bo glioni Beaulieu Guillaume 2016 memoire.pdf?sequence=2
work page 2016
-
[8]
Caba˜ na, E. M. (1997). Contiguidad, pruebas de ajuste y Procesos Emp´ ı ricos Trans- formados. D´ ecima escuela venezolana de Matem´ aticas
work page 1997
Show all 27 references
-
[9]
P, Oliffson Kamphorst, S, Ruelle, D
Eckmann, J. P, Oliffson Kamphorst, S, Ruelle, D. (1987). Recurrence plots of dy- namical systems. Europhys. Lett. 4, 973-977
1987
-
[10]
Galton, F. (1888). Co-relations and their measurement, chiefly from anthropometric data. Proceedings of the Royal Society of London. 45(273-279):135–145
-
[11]
and Zinn, J
Gin´ e, E. and Zinn, J. (1986). Lectures on the central limit theorem for empirical processes. Lecture Note in Math. 1221. 50-113. Springer, New York
1986
-
[12]
and Sch¨ olkopf, B
Gretton, A., Bousquet, O., Smola, A. and Sch¨ olkopf, B. (2005). Measuring statisti- cal dependence with hilbert-scmidts norms. International Conference on algorithmic learning theory. 63-77. Springer
2005
-
[13]
Heller, R., Heller, Y., and Gorfine, M. (2012). A consistent multivariate test of association based on ranks of distances. Biometrika. 100(2):503–510
2012
-
[14]
Hoeffding, W. (1961). The strong law of large numbers for U-statistics. http://www.lib.ncsu.edu/resolver/1840.4/2128
1961
-
[15]
Hoefdding, W. (1948). A non-parametric test of independence. The Annals of Math- ematical Statistics. 546-557
1948
-
[16]
and R´ emillard, B
Genest, G. and R´ emillard, B. (2004). Test of independence and randomness based on the empirical copula process. Test, 13(2): 335-369
2004
-
[17]
Kalemkerian, J. (2017). Fractional Iterated Ornstein-Uhlenbeck Processes. arXiv preprint arXiv:1709.07143v1[math.ST]
2017 arXiv
-
[18]
and Holmes, M
Kojadinovic, I. and Holmes, M. (2009). Tests of independence among continuous random vectors based on cram´ er-von mises functional of the empirical copula process. Journal of Multivariate Analysis, 100(6): 1137-1154
2009
-
[19]
Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1/2):81–93
1938
-
[20]
& Yang, G
Le Cam, L. & Yang, G. L. (1990). Asymptotics in Statistics. Some Basic Concepts. Springer, New York
1990
-
[21]
Marwan, N. (2008). A Historical Review of Recurrence Plots. European Physical Journal—Special Topics, 164, 3-12. 25
2008
-
[22]
& Van Zwet, W
Oosterhoof, J. & Van Zwet, W. R. (1979). A note on contiguity and Hellinger distance. Contributions to Statistics. Jaroslav H´ ajek Memorial Volume (J. Jorecov´ a, ed) Reidel Dordrecht. 157-166
1979
-
[23]
Pearson, K. (1898). Mathematical contributions to the theory of evolution on a form of spurious correlation which may arise when indices are used in the measurement of organs. Proceedings of the Royal Society of London. 60(359-367):489–498
-
[24]
Spearman, C. (1904). The proof and measurement of association between two things. The American journal of psychology. 15(1):72–101
1904
-
[25]
J., Rizzo, M
Sz´ ekely, G. J., Rizzo, M. L., Bakirov, N. K., et al. (2007). Measuring and testing dependence by correlation of distances. The Annals of Statistics. 35(6):2769– 2794
2007
-
[26]
J., Rizzo, M
Sz´ ekely, G. J., Rizzo, M. L., et al. (2009). Brownian distance covariance. The annals of applied statistics, 3(4):1236–1265
2009
-
[27]
Wilks, S. (1935). On the independence of k sets of normality distributed statistical variables. Econometrica, Journal of the Econometric Society. 309-326. 26 Figure 2: Parabola, Two parabolas, circle, diamond, wshape and four independent clouds. 27
1935
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.