Pith. sign in

REVIEW 2 major objections 6 minor 1 cited by

Coordinate-wise Elephant Random Walk

T0 review · 2 major / 6 minor · reviewed 2026-07-09 · glm-5.2

Pith's one-line read Memory washes out of elephant walks on the hypercube

desk verdict CERW on the hypercube: clean perturbation argument, universal limiting variance independent of memory parameters read the letter →

arxiv 2607.07022 v1 pith:LHQW6ZW3 submitted 2026-07-08 math.PR

classification math.PR
keywords coordinate-wisememorywalkhypercuberandomboundedcentralcoordinate
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces a coordinate-wise elephant random walk (CERW) on the finite hypercube Q_k = {-1,1}^k. At each step, one coordinate is chosen uniformly and updated by sampling a past value of that coordinate and repeating it with probability p_i or flipping it with probability 1-p_i. The process is non-Markovian because transition probabilities depend on the full empirical history of each coordinate. The central claim is that when every p_i < 1, the per-coordinate memory biases converge to zero almost surely, so the time-dependent transition kernels become asymptotically indistinguishable in total variation from a memoryless refresh kernel that picks a coordinate uniformly and sets it to a fresh random sign. This reduction to a perturbation of a uniformly ergodic Markov chain yields a weak law of large numbers for bounded observables and, via the Doob martingale, a martingale central limit theorem and functional central limit theorem. The limiting variance sigma^2_f = pi(v_f) depends only on the observable f, the refresh kernel, and the uniform measure pi on the hypercube, and is completely independent of the memory parameters p_1, ..., p_k.

What carries the argument

The argument rests on three pillars. First, a stochastic-approximation recursion for the empirical proportion x^(i)_r of +1 values in coordinate i's embedded history: x^(i)_{r+1} - 1/2 = (1 - c_i/(r+2))(x^(i)_r - 1/2) + noise/(r+2), where c_i = 2(1-p_i). When c_i > 0 (i.e., p_i < 1), the Robbins-Siegmund theorem gives almost-sure convergence of the bias to zero. Second, a total-variation perturbation bound (Lemma 4.1): the TV distance between the CERW kernel K_n and the refresh kernel K* is at most (1/2) max_i |bm^(i)_n|. Third, a finite-horizon comparison estimate (Lemma 4.2) proved by induction, controlling the accumulated discrepancy over T steps, combined with the uniform ergodicity of K

What would settle it

If one could exhibit a coordinate i with p_i < 1 for which the empirical bias m^(i)_r does not converge to zero almost surely, or if the total variation bound delta_n = sup_x ||K_n(x,.) - K*(x,.)||_TV did not vanish, then the perturbation argument collapses and the CLT with universal variance would fail. The load-bearing step is the Robbins-Siegmund convergence in Proposition 3.3.

Watch

Extended reading notes

Core claim

The core mechanism is the almost-sure vanishing of the coordinate-wise empirical memory bias m^(i)_r, proved via a stochastic approximation recursion (Lemma 3.2) and the Robbins-Siegmund almost-supermartingale theorem. When p_i < 1, the contraction factor c_i = 2(1-p_i) > 0 forces the proportion of +1 values in each coordinate's history toward 1/2, so the bias m^(i)_r = 2x^(i)_r - 1 converges to zero. This makes the random kernel K_n (which uses the current bias to set the refresh probability) converge in total variation to the memoryless kernel K*. A key technical step is Lemma 4.2, which uses induction over time horizons to control the cumulative path-dependent discrepancy between the CERW

Load-bearing premise

The entire argument requires p_i < 1 for every coordinate. When p_i = 1, the contraction constant c_i = 2(1-p_i) vanishes, the Robbins-Siegmund supermartingale argument fails, the coordinate's empirical bias need not converge to zero, and the kernel K_n need not approach the refresh kernel K*. The paper does not analyze what happens at or beyond this boundary.

Editorial extensions

If this is right

  • On a finite state space, long-range memory that operates per-coordinate washes out in the large-time limit whenever each coordinate's memory parameter is sub-critical (p_i < 1), so the process inherits the stationary behavior of a simple memoryless chain.
  • The phase transition at p_i = 3/4, known from the classical 1D elephant random walk, appears here only in the rate of bias decay (Proposition 3.5) but does not affect the final CLT variance, suggesting that on finite spaces the trichotomy structure collapses.
  • The universality of the limiting variance means that for any choice of memory parameters p_1,...,p_k in [0,1), the fluctuation scale of observables like height, parity, or Hamming distance is identical and computable in closed form from the hypercube geometry alone.
  • The perturbation-plus-ergodicity template used here could apply to other non-Markovian processes on finite state spaces where memory biases can be shown to vanish, reducing the analysis to a comparison with a known ergodic Markov chain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The boundary case p_i = 1 for some coordinate i is not studied but is the natural next question: if one coordinate has perfect memory (always repeats its past), its bias does not vanish, the kernel does not converge to K*, and the universality result should break. The limiting behavior in this mixed regime may involve a non-trivial dependence on p_i = 1 coordinates.
  • The rate of bias decay (Proposition 3.5) transitions at p_i = 3/4, mirroring the classical ERW trichotomy, but since the bias vanishes for all p_i < 1, this rate only controls finite-time corrections and not the asymptotic variance. One could ask whether the convergence rate to the CLT limit is noticeably slower near p_i = 3/4, even though the limit itself is unchanged.
  • If the state space were extended to an infinite graph (e.g., Z^k instead of Q_k), the refresh chain would not be uniformly ergodic and the perturbation argument would not directly apply, suggesting that the finiteness of the hypercube is essential to the universality phenomenon.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper introduces the Coordinate-wise Elephant Random Walk (CERW) on the $k$-dimensional hypercube $Q_k = {-1,1}^k$. At each step, a coordinate is selected uniformly and updated via an elephant-type memory rule using only that coordinate's past values. The process is non-Markovian on $Q_k$ due to path-dependent transition probabilities. The author shows that when all memory parameters satisfy $p_i < 1$, the coordinate-wise empirical biases vanish almost surely (Proposition 3.3, via Robbins-Siegmund). This yields total-variation closeness of the random transition kernels to a memoryless refresh kernel $K^*$ (Lemma 4.1). Using a finite-horizon comparison estimate (Lemma 4.2) and uniform ergodicity of $K^*$, a weak law of large numbers for bounded observables is established (Theorem 5.3). The Doob martingale associated with a bounded observable is then analyzed, yielding a martingale CLT (Theorem 6.5) and functional CLT (Theorem 6.9). The limiting variance $sigma_f^2 = pi(v_f)$ depends only on the refresh kernel and uniform measure, and is independent of the memory parameters. Explicit variance formulas are computed for eight natural observables in Section 7.

Significance. The paper makes a solid contribution by introducing a genuinely new variant of the elephant random walk on a finite state space with anisotropic, coordinate-wise memory. The key technical achievement is the perturbation argument reducing the non-Markovian dynamics to the memoryless refresh chain, culminating in the universality of the limiting variance $sigma_f^2 = pi(v_f)$. The explicit computation of variances for eight observables (Table 1) provides concrete, falsifiable predictions that illustrate the abstract limit theorems. The proofs are self-contained and rely on standard, external tools (Robbins-Siegmund, Hall-Heyde martingale CLT, Whitt's FCLT, Markov chain minorization), which is a strength. The identification of the $p_i = 3/4$ threshold in the rate of bias decay (Proposition 3.5), mirroring the classical ERW trichotomy, adds further interest, though its role in the global limit theorems is limited to rate considerations.

major comments (2)
  1. Section 3.3, Proposition 3.5: The phase transition at $p_i = 3/4$ in the rate of decay of $E[(m_r^{(i)})^2]$ is noted, but its implications for the global limit theorems are not discussed. Since the main results (Theorems 5.3, 6.5, 6.9) hold for all $p_i < 1$ regardless of whether $p_i$ is above or below $3/4$, it would strengthen the paper to briefly clarify that this phase transition affects only the rate of convergence of the biases, not the validity of the asymptotic limit theorems themselves. This is not a load-bearing issue for correctness, but the current presentation may leave readers wondering about the relationship between the $3/4$ threshold and the global results.
  2. The boundary case $p_i = 1$ is excluded from all main results (Corollary 3.4, Theorems 5.3, 6.5, 6.9). The paper notes that $c_i = 2(1-p_i) = 0$ causes the Robbins-Siegmund argument to fail. A brief remark on what happens when $p_i = 1$ for some coordinate (e.g., the bias does not vanish, the kernel does not converge to $K^*$, and the limit theorems break down) would clarify the sharpness of the assumption and the boundary of the theory. This is a natural question that readers familiar with the ERW literature will ask.
minor comments (6)
  1. Section 2, definition of $H_n^{(i)}$: The history is defined as a finite sequence including $X_0^{(i)}$. It would help to explicitly state that the sampling in step (2) of the update rule is uniform over the $N_n^{(i)}+1$ elements of $H_n^{(i)}$, to avoid ambiguity about whether the initial value is included in the sampling.
  2. Section 4, Lemma 4.1: The bound $delta_n leq (1/2) max_i |bm_n^{(i)}|$ uses $|alpha_i| leq 1$. It would be clearer to state this explicitly in the proof, since $alpha_i = 2p_i - 1 in [-1,1]$.
  3. Section 7.5, Observable 5 (Pair correlation): The computation of $K^* H^4(x)$ involves expanding $(H(x) - x^{(i)} + xi)^4$ and averaging over $xi$. The intermediate steps are omitted; including a brief derivation or stating the key identity $E[xi^4] = 1$, $E[xi^2] = 1$ would aid verification.
  4. Table 1, row 6 (Occupation level): The formula reference '(7.2)' is correct, but the table entry 'See formula (7.2)' could be made more self-contained by at least indicating the dependence on $r$ and $k$.
  5. Section 6.3, Theorem 6.9: The statement writes $M^{(f,n)} Rightarrow sigma_f B$ but does not explicitly state the space of convergence. The proof mentions $D([0,infty), mathbb{R})$, which should be stated in the theorem itself for completeness.
  6. Typographical: In Section 7.6, the set $L_r$ is defined as ${x in Q_k : H(x) = k - 2r}$, which represents vertices with exactly $r$ negative coordinates. This is correct but could be stated more directly as ${x: R(x) = r}$ for consistency with the subsequent notation $R(x)$.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: fully self-contained mathematical derivation

full rationale

This is a pure mathematics paper with a self-contained derivation chain. The main results (WLLN, martingale CLT, functional CLT, universality of limiting variance) are derived from standard external theorems: Robbins-Siegmund almost supermartingale convergence (Prop 3.3), Borel-Cantelli (Cor 3.4), standard Markov chain minorization (Lemma 4.3), Hall-Heyde martingale CLT (Theorem 6.5), and Whitt's FCLT (Theorem 6.9). No parameters are fitted to data, no predictions reduce to fitted values, and no self-citation is load-bearing. The universality claim (σ²_f = π(v_f) independent of p_i) follows from the structural fact that v_f is defined purely in terms of the refresh kernel K* and uniform measure π (Eq. 6.2), combined with the perturbation argument showing the memory-dependent kernels K_n converge to K* in total variation. The variance formula contains no p_i by construction of the definition, but this is a genuine mathematical consequence of the bias vanishing (Prop 3.3 → Cor 3.4 → Lemma 4.1 → Lemma 6.2 → Prop 6.3), not a circular restatement. The p_i < 1 assumption is a stated boundary condition, not an internal inconsistency. The derivation chain is non-circular throughout.

Assumptions & free parameters 1 free parameters · 5 assumptions · 2 invented entities

No physical entities are invented. The CERW and K* are mathematical constructions with well-defined properties.

free parameters (1)
  • p_1, ..., p_k = not fitted; model parameters in [0,1]
    These are model parameters defining the repeat/reverse probability for each coordinate's memory. They are not fitted to data or chosen ad hoc to make the derivation work; they are part of the model definition. The main results hold for all p_i < 1.
assumptions (5)
  • standard math Robbins-Siegmund almost supermartingale convergence theorem
    Invoked in the proof of Proposition 3.3 to conclude u_r^(i) → 0 a.s. from the recursion (3.7). This is a standard result from [10].
  • standard math Hall-Heyde martingale central limit theorem (Corollary 3.1 of [8])
    Invoked in Theorem 6.5 to establish the CLT from the conditional variance and Lindeberg conditions.
  • standard math Whitt's functional CLT (Theorem 2.1 of [15])
    Invoked in Theorem 6.9 to establish the FCLT from jump negligibility and quadratic variation convergence.
  • standard math Standard Markov chain minorization / Doeblin condition
    Invoked in Lemma 4.3 to obtain the geometric ergodicity bound (4.8) for the refresh kernel K*.
  • standard math Second Borel-Cantelli lemma
    Invoked in Corollary 3.4 to show each coordinate is selected infinitely often almost surely.
invented entities (2)
  • Coordinate-wise Elephant Random Walk (CERW) independent evidence
    purpose: The stochastic process studied throughout the paper
    The CERW is a well-defined stochastic process. Its limiting behavior is derived from first principles and verified through explicit variance calculations for eight observables. No new physical entity is postulated; this is a mathematical construction.
  • Refresh kernel K* independent evidence
    purpose: Memoryless benchmark chain used for comparison
    K* is a standard random-scan Gibbs-type kernel on the hypercube. It is a mathematical tool, not a postulated physical entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Coordinate-wise Elephant Random Walk." pith.science (2026). https://pith.science/paper/LHQW6ZW3

@misc{pith2026260707022,
  author       = {Pith},
  title        = {Pith review of: Coordinate-wise Elephant Random Walk},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LHQW6ZW3}},
  note         = {Machine review of arXiv:2607.07022}
}
abstract

We introduce a coordinate-wise version of the elephant random walk on the $k$-dimensional discrete hypercube $Q_k=\{-1,1\}^k$. At each global time step, one coordinate is selected uniformly at random and updated according to an elephant-type memory rule using only the past values of that coordinate. The resulting process is a nearest-neighbor walk with possible holding on the hypercube, but it is not Markovian on $Q_k$ because the transition probabilities depend on coordinate-wise empirical histories. We show that, when all memory parameters satisfy $p_i<1$, the coordinate-wise memory biases vanish almost surely. Consequently, the time-dependent transition kernels of the walk are asymptotically close in total variation to the memoryless coordinate-refresh kernel. Using this perturbation argument and the uniform ergodicity of the refresh chain, we prove a weak law of large numbers for bounded observables. We then study the Doob martingale associated with a bounded observable and prove a martingale central limit theorem and functional central limit theorem. The limiting variance is determined by the refresh kernel and the uniform measure on the hypercube, and is completely independent of the memory parameters $p_1,p_2,\dots,p_k.$

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A model of opinion dynamics evolving via a preferential attachment mechanism involving multiple extractions

    math.PR 2026-08 conditional novelty 6.0 of 10

    For a two-opinion preferential-attachment network with multiple sampling and general reinforcement, the normalized opinion count, influence capital and activity converge almost surely to invariant sets of a mean-field...

Reference graph

Works this paper leans on

15 extracted references · 15 canonical work pages · cited by 1 Pith paper

  1. [1]

    Elephant random walks and their connection to Pólya- type urns

    Erich Baur and Jean Bertoin. “Elephant random walks and their connection to Pólya- type urns”. In:Physical Review E94.5 (2016), p. 052134.doi:10.1103/PhysRevE.94. 052134.url:https://link.aps.org/doi/10.1103/PhysRevE.94.052134

  2. [2]

    Harmonic rigidity at fixed spectral gap in one dimension

    Bernard Bercu. “A martingale approach for the elephant random walk”. In:Journal of Physics A: Mathematical and Theoretical51.1 (2018), p. 015201.doi:10.1088/1751- 8121/aa95a6.url:https://doi.org/10.1088/1751-8121/aa95a6

  3. [3]

    On the multi-dimensional elephant random walk

    Bernard Bercu and Lucile Laulin. “On the multi-dimensional elephant random walk”. In:Journal of Statistical Physics175.6 (2019), pp. 1146–1163.doi:10.1007/s10955- 019-02282-8.url:https://doi.org/10.1007/s10955-019-02282-8

  4. [4]

    Functional limit theorems for the multi-dimensional elephant ran- dom walk

    Marco Bertenghi. “Functional limit theorems for the multi-dimensional elephant ran- dom walk”. In:Stochastic Models38.1 (2022), pp. 37–50.doi:10.1080/15326349. 2021.1971092.url:https://doi.org/10.1080/15326349.2021.1971092

  5. [5]

    Elephant random walk on triangular lattice

    Rohit Chaudhuri. “Elephant random walk on triangular lattice”. In:arXiv preprint (2026). arXiv:2603.14402

  6. [6]

    Central limit theorem and related results for the elephant random walk

    Cristian F. Coletti, Renato Gava, and Gunter M. Schütz. “Central limit theorem and related results for the elephant random walk”. In:Journal of Mathematical Physics58.5 (2017), p. 053303.doi:10.1063/1.4983566.url:https://doi.org/10.1063/1. 4983566

  7. [7]

    On multidimensional elephant random walk with stops and random step sizes

    Shyan Ghosh, Manisha Dhillon, and Kuldeep Kumar Kataria. “On multidimensional elephant random walk with stops and random step sizes”. In:arXiv preprint(2026). arXiv:2601.07502

  8. [8]

    Heyde.Martingale Limit Theory and Its Application

    Peter Hall and Christopher C. Heyde.Martingale Limit Theory and Its Application. Academic Press, 1980

Show all 15 references
  1. [9]

    Elephants Explore in Spirals Sometimes

    Lucile Laulin and Bastien Mallein. “Elephants Explore in Spirals Sometimes”. In: Stochastics and Quality Control().doi:doi:10.1515/eqc-2026-0017.url:https: //doi.org/10.1515/eqc-2026-0017

  2. [10]

    A convergence theorem for non negative al- most supermartingales and some applications

    Herbert Robbins and David Siegmund. “A convergence theorem for non negative al- most supermartingales and some applications”. In:Optimizing Methods in Statistics. Academic Press, 1971, pp. 233–257

  3. [11]

    Minorization conditions and convergence rates for Markov chain MonteCarlo

    Jeffrey S. Rosenthal. “Minorization conditions and convergence rates for Markov chain MonteCarlo”.In:Journal of the American Statistical Association90.430(1995),pp.558– 566.doi:10.2307/2291067

  4. [12]

    Elephants can always remember: Exact long- range memory effects in a non-Markovian random walk

    Gunter M. Schütz and Steffen Trimper. “Elephants can always remember: Exact long- range memory effects in a non-Markovian random walk”. In:Physical Review E70.4 (2004), p. 045101.doi:10.1103/PhysRevE.70.045101.url:https://link.aps. org/doi/10.1103/PhysRevE.70.045101

  5. [13]

    Functional limit theorems for elephant random walks on general pe- riodic structures

    Shuhei Shibata. “Functional limit theorems for elephant random walks on general pe- riodic structures”. In:arXiv preprint(2025). arXiv:2511.10347

  6. [14]

    Stable functional CLTs for scaled elephant random walks

    Go Tokumitsu. “Stable functional CLTs for scaled elephant random walks”. In:arXiv preprint(2026). arXiv:2603.13690

  7. [15]

    ProofsofthemartingaleFCLT

    WardWhitt.“ProofsofthemartingaleFCLT”.In:Probability Surveys4(2007),pp.268– 302.doi:10.1214/07-PS122.url:https://doi.org/10.1214/07-PS122

Pith tools

Reviewed July 9, 2026 · model on record in the stance chip above.