Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

A minimal agent that trains on its own predicted labels, with no ground truth or reward, can spontaneously learn the true latent structure of its environment, and this capacity appears as a sharp phase transition.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A self-labeling linear classifier can bootstrap alignment with the latent structure of Gaussian data when regularization is strong enough, and label-exchange between agents can produce collective consensus.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection A clean analytically tractable phase transition for label-only self-consistent learning, held together by a fresh-data idealization that the paper states plainly but does not quantify. the 3 major comments →

arxiv 2509.23212 v2 pith:IB5D3HQ7 submitted 2025-09-27 cond-mat.dis-nn cond-mat.stat-mechnlin.AO

Replication and Information Extraction in a Minimal Agent-Environment Model

classification cond-mat.dis-nn cond-mat.stat-mechnlin.AO MSC 68T0562H3082B26 PACS 05.20.-y64.60.-i07.05.Mh
keywords functional replicatorsself-consistency learningunsupervised learningphase transitionreplica methodGaussian mixturecollective learninglabel exchange
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a classifying agent can extract genuine latent structure from data without any ground-truth labels, rewards, or teaching signal, purely by repeatedly labeling fresh data and then training on its own labels under a simplicity bias. The authors call the resulting self-sustaining labeling regimes 'functional replicators' and show, with replica-method statistical mechanics, that their onset is a phase transition: when the regularization is strong enough, the zero-magnetization fixed point becomes unstable and the agent aligns with the true data centroids. In a Gaussian-mixture environment a single agent reaches the same mutual information as a supervised classifier in most noise regimes, sometimes better. A population of such agents exchanging labels can reach consensus, and interaction can either enable or suppress spontaneous learning depending on a stubbornness parameter. If correct, this provides a minimal principled setting where learning emerges from self-consistency alone, without selection or external objectives.

Core claim

On the paper's own terms: a linear classifier updated to be self-consistent with its own hard labels on a fresh batch of Gaussian-mixture data, plus an L2 regularization, has its whole high-dimensional dynamics reduced to a scalar recursion for the magnetization mt = (Wt·μ)/||Wt||. The paper shows that the trivial fixed point m*=0 loses stability when the derivative r = lim_{mt→0} f'(mt) crosses 1, which happens as the regularization λ grows relative to the noise σ; the stable fixed point then has m*>0, meaning the self-made labels are correlated with the environment's true centroids. Long-time correlations between agent generations equal (m*)², so the only persistent memory is the signal pr

What carries the argument

The central object is the magnetization mt = Wt·μ/||Wt||, the agent's alignment with the environment's centroid. The argument is carried by the scalar recursion mt+1 = f(mt; α, σ, λ), obtained via a replica calculation (replica-symmetric ansatz, β→∞) of the convex self-consistency loss with fresh data at each step; and by the stability condition r = lim_{mt→0} f'(mt), whose crossing of 1 marks the learning phase transition. The regularization λ acts as a simplicity bias that concentrates weights along informative directions, and the fresh-data assumption removes inter-step cross correlations, making the one-step replica computation exact in the thermodynamic limit. The population extension u

Load-bearing premise

The load-bearing premise is that every update sees a fresh, independent batch of data—so that inter-step correlations vanish and the replica computation of the one-step map is exact in the thermodynamic limit; if data are finite or reused, the transition boundary changes or disappears, as the paper's own uniform-splitting example shows.

What would settle it

Initialize agents at a small nonzero magnetization m0 in a parameter region where the replica calculation predicts r<1 (low λ, high σ) and confirm the magnetization relaxes to zero; and in a region where r>1 (high λ, low σ) confirm a small seed grows to m*>0. The replica prediction gives exact r in the thermodynamic limit, so a discrepancy between the measured slope of mt+1 vs mt at mt=0 and the replica r would settle the matter directly.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If correct, self-supervised bootstrap is a genuine phase transition: there is a sharp boundary in (σ, λ) separating agents that never align from agents that reach a stable informative labeling; small initial fluctuations get amplified.
  • The steady-state performance of the unsupervised turnover is comparable to supervised training (sometimes superior at intermediate noise and strong regularization), so self-consistency alone can be as good as ground truth in simple environments.
  • Functional replicators are the only long-time attractors in the binary case: the only persistent memory channel is the signal projection, since orthogonal components decorrelate exponentially.
  • In populations, label exchange alone (without access to other agents' weights) can drive consensus; interaction can enable learning in regimes where isolated agents would fail, but can also suppress it for high stubbornness.
  • Multiclass environments support information-bearing replicators, but full latent-structure recovery becomes harder as C grows; output overparameterization (K > C) increases the chance of full coverage.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Testable extension: for a fixed finite dataset, the phase boundary should renormalize as a function of the resampling rate, shrinking toward zero as the rate goes to zero; the paper's own 0% resampling result suggests this, and a quantitative finite-budget replica theory with temporal correlations would settle it.
  • An inference from the η=1/2 crossover: decentralized learning systems may have an optimal peer-distrust level where consensus is fastest even though individual learning is slowest—a possible design principle for swarm or federated learning that the paper does not state.
  • Because the only persistent channel is the signal projection, the model predicts that any self-consistency algorithm with strong regularization will discard all spurious directions; if a real self-distillation method shows persistent non-signal correlations, the model's minimal description is incomplete.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies a minimal unsupervised learning loop in which a linear classifier labels fresh samples from a symmetric two-component Gaussian mixture and then updates its weights by minimizing cross-entropy on its own labels with L2 regularization. In the thermodynamic limit and under a replica symmetric ansatz, the dynamics is reduced to a scalar recursion for the magnetization m_t, Eq. (5). The stability of the trivial fixed point, Eq. (6), yields a phase boundary r=1 in the (σ, λ) plane: above a critical regularization strength the agent spontaneously develops nonzero alignment with the latent centroid, despite never seeing ground-truth labels. The authors validate the boundary against simulations without fitted parameters, compare performance with supervised baselines, extend the phenomenology numerically to C>2 Gaussian mixture classes, and study a population of agents exchanging labels, finding consensus regimes with distinct O(M^2) and O(M) timescales.

Significance. If the central claim holds, the paper provides a clean statistical-mechanics example of a genuine learning phase transition in a self-supervised, teacher-free setting: self-consistency plus a simplicity bias is sufficient to bootstrap latent-structure extraction. The strength of the work is that the binary case is derived from the model definition via a replica calculation, with no fitted parameters in the phase boundary, and the theoretical transition line is checked against direct simulations. The data-constraint experiments in the SM, while highlighting an important limitation, are a useful and honest addition. The population part is more exploratory but connects the model to collective-learning and consensus phenomenology.

major comments (3)
  1. [Theoretical analysis, Eq. (5)-(6), and SM I.D] The scalar recursion and the stability line r=1 rely on drawing a fresh, independent batch at every timestep; the SM states this explicitly ('Since we assume fresh data at each timestep, we conveniently do not have to deal with inter-step cross correlations'). This is not a harmless technical convenience: SM Fig. S7 shows that under uniform splitting (0% resampling) the dynamics typically fails to reach the structured state that the fresh-stream theory predicts for the same parameters, and alternating/partial-resampling protocols change the trajectories. The phase transition is therefore established only for the infinite fresh-stream protocol. The paper should either quantify a finite-data/reuse regime in which the transition survives, or explicitly restrict the main claims to this idealization. As written, the abstract's 'without explicit guidance' and the general framing overstate the
  2. [Theoretical analysis, replica computation (SM II.B.1)] The phase boundary is derived under the Replica Symmetric (RS) ansatz. The paper justifies RS by 'the convex nature of the optimization problem', but the replicated free entropy involves the sign function and the iterated dynamics is not convex; RS is an assumption, not a consequence of per-step convexity. The excellent numerical agreement for m* and for the boundary is encouraging, but it does not by itself rule out replica-symmetry-breaking corrections to the free energy or to the stability condition. The authors should state more carefully that the boundary is obtained under RS and, if possible, provide a stability check of the RS saddle point or a direct test of RS order parameters.
  3. [Collective learning in interacting agents and Abstract] The abstract claims that 'interaction reshapes the learning phase boundary', but the population section does not compute a learning phase boundary. It reports the evolution of π_t and φ_t, the η=1/2 crossover, and O(M^2) vs O(M) timescales, but no boundary in (σ, λ, η) is derived or mapped. The claim that interaction reshapes the phase boundary is therefore unsupported by the presented evidence. Either compute the population-level boundary (even approximately) or soften the claim to describe the observed consensus/coordination crossover.
minor comments (4)
  1. [SM I.A, Fig. S1 caption] The caption says 'for increasing noise levels α = 0.15, 0.45, 0.75'; α is the load P/N, not a noise level. The variables should be labeled consistently, and the color code (λ? training time?) is unclear.
  2. [Main text, Fig. 2(c) and footnote [38]] The average NMI is not self-averaging: footnote [38] states that low averages can be dominated by rare lucky initializations, and error bars are omitted from Fig. 2(c). Since this is the only quantitative evidence for the multiclass claim, the distribution or best-case curves should be shown alongside the average.
  3. [SM I.D, Fig. S7] The figure compares protocols but the marker for the replica solution is described only as starting 'from the experimental mean alignment at t=0'. Clarify whether the replica curve is the deterministic recursion or an average, and report variability across the 50 runs.
  4. [Eq. (6)] The notation r = lim_{m_t -> 0} f'(m_t; α, σ, λ) is fine, but f itself is not given in closed form and is only defined implicitly through the replica saddle-point calculation. A sentence pointing to the explicit saddle-point equations in the SM would help reproducibility.

Circularity Check

0 steps flagged

Derivation is self-contained; only minor methodological self-citation, no circular reduction.

full rationale

The central claim—that iterated self-consistency with regularization yields a spontaneous alignment transition—is derived by an explicit replica computation (SM II.B) from the model definition (Eqs. (1)-(3)), leading to the scalar recursion m_{t+1}=f(m_t; alpha, sigma, lambda) (Eq. (5)) and the stability condition r=1 (Eq. (6)). These results are not fitted: the analytical phase boundary is compared to direct numerical simulations (Fig. 2b) and to supervised baselines (Fig. 2c, Fig. S3), and the order-parameter map is tested against transient dynamics (Fig. S1). The only stated special assumption is the fresh-data-per-timestep idealization; the paper explicitly says 'Since we assume fresh data at each timestep, we conveniently do not have to deal with inter-step cross correlations in the dataset'. The SM's data-constraint experiments (Fig. S7) show that finite or reused data can change the outcome, but that is a domain-of-validity limitation, not a circular reduction of the predicted transition to its own inputs. The main self-citation is methodological: the SM states 'The analytical computation is adapted from the procedure outlined in [1]', where [1] is a coauthor paper (Mannelli et al.), but the full replica calculation is reproduced in the SM and the central claim does not rest on an unverified uniqueness or existence theorem imported from that citation. Reference [51] is a contextual citation for Schelling-type dynamics and is not load-bearing. No step of the derivation reduces by construction to a fitted parameter, to a renaming, or to an equation supplied solely by the authors.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

No fitted free parameters are involved; alpha, sigma, lambda, C, eta, M are tunable inputs. The paper relies on six stated assumptions, chiefly fresh independent batches per step, exact per-step optimization, and the Replica Symmetric ansatz. No new physical entity or hidden mechanism is introduced; 'functional replicator' is a descriptive label for a fixed point, not a postulated entity.

axioms (6)
  • domain assumption Gaussian mixture environment with two symmetric centroids of unit norm and isotropic noise.
    Eq. (1) defines the environment; all exact single-agent results are for this distribution.
  • domain assumption Fresh independent data batch at every timestep, so no inter-step dataset correlations.
    Main text: 'Since we assume fresh data at each timestep...'; this is what makes the replica reduction a static problem.
  • domain assumption Each per-step optimization reaches its global minimizer (Bayesian optimum).
    SM I.A requires 'that each turnover update reaches its Bayesian optimum' for the scalar recursion to be exact.
  • domain assumption Replica Symmetric ansatz is exact for the saddle point.
    Main text and SM Eq. (S32)-(S35): RS is 'justified here by the convex nature of the optimization problem', but no rigorous proof is given.
  • standard math Weight norm is irrelevant because labels are hard labels, allowing normalization and spherical perceptron treatment.
    Main text: 'Since the agent works with hard labels, its norm is irrelevant'; used throughout the replica computation.
  • domain assumption Thermodynamic limit with N -> infinity and alpha = P/N fixed.
    Main text 'Theoretical analysis': exact statements are asymptotic in this proportional regime.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Replication and Information Extraction in a Minimal Agent-Environment Model." pith.science (2026). https://pith.science/paper/IB5D3HQ7

@misc{pith2026250923212,
  author       = {Pith},
  title        = {Pith review of: Replication and Information Extraction in a Minimal Agent-Environment Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IB5D3HQ7}},
  note         = {Machine review of arXiv:2509.23212}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

How can information be extracted from data without explicit guidance or rewards? We investigate this question in a minimal setting where a classifying agent is exposed to a stream of structured data produced by a generative environment, and evolves by seeking consistency of its own labels in time. We find that imposing a simplicity bias on the classification rule can drive the dynamics toward label-coherent steady states, which we coin functional replicators. Remarkably, these persistent labeling rules align with the latent structure of the data. Using analytical tools from statistical mechanics, we characterize this spontaneous learning phase transition. Extending the analysis to a population of agents that pool labels from one another, we show that interaction reshapes the learning phase boundary, in some regimes enabling spontaneous learning that no isolated agent can achieve, while suppressing it in others. Our minimal framework thus opens a route to decentralized learning through label exchange alone, requiring no access to the internal weights of other agents.

Figures

Figures reproduced from arXiv: 2509.23212 by Davide Straziota, Jerome Garnier-Brun, Luca Saglietti, Sebastiano Ariosto.

Figure 1
Figure 1. Figure 1: FIG. 1. Learning loop: the agent (linear in this example) assigns [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. (a) Evolution of the agent’s alignment with the environment signal [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Evolution of the average absolute value of the alignment [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Theory of collective learning in populations of adaptive agents

    cond-mat.stat-mech 2026-07 unverdicted novelty 5.0

    An effective reward function emerges that fully governs the evolution of policy distributions across the population, yielding closed equations for mean and variance under Gaussian assumptions.

Reference graph

Works this paper leans on

73 extracted references · 6 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Random Init ( t = 0 only) Environment E Agent W0

  2. [2]

    Sample Mini Batch Environment E Agent W0

  3. [3]

    Self Labelling Environment E Agent W0

  4. [4]

    Update Weights Environment E Agent W1 FIG. 1. Learning loop: the agent (linear in this example) assigns labels to fresh data and uses those predictions as training targets. portantly, such a parsimony constraint prevents trivial self- agreement and promotes larger classification margins. The resulting dynamics, driven by the stochasticity of fresh samples,...

  5. [5]

    0 0. 2 0. 4 0. 6 0. 8 mt

  6. [6]

    8 mt+1 (a) 10− 4 10− 2 100 102 λ 10− 4 10− 2 100 102 λ

  7. [7]

    0 NMI (c) C = 2 C = 3 C = 4 C = 5 C = 6 FIG. 2. (a) Evolution of the agent’s alignment with the environment signal mt ↦→mt+1 yielded by the replica analysis (SM). Star symbols indicate the fixed point solution the iterative map will reach for sufficiently large t. (b) Fixed point alignment reached by numerical simulations run with α = 1 , N = 10 3, and aver...

  8. [8]

    0 πt η = 0. 0 η = 0. 25 η = 0. 5 η = 0. 75 η = 1. 0 0 10 20 t/M

  9. [9]

    Easy” learning problem, σ = 0.25. Right: “Difficult

    0 ϕ t 0 10 20 t/M FIG. 3. Evolution of the average absolute value of the alignment to the signal πt describing individual performance (top) and of the absolute value of the average alignment φt describing consensus within the population (bottom) in a fully connected population of size M = 128 , N = 500 , α = 1 , λ = 10 , averaged over 32 trials, faint lin...

  10. [10]

    Kushner, Self-assembly of biological structures, Bacterio- logical reviews 33, 302 (1969)

    D. Kushner, Self-assembly of biological structures, Bacterio- logical reviews 33, 302 (1969)

  11. [11]

    Stradner, H

    A. Stradner, H. Sedgwick, F. Cardinaux, W. C. Poon, S. U. Egelhaaf, and P . Schurtenberger, Equilibrium cluster forma- tion in concentrated protein solutions and colloids, Nature 432, 492 (2004)

  12. [12]

    P . G. Higgs and N. Lehman, The RNA world: molecular co- operation at the origins of life, Nature Reviews Genetics 16, 7 (2015)

  13. [13]

    Ball, The self-made tapestry: pattern formation in nature (Oxford University Press, Inc., 1999)

    P . Ball, The self-made tapestry: pattern formation in nature (Oxford University Press, Inc., 1999)

  14. [14]

    G. M. Whitesides and B. Grzybowski, Self-assembly at all scales, Science 295, 2418 (2002)

  15. [15]

    Vicsek, A

    T. Vicsek, A. Czir ´ok, E. Ben-Jacob, I. Cohen, and O. Shochet, Novel type of phase transition in a system of self-driven parti- cles, Physical review letters 75, 1226 (1995)

  16. [16]

    Pross and V

    A. Pross and V . Khodorkovsky, Extending the concept of ki- netic stability: toward a paradigm for life, Journal of physical organic chemistry 17, 312 (2004)

  17. [17]

    Wagner and A

    N. Wagner and A. Pross, The nature of stability in replicating systems, Entropy 13, 518 (2011)

  18. [18]

    Pascal and A

    R. Pascal and A. Pross, Stability and its manifestation in the chemical and biological worlds, Chemical Communications 51, 16160 (2015)

  19. [19]

    Gardner, Mathematical games, Scientific american 222, 132 (1970)

    M. Gardner, Mathematical games, Scientific american 222, 132 (1970)

  20. [20]

    C. G. Langton, Studying artificial life with cellular automata, Physica D: nonlinear phenomena 22, 120 (1986)

  21. [21]

    the arrival of the fittest

    W. Fontana and L. W. Buss, “the arrival of the fittest”: Toward a theory of biological organization, Bulletin of mathematical biology 56, 1 (1994)

  22. [22]

    J. H. Holland, Hidden order, Business Week-Domestic Edition 21, 93 (1995)

  23. [23]

    B. W.-C. Chan, Lenia-biology of artificial life, arXiv preprint arXiv:1812.05433 (2018)

  24. [24]

    Ackley and M

    D. Ackley and M. Littman, Interactions between learning and evolution, Artificial life II 10, 487 (1991)

  25. [25]

    Littman, Simulations combining evolution and learning, Adaptive Individuals in Evolving Populations: Models and Algorithms 26, 465 (1996)

    M. Littman, Simulations combining evolution and learning, Adaptive Individuals in Evolving Populations: Models and Algorithms 26, 465 (1996)

  26. [26]

    M. L. Wong, C. E. Cleland, D. Arend Jr, S. Bartlett, H. J. Cleaves, H. Demarest, A. Prabhu, J. I. Lunine, and R. M. Hazen, On the roles of function and selection in evolving sys- tems, Proceedings of the National Academy of Sciences 120, e2310223120 (2023)

  27. [27]

    Ag ¨uera y Arcas, J

    B. Ag ¨uera y Arcas, J. Alakuijala, J. Evans, B. Laurie, A. Mordvintsev, E. Niklasson, E. Randazzo, and L. V er- sari, Computational life: How well-formed, self-replicating programs emerge from simple interaction, arXiv preprint arXiv:2406.19108 (2024)

  28. [28]

    Engel and C

    A. Engel and C. van den Broeck, Statistical mechanics of learning (Cambridge University Press, 2001)

  29. [29]

    Advani, S

    M. Advani, S. Lahiri, and S. Ganguli, Statistical mechanics of complex neural systems and high dimensional data, Journal of Statistical Mechanics: Theory and Experiment 2013, P03014 (2013)

  30. [30]

    Gabri ´e, S

    M. Gabri ´e, S. Ganguli, C. Lucibello, and R. Zecchina, Neural networks: From the perceptron to deep nets, in Spin Glass Theory and Far Beyond: Replica Symmetry Breaking After 40 Years (World Scientific, 2023) pp. 477–497

  31. [31]

    Cui, High-dimensional learning of narrow neural networks, Journal of Statistical Mechanics: Theory and Experiment 2025, 023402 (2025)

    H. Cui, High-dimensional learning of narrow neural networks, Journal of Statistical Mechanics: Theory and Experiment 2025, 023402 (2025)

  32. [32]

    G. J. Hollich, K. Hirsh-Pasek, R. M. Golinkoff, R. J. Brand, E. Brown, H. L. Chung, E. Hennon, C. Rocroi, and L. Bloom, Breaking the language barrier: An emergentist coalition model for the origins of word learning, Monographs of the society for research in child development (2000)

  33. [33]

    Pinker, The language instinct: How the mind creates lan- guage (Penguin UK, 2003)

    S. Pinker, The language instinct: How the mind creates lan- guage (Penguin UK, 2003)

  34. [34]

    Dupoux, Cognitive science in the era of artificial in- telligence: A roadmap for reverse-engineering the infant language-learner, Cognition 173, 43 (2018)

    E. Dupoux, Cognitive science in the era of artificial in- telligence: A roadmap for reverse-engineering the infant language-learner, Cognition 173, 43 (2018)

  35. [35]

    Zaadnoordijk, T

    L. Zaadnoordijk, T. R. Besold, and R. Cusack, The next big thing (s) in unsupervised machine learning: Five lessons from infant learning, arXiv preprint arXiv:2009.08497 (2020)

  36. [36]

    Fontana, Algorithmic chemistry , Tech

    W. Fontana, Algorithmic chemistry , Tech. Rep. (Los Alamos National Lab., NM (USA), 1990)

  37. [37]

    Y . Y ao, L. Rosasco, and A. Caponnetto, On early stopping in gradient descent learning, Constructive approximation 26, 289 (2007)

  38. [38]

    Pross, Dynamic kinetic stability, in Encyclopedia of Astro- biology (Springer, 2022) pp

    A. Pross, Dynamic kinetic stability, in Encyclopedia of Astro- biology (Springer, 2022) pp. 1–2

  39. [39]

    Lee et al

    D.-H. Lee et al. , Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks, in Workshop on challenges in representation learning, ICML , V ol. 3 (Atlanta, 2013) p. 896

  40. [40]

    Arazo, D

    E. Arazo, D. Ortego, P . Albert, N. E. O’Connor, and K. McGuinness, Pseudo-labeling and confirmation bias in deep semi-supervised learning, in 2020 International joint conference on neural networks (IJCNN) (IEEE, 2020) pp. 1–8

  41. [41]

    Zhang, J

    L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, and K. Ma, Be your own teacher: Improve the performance of convolu- tional neural networks via self distillation, in Proceedings of the IEEE/CVF international conference on computer vision 6 (2019) pp. 3713–3722

  42. [44]

    We leave this for future work

    In principle, one could adapt the toolbox of Takahashi [34] to our setting to consider variations of our dynamics, the main difference being the bootstrapping of the learning process at short times. We leave this for future work

  43. [46]

    Note that, while the effect on the agent parameters of the turnover step can be framed as an optimization problem and analyzed with equilibrium statistical mechanics, the math- ematical description of the agent-environment interaction would require an out-of-equilibrium analysis, since the agent needs an energy input to sustain the learning process at each step

  44. [47]

    Note that in this setting the NMI is not self-averaging with respect to different initial conditions, but develops different modes in correspondence of different levels of centroid cov- erage. Decreasing values of the average NMI < 1 denote set- tings where the maximum achievable NMI is attained only in few lucky initializations, while other cases converg...

  45. [48]

    Frankle and M

    J. Frankle and M. Carbin, The lottery ticket hypothesis: Finding sparse, trainable neural networks, arXiv preprint arXiv:1803.03635 (2018)

  46. [49]

    Boyd and P

    R. Boyd and P . J. Richerson, Culture and the Evolutionary Process (University of Chicago Press, 1985)

  47. [50]

    Henrich and R

    J. Henrich and R. McElreath, The evolution of cultural evolu- tion, Evolutionary Anthropology 12, 123 (2003)

  48. [51]

    K. N. Laland, Social learning strategies, Learning & Behavior 32, 4 (2004)

  49. [52]

    Whiten and C

    A. Whiten and C. P . van Schaik, The evolution of animal ’cul- tures’ and social intelligence, Philosophical Transactions of the Royal Society B 362, 603 (2007)

  50. [53]

    P . Long, T. Fan, X. Liao, W. Liu, H. Zhang, and J. Pan, To- wards optimally decentralized multi-robot collision avoidance via deep reinforcement learning, in 2018 IEEE international conference on robotics and automation (ICRA) (IEEE, 2018) pp. 6252–6259

  51. [54]

    Bredeche and N

    N. Bredeche and N. Fontbonne, Social learning in swarm robotics, Philosophical Transactions of the Royal Society B 377, 20200309 (2022)

  52. [55]

    M. Y . Ben Zion, J. Fersula, N. Bredeche, and O. Dauchot, Morphological computation and decentralized learning in a swarm of sterically interacting robots, Science Robotics 8, eabo6140 (2023)

  53. [56]

    P . A. P . Moran, Random processes in genetics, in Mathe- matical proceedings of the cambridge philosophical society , V ol. 54 (Cambridge University Press, 1958) pp. 60–71

  54. [57]

    R. A. Blythe and A. J. McKane, Stochastic models of evolu- tion in genetics, ecology and linguistics, Journal of Statistical Mechanics: Theory and Experiment 2007, P07018 (2007)

  55. [58]

    T. C. Schelling, Micromotives and macrobehavior (WW Nor- ton & Company, 2006)

  56. [59]

    Grauwin, E

    S. Grauwin, E. Bertin, R. Lemoy, and P . Jensen, Competition between collective and individual dynamics, Proceedings of the National Academy of Sciences 106, 20622 (2009)

  57. [60]

    Garnier-Brun, R

    J. Garnier-Brun, R. Zakine, and M. Benzaquen, Hydrodynam- ics of cooperation and self-interest in a two-population occu- pation model, Phys. Rev. Lett. 135, 107402 (2025)

  58. [61]

    Replication and Information Extraction in a Minimal Agent-Environment Model

    G. Jung, M. Ozawa, and E. Bertin, Kinetic theory of decen- tralized learning for smart active matter, Phys. Rev. Lett. 134, 248302 (2025). Supplemental Material for “Replication and Information Extraction in a Minimal Agent-Environment Model” Sebastiano Ariosto, ∗ J´erˆome Garnier-Brun, Luca Saglietti, and Davide Straziota Bocconi Institute for Data Scien...

  59. [62]

    ADDITIONAL RESULTS A

    Mixture of C > 2 centroids 11 References 15 I. ADDITIONAL RESULTS A. Learning Dynamics In the main text we focused on the stationary properties of the agent, showing how the long-term behavior is captured by the fixed point of the recursion equation. It is worth emphasizing, however, that this map is not restricted to the stationary regime, as it provides ...

  60. [63]

    Alternating fixed batches , where two (or more) disjoint mini-batches are defined once and then reused across successive time steps

  61. [64]

    Uniform splitting (or 0% resampling), where the budget is partitioned into T disjoint subsets of size P = Ptot/T , each of which is presented only once

  62. [65]

    The comparison, reported in Fig

    Partial X% resampling, where at each step a fraction P/P tot = X/100 is sampled uniformly at random from the entire dataset, so that examples may be reused across time with a tunable frequency. The comparison, reported in Fig. S7, reveals that replicators can still emerge under strong data limitations, provided that each step offers sufficiently rich stati...

  63. [66]

    5 0 20 40 t/M

    0 πt “Easy” η < 0. 5 0 20 40 t/M

  64. [67]

    Easy” η ≥ 0. 5 0 1 2 t/M 2 “Hard

    0 πt “Easy” η ≥ 0. 5 0 1 2 t/M 2 “Hard” η = 0 0 20 40 t/M “Hard” η > 0 FIG. S8. Evolution of the absolute value of the average alignment describing consensus within the population in fully connected populations of size M = [32 , 64, 128], N = 500 , α = 1 , λ = 10 , averaged over 32 trials. Colors indicate the value of η, matching the legend of the Main te...

  65. [68]

    Decompose the weight vector as Wt = mtµ + ξt, with ⟨µ, ξt⟩ = 0

    Lower bound. Decompose the weight vector as Wt = mtµ + ξt, with ⟨µ, ξt⟩ = 0. For a slow dynamics it is natural to assume that fluctuations retain some positive correlation across nearby times (i.e. ⟨ξt, ξt+1⟩ > 0), which is consistent with our learning dynamics. Then C(∆t) = ⟨Wt, Wt+∆t⟩ = mtmt+∆t + ⟨ξt, ξt+∆t⟩. The first term tends to (m∗)2, while the secon...

  66. [69]

    Under dynamical stability, mt → m∗

    Long-time limit. Under dynamical stability, mt → m∗. Meanwhile, the orthogonal components ξt evolve stochastically within the high-dimensional subspace orthogonal to µ. In the long-time limit, these fluctuations decorrelate, so that lim∆t→∞⟨ξt, ξt+∆t⟩ = 0. Hence the correlation reduces to lim ∆t→∞ C(∆t) = ( m∗)2. Intuitively, the only persistent contributi...

  67. [70]

    The analytical computation is adapted from the procedure outlined in [1]

    Two centroid case The analysis of the replica investigation begins by considering the simplest case, when data is generated from a two-centroid distribution, C = 2. The analytical computation is adapted from the procedure outlined in [1]. The case of our interest is the one in which the agent, at the iteration t + 1, can only observe its labels generated ...

  68. [71]

    The core steps for the calculation are identical to those presented in the previous paragraphs

    Mixture of C > 2 centroids In this section, in analogy to the results presented in the C = 2 case, we propose a replica computation for C > 2 centroids. The core steps for the calculation are identical to those presented in the previous paragraphs. The main difficulty comes from the intrinsic nature of the labels, which are not binary anymore. The key step...

  69. [72]

    S. S. Mannelli, F. Gerace, N. Rostamzadeh, and L. Saglietti, Bias-inducing geometries: An exactly solvable data model with fairness implications, in Proceedings of the Geometry-Grounded Representation Learning and Generative Modeling Workshop (ICML 2024) (2024) arXiv:2205.15935

  70. [73]

    Engel and C

    A. Engel and C. V . den Broeck, Statistical Mechanics of Learning (Cambridge University Press, Cambridge, UK, 2001)

  71. [74]

    Takanami, T

    K. Takanami, T. Takahashi, and A. Sakata, The effect of optimal self-distillation in noisy gaussian mixture model, arXiv preprint arXiv:2501.16226 (2025)

  72. [75]

    Takahashi, The role of pseudo-labels in self-training linear classifiers on high-dimensional gaussian mixture data, arXiv preprint arXiv:2205.07739 (2022)

    T. Takahashi, The role of pseudo-labels in self-training linear classifiers on high-dimensional gaussian mixture data, arXiv preprint arXiv:2205.07739 (2022)

  73. [76]

    Loureiro, G

    B. Loureiro, G. Sicuro, C. Gerbelot, A. Pacco, F. Krzakala, and L. Zdeborov ´a, Learning gaussian mixtures with generalised linear models: Precise asymptotics in high-dimensions, in Advances in Neural Information Processing Systems 34 (NeurIPS 2021) (2021) spotlight presentation

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.