{"id":"a289b659-de0d-41db-be0d-907095a47742","arxiv_id":"2608.08922","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For normalized self-attention in the d ~ N limit, the overlap gap controls a manifold of clustered fixed points and a finite-sharpness dynamical attention-condensation transition.","lead":"What happens when token representations in a transformer are both the players and the board of the game? This paper studies a minimal version of that feedback loop and finds a manifold of clustered attractor states governed by each cluster's overlap gap to its nearest rival, plus a finite-sharpness threshold for condensation out of random initial states.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The finite-beta condensation onset from Gaussian initial conditions is supported only by T=N finite-size data; without a beta_c(N) extrapolation, the paper's own lifetime bound suggests the apparent onset may drift to zero, leaving the central claim unestablished.","rationale":"The reader's weakest assumption identifies the amplification of O(1) fluctuations into O(1) overlap gaps before rank collapse as the load-bearing mechanism, supported only by heuristic argument and T=N finite-size simulations. My concern sharpens this: the missing scaling collapse of beta_c(N) is not merely a technical omission, because the paper's own lifetime analysis (Eq. 24, Eq. 118) demonstrates that finite-time stability boundaries in the same model shift to zero with N. If the Gaussian-onset beta_c(N) also decays as (log N)/sqrt(N), then the claimed finite-beta transition is not a thermodynamic transition but a finite-observation-time crossover. The fixed-point manifold and linear stability results (Leakage bound Eq. 10, linear contraction in Supp. III B) are direct and correct; the weakness is exclusively in the dynamical-accessibility half of the paper, which is the headline claim. The concrete test of extrapolating beta_c(N) and extending observation times will settle the issue, so the CONDITIONAL verdict remains appropriate pending that evidence. I do not see an internal inconsistency in the fixed-point analysis; the concern is the unsupported asymptotic claim about nucleation, which the paper itself flags as heuristic in the 'Dynamical accessibility' section.","tokens_in":21504,"tokens_out":6459,"duration_ms":68592,"concrete_test":"For Gaussian initial conditions, define beta_c(N) as the smallest beta at which a chosen condensation diagnostic (e.g., mean attention IPR YA(T=N) >= 0.1, or the peak of V/N) first exceeds its diffuse-regime value, for N = 512, 1024, 2048, 4096, 8192, and 16384. Plot beta_c(N) against 1/N, 1/sqrt(log N), and (log N)/sqrt(N). If the extrapolated limit is positive, the finite-beta claim is supported; if it approaches zero, the transition is a finite-size/finite-time crossover and the central claim fails. As a complementary check, for a beta below the apparent onset at N = 4096 (e.g., beta = 0.2), extend the simulation to T = 10N and T = 100N; if condensation or fragmentation appears at longer times, the T = N onset is an artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that clustered states nucleate from an unstructured Gaussian cloud only above a finite threshold in attention sharpness beta, giving a genuine dynamical condensation transition at beta = O(1) in the joint limit d ~ N -> infinity. The evidence is Figure 3, measured at the single observation time T = N for N up to 4096, plus a heuristic feedback-amplification argument. This evidence does not establish a thermodynamic transition for two reasons. First, the paper itself shows (Eq. 24 and Supplemental Sec. III D, Eq. 118) that for any fixed beta > 0, fragmented states have lifetimes diverging faster than any power of N, and the apparent finite-time stability boundary satisfies beta_surv(N,T) ~ (log N + log T)/sqrt(N) -> 0. The Gaussian-onset data are collected at the same T = N, so the apparent onset could be a finite-time artifact: for any beta > 0, the initial O(1) logit fluctuations might nucleate clusters on a time scale that grows with N, making the measured beta_c(N) decay rather than converge to a positive constant. The paper provides no scaling collapse of beta_c and no extrapolation to N -> infinity. Second, the only quantitative benchmark is the static REM result beta_REM ~ sqrt(log N) (Supplemental Sec. I); the dynamical feedback must be shown to reduce this to O(1) rigorously, not just heuristically. Because the fixed-point and stability analyses are sound, the gap is specifically in the dynamical-accessibility claim, which is exactly the paper's headline result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a minimal normalized self-attention dynamics in which token representations determine the attention matrix and the attention matrix updates the representations, with token dimension d and number N taken together to infinity (d ~ N). The authors define normalized token overlaps q_ij and scaled logits z_ij = sqrt(d) q_ij, with softmax attention at inverse temperature beta. They show that, when tokens form clusters with a nonvanishing same-cluster overlap advantage, inter-cluster attention is exponentially suppressed, producing a manifold of clustered fixed points ranging from few macroscopic clusters to extensive microscopic fragmentation. They analyze the linear stability of these states, the noise-induced fragmentation of a macroscopic parent cluster characterized by effective sharpness alpha = beta epsilon^2, the lifetime of fragmented states under finite-N coarsening, and the order-of-limits distinction between t -> infinity first and N -> infinity first. The final part addresses dynamical accessibility from independent Gaussian initial conditions and claims a finite-beta attention-condensation transition, with three regimes: diffuse rank collapse, macroscopic-clustered condensation, and fragmented condensation.","tokens_in":21706,"tokens_out":4566,"duration_ms":51315,"significance":"If established, the paper would provide a clean statistical-mechanical picture of self-attention dynamics: the overlap gap as the organizing quantity, a normally attracting clustered manifold, exponentially long-lived fragmentation protected by sqrt(N) logit advantages, and a dynamical nucleation transition at O(1) beta from unstructured initial states. The fixed-point characterization, the leakage bound, the linear stability calculation, and the lifetime scaling (Eqs. 10, 23, 24 and Supplemental Sec. III) are internally consistent and are genuine contributions. The alpha-collapse of the fragmentation data and the explicit order-of-limits discussion are also valuable. However, the central dynamical-accessibility claim—that a finite beta_c survives in the joint limit—is not established by the evidence presented; the paper's own finite-size survival criterion in the Supplemental Material (Eq. 118) points to a possible drift of the apparent onset to zero. This gap is load-bearing because the finite-beta condensation transition is the headline result.","major_comments":[{"comment":"The claim that \"the condensation onset remains at finite beta as N -> infinity\" is supported only by simulations at the single observation time T = N, with no beta_c(N) extrapolation, no scaling collapse, and no quantitative definition of the onset. This is exactly the regime where the paper's own survival criterion applies: Supplemental Eq. (118) gives beta_surv(N,T) ~ (log N + log T)/sqrt(N), which at T = N behaves as ~ 2 log N / sqrt(N) -> 0. The apparent onset seen in Fig. 3 could therefore be a finite-time artifact: for any fixed beta > 0, nucleation from O(1) initial logit fluctuations might occur on a time scale that grows with N, making beta_c(N) decay rather than converge to a positive constant. The time-window plateau shown in Supplemental Fig. 7 (T = N to 5N) does not resolve this, because the predicted coarsening time diverges faster than any power of N. The authors must provide either a scaling collapse of the onset with an N -> infinity extrapolation, or an asymptotic derivation of the nucleation time showing that it remains O(poly(N)) only for beta above a positive constant.","section":"Dynamical accessibility from Gaussian initial conditions; Fig. 3"},{"comment":"The mechanism invoked to justify the finite-beta onset is the statement that \"feedback between attention and token geometry amplifies overlap fluctuations and generates finite overlap gaps.\" This is only a heuristic assertion; no linearized instability analysis of the diffuse Gaussian state, no mean-field nucleation calculation, and no bound on the growth rate of the overlap gap is provided. The initial logits are O(1), the same order as the static REM benchmark logits, so the claim that dynamical feedback reduces the threshold from beta_REM ~ sqrt(log N) to beta = O(1) requires a quantitative demonstration. Without such an analysis, the comparison to the static REM benchmark does not by itself support the dynamical transition.","section":"Dynamical accessibility from Gaussian initial conditions; Eqs. (26) and feedback-amplification paragraph"},{"comment":"The manuscript itself states that the apparent small-beta stability boundary at finite size and finite time moves to zero in the thermodynamic limit, beta_surv(N,T) ~ (log N + log T)/sqrt(N). This is in direct tension with the main-text assertion that the Gaussian-initial-condition condensation onset remains finite as N -> infinity. The two statements could be consistent if the Gaussian nucleation process is qualitatively different from the survival of an already formed fragmented state, but the paper does not show this. The authors need to explain how the finite-time Gaussian-onset data avoid the drift predicted by Eq. (118), or provide numerical evidence that beta_c(N) converges to a positive value, in which case the central claim would be supported.","section":"Stability of extensive fragmentation; Supplemental Sec. III D, Eq. (118)"}],"minor_comments":[{"comment":"The abstract and introduction state the finite-beta condensation transition as an established result, but the body of the paper only provides finite-time numerical evidence; the wording should be softened until the scaling analysis is supplied.","section":"Abstract and Introduction"},{"comment":"The cluster statistics depend on a transitive overlap threshold q_th; while Supplemental Sec. V B shows robustness for 1 - q_th = 10^-4 and 10^-5, the main text should mention this threshold choice and its insensitivity explicitly in the caption or text.","section":"Figure 1 and cluster diagnostics"},{"comment":"The lifetime formula tau_frag ~ (1/gamma N) exp(beta sqrt(N) Delta) is stated for bounded-size clusters, but the geometric factor g_a from the Supplemental derivation is omitted; a brief statement that g_a = O(1) for finite angular gaps would improve clarity.","section":"Equation (24) and surrounding text"},{"comment":"The paper cites Refs. [13-20] for clustering and mean-field results, but the distinction between those results and the present overlap-gap mechanism could be made more explicit, especially regarding which results are new and which are refinements of prior work.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The fixed-point and stability analysis is solid and likely publishable, but the headline claim of a finite-beta dynamical condensation transition from Gaussian initial conditions is not yet established. I would need to see a proper finite-size scaling of the onset or an asymptotic nucleation argument before accepting the central claim. The paper's own Supplemental Eq. (118) is a serious red flag that the apparent onset may drift to zero; this must be addressed head-on."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is what I think you should know. The overlap-gap mechanism is real and clean. The paper proves that if tokens form clusters with a nonvanishing overlap advantage, inter-cluster attention is exponentially suppressed in sqrt(N), giving a manifold of clustered fixed points that are locally stable. The leakage bound (Eq. 10) and the linear contraction of internal modes are correct and transferable. The concentration argument for extensive fragmentation is also solid: for K=Theta(N) random centers, the max overlap is O(sqrt(log N/N)), so the gap tends to 1 and every fixed beta>0 gives vanishing leakage. This part is a genuine contribution, and the diagnostics YA, Reff, V/N are sensible.\n\nThe soft spot is exactly where the reader put it: the dynamical condensation transition. The claim that Gaussian initial conditions nucleate clusters only above a finite beta in the joint limit is supported by Fig. 3, measured at T=N for N up to 4096, plus a heuristic feedback-amplification story. No beta_c(N) extrapolation is shown. The paper's own Eq. (24) and Supp. Eq. (118) say that for any fixed beta>0, fragmented states have lifetimes ~ exp(beta sqrt(N))/N, and the apparent finite-time stability boundary is beta_surv(N,T) ~ (log N + log T)/sqrt(N) -> 0. Since the numerical onset is at T=N, the data are exactly consistent with beta_c(N) ~ 2 log N/sqrt(N) decaying to zero. So the central claim that the onset is at finite beta is unestablished. The feedback amplification may be enough to create O(1) gaps before coarsening, but that is not derived; it is a conjecture. The coexistence regime (one macroscopic cluster plus many microscopic ones) is also defined by finite-size peaks and may be a crossover.\n\nThe paper deserves a serious referee. The rigorous core is publishable, and the dynamical claim is important enough that we need someone to try to prove it or kill it. A referee should ask for a scaling collapse of beta_c with N, an asymptotic treatment of the nucleation step, or at least an honest reframing as a conjecture. If the authors can show beta_c saturates to a positive constant, this becomes a strong paper; if it drifts to zero, the three-regime picture collapses into a finite-time crossover.\n\nI would bring it to the group, and I would cite the overlap-gap fixed-point results. The thinking is serious; the gap is a missing proof, not a confused argument.","headline":"Solid overlap-gap fixed-point analysis, but the headline finite-beta condensation transition from Gaussian initial conditions is not supported by the T=N numerics, which are consistent with the paper's own lifetime bound pushing the onset to zero.","tokens_in":22311,"tokens_out":3517,"would_cite":true,"duration_ms":36718,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a positive overlap gap makes clustered token states stable attractors of self-attention dynamics, with a finite sharpness threshold for their nucleation from random initial conditions.","keywords":["self-attention dynamics","overlap gap","attention condensation","rank collapse","clustered attractor manifold","fragmentation","random-energy model","participation ratio"],"falsifier":"Simulate the model with $d=N$, $\\gamma=0.3$, Gaussian initial conditions, and a fixed $\\beta$ such as $0.8$, measuring the mean attention IPR $Y_A$ at time $T=N$ for increasing $N$. If $Y_A$ decays to zero with $N$, or if the apparent condensation onset $\\beta_c(N)$ does not converge to a positive constant, the claimed finite-$\\beta$ dynamical transition is false; a complementary check is whether the normalized overlap variance grows from $O(N^{-1})$ to $O(1)$ gaps rather than decaying.","tokens_in":21211,"feed_emoji":"🌀","tokens_out":11403,"duration_ms":109370,"temperature":0.7,"pith_summary":"The paper studies a minimal self-attention dynamics in which token representations set the attention matrix and the attention matrix updates the representations, in the joint limit where the number of tokens $N$ and the embedding dimension $d$ grow together. It claims that the overlap gap—the surplus similarity of a token to its own cluster over every other cluster—is the quantity that decides the attractor structure. A nonvanishing gap produces a $\\sqrt{N}$ logit advantage, so inter-cluster attention is exponentially suppressed at fixed sharpness $\\beta$, making clustered configurations stable fixed points across a whole manifold from a few macroscopic clusters to extensive microscopic fragmentation. Starting from unstructured Gaussian tokens, clusters form only above a finite $\\beta$; below it, averaging erases diversity and the system rank-collapses, while above it, feedback amplifies fluctuations and condenses attention. The result matters because it locates a genuine dynamical transition at $\\beta=O(1)$, whereas static random logits would need $\\beta\\sim\\sqrt{\\log N}$.","feed_headline":"Attention condenses past a finite sharpness threshold","feed_subtitle":"With token count and dimension growing together, clusters appear at beta = O(1), not at the sqrt(log N) static onset.","key_machinery":"The paper works with the self-masked softmax attention matrix $A_{ij}\\propto\\exp(\\beta\\sqrt{d}\\,q_{ij})$ (diagonal excluded) and the normalized residual update $x_i(t+1)\\propto(1-\\gamma)x_i(t)+\\gamma\\sum_{j\\neq i}A_{ij}x_j(t)$. The object that carries the argument is the overlap gap $\\Delta_a=1-\\max_{b\\neq a}Q_{ab}$, the amount by which a token's overlap with members of its own cluster exceeds its largest overlap with any other cluster. When $\\Delta_a$ is bounded away from zero, the total attention a source cluster sends to all other clusters is bounded by $N e^{-\\beta\\sqrt{N}\\Delta_a}$, which vanishes for every fixed $\\beta>0$ in the $d=N\\to\\infty$ limit. This exponential leakage bound simultaneously yields the fixed-point condition, the local stability of internal modes (contraction factor $1-\\gamma n_a/(n_a-1)$), and the fragmentation lifetime $e^{\\beta\\sqrt{N}\\Delta}$. A companion control parameter is the effective sharpness $\\alpha=\\beta\\epsilon^2$ for a perturbed macroscopic cluster, which decides whether the perturbation heals or fragments into a narrow cone of microscopic descendants.","core_discovery":"The central claim is that the overlap gap, not the average similarity, organizes the attractor phase diagram of minimal normalized self-attention. In the joint limit $d\\sim N\\to\\infty$, a token whose overlap with its own cluster exceeds its overlap with every other cluster by a nonvanishing amount $\\Delta$ gives that cluster a logit advantage $\\beta\\sqrt{N}\\Delta$; because the competing set has only $O(N)$ targets and hence $O(\\log N)$ entropy, the inter-cluster attention fraction decays like $N e^{-\\beta\\sqrt{N}\\Delta}$, which vanishes for every fixed $\\beta>0$. Consequently, every clustered configuration with a uniform positive gap is an exact fixed point of the limiting dynamics, and the manifold of such states spans diffuse macroscopic clusters with $Y_A=O(1/N)$ and condensed microscopic fragmentation with $Y_A=O(1)$, with narrow-cone fragments remaining low-rank in representation. These fixed points are normally attracting: internal deformations contract, while collective rotations are neutral. From Gaussian initial conditions the gap must be created by the dynamics, and the paper finds this happens only above a finite sharpness $\\beta$; below it, attention averaging collapses the cloud to rank one, so the finite-$O(1)$ onset is a genuinely dynamical condensation transition rather than the static random-energy threshold $\\beta\\sim\\sqrt{\\log N}$.","pith_inferences":["A practical diagnostic follows for trained transformers: measure the gap between the mean intra-cluster and maximum inter-cluster overlap of token representations; heads with a finite gap should show exponentially suppressed attention leakage, and head specialization could be predicted from this gap.","The kernel-threshold result suggests a testable architectural prediction: replacing softmax with a quadratic nonlinearity should still produce condensation in the $d\\sim N$ regime, while linear attention should not, separating the role of exponential selection from mere feedback.","The three dynamical regimes imply that depth acts as a control parameter: repeated layers at fixed $\\beta$ below onset should drive rank collapse, while above onset they should drive further fragmentation or slow coarsening depending on the order of limits, which could be tested in deep attention stacks.","The coexistence regime of one macroscopic cluster plus many microscopic fragments resembles attention-sink phenomenology; testing whether the macroscopic cluster absorbs most inter-cluster attention could connect the model to observed training instabilities."],"forward_implications":["Clustered configurations with a uniform positive overlap gap are exact fixed points of the limiting dynamics at any fixed $\\beta>0$, so self-attention alone can sustain a high-dimensional manifold of stable clustered states without external regularization.","In the thermodynamic limit taken first, extensive microscopic fragmentation is locally stable; at finite $N$ the same states coarsen on a time scale growing as $\\exp(\\beta\\sqrt{N}\\Delta)$, so the apparent stability boundary shifts toward $\\beta=0$ as $N$ grows.","From Gaussian initial data the dynamics shows three regimes—diffuse rank collapse, condensed coexistence with one macroscopic cluster, and fragmented condensation—separated by transitions in the attention inverse participation ratio, the overlap variance, and the representation participation rank.","The finite-$\\beta$ condensation onset contrasts with frozen random logits, which would require $\\beta\\sim\\sqrt{\\log N}$ according to the random-energy-model benchmark.","If the dimension is held fixed while $N\\to\\infty$, or if the attention kernel is linear or subquadratic, fixed-$\\beta$ condensation disappears, so the exponential softmax together with $d\\sim N$ scaling is the mechanism behind the transition."],"supporting_citations":[{"why":"Establishes that repeated pure attention loses rank doubly exponentially, the baseline the diffuse regime extends.","marker":"[9]"},{"why":"Provides the signal-propagation and rank-collapse analysis behind the small-$\\beta$ erasure mechanism.","marker":"[10]"},{"why":"Introduces the emergence of clusters in self-attention dynamics, the starting point for the exact clustered fixed-point manifold.","marker":"[13]"},{"why":"Supplies the mean-field polarization dynamics whose logistic growth sets the $O(\\log N)$ macroscopic cluster formation time.","marker":"[16]"},{"why":"Connects random attention logits at initialization to the random-energy model, defining the static benchmark.","marker":"[21]"},{"why":"Derives the random-energy-model thermodynamics used for the static condensation threshold $\\beta\\sim\\sqrt{\\log N}$.","marker":"[22]"},{"why":"Analyzes dynamic metastability in self-attention, supporting the exponentially long-lived fragmented states.","marker":"[24]"},{"why":"Shows meta-stable clustering in mean-field transformer models, a companion to the narrow-cone fragmented states.","marker":"[25]"},{"why":"Provides the linear-attention construction used as the non-softmax kernel comparison that does not condense at fixed $\\beta$.","marker":"[26]"}],"fun_headline_variants":["Overlap gap sets attention attractor phase","Finite sharpness trigger for attention condensation","Dynamical condensation in self-attention at finite beta","Attention clusters only past a finite sharpness","Self-attention condensation: gap-driven phase transition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that, from a random start, the feedback loop amplifies small accidental similarities into real clusters before averaging erases them; the paper supports this amplification by heuristic reasoning and finite-size simulations, not by an asymptotic proof in the thermodynamic limit.","fun_headline_variants_meta":{"raw":{"variants":["Overlap gap sets attention attractor phase","Finite sharpness trigger for attention condensation","Dynamical condensation in self-attention at finite beta","Attention clusters only past a finite sharpness","Self-attention condensation: gap-driven phase transition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000422,"raw_usage":{"total_tokens":2176,"prompt_tokens":961,"completion_tokens":1215,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":1145}},"tokens_in":577,"tokens_out":1215,"duration_ms":9810,"temperature":1.0,"reasoning_tokens":1145,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:20:47.527508+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the model with $d=N$, $\\gamma=0.3$, Gaussian initial conditions, and a fixed $\\beta$ such as $0.8$, measuring the mean attention IPR $Y_A$ at time $T=N$ for increasing $N$. If $Y_A$ decays to zero with $N$, or if the apparent condensation onset $\\beta_c(N)$ does not converge to a positive constant, the claimed finite-$\\beta$ dynamical transition is false; a complementary check is whether the normalized overlap variance grows from $O(N^{-1})$ to $O(1)$ gaps rather than decaying.","supporting_citations":[{"cited_title":"Dong, J.-B","cited_arxiv_id":null,"evidence_quote":"Establishes that repeated pure attention loses rank doubly exponentially, the baseline the diffuse regime extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the signal-propagation and rank-collapse analysis behind the small-$\\beta$ erasure mechanism."},{"cited_title":"Rigollet, The mean-field dynamics of transformers, in Proceedings of the International Congress of Mathemati- cians 2026—Volume 7: Invited Lectures (Sections 15–20), edited by S","cited_arxiv_id":null,"evidence_quote":"Supplies the mean-field polarization dynamics whose logistic growth sets the $O(\\log N)$ macroscopic cluster formation time."},{"cited_title":"Katharopoulos, A","cited_arxiv_id":null,"evidence_quote":"Provides the linear-attention construction used as the non-softmax kernel comparison that does not condense at fixed $\\beta$."}],"review_version":1}