{"id":"fde2e147-7ff4-4053-b560-c695326ba30f","arxiv_id":"2506.06134","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A continuous-time Hebbian/anti-Hebbian similarity matching network is shown to converge layer by layer to the principal subspace solution, with the slow layer convergence relying on two unproven conjectures.","lead":"This paper studies a brain-inspired network that compresses data by matching input and output similarities, and proves that its fast and middle layers converge to optimal values. It provides a conditional proof and simulations that the slow learning layer finds the principal subspace projection, but the full network proof is not complete.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full-network convergence is never proved: only the ε→0 reduced flows are analyzed, and no singular-perturbation theorem links (7) to (20), (24), and (26).","rationale":"The reader identified the same load-bearing concern: the reduction from the coupled system (7) to the three sequential reduced flows is the only mechanism connecting the rigorous Theorems 1–3 to the paper's central claim about the full network. The paper proves strong convexity/concavity and contractivity for each isolated level, gives an explicit global-minimizer characterization for the slow cost, and honestly labels the slow-level results as dependent on two empirical conjectures. Those partial results are sound and useful. The missing piece is exactly a singular-perturbation/tracking theorem that would justify replacing (7) by (20), (24), and (26) for small finite ε1, ε2. Because this gap affects the headline claim about the full network but not the validity of the per-level theorems, the appropriate verdict remains CONDITIONAL as the reader concluded; no adjustment is needed.","tokens_in":28750,"tokens_out":7873,"duration_ms":85122,"concrete_test":"Derive a quantitative singular-perturbation tracking bound: prove that under Conjectures 1–2 there exist C, T0, ε* > 0 such that for all 0 < ε1, ε2 < ε*, sup_{0≤t≤T0} [ ||Y(t) − M(t)^{-1}W(t)X||_F + ||M(t) − (W(t)C_XW(t)^T)^{1/3}||_F + ||W(t) − W_red(t)||_F ] ≤ C(ε1+ε2), where W_red solves (26). As a specific check, verify whether the frozen-M contraction rate (4/T)λmin(M) admits a uniform positive lower bound along solutions of (7); if the bound degenerates, construct X and W0 for which the fast subsystem's convergence time is O(1/λmin(M)) and show the O(ε) tracking error bound fails. If such a bound cannot be established, the paper should state convergence only for the singular-limit reduced flows, not for the finite-ε network (7).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The advertised claim—that trajectories of the similarity matching network (7) converge to a global optimum (Informal statement, Section 5)—is not established. What is proved is convergence of three decoupled gradient flows: (20) for Y with W,M frozen, (24) for M with W frozen and Y at its fast equilibrium, and (26) for W with M,Y at their equilibria. The reduction from (7) to these flows is made informally in Section 4 via 'as ε1→0 ... the reduced slower subsystem is therefore ...' and 'when ε2→0 ... the slow system becomes ...' (eqs. (8)–(11)). No singular-perturbation theorem (Tikhonov, center-manifold, or contraction-based bounded-error tracking) is stated or proved. The gap is not cosmetic: the fast contraction rate ν_Y=(4/T)λmin(M) depends on M and can degenerate if M approaches the boundary of S^m_{>0}, so the uniform-in-parameters contractivity needed for a singular-perturbation argument is not verified. Similarly, Conjecture 1 only asserts forward invariance of the full-rank set, not a uniform lower bound on σ_min(W(t)) over time, so the slow subsystem's Lipschitz constants are not controlled uniformly. Thus for any fixed small ε1, ε2, convergence of the full coupled system (7) remains unproven. The numerical experiments in Section 6 show small SM(Y)-SM(Y*) for one data distribution, but they do not test the tracking bound that would fill this gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes a continuous-time Hebbian/anti-Hebbian network derived from an embedded similarity matching objective, with neural, lateral-synaptic, and feedforward-synaptic variables evolving on three time scales. Using a multilevel optimization viewpoint, it studies three decoupled gradient flows: fast neural dynamics (20), intermediate lateral dynamics (24), and slow feedforward dynamics (26). For the first two, it proves strong convexity/concavity and global exponential convergence, and it gives explicit equilibrium formulas. For the slow level, it characterizes stationary points and global minima and proves, under two (plus one auxiliary) empirically motivated conjectures, almost sure convergence of the slow flow to a global minimizer. Projecting the composed optimum back to the output space yields the principal subspace projection. The paper's informal statement claims convergence of the full coupled network (7), but no theorem analyzes the finite-epsilon coupled system directly.","tokens_in":29148,"tokens_out":15151,"duration_ms":138864,"significance":"If the results hold, the paper would provide the first continuous-time convergence analysis of a biologically plausible similarity matching network, with clean contraction-based proofs, explicit closed-form equilibria, a formal proof of positive-definiteness invariance for the lateral dynamics, and a plausible landscape analysis of the non-convex slow level. The conjectures are stated clearly and tested numerically, and the code is provided. Theorems 1 and 2 appear sound and are useful in their own right. However, the gap between the analyzed reduced flows and the actual multi-time-scale system prevents the paper from delivering the advertised complete convergence analysis of the network (7); the main convergence claim is therefore conditional on an unproved singular-perturbation step and on two unproved conjectures.","major_comments":[{"comment":"","section":"Section 4, eqs. (8)–(11); Section 5, Informal statement"},{"comment":"","section":"Section 5.3, Theorem 3; Conjectures 1 and 2"},{"comment":"","section":"Appendix B.3, proof of Lemma 3"}],"minor_comments":[{"comment":"","section":"Section 4, eq. (10)"},{"comment":"","section":"Section 5.3 and Appendix D"},{"comment":"","section":"Theorem 2(iii)"},{"comment":"","section":"Section 6, simulations"},{"comment":"","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its conjectures, which is a positive feature, but the combination of an unproved singular-perturbation step and unproved landscape conjectures makes the headline claim substantially stronger than the proven content. The authors should be pushed to either supply the missing analysis for the finite-ε system or reframe the contribution as a convergence analysis of the multilevel reduced dynamics. If the latter route is taken, the paper would still be a solid contribution to the similarity matching literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about 2506.06134. First, the first two levels of the analysis, neural dynamics and lateral synaptic dynamics, are proved cleanly and are genuinely useful. Second, the paper's headline claim, that trajectories of the full coupled network converge to the global optimum, is not established. What is proved is convergence of three decoupled gradient flows in the singular limits, and no singular perturbation theorem connects those to the finite-ε system (7).\n\nThe real new content: a continuous-time three-time-scale version of the similarity matching network from Pehlevan et al., a formal proof of the lateral SPD invariance that the discrete-time literature had asserted without proof, and explicit expressions for the global minima of the slow cost along with an almost-sure convergence theorem for the slow flow. Theorems 1 and 2 are straightforward contraction-theory arguments, but they are correct and properly quantified. The SPD invariance proof is short and elegant. The slow-level analysis is honest: Theorem 3 is explicitly conditional on two conjectures, and the appendix gives numerical evidence for them. That is the right way to present an incomplete result.\n\nThe soft spot is structural. The paper never analyzes the coupled system (7). The reduction in Section 4 just says \"as ε1→0 ... the reduced slower subsystem is therefore ...\" and \"when ε2→0 the slow system becomes ...\". That is an informal time-scale argument, not a theorem. A proper singular perturbation or bounded-error tracking statement is missing, and it is not a cosmetic gap. The fast contraction rate νY = (4/T)λmin(M) depends on M, and M can approach the boundary of the SPD cone, so uniform contractivity over the trajectory is not shown. Likewise, Conjecture 1 asserts forward invariance of the full-rank set but not a uniform lower bound on σmin(W(t)), so the slow subsystem's Lipschitz constants are not controlled uniformly in time. For fixed small ε1, ε2, convergence of (7) remains open. The numerics in Section 6 support the conjecture but do not test the tracking bound that would close the gap.\n\nThat said, the paper is not sloppy. It explicitly labels the conjectures as conjectures, states the informal claim as informal, and ships code and data. The gap is in the advertised completeness, not in the internal logic.\n\nWho is this for? Researchers working on Hebbian/anti-Hebbian networks, normative models of dimensionality reduction, or multiple-time-scale neural dynamics. They will get real value from Theorems 1 and 2 and the SPD result, and a clear cautionary tale about time-scale reduction. It deserves a serious referee: the referee's job should be to require either a singular perturbation theorem, a bound on the tracking error, or a reframing of the claims to match what is proved. I would send it out.","headline":"A continuous-time similarity matching network with two solid early theorems and an honest but incomplete convergence story: the full coupled-network convergence is not proven because the time-scale reduction is informal.","tokens_in":29557,"tokens_out":2490,"would_cite":true,"duration_ms":24182,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["92B20","37N25","68T07","65K10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under two empirically supported conjectures, the three-timescale similarity matching network provably converges to the principal subspace projection of the input data, with the composed solution a global minimizer of the similarity…","keywords":["similarity matching","Hebbian learning","anti-Hebbian learning","principal subspace projection","multi-time-scale dynamics","contraction theory","gradient flow convergence","dimensionality reduction"],"falsifier":"Run the slow flow (26) from a full-rank $W_0$ and monitor the smallest singular value of $W(t)$; if it reaches zero at any finite time, Conjecture 1 fails and the dynamics is no longer defined. Alternatively, compare a trajectory of the full network (7) with fixed small $\\epsilon_1,\\epsilon_2$ to the corresponding reduced flow (26) over the same horizon: if the error does not vanish as $\\epsilon_1,\\epsilon_2\\to 0$, the time-scale reduction that carries the proof breaks down.","tokens_in":28534,"feed_emoji":"🧠","tokens_out":6611,"duration_ms":60841,"temperature":0.7,"pith_summary":"Similarity matching, which asks that pairwise similarities in a low-dimensional output match those of the input, is a biologically motivated way to do dimensionality reduction, but its neural implementations have lacked a full convergence proof. This paper supplies one for a continuous-time network whose fast neural activity, intermediate lateral synapses, and slow feedforward weights evolve on three separate time scales. Level by level, the paper shows a strongly convex problem with exponential convergence, a strongly concave problem with exponential convergence in the positive-definite cone, and a nonconvex slow problem whose gradient flow converges almost surely to a global minimizer under two technical conjectures. Composing the three optima yields the projection onto the principal subspace of the input covariance, the exact solution of the similarity matching problem. The result would close the gap between the biological plausibility of Hebbian and anti-Hebbian rules and rigorous guarantees for offline convergence.","feed_headline":"Hebbian/anti-Hebbian network provably finds the principal subspace","feed_subtitle":"Three time scales—fast neural, lateral, slow Hebbian—collapse onto the exact PCA projection, closing a proof gap in similarity matching.","key_machinery":"The machinery is the embedded min-max-min objective (6), which lifts the similarity matching problem into a space of neural activities $Y$, lateral weights $M$, and feedforward weights $W$, together with the three-level optimization scheme (12) that mirrors the three time scales. At the first two levels, strong convexity and strong concavity, combined with contraction theory, give global exponential convergence to uniquely defined optima. At the third level, the cost is nonconvex and nonsmooth; the proof characterizes stationary points via singular value decomposition, establishes coercivity, and invokes a strict-saddle argument to conclude almost-sure convergence to $W^\\star$ from random initialization. The load-bearing identities are $Y^\\star = M^{-1}WX$, $M^\\star = (W C_X W^\\top)^{1/3}$, and $W^\\star = U\\Lambda_C^m(V_C^m)^\\top$.","core_discovery":"The paper's central claim is that the three-timescale similarity matching network (7) solves the principal subspace projection problem. Formally, under Conjectures 1 and 2, the slow feedforward gradient-flow dynamics (26) converge almost surely from random full-rank initialization to $W^\\star = U\\Lambda_C^m (V_C^m)^\\top$, where $\\Lambda_C^m$ contains the $m$ largest eigenvalues of the input covariance $C_X$ and $V_C^m$ the corresponding eigenvectors. The neural and lateral dynamics converge, under the same separation, to $Y^\\star = M^{-1} W X$ and $M^\\star = (W C_X W^\\top)^{1/3}$. Substituting the converged values gives $Y^\\star(M^\\star,W^\\star) = U\\Sigma_X^m (V_X^m)^\\top$, which is a global minimizer of the similarity matching cost (3) and exactly the projection of the data onto the principal subspace. The paper also claims, informally, that the full coupled network inherits this convergence through the time-scale separation, a step it does not prove at finite $\\epsilon_1,\\epsilon_2$.","pith_inferences":["If the convergence result extends to finite time-scale parameters, the same network can serve as a local, biologically plausible PCA algorithm whose offline behavior is certified; the paper's simulations with $\\epsilon_1=0.01$, $\\epsilon_2=0.5$ suggest that the separation is not a serious obstacle in practice.","The three-level proof pattern, contractive fast layers paired with a strict-saddle slow layer, may transfer to other similarity-matching variants such as non-negative or sparse matching, though each variant would require its own convexity and saddle analysis.","The unresolved singular-perturbation step means that the formal guarantee applies to the reduced flows rather than directly to the full coupled system at finite $\\epsilon_1,\\epsilon_2$; a rigorous tracking theorem would complete the proof."],"forward_implications":["The fast neural dynamics globally exponentially converge to $Y^\\star = M^{-1}WX$, so the network's activity settles at the unique best output for the current weights.","The lateral dynamics preserve symmetry and positive definiteness and converge exponentially to $M^\\star = (W C_X W^\\top)^{1/3}$, giving the first formal proof of this invariance property.","The slow Hebbian weights converge almost surely to $W^\\star = U\\Lambda_C^m(V_C^m)^\\top$, whose singular vectors are the top $m$ eigenvectors of the input covariance.","Composed, these limits give $Y^\\star(M^\\star,W^\\star) = U\\Sigma_X^m(V_X^m)^\\top$, a global minimizer of the similarity matching problem, so the network provably computes a principal subspace projection rather than merely a stationary point.","Because the first two levels are contracting, the fast and intermediate dynamics are robust to noise and insensitive to initial conditions."],"supporting_citations":[{"why":"Supplies the offline algorithm and the min-max-min embedding of the similarity matching objective that the network is derived from.","marker":"[37]"},{"why":"Provides the characterization of global minimizers of the similarity matching problem, used as the target solution in Lemma 2.","marker":"[27]"},{"why":"Supplies the strict-saddle escape result, Corollary 4, used in Theorem 3 for almost-sure convergence from random initialization.","marker":"[12]"},{"why":"Provides the contraction-theory tools used to prove exponential convergence of the fast and intermediate dynamics.","marker":"[3]"},{"why":"Introduces the similarity matching cost and the Hebbian/anti-Hebbian framework underlying the network.","marker":"[22]"},{"why":"Presents the earlier Hebbian/anti-Hebbian network for linear subspace learning that the present model builds on.","marker":"[35]"}],"fun_headline_variants":["Three-timescale Hebbian network provably reaches PCA subspace","Hebbian learning nets: proof of PCA subspace convergence","Similarity matching tackles PCA with proven convergence","Fast neural, slow Hebbian: PCA proof via time-scale separation","Provable principal subspace from Hebbian anti-Hebbian learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof's load-bearing premise is that the three coupled dynamics can be reduced, in the limit where the time-scale parameters go to zero, to three sequential gradient flows, so that convergence of the reduced flows transfers to the full network; no singular-perturbation theorem is proved for finite small time-scale parameters.","fun_headline_variants_meta":{"raw":{"variants":["Three-timescale Hebbian network provably reaches PCA subspace","Hebbian learning nets: proof of PCA subspace convergence","Similarity matching tackles PCA with proven convergence","Fast neural, slow Hebbian: PCA proof via time-scale separation","Provable principal subspace from Hebbian anti-Hebbian learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1479,"prompt_tokens":1093,"completion_tokens":386,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":303}},"tokens_in":709,"tokens_out":386,"duration_ms":3972,"temperature":1.0,"reasoning_tokens":303,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T06:01:04.465471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the slow flow (26) from a full-rank $W_0$ and monitor the smallest singular value of $W(t)$; if it reaches zero at any finite time, Conjecture 1 fails and the dynamics is no longer defined. Alternatively, compare a trajectory of the full network (7) with fixed small $\\epsilon_1,\\epsilon_2$ to the corresponding reduced flow (26) over the same horizon: if the error does not vanish as $\\epsilon_1,\\epsilon_2\\to 0$, the time-scale reduction that carries the proof breaks down.","supporting_citations":[{"cited_title":"Pehlevan, A","cited_arxiv_id":null,"evidence_quote":"Supplies the offline algorithm and the min-max-min embedding of the similarity matching objective that the network is derived from."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the characterization of global minimizers of the similarity matching problem, used as the target solution in Lemma 2."},{"cited_title":"Convergence Analysis of Gradient Flow for Overparameterized LQR Formulations","cited_arxiv_id":"2408.15456","evidence_quote":"Supplies the strict-saddle escape result, Corollary 4, used in Theorem 3 for almost-sure convergence from random initialization."},{"cited_title":"Bullo.Contraction Theory for Dynamical Systems","cited_arxiv_id":null,"evidence_quote":"Provides the contraction-theory tools used to prove exponential convergence of the fast and intermediate dynamics."},{"cited_title":"Pehlevan, T","cited_arxiv_id":null,"evidence_quote":"Presents the earlier Hebbian/anti-Hebbian network for linear subspace learning that the present model builds on."}],"review_version":1}