{"id":"4bb83f78-2215-4f7b-8068-d3a006b85c67","arxiv_id":"2607.08843","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Exact solutions show abstraction is set by input/target geometry, rises with depth, peaks under small init, and is attenuated by nonlinearities—improving LLM probes via GELU ablation.","lead":"This paper derives exact training-time trajectories for how concept directions align in simple neural networks, and proves that nonlinearities weaken that alignment in activations versus preactivations. The results give concrete rules for depth, initialization, and probe design that improve concept extraction in open models.","discovery_kind":"first_principles","skeptic_critique":{"model":"grok-4.5","headline":"The attenuation law and GELU-ablation application rest on infinite-width 2FS NNGP maps; finite-width residual streams may not inherit the same attenuation, so the probe-generalization claim is under-supported.","rationale":"The reader correctly flags Assump. 2 as the price of exact scalar ODEs for the linear theory; that concern is real but already scoped by the paper’s theorems. The more load-bearing issue for the paper’s strongest applied claim is the gap between the infinite-width attenuation proof (Thm. 8 / C.2) and the finite-width transformer interventions. The linear terminal/depth/init laws are clean under stated axioms and have qualitative CNN support (Fig. 3). The nonlinear attenuation result is mathematically careful for erf/L-ReLU NNGPs, but Fig. 5 and the LLM probe improvement are presented as applications of that law without verifying the pre/post inequality or 2FS structure in residual streams. A single pre/post measurement on the existing evaluation sets would settle whether the theory actually underwrites the intervention. Verdict remains CONDITIONAL; confidence in the applied claim should be lower than in the linear exact solutions.","tokens_in":61462,"tokens_out":692,"duration_ms":7024,"concrete_test":"On the same Gemma 4 bilingual pairs and DINOv3 3dshapes subset used in Fig. 5, extract both pre-activation (MLP input) and post-activation (MLP output) residual-stream features at the ablated layer, compute αQ vs αK without any ablation, and test whether |αK| ≤ |αQ| holds for ≥15/18 concepts (or every ViT layer). If the inequality fails systematically, the attenuation-law justification for GELU ablation does not transfer and the probe-generalization claim should be reclassified as empirical only.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorem 8 (and formal Theorem C.2) proves |αK| ≤ |αQ| only for infinite-width Gaussian preactivations whose kernel is exactly 2FS, with the feature kernel given by the NNGP maps ψeta/ψω of Eq. (15). The central applied claim—that local GELU ablation improves abstraction and probe transfer in DINOv3/Gemma 4 (Fig. 5, §5)—treats residual-stream activations as if they obey the same instantaneous attenuation. Residual streams are finite-width, non-Gaussian, multi-layer, and not 2FS; GELU is also not among the two nonlinearities for which the secant/power-series conditions of §C.5 are verified. Thus the theory-to-application bridge for the strongest practical claim is an unproven extrapolation rather than a consequence of Theorem 8. The linear exact laws (Theorems 4, 7) remain internally solid under Assumps. 1–2; the load-bearing soft spot is the nonlinear/empirical leap.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper develops a dynamical theory of “abstraction” (cosine alignment of context-specific concept vectors) under the linear representation hypothesis. In a two-layer linear network with variable-projected readout and two-factor symmetry (2FS) kernels, the authors reduce gradient flow to scalar mode ODEs and obtain an exact implicit trajectory for abstraction (Theorem 3). From this they derive a terminal law α∞ determined by the geometric mean of input and target inverse-SNRs (Theorem 4), overshoot and initialization-scale control of peak abstraction (Theorem 5, Proposition 6), and depth laws for layerwise interpolation and terminal abstraction under balanced deep linear networks (Theorem 7). They extend to infinite-width two-layer erf and leaky-ReLU networks, prove an attenuation law |αK| ≤ |αQ| for features vs preactivations (Theorem 8 / C.2), and characterize how ReLU terminal abstraction depends more on input than target geometry. Empirical checks include ResNets on 3dshapes, local GELU ablation in DINOv3 and Gemma 4, and macaque V4 vs IT.","tokens_in":61845,"tokens_out":1192,"duration_ms":16614,"significance":"If the linear results hold under the stated axioms, they give a rare closed-form account of how linear concept geometry evolves during training—not only at convergence—linking data/target geometry, depth, and rich/lazy initialization to abstraction trajectories. That is a clear advance over asymptotic perfect-abstraction results. The attenuation law and ReLU vs erf phase diagrams offer a mechanistic explanation for known nonlinearity effects on abstract geometry. Strengths include exact solutions (Riccati reduction, eigenvalue ODEs, Theorems 3–7), careful infinite-width nonlinear maps and proofs (Appendix C), and falsifiable qualitative predictions tested on ResNets, open transformers, and neural data. The applied GELU-ablation result is modest but practically relevant for probing/steering if the effect is robust.","major_comments":[{"comment":"§5 and abstract claim evidence for the attenuation law and improved probe generalization via local GELU ablation (Fig. 5). Theorem 8 / Theorem C.2 only prove |αK| ≤ |αQ| for infinite-width Gaussian preactivations with exact 2FS kernels under erf or leaky ReLU (Eq. 15; §C.5). Residual streams in DINOv3/Gemma are finite-width, multi-layer, non-Gaussian, and not 2FS; GELU is not among the verified nonlinearities. The manuscript should explicitly separate the theorem’s hypotheses from the empirical intervention, avoid language that presents Fig. 5 as a consequence of Theorem 8, and add controls (e.g., ablating other MLP components, random linear maps, or non-concept baselines) so the ~1pp probe gain is not over-attributed to the attenuation mechanism.","section":null},{"comment":"Assumption 2 (2FS; §2.2, §A.3) is load-bearing for every exact scalar ODE and closed-form law (Proposition 2, Theorems 3–7). It is imposed via group invariance rather than derived for realistic data. The ResNet/3dshapes checks (§3.3, Fig. 3) are qualitative and still use a 2×2 factor design. The paper should state more sharply which predictions are expected to survive without 2FS (e.g., qualitative depth/init trends) versus which are 2FS-specific (exact α∞ formula, arctanh layerwise interpolation), and ideally report a controlled symmetry-breaking experiment (e.g., unbalanced class sizes or non-orthogonal Fourier components) to bound sensitivity.","section":null}],"minor_comments":[{"comment":"Assump. 1 (variable-projected readout) is better motivated than a frozen readout, and §A.4/Fig. 7 help, but early-training discrepancies should be flagged when interpreting non-monotonic α(t) near initialization.","section":null},{"comment":"Setting 2 (layerwise balancing) for depth results (§3.2, §B.8) is standard but strong; a short remark on how unbalanced init or SGD noise would perturb Eq. (13)–(14) would help readers.","section":null},{"comment":"Proposition 1’s Gaussian score approximation is used to link α to probe transfer; state when the approximation fails (heavy-tailed residual streams, multi-token concepts).","section":null},{"comment":"Fig. 1E / 2FS entry notation (ad, a2, a1s, a1c, a0) is dense; a small table mapping entries to modes S/SC would improve readability.","section":null},{"comment":"Related work on Word2Vec analogies and LRH is good; a one-sentence contrast with Jiang et al. [4] on dynamics vs asymptotics in the main text (not only appendix) would help orientation.","section":null},{"comment":"Gemma results: report confidence intervals or paired tests for the 17/18 probe improvements; French–Spanish decrease should be discussed briefly.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Core linear theory is high quality and suitable for a strong ML theory venue. The main risk is overselling the transformer application relative to Theorem 8’s hypotheses; if the authors tighten that bridge and clarify 2FS scope, I would support acceptance. Scope fits cs.LG / theory of deep learning well."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that they actually solve the training-time trajectory of concept alignment, not just the end point. Under a two-factor symmetry (2FS) kernel and a variable-projected readout, they get closed-form eigenvalue ODEs, an exact implicit solution for α(t), a terminal law α∞ = (1 − √(νx νy))/(1 + √(νx νy)), depth interpolation in arctanh space, and rich-limit metastability of near-perfect abstraction. That package is new relative to Jiang et al. and Wang et al., who mostly treat existence or global optima.\n\nThe linear math is careful and complete under the stated axioms. The Riccati reduction, spectral filtering by M(Q), and the proofs of Theorems 3–7 and the residence-time result look solid. The infinite-width nonlinear extension is also done properly: NNGP maps preserve 2FS, erf stays close to the linear terminal law, ReLU is shown to depend more on input than target geometry, and the attenuation law |αK| ≤ |αQ| is proved for erf and leaky ReLU via a clean secant/power-series argument. Qualitative checks on small ResNets, DINOv3 layerwise trends, and macaque IT > V4 line up with the depth prediction.\n\nSoft spots, in proportion. 2FS and balanced depth are modeling choices that buy solvability; without them the scalar ODEs disappear. That is fine if you treat the results as exact laws inside a minimal model, not as universal theorems. The bigger gap is the applied claim. Theorem 8 is infinite-width, Gaussian, exactly 2FS, and only verified for erf/leaky-ReLU. Local GELU ablation in finite-width residual streams of DINOv3/Gemma is an empirical intervention that works in their figures, but it is not a corollary of the theorem. The probe-generalization gains are small (~1 pp) and the theory-to-practice bridge is an extrapolation. Still, the linear core does not depend on that leap.\n\nThis is for people who care about learning dynamics of representation geometry, LRH foundations, or hierarchy predictions in systems neuroscience. It deserves a serious referee. I would engage with the linear laws and the attenuation idea; I would not treat the GELU recipe as theory-backed until someone checks finite-width residual streams more carefully.","headline":"Exact abstraction trajectories under 2FS linear nets are the real contribution; the GELU-ablation application is a useful but under-proved leap from the infinite-width attenuation law.","tokens_in":62418,"tokens_out":587,"would_cite":true,"duration_ms":8692,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Exact solutions show how concept directions align during training, not just after it.","keywords":["linear representation hypothesis","abstraction dynamics","deep linear networks","kernel Riccati dynamics","attenuation law","linear probes","concept geometry","infinite-width NNGP"],"falsifier":"Train a deep ReLU or transformer stack on data that clearly violates two-factor symmetry and check whether terminal abstraction still tracks the geometric-mean formula, whether deeper layers are more abstract, and whether ablating the local nonlinearity still raises feature abstraction and probe transfer.","tokens_in":62374,"feed_emoji":"↗️","tokens_out":566,"duration_ms":6965,"temperature":0.7,"pith_summary":"This paper asks how linear concept directions form while a network trains, not merely whether they exist at the end. In a minimal linear network with two binary factors, the authors solve the full trajectory of abstraction—the cosine alignment of concept vectors across contexts. Terminal abstraction is fixed by the geometric mean of input and target noise ratios; depth improves it when targets are cleaner than inputs; small initialization lets abstraction overshoot and stay near-perfect for a long time. Nonlinearities (erf, ReLU) change the dynamics and, at any fixed time, weaken abstraction in post-activation features relative to preactivations. Evidence in vision and language models, plus a simple local ablation that improves probe transfer, shows the laws still matter outside the solvable setting.","feed_headline":"How concept directions line up during training, solved exactly","feed_subtitle":"Terminal abstraction is set by data and targets; depth and init scale control the path","key_machinery":"Abstraction score α: the cosine similarity of a concept direction measured in two contexts, rewritten as a function of the signal-to-noise ratio of two eigenmodes (S versus SC) of a five-entry two-factor-symmetric kernel. Exact scalar ODEs for those modes yield the terminal law, depth interpolation, and attenuation factor.","core_discovery":"Under a minimal linear network with two-factor symmetry and optimal readout, abstraction has an exact closed-form trajectory. Its terminal value is set only by the geometric mean of input and target inverse signal-to-noise ratios; depth and initialization separately control layerwise growth and peak overshoot. In infinite-width nonlinear nets the same geometry is reshaped by the nonlinearity, and both erf and leaky ReLU attenuate abstraction so that feature abstraction never exceeds preactivation abstraction.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Exact closed-form trajectory of concept direction alignment","Data-target geometry sets terminal abstraction; depth grows it","Init scale controls peak overshoot of abstraction during training","Nonlinearity attenuates abstraction from preacts to features","Linear theory of abstraction yields exact dynamics under symmetry"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"The whole exact theory needs the data, targets, and initial features to obey a strong two-factor symmetry so that every kernel collapses to five numbers and the same five modes; without that symmetry the closed-form trajectory disappears.","fun_headline_variants_meta":{"raw":{"variants":["Exact closed-form trajectory of concept direction alignment","Data-target geometry sets terminal abstraction; depth grows it","Init scale controls peak overshoot of abstraction during training","Nonlinearity attenuates abstraction from preacts to features","Linear theory of abstraction yields exact dynamics under symmetry"]},"model":"grok-4.5","effort":"low","cost_usd":0.004778,"raw_usage":{"total_tokens":1413,"prompt_tokens":829,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":47780000,"prompt_tokens_details":{"text_tokens":829,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":507,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":829,"tokens_out":77,"duration_ms":4422,"temperature":1.0,"reasoning_tokens":507,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T06:18:33.353821+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train a deep ReLU or transformer stack on data that clearly violates two-factor symmetry and check whether terminal abstraction still tracks the geometric-mean formula, whether deeper layers are more abstract, and whether ablating the local nonlinearity still raises feature abstraction and probe transfer.","supporting_citations":[],"review_version":1}