{"id":"05797765-a47a-4d54-b9af-d408a5d80f65","arxiv_id":"2608.12597","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For quadratic loss wells, the paper derives an orientation-resolved formula predicting the random-subspace dimension where training succeeds, and shows it tracks measured transitions.","lead":"Neural networks can often be trained by optimizing only a small random projection of their parameters, but picking the projection size has been a guessing game. This paper derives a formula that predicts the needed size from the loss landscape's curvature and displacement, and wraps it in a memory-efficient training method.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ridge-leverage surrogate for E[P_W] is unproven for general spectra; a dominant-outlier spectrum could break the master-formula predictor.","rationale":"The reader's weakest_assumption correctly identifies Approximation 1 as the load-bearing step for the master-formula predictor. The leave-one-out derivation is transparent about the surrogate, but the paper provides no rigorous bounds for general spectra, only empirical support on specific synthetic and neural-curvature cases and exactness in the block model. Because the central claim is that bdMF predicts the empirical transition, the validity of E[P_W] ≈ H(H+κI)^{-1} is essential; if it fails for a realistic outlier-plus-bulk spectrum, the predictor's generality collapses. This is a genuine concern, but it is not a fatal flaw: the paper is honest about the heuristic status, the experiments are well designed, and the block-model exact case plus extensive empirical validation provide substantial support. The appropriate disposition remains conditional acceptance, matching the reader's verdict: the concern should be addressed by either deriving quantitative approximation bounds or testing the surrogate on deliberately adversarial spectra before the predictor is claimed as a generally reliable tool. No additional concern outweighs this one; the mean-level versus high-probability distinction and the localization issue are acknowledged and partially controlled in the paper, and the end-to-end transitions are explicitly protocol-dependent rather than presented as intrinsic constants.","tokens_in":40085,"tokens_out":4369,"duration_ms":51072,"concrete_test":"Run a Monte Carlo check of Approximation 1 on a deep-network-like spectrum: P=10^4, H=diag(10^6, 1×100, 0×9899) with d=10, and displacement Δ0=R·v_1 aligned with the dominant eigenvector. Compute the exact expected residual E[ρ(A)] from (B.1) by averaging over at least 10^4 Gaussian matrices A, and compare to bρMF(d; Δ0) from (B.4) with κ solving (B.3). If the relative error exceeds 10%, the surrogate fails in the outlier-plus-bulk regime; if the error is small, repeat for d near the rank of the bulk to probe the transition region. This isolates Approximation 1 from all other approximations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central predictor bρMF in Definition 1 rests on Approximation 1 (Appendix B.2): E[P_W] ≈ H(H+κI)^{-1}. The leave-one-out derivation approximates t_i ≈ τ uniformly across i, but for spectra with large eigenvalue disparity, tr(Q_{-i}^{-1}) can vary substantially with i. In particular, omitting the dominant eigenvector's row from Q changes the resolvent trace by a large amount, so for the stiffest direction t_1 can be much smaller than τ. The approximation is exact only in the block-constant model of Section B.5, where all stiff eigenvalues are equal; no quantitative error bound is given for general spectra. Since the paper's central empirical claim is that bdMF tracks the measured transition across spectra (Tables 2, 4, 5, 14), a spectrum where the surrogate is inaccurate would directly undermine the claimed generality. The controlled experiments support the surrogate in tested regimes, but its domain of validity is uncharacterized, leaving predictions for new architectures and spectral tails (e.g., the ViT-scale surrogate) uncertain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies random low-dimensional reparameterizations of neural networks, where a small latent vector z is mapped to a parameter update by a frozen random map. It first recasts the known accessibility transition for compact convex targets in conic form, centering it at the statistical dimension of the polar cone (Theorem 1/2). Its main claimed contribution is an orientation-resolved quadratic master formula bρMF(d; Δ0) = Σ_i κ(d) λ_i / (λ_i + κ(d)) Δ_i^2, with κ(d) fixed by the trace equation tr[H(H+κI)^{-1}] = d, predicting the expected random-slice residual from the curvature spectrum and the displacement profile. Two specializations are derived: an isotropic-orientation predictor and an orientation-uniform predictor that recovers the earlier radius-only bound of Larsen et al. The paper then introduces RaMaN, a framework using structured Hadamard or seed-regenerated Gaussian frozen maps to avoid O(dP) storage, and reports twelve experiments: controlled quadratic transitions, neural-curvature probes, end-to-end training transitions, map-family and optimizer/tolerance ablations, and a ViT-scale GGN–Ritz probe.","tokens_in":40325,"tokens_out":4657,"duration_ms":52203,"significance":"If the master formula is taken as a validated operational predictor, the paper makes a useful contribution: it turns a geometric transition into a computable dimension-selection rule and identifies displacement orientation, not just the Hessian spectrum, as a control parameter. The rigorous conic part (Lemma 1 and Theorem 2) is a correct application of known results, and the paper is commendably explicit about several limitations, notably stating that Experiment 12 is an internal-consistency test rather than independent validation. The controlled experiments (Tables 2, 4, 5) give strong in-model support for the predictor, and the RaMaN framework addresses a real memory bottleneck. However, the central predictor rests on an unproved approximation whose domain of validity is not characterized; this limits the strength of the paper's main claim until that gap is closed or the claim is appropriately narrowed.","major_comments":[{"comment":"The master formula bρMF(d; Δ0) in Definition 1 rests entirely on Approximation 1, E[P_W] ≈ H(H+κI)^{-1} with κ fixed by the trace equation (B.3). The derivation replaces the per-row quantities t_i = a_i^T Q_{-i}^{-1} a_i by a common τ, but this uniform approximation is not controlled. For a spectrum with a dominant outlier, tr(Q_{-i}^{-1}) varies substantially across i, so the stiffest direction can have t_1 far from τ, and the ridge-leverage surrogate can break down. No quantitative error bound is given outside the block-constant model of Section B.5, and the paper's own Experiment 12 is labeled an internal-consistency test. Since Tables 4, 5, and 14 use bdMF as the central predictor, I ask for either (i) a rigorous or quantitative bound on the approximation error under stated conditions (e.g., stable-rank and d/r regime), or (ii) an explicit stress-test experiment with spectra designed to challenge the surrogate, such as a dominant-outlier spectrum, reporting where the predicted critical dimension fails. Without one of these, the domain of validity of the central claim is uncharacterized.","section":"Appendix B.2, Approximation 1; Definition 1"},{"comment":"The paper presents the master formula as the main theoretical contribution, but the rigorous conic transition of Theorem 1 applies to the localized target S = Sε ∩ B(θ0, Rloc), whereas the derivation in Appendix B.1 analyzes the unlocalized quadratic residual and ignores Rloc. This gap is acknowledged in Remark 11, but it means that bdMF is not proven to approximate δ(C°), the quantity identified by the theorem. The experiments test the unlocalized least-squares residual rather than the localized conic intersection. I recommend closing the gap by stating explicitly throughout the abstract and introduction that bdMF is a heuristic surrogate for the conic threshold, and by reporting, for at least the quadratic experiments, the quantity ∥Az⋆∥2 relative to Rloc to justify that localization is inactive. As written, the connection between the rigorous theorem and the paper's flagship predictor is not fully closed.","section":"Theorem 1, Definition 1, Remark 11"}],"minor_comments":[{"comment":"The caption labels the √P curve a 'normalized reference' rather than a prediction, but a reader may misread it as a theoretical envelope; I suggest adding one sentence stating explicitly that no claim of O(√P) width is being made.","section":"Figure 3, middle panel"},{"comment":"The bert-tiny row lists P = 4,386,178 and then a separate row 'Prep = 413,314' that appears detached; align the reparameterized count with the model row or explain the relationship in the caption.","section":"Table 1"},{"comment":"The calibration parameters γ and b are introduced without any sensitivity analysis; a sentence on how the selected dimension depends on their choice, or a reference to an ablation if one exists, would help practitioners.","section":"Algorithm 1"},{"comment":"The 'resolved fraction' of 0.2–6% and the exact-Hessian-vector spot-check failure are material qualifications of the sub-percent agreement and should be repeated in the table caption or its immediately surrounding text, not only in the prose paragraph.","section":"Section 4.9, Table 14"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations and the experimental work is thorough. My main concern is whether a heuristic approximation, however well validated on a set of controlled spectra, is sufficient to carry the title-level claim of 'predicting when' random low-dimensional training works. If the authors can add a quantitative error analysis or an explicit failure characterization for Approximation 1, the paper would be substantially stronger; alternatively, the claims need to be carefully narrowed so that the central contribution is presented as a validated empirical predictor rather than a proven theoretical law."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. It's a real step forward in predicting when random low-dimensional reparameterization can train a network. The orientation-resolved master formula, rho_hat(d; Delta_0) = sum_i kappa(d) lambda_i / (lambda_i + kappa(d)) Delta_i^2, is new relative to Larsen et al.'s radius-only bound, and the paper demonstrates that displacement orientation can shift the transition by ~37x while spectrum and radius stay fixed. Theorem 2's conic transition is rigorous, and the controlled quadratic experiments are solid: median relative error 0.45% over 30 orientation-spectrum combos, with a clean solvable block model backing the main approximation. The RaMaN framework's memory savings are real, and the benchmark is thoughtfully designed, including honest separation of training accessibility from generalization. The paper is also unusually upfront about its own limitations.\n\nThe soft spot is exactly where the stress-test points: the predictor relies on Approximation 1 (Appendix B.2), E[P_W] ≈ H(H+kappa I)^{-1}, a finite-dimensional heuristic. The leave-one-out derivation sets t_i ≈ tau uniformly, which can fail when the spectrum has dominant outliers—removing the stiffest row changes the resolvent trace substantially. It's only exact in the block-constant model, and no quantitative error bound is given for general spectra. So the claim that d_hat_MF tracks the transition 'across spectra' is stronger than what's proven. That said, the experiments cover several spectra and the approximation holds numerically in all tested cases. This is a gap in characterized generality, not a demonstrated breakdown. At minimum, I'd want a bound under stable-rank conditions or an explicit description of regimes where the surrogate could be inaccurate, plus code/data for reproduction.\n\nProportionately, these are real caveats but not fatal. The paper is careful, honest, and the central argument holds up in its tested domain. It deserves a serious referee, and I'd take a revise-and-resubmit path. Worth citing if you work on intrinsic dimension, random subspace training, or parameter-efficient fine-tuning.","headline":"An honest, well-tested paper with a genuinely new orientation-resolved predictor; the unproved ridge-leverage surrogate is a real gap but not fatal—worth refereeing.","tokens_in":40838,"tokens_out":3326,"would_cite":true,"duration_ms":34512,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the latent dimension at which random low-dimensional reparameterization can reach a low-loss region is not an empirical accident: in a quadratic model it is set by a master formula combining the curvature spectrum…","keywords":["random low-dimensional reparameterization","accessibility transition","orientation-resolved master formula","statistical dimension","latent dimension selection","structured Hadamard mapping","seed-regenerated Gaussian maps","random subspace training"],"falsifier":"In a controlled quadratic experiment with $P=1024$ and a slowly decaying spectrum such as $\\lambda_i \\propto i^{-1}$, take a displacement of fixed norm aligned with a mid-curvature eigenvector, run many random Gaussian slices, and compare the measured $d_{50}$—the dimension at which 50% of slices hit the $\\varepsilon$-sublevel set—with the master-formula threshold $\\hat d_{\\mathrm{MF}}$. If the discrepancy exceeds the sharpened conic window $O(\\sqrt{d_{\\mathrm{conic}} \\log(1/\\eta)})$ or a few percent relative error, the ridge-leverage surrogate is falsified for that spectrum.","tokens_in":39871,"feed_emoji":"🎯","tokens_out":5599,"duration_ms":50765,"temperature":0.7,"pith_summary":"The paper claims that the latent dimension at which random low-dimensional reparameterization can reach a low-loss region is predictable from local curvature and displacement geometry, rather than discoverable only by sweeping candidate dimensions. It recasts the known random-slice accessibility transition as a conic phase transition centered at the statistical dimension of a polar cone, then derives an orientation-resolved quadratic master formula for the expected residual of a random slice. If the claim is right, practitioners can choose a latent dimension prospectively from measurable quantities. The paper also builds scalable random mappings at the predicted dimension, cutting frozen-map storage and optimizer-state memory, and reports sharp, protocol-dependent training transitions across image and language models.","feed_headline":"Latent dimension for random-map training is now predictable","feed_subtitle":"A curvature-and-orientation master formula locates the random-slice transition; RaMaN picks the latent size without a sweep.","key_machinery":"The orientation-resolved master formula $\\hat\\rho_{\\mathrm{MF}}(d; \\Delta_0) = \\sum_i \\frac{\\kappa(d)\\lambda_i}{\\lambda_i+\\kappa(d)} \\Delta_i^2$, with $\\kappa(d)$ fixed by $d = \\operatorname{tr}[H(H+\\kappa I)^{-1}]$, is the central object. It acts as a ridge-leverage surrogate for the expected projection onto the random slice's curvature image, turning an intractable geometric threshold into a computable mean-level predictor. The rigorous conic theorem supplies the transition center as the statistical dimension of the polar cone, $d_{\\mathrm{conic}} = \\delta(C^\\circ)$; the master formula is the operational replacement that uses only the curvature spectrum and the displacement profile.","core_discovery":"For a localized quadratic loss with positive-semidefinite curvature $H$ and displacement $\\Delta_0$ from the reference point to the minimizer, the expected residual of a random $d$-dimensional slice is predicted as $\\hat\\rho_{\\mathrm{MF}}(d; \\Delta_0) = \\sum_i \\frac{\\kappa(d)\\lambda_i}{\\lambda_i+\\kappa(d)} \\Delta_i^2$, where $\\kappa(d)$ solves $d = \\operatorname{tr}[H(H+\\kappa I)^{-1}]$. The predicted critical dimension is the smallest $d$ with $\\hat\\rho_{\\mathrm{MF}}(d; \\Delta_0) \\le \\varepsilon_q = 2\\varepsilon$. A key consequence is that two displacements with the same spectrum and same norm can require different latent dimensions when their energy is distributed differently across stiff versus flat curvature directions. The paper reports that this predictor tracks measured quadratic transitions with median relative error about 0.45% in synthetic orientation sweeps and at most 0.91% on neural-curvature probes, while an isotropic-orientation specialization is sharply correct under equal-energy displacement and a radius-only specialization is conservative, recovering the earlier quadratic bound.","pith_inferences":["The same master formula could be used to set ranks prospectively in parameter-efficient fine-tuning methods that share frozen random matrices, by treating the rank as the predicted latent dimension for a Fisher or GGN curvature surrogate.","Because the equal-tail approximation depends on the unresolved spectrum mostly through its sum, coarse spectral estimates with a few leading Ritz pairs plus a trace may suffice for dimension selection; this is testable by comparing predictions made with partial spectra against full-spectrum predictions.","A natural extension is to predict $d_{\\mathrm{gen}}$, the dimension at which generalization approaches the full-parameter model, by applying the same residual analysis to a data-dependent curvature surrogate; the paper separates $d_{\\mathrm{train}}$ from $d_{\\mathrm{gen}}$ but does not predict the latter.","For nonconvex losses, the one-sided conic bounds suggest that miss probabilities survive through Gaussian width even when the full transition is not rigorous; a testable extension is whether $\\hat d_{\\mathrm{MF}}$ still predicts the empirical midpoint when monitored least-squares displacements stay within the localization ball."],"forward_implications":["A latent dimension can be selected without sweeping: given leading curvature eigenvalues and a displacement estimate, $\\hat d_{\\mathrm{MF}}$ gives an operating dimension for random low-dimensional training.","Displacement orientation is load-bearing: two tasks with the same Hessian spectrum and the same distance to the solution can differ in required latent dimension, empirically by up to a factor of 37 in the controlled orientation sweeps.","When only a displacement radius is known, the orientation-uniform predictor $r_{\\mathrm{eff}}(\\varepsilon, R) = \\sum_i \\frac{\\lambda_i R^2}{\\lambda_i R^2 + \\varepsilon_q}$ is conservative and reproduces the earlier radius-only quadratic bound.","Matrix-free implementations with seeded Gaussian maps or structured Hadamard mappings remove $O(dP)$ frozen-map storage, reduce optimizer-state memory to $O(d)$, and still exhibit sharp end-to-end training transitions across MLPs, CNNs, transformers, ResNets, and a pretrained language model.","The predicted transition is a training-accessibility threshold, not a generalization guarantee: end-to-end midpoints depend on the optimizer, learning-rate protocol, and loss tolerance, and crossing the transition does not by itself ensure full-model predictive performance."],"supporting_citations":[{"why":"Supplies the conic integral-geometry phase-transition results and statistical-dimension framework behind Theorem 2.","marker":"[3]"},{"why":"Gives the Gaussian-width geometric transition and the radius-only quadratic bound that the orientation-uniform specialization recovers.","marker":"[11]"},{"why":"Provides the leave-one-out resolvent and deterministic-equivalent argument underlying the ridge-leverage surrogate for the projector expectation.","marker":"[4]"},{"why":"Reports the empirical sharp transition and intrinsic-dimension experiments that motivate predicting the transition location.","marker":"[14]"},{"why":"Shows pretrained language models adapt in very low-dimensional random subspaces, the practical setting the predictors target.","marker":"[1]"},{"why":"Supplies the fast structured projection ideas that the SHM construction adapts for matrix-free maps.","marker":"[12]"},{"why":"Provides the escape-through-a-mesh bound used for one-sided statements on nonconvex sublevel sets.","marker":"[6]"}],"fun_headline_variants":["Random-map training dimension now predictable via curvature","RaMaN's formula predicts latent size for neural nets","Orientation matters: predictor for random slice training","Quadratic master formula locates random-map transition","Latent dimension for random maps: a predictive theory"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole quadratic predictor rests on the ridge-leverage surrogate $\\mathbb{E}[P_W] \\approx H(H+\\kappa I)^{-1}$, which is asserted as a finite-dimensional heuristic with no proven accuracy bounds for general spectra; if that surrogate fails for a given spectrum, the predicted critical dimension $\\hat d_{\\mathrm{MF}}$ is unreliable even though the conic transition theorem itself is rigorous.","fun_headline_variants_meta":{"raw":{"variants":["Random-map training dimension now predictable via curvature","RaMaN's formula predicts latent size for neural nets","Orientation matters: predictor for random slice training","Quadratic master formula locates random-map transition","Latent dimension for random maps: a predictive theory"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000628,"raw_usage":{"total_tokens":2945,"prompt_tokens":1030,"completion_tokens":1915,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":1842}},"tokens_in":646,"tokens_out":1915,"duration_ms":13390,"temperature":1.0,"reasoning_tokens":1842,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:04:10.670215+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In a controlled quadratic experiment with $P=1024$ and a slowly decaying spectrum such as $\\lambda_i \\propto i^{-1}$, take a displacement of fixed norm aligned with a mid-curvature eigenvector, run many random Gaussian slices, and compare the measured $d_{50}$—the dimension at which 50% of slices hit the $\\varepsilon$-sublevel set—with the master-formula threshold $\\hat d_{\\mathrm{MF}}$. If the discrepancy exceeds the sharpened conic window $O(\\sqrt{d_{\\mathrm{conic}} \\log(1/\\eta)})$ or a few percent relative error, the ridge-leverage surrogate is falsified for that spectrum.","supporting_citations":[{"cited_title":"How Many Degrees of Freedom Do We Need to Train Deep Networks: A Loss Landscape Perspective","cited_arxiv_id":null,"evidence_quote":"Gives the Gaussian-width geometric transition and the radius-only quadratic bound that the orientation-uniform specialization recovers."},{"cited_title":"Exact expressions for double descent and implicit regularization via surrogate random design","cited_arxiv_id":null,"evidence_quote":"Provides the leave-one-out resolvent and deterministic-equivalent argument underlying the ridge-leverage surrogate for the projector expectation."},{"cited_title":"Measuring the intrinsic dimension of objective landscapes","cited_arxiv_id":null,"evidence_quote":"Reports the empirical sharp transition and intrinsic-dimension experiments that motivate predicting the transition location."},{"cited_title":"Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning","cited_arxiv_id":null,"evidence_quote":"Shows pretrained language models adapt in very low-dimensional random subspaces, the practical setting the predictors target."},{"cited_title":"Fastfood - Computing Hilbert Space Ex- pansions in Loglinear Time","cited_arxiv_id":null,"evidence_quote":"Supplies the fast structured projection ideas that the SHM construction adapts for matrix-free maps."},{"cited_title":"On Milman’s Inequality and Random Subspaces which Escape through a Mesh in Rn","cited_arxiv_id":null,"evidence_quote":"Provides the escape-through-a-mesh bound used for one-sided statements on nonconvex sublevel sets."}],"review_version":1}