DynMuon dynamically schedules the spectral exponent p in Muon-style updates according to curvature, noise, and training stage, yielding lower validation loss with 10-26% fewer steps than fixed Muon.
Canonical reference
Springer, 2006
Canonical reference. 83% of citing Pith papers cite this work as background.
citation-role summary
citation-polarity summary
representative citing papers
Spectra defines and controls effective capacity in graph embeddings via the Shannon effective rank of a trace-normalized kernel spectrum, making capacity a post-fit property rather than a pre-training hyperparameter.
A framework learns constitutive priors from noisy data to enable PDE-constrained inverse design of elastic networks using latent variables, homotopy continuation, Chamfer distance matching, and neural smoothness constraints.
AS-LoRA adaptively chooses which LoRA factor to update per layer and round using a curvature-aware second-order score, eliminating reconstruction error floors and improving performance in DP federated learning.
Joint-DP co-optimizes sensing geometry with Bellman-optimal adaptive policies via differentiable dynamic programming and relaxations, scaling to photonic designs exceeding 10^5 pixels.
KSOS-BO improves acquisition function optimization in Bayesian optimization by casting it as a kernel sum of squares semidefinite program, outperforming Sobol, DE, and CMA-ES baselines on 10/15 benchmarks with 81% average regret reduction.
Gauss-Newton descent whitens errors by projecting Newton directions or gradients onto the tangent space, replacing JJ^T with the identity and removing parameterization distortions that affect Newton descent.
A new regularized Hessian-free Newton-type method for smooth convex optimization achieves global O(k^{-2}) convergence and local quadratic convergence in a variant, with practical speedups over prior methods.
Quantum probes with squeezing or entanglement, routed via an identifiability-guaranteeing algorithm, enable improved estimation of link transmissivities in general optical networks using det(FIM) and tr(FIM^{-1}) metrics.
Fisher Decorator refines flow policies in offline RL via a local transport map and Fisher-matrix quadratic approximation of the KL constraint, yielding controllable error near the optimum and SOTA benchmark results.
Swarm-based inertial methods derived from the Onsager principle achieve provable O(1/δ(t)) asymptotic convergence rates for global minimization with structure-preserving discretizations.
AdamFLIP treats PDE constraint residuals in PINNs as a controlled dynamical system, computes Lagrange multipliers via feedback linearization to drive residuals to zero, and applies Adam-style adaptation to the resulting gradient for scalable hard-constrained training.
MDPO improves differentiable planning by injecting gradient-sensitivity-adapted noise into the action space, outperforming both deterministic variants and PPO on nonlinear and hybrid benchmarks.
PINNACLE is an open-source framework for classical and quantum PINNs that supplies modular training methods and benchmarks showing high sensitivity to architecture choices plus parameter-efficiency gains in some hybrid quantum regimes.
A research roadmap analyzing the current state of search-based software engineering with foundation models, outlining challenges and directions across three integration aspects.
citing papers explorer
-
DynMuon: A Dynamic Spectral Shaping View of Muon
DynMuon dynamically schedules the spectral exponent p in Muon-style updates according to curvature, noise, and training stage, yielding lower validation loss with 10-26% fewer steps than fixed Muon.
-
Rank Is Not Capacity: Spectral Occupancy for Latent Graph Models
Spectra defines and controls effective capacity in graph embeddings via the Shannon effective rank of a trace-normalized kernel spectrum, making capacity a post-fit property rather than a pre-training hyperparameter.
-
Constitutive Priors for Inverse Design
A framework learns constitutive priors from noisy data to enable PDE-constrained inverse design of elastic networks using latent variables, homotopy continuation, Chamfer distance matching, and neural smoothness constraints.
-
Adaptive Selection of LoRA Components in Privacy-Preserving Federated Learning
AS-LoRA adaptively chooses which LoRA factor to update per layer and round using a curvature-aware second-order score, eliminating reconstruction error floors and improving performance in DP federated learning.
-
Adaptive Sensing beyond Non-Adaptive Information Limits: End-to-End Co-Design of Geometry, Policy, and Inference
Joint-DP co-optimizes sensing geometry with Bellman-optimal adaptive policies via differentiable dynamic programming and relaxations, scaling to photonic designs exceeding 10^5 pixels.
-
KSOS-BO: Improving Sampling in Bayesian Optimization via Kernel Sum of Squares
KSOS-BO improves acquisition function optimization in Bayesian optimization by casting it as a kernel sum of squares semidefinite program, outperforming Sobol, DE, and CMA-ES baselines on 10/15 benchmarks with 81% average regret reduction.
-
Error whitening: Why Gauss-Newton outperforms Newton
Gauss-Newton descent whitens errors by projecting Newton directions or gradients onto the tangent space, replacing JJ^T with the identity and removing parameterization distortions that affect Newton descent.
-
A Regularized Hessian-Free Inexact Newton-Type Method with Global $\mathcal{O}(k^{-2})$ Convergence
A new regularized Hessian-free Newton-type method for smooth convex optimization achieves global O(k^{-2}) convergence and local quadratic convergence in a variant, with practical speedups over prior methods.
-
Quantum-enhanced Network Tomography
Quantum probes with squeezing or entanglement, routed via an identifiability-guaranteeing algorithm, enable improved estimation of link transmissivities in general optical networks using det(FIM) and tr(FIM^{-1}) metrics.
-
Fisher Decorator: Refining Flow Policy via a Local Transport Map
Fisher Decorator refines flow policies in offline RL via a local transport map and Fisher-matrix quadratic approximation of the KL constraint, yielding controllable error near the optimum and SOTA benchmark results.
-
Swarm-Based Inertial Methods for Optimization
Swarm-based inertial methods derived from the Onsager principle achieve provable O(1/δ(t)) asymptotic convergence rates for global minimization with structure-preserving discretizations.
-
AdamFLIP: Adaptive Momentum Feedback Linearization Optimization for Hard Constrained PINN Training
AdamFLIP treats PDE constraint residuals in PINNs as a controlled dynamical system, computes Lagrange multipliers via feedback linearization to drive residuals to zero, and applies Adam-style adaptation to the resulting gradient for scalable hard-constrained training.
-
Model-Driven Policy Optimization in Differentiable Simulators via Stochastic Exploration
MDPO improves differentiable planning by injecting gradient-sensitivity-adapted noise into the action space, outperforming both deterministic variants and PPO on nonlinear and hybrid benchmarks.
-
PINNACLE: An Open-Source Computational Framework for Classical and Quantum PINNs
PINNACLE is an open-source framework for classical and quantum PINNs that supplies modular training methods and benchmarks showing high sensitivity to architecture choices plus parameter-efficiency gains in some hybrid quantum regimes.
-
Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap
A research roadmap analyzing the current state of search-based software engineering with foundation models, outlining challenges and directions across three integration aspects.