Riemannian networks are introduced for the full-rank correlation matrix manifold by extending MLR, FC, and convolutional layers to five geometries with backpropagation methods for two, showing effectiveness over SPD and Grassmannian baselines.
hub
The annals of mathematical statistics , pages=
13 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
HPO enables unbiased policy optimization in hybrid action spaces by mixing differentiable simulation gradients with score-function estimates, outperforming PPO as continuous dimensions increase.
A convex data-driven inverse RL framework for linear systems with uncertainty that uses a generalized LQR cost with cross terms, kernel regression from data, and differentiable SDPs for robust cost design over perturbations.
LLM.int8() performs 8-bit inference for transformers up to 175B parameters with no accuracy loss by combining vector-wise quantization for most features with 16-bit mixed-precision handling of systematic outlier dimensions.
Neural networks trained on scalar diffusion data predict adaptive coarse basis functions for Schwarz methods, transferring without retraining to linear elasticity and nonlinear p-Laplace problems.
SGD on multiclass cross-entropy loss alternates between curvature-driven oscillations and stable regimes but self-stabilizes to enable best-iterate convergence with large learning rates for linear and two-layer models.
SGD is reformulated via a master equation from discrete updates, producing a discrete Fokker-Planck equation that predicts non-stationary variance growth proportional to learning rate in flat Hessian directions.
Spectral decomposition of the logit Jacobian yields an adaptive MSA with linear convergence and a tractable Newton method for path-based SUE, with reported speedups on networks up to Chicago Regional size.
New Berry-Esseen bounds for multivariate martingale difference sequences achieve n^{-1/4} rate and polylog(d) dimension dependence in Kolmogorov distance.
AdamO modifies Adam with an orthogonality correction to ensure the spectral radius of the TD update operator stays below one, providing a theoretical stability guarantee for offline RL.
A recursive cubing framework identifies stable hyperparameter regions for MC dropout uncertainty quantification in spatial deep learning and produces competitive or superior predictive intervals versus a statistical baseline on simulations and land-surface temperature data.
A survey equating offline Monte-Carlo/SAA and online stochastic-approximation sample complexities for convex stochastic optimization arising in statistics and ML.
citing papers explorer
-
Berry-Esseen bounds for multivariate martingale difference sequences in the Kolmogorov distance
New Berry-Esseen bounds for multivariate martingale difference sequences achieve n^{-1/4} rate and polylog(d) dimension dependence in Kolmogorov distance.