HORST uses non-commutative operator composition and a hyperbolic mirror map to combine stability from adaptive optimizers with L1 sparsity bias, outperforming AdamW across sparsity levels on vision and language tasks.
Proceedings of the 38th International Conference on Machine Learning , pages =
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2representative citing papers
Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.
citing papers explorer
-
HORST: Composing Optimizer Geometries for Sparse Transformer Training
HORST uses non-commutative operator composition and a hyperbolic mirror map to combine stability from adaptive optimizers with L1 sparsity bias, outperforming AdamW across sparsity levels on vision and language tasks.
-
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.