Pith. sign in

REVIEW 8 cited by

Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.01113 v1 pith:NQNC2HU6 submitted 2020-02-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords cayleymanifoldoptimizationstiefeladamalgorithmstransformconvergence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Strictly enforcing orthonormality constraints on parameter matrices has been shown advantageous in deep learning. This amounts to Riemannian optimization on the Stiefel manifold, which, however, is computationally expensive. To address this challenge, we present two main contributions: (1) A new efficient retraction map based on an iterative Cayley transform for optimization updates, and (2) An implicit vector transport mechanism based on the combination of a projection of the momentum and the Cayley transform on the Stiefel manifold. We specify two new optimization algorithms: Cayley SGD with momentum, and Cayley ADAM on the Stiefel manifold. Convergence of Cayley SGD is theoretically analyzed. Our experiments for CNN training demonstrate that both algorithms: (a) Use less running time per iteration relative to existing approaches that enforce orthonormality of CNN parameters; and (b) Achieve faster convergence rates than the baseline SGD and ADAM algorithms without compromising the performance of the CNN. Cayley SGD and Cayley ADAM are also shown to reduce the training time for optimizing the unitary transition matrices in RNNs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization

    cs.DC 2026-06 unverdicted novelty 7.0 of 10

    TwinQuant learns quantization-friendly subspaces for 4-bit LLM weights via manifold optimization and a fused kernel, preserving near-FP16 accuracy with up to 1.8x speedup on LLaMA3 and Qwen3 models.

  2. SpinQuant: LLM quantization with learned rotations

    cs.LG 2024-05 conditional novelty 7.0 of 10

    SpinQuant learns optimal rotations to enable accurate 4-bit quantization of LLM weights, activations, and KV cache, reducing the zero-shot gap to full precision to 2.9 points on LLaMA-2 7B.

  3. Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Quasi-SVD learns a Lie-constrained approximate SVD whose one-sided orthogonal factor enables GPU-parallel medical imaging decompositions above 25 FPS with SSIM 0.89–0.94.

  4. ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    Layer-wise LLM rotations can be fused offline and residual basis mismatch approximated by a low-rank subspace, matching expensive layer-wise PTQ accuracy with near-global overhead.

  5. Tensor-Programmable Quantum Circuits for Solving Differential Equations

    quant-ph 2025-02 unverdicted novelty 6.0 of 10

    A quantum solver for PDEs is introduced via flexible matrix product operator representations with mid-circuit measurements and state-dependent norm correction to handle non-unitary dynamics.

  6. Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Pion is an optimizer that preserves the singular values of weight matrices in LLM training by applying orthogonal equivalence transformations.

  7. ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    ReSpinQuant achieves state-of-the-art accuracy in W4A4 and W3A3 LLM quantization by using efficient residual subspace rotation approximations that match layer-wise performance while retaining the inference speed of gl...

  8. Optimizing LOCC Protocols on Product Stiefel Manifold

    quant-ph 2025-10 conditional novelty 5.0 of 10

    Fixed-round LOCC protocols are parameterized by a product Stiefel manifold and optimized with Riemannian gradient methods, yielding achievable distillation and state-merging fidelities that sometimes match PPT upper bounds.

Pith tools