Optimizing training data via a differentiable SCM yields climate emulators that outperform those trained on six standard ScenarioMIP pathways while using less data and isolating distinct forcing responses.
Zhiyuan Li and Sanjeev Arora
4 Pith papers cite this work, alongside 44 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Manifold constraints via the new MACRO optimizer independently bound activation scales and enforce rotational equilibrium in LLM pre-training, subsuming RMS normalization and decoupled weight decay while delivering competitive performance with convergence guarantees.
NDSI-BWE deploys seven nonlinear-dynamics discriminators and a dual-stream ConformerNeXt generator to claim new state-of-the-art results in speech bandwidth extension.
Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.
citing papers explorer
-
Optimal scenario design for climate emulation
Optimizing training data via a differentiable SCM yields climate emulators that outperform those trained on six standard ScenarioMIP pathways while using less data and isolating distinct forcing responses.
-
Demystifying Manifold Constraints in LLM Pre-training
Manifold constraints via the new MACRO optimizer independently bound activation scales and enforce rotational equilibrium in LLM pre-training, subsuming RMS normalization and decoupled weight decay while delivering competitive performance with convergence guarantees.
-
CIS-BWE: Chaos-Informed Speech Bandwidth Extension
NDSI-BWE deploys seven nonlinear-dynamics discriminators and a dual-stream ConformerNeXt generator to claim new state-of-the-art results in speech bandwidth extension.
-
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors
Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.