REVIEW 2 cited by
Universal Neural Functionals
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A challenging problem in many modern machine learning tasks is to process weight-space features, i.e., to transform or extract information from the weights and gradients of a neural network. Recent works have developed promising weight-space models that are equivariant to the permutation symmetries of simple feedforward networks. However, they are not applicable to general architectures, since the permutation symmetries of a weight space can be complicated by recurrence or residual connections. This work proposes an algorithm that automatically constructs permutation equivariant models, which we refer to as universal neural functionals (UNFs), for any weight space. Among other applications, we demonstrate how UNFs can be substituted into existing learned optimizer designs, and find promising improvements over prior methods when optimizing small image classifiers and language models. Our results suggest that learned optimizers can benefit from considering the (symmetry) structure of the weight space they optimize. We open-source our library for constructing UNFs at https://github.com/AllanYangZhou/universal_neural_functional.
Forward citations
Cited by 2 Pith papers
-
Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training
LLaNA processes NeRF network weights directly with a frozen meta-encoder and a LLaMA 2 backbone, beating image- and point-cloud-based baselines on NeRF captioning and Q&A, and is trained on a new 280K-object ObjaNeRF-...
-
Improving Learning to Optimize Using Parameter Symmetries
A new lemma shows the symmetry part of Newton's direction increases gradient norm for convex objectives, but the paper's own experiments show teleportation-augmented L2O underperforms vanilla L2O, with momentum yieldi...
Discussion (0). Continue with ORCID to comment.