Pith. sign in

REVIEW 4 major objections 4 minor 82 references

All prominent permutation-equivariant weight-space networks have identical expressive power, and under a natural distinct-bias condition they become universal approximators for invariant and equivariant maps.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 05:47 UTC pith:YSP33AA6

load-bearing objection Strong equivalence and GP universality theory for weight-space nets, but Theorem 6.1 as stated is false for linear activations and the abstract's 34% number has no support in the manuscript. the 4 major comments →

arxiv 2602.01083 v2 pith:YSP33AA6 submitted 2026-02-01 cs.LG

On the Expressive Power of Permutation-Equivariant Weight-Space Networks

classification cs.LG
keywords weight-space learningpermutation equivarianceexpressive equivalenceuniversal approximationgeneral positionfunction-space functionalscanonizationMLP weights
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper aims to settle how expressive permutation-equivariant networks over MLP weight spaces are. It proves that five prominent architectures—Deep Weight Space networks, Neural Functional Networks, Graph Meta-Networks, Neural Graph GNNs, and, under a distinct-bias condition, Neural Functional Transformers—approximate exactly the same class of maps on any compact weight set. It then proves universality in four settings: function-space functionals are always approximable; permutation-invariant functionals and permutation-equivariant operators are approximable exactly when hidden-layer bias values within each layer are pairwise distinct; and function-space operators are not approximable by fixed-architecture models but become approximable when the input architecture is large enough.

Core claim

The central claim is a unification: partial expressivity results for weight-space networks are replaced by one account. On compact sets, all prominent permutation-equivariant weight-space architectures are expressively equivalent; under a general-position assumption on the input weights, they are universal for permutation-invariant functionals and permutation-equivariant operators, and function-space functionals are universal on any compact set. The mechanism is a continuous canonization map available exactly when hidden-layer biases are distinct: sorting neurons by bias values gives each permutation orbit a canonical representative, and every equivariant operator factors through this canoni

What carries the argument

The load-bearing object is the weight-space canonization map defined on inputs with pairwise distinct hidden-layer biases. Because the bias ordering is unique and locally constant, it continuously selects one representative from each permutation orbit; the paper shows this map can be approximated by the architectures under study. Once canonization is available, any permutation-equivariant operator factors through a single continuous function applied pointwise to weight entries, reducing universality to standard pointwise MLP approximation. A separate expressive-equivalence cycle shows that the prominent architectures can all implement the same update primitives, so results proved for one arc

Load-bearing premise

The broadest universality claim, for function-space functionals, rests on a quoted separation-to-approximation theorem whose exact hypotheses are not restated for these equivariant architectures, so if that theorem does not apply, universality is unproven even though separation itself is established.

What would settle it

Exhibit two MLP weight vectors with all hidden-layer biases pairwise distinct, not related by neuron permutation, such that every permutation-equivariant weight-space network of every depth outputs the same value on both. The paper's general-position universality theorem predicts that no such pair can exist; a concrete construction with a proof of indistinguishability would refute it.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Architecture choice among the major permutation-equivariant weight-space models does not change expressive power, so empirical differences between them must come from optimization or inductive bias rather than representational limits.
  • Any continuous target that depends only on the realized function, such as accuracy prediction or INR classification, can be approximated by existing weight-space networks on any compact weight set.
  • Parameter-dependent tasks like loss-landscape curvature prediction or pruning-mask prediction are approximable once hidden biases are distinct, but fail on degenerate inputs such as heavily quantized or binary weights.
  • Function-space operators that increase the geometric complexity of the output function, such as zoom-out edits or domain adaptation, cannot be handled by fixed-architecture weight-to-weight maps; the input architecture must be enlarged.
  • The theoretical guidance translates into a concrete practical payoff: small modifications to existing weight-space models reportedly yield a 34 percent improvement over prior state-of-the-art results.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The distinct-bias condition is effectively a way to give each hidden neuron a unique identifier from the weights themselves; one testable extension is to add explicit learned neuron identifiers, which might restore universality even on the exclusion set where biases collide.
  • The fixed-architecture obstruction for function-space operators suggests a new role for overparameterization in weight-space learning: increasing input width is not just an optimization convenience but a representational resource needed to express certain function-level transformations.
  • The NFT gap on degenerate inputs suggests that attention-based weight-space models read bias values as tokens and can separate weight pairs that message-passing models cannot; this points to a concrete experimental probe comparing architectures on zero-bias initialization regions.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper develops a unified expressivity theory for permutation-equivariant weight-space networks operating on MLP parameters. It claims that the prominent architectures—DWS, NP-NFN, HNP-NFN, GMN, and NG-GNN—are all equivalent in expressive power, and that NFTs join this equivalence class under a general-position assumption (distinct hidden-layer biases). It then analyzes four approximation settings: function-space functionals, permutation-invariant functionals, function-space operators, and permutation-equivariant operators. The main results are universality for function-space functionals, non-universality for general invariant/equivariant maps, universality under general position for the latter, and universality for function-space operators only when the input architecture is sufficiently large. The abstract also claims a 34% empirical improvement over prior state-of-the-art.

Significance. If the main results hold, this is a useful unification of a fragmented literature: it reduces five active architectural families to one expressivity class, introduces a clean taxonomy of approximation settings, and gives concrete non-universality witnesses based on WL indistinguishability. The appendix contains substantial constructive proofs, and the general-position canonization technique is a plausible reusable tool. However, the paper's headline universality theorem for function-space functionals depends on an external separation-to-approximation theorem whose hypotheses are not verified, and the claimed 34% empirical improvement appears nowhere in the manuscript. These issues need resolution before the paper's central claims can be relied upon.

major comments (4)
  1. [Abstract / §1] The abstract and introduction state that "slight modifications to existing weight-space models yield a 34% improvement over prior SOTA." I could not find any experiment, table, dataset, or method description in the manuscript supporting this number. The paper is otherwise purely theoretical, and this unsupported empirical claim should be either fully documented or removed from the abstract and introduction.
  2. [Appendix E.2 / Theorem 6.1] The proof of universality for function-space functionals reduces to [50, Thm 1] without stating or verifying its hypotheses. The manuscript proves a separation property (Prop. E.6) and then asserts that a bounded-depth separation-to-approximation theorem applies. This is not established. Definition 3.1 allows any continuous activation σ, and the DWS networks in Definition D.7 use σ as the interleaved pointwise nonlinearity. If σ(x)=x, DWS networks are affine equivariant maps and cannot approximate arbitrary continuous function-space functionals—e.g., the square of a point-evaluation functional—so Theorem 6.1 is false as stated. The authors should add explicit assumptions on the activation (e.g., non-polynomial), verify the hypotheses of [50]/[51] for the DWS family, and justify the claim that the separating networks have bounded depth independent of the pair being separated.
  3. [Appendix D.5 / Proposition D.22] The proof that NFTs match DWS expressivity under general position requires NFTs to implement the listed DWS primitives. For the first-layer and last-layer per-neuron operators, the proof states "we recall that NFT uses the positional encoding PE of NP-NFN" and invokes Definition D.9. However, the NFT definition in Definition D.12 includes only layer encodings (LAYERENC), not per-neuron positional encodings. Without such identifiers, the attention mechanism cannot select an arbitrary input or output neuron for these operators when weight features are uniform. Either the NFT definition must explicitly include the positional encoding, or the authors must supply an implementation using the attention primitives actually defined. As written, the second part of Proposition 5.3 is unsupported.
  4. [Appendix F.2 / Theorem F.10] Theorem F.10 states that for any δ-approximating compact set K′⊂V_A there exists a DWS network approximating Ψ. The proof, however, uses Lemma F.16 and then invokes Theorem 7.4, which requires the relevant compact set to lie in V_A\E_A. The theorem does not require K′ to avoid the exclusion set, and arbitrary δ-approximators may intersect E_A, where Theorem 7.4 does not apply. The statement should be corrected to require K′⊂V_A\E_A, or the proof should construct a δ-approximator with this property.
minor comments (4)
  1. [Appendix E.3] The appendix remark says the proof of Theorem 6.3 is conceptually different from the proof sketch in §6.2. The main text should clearly label §6.2 as intuition, or align the two proofs, to avoid the appearance of a missing formal argument.
  2. [Appendix F.7 / Lemma F.7] There is a dimension mismatch in the lemma statement: Ψ is defined as a map into V^{n+5}_A, but the approximating network Φ is stated as a map into V^n_A. This should be V^{n+5}_A unless n is defined differently.
  3. [Appendix E.1 / Proposition E.6] The proof asserts that the depth of the separating DWS network "depends only on the architecture and not on the particular choice of v,v′." Since [49]'s forward-pass simulation is usually stated as an approximation guarantee, the exact separation needs a short argument or a precise citation; otherwise the bounded-depth claim is not justified.
  4. [Definition 5.1 / Definition D.3] The relationship between the main-text definition of expressive power N^π(K;Y) and the appendix's expressive-containment relation should be made more prominent, since later proofs switch freely between the two.

Circularity Check

0 steps flagged

No circular derivation found: the main universality and equivalence results are built from explicit constructions and external published theorems, not from their own conclusions.

full rationale

The paper's central claims are derived non-circularly. Theorem 5.2/D.6 (architectural equivalence) is proved by explicit layer-wise simulations among DWS, NP-NFN+PE, GMN, and NG-GNN in Appendix D, using prior published characterizations of DWS layers and GMN expressiveness ([49], [41]); these are parameter-free external results with stated assumptions, so they count as independent evidence rather than circularity. The GP universality results (Theorems 6.3 and 7.4) are proved by constructing a bias-sorting canonization map and then approximating the induced pointwise representations with DWS/DeepSets primitives (Appendix F.1); no target conclusion is assumed, and the formal proof of Theorem 6.3 explicitly routes through Theorem 7.4 in an acyclic, structural way. The function-space operator results (Theorem 7.2/F.10) reduce to a continuous lifting and then to Theorem 7.4, again without assuming the conclusion. Theorem 6.1/E.9 does rely on an external separation-to-approximation theorem from [50]/[51] without restating or verifying its hypotheses in Appendix E.2. If that theorem does not apply, the function-space functional universality proof would fail; however, this is a correctness/evidence gap, not circularity: the paper proves separation (Proposition E.6) and then invokes an external conversion result, without equating the conclusion to the input by construction. No fitted parameter is renamed as a prediction, no self-citation chain is used to forbid alternatives, and no ansatz is smuggled in via prior work. The load-bearing self-citations ([49], [41]) are independent prior proofs about the relevant architectures, not restatements of the present claims. Therefore no specific circular step can be exhibited, and the correct circularity score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 7 axioms · 0 invented entities

The paper's theoretical core fits within standard approximation theory; it imports four external results (forward-pass simulation, separation-to-approximation, DeepSets universality, WL bound). No free parameters are fitted, and no new physical or architectural entities are postulated; canonization and neuron-identification maps are mathematical constructions on existing weight space. The general-position assumption is the principal domain restriction.

axioms (7)
  • domain assumption Input weights are in general position: within each hidden layer, all bias entries are pairwise distinct (v ∉ E_A).
    Used to define a continuous canonization map; Theorems 6.3 and 7.4 fail on E_A by Propositions 6.2 and 7.3. Stated in Definition 3.4.
  • domain assumption Prior result [49]: DWS networks can simulate (approximate) the forward pass of the input MLP on compact weight sets.
    Used in the proof of Theorem 6.1 and Proposition E.6 to separate functionally distinct weights; central to universality of function-space functionals.
  • domain assumption Separation-to-approximation theorem of [50,51]: continuous invariant functions constant on equivalence classes of bounded-depth equivariant affine networks can be uniformly approximated on compact sets by such networks.
    Invoked in Appendix E.2 to convert DWS-separation into universality; exact hypotheses are not restated in the paper.
  • standard math Continuous activation σ and MLP universal approximation [10,32].
    Used throughout to approximate continuous maps by MLPs in DWS/GMN/NG simulations.
  • standard math DeepSets/equivariant set networks are universal for continuous permutation-equivariant functions on compact domains [59].
    Used in Lemma F.4 to approximate ranking/canonization maps by DWS bias-to-bias layers.
  • standard math Fixed ReLU MLP architectures have uniformly bounded number of linear regions and region vertices [46,54,60].
    Used in Theorem F.9 to establish the expressivity ceiling for fixed architectures.
  • standard math The 1-WL test bounds message-passing GNN expressivity [47].
    Used in Proposition 6.2 and Proposition E.8 to show NG-GNN cannot separate the two constructed neural graphs.

pith-pipeline@v1.3.0-alltime-deepseek · 59664 in / 15903 out tokens · 172030 ms · 2026-08-03T05:47:53.497060+00:00 · methodology

0 comments
read the original abstract

Weight-space learning studies neural architectures that operate directly on the parameters of other neural networks. Motivated by the growing availability of pretrained models, recent work has demonstrated the effectiveness of weight-space networks across a wide range of tasks. SOTA weight-space networks rely on permutation-equivariant designs to improve generalization. However, this may negatively affect expressive power, warranting theoretical investigation. Importantly, unlike other structured domains, weight-space learning targets maps operating on both weight and function spaces, making expressivity analysis particularly subtle. While a few prior works provide partial expressivity results, a comprehensive characterization is still missing. In this work, we address this gap by developing a systematic theory for expressivity of weight-space networks. We first prove that all prominent permutation-equivariant networks are equivalent in expressive power. We then establish universality in both weight- and function-space settings under mild, natural assumptions on the input weights, and characterize the edge-case regimes where universality no longer holds. Guided by our theoretical results, we show that slight modifications to existing weight-space models yield a 34% improvement over prior SOTA, demonstrating the practical relevance of our framework.

Figures

Figures reproduced from arXiv: 2602.01083 by Adir Dayan, Haggai Maron, Yam Eitan.

Figure 1
Figure 1. Figure 1: Expressivity landscape for permutation-equivariant weight-space networks on MLPs (blue arrows: new [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Real-world examples of target functions for all approximation settings in weight-space learning. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Computational graphs induced by weights v and v ′ used in the proof of Proposition 6.2. The weight matrices have binary entries, where 1 corresponds to an edge and 0 to the absence of an edge, and bias terms are equipped with features encoding their layer index. Both graphs admit identical 1-WL colorings (node colors) and are therefore indistinguishable by message-passing GNNs (NG-GNNs). Formal definitions… view at source ↗
Figure 4
Figure 4. Figure 4: Diagram illustrating Theorem 7.2. A function-space operator Ψ : K ⊂ C(X, R m) → C(X, R m) is approximated by a permutation-equivariant weight-space operator Φ : K′ ⊂ VA \ EA → VA. The set K′ serves as an approximation of K under the realization map R. The diagram is approximately commutative, showing that Φ approximates Ψ via R. sufficiently large architecture A and any compact set K′ ⊂ VA \ EA whose reali… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

82 extracted references · 23 linked inside Pith

  1. [1]

    A convergence theory for deep learning via over-parameterization

    Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. A convergence theory for deep learning via over-parameterization. InInternational conference on machine learning, pages 242–252. PMLR, 2019

  2. [2]

    Expressive power of invariant and equivariant graph neural networks.arXiv preprint arXiv:2006.15646, 2020

    Waiss Azizian and Marc Lelarge. Expressive power of invariant and equivariant graph neural networks.arXiv preprint arXiv:2006.15646, 2020

  3. [3]

    A flexible, equivariant framework for subgraph gnns via graph products and graph coarsening.Advances in Neural Information Processing Systems, 37:101168– 101222, 2024

    Guy Bar-Shalom, Yam Eitan, Fabrizio Frasca, and Haggai Maron. A flexible, equivariant framework for subgraph gnns via graph products and graph coarsening.Advances in Neural Information Processing Systems, 37:101168– 101222, 2024

  4. [4]

    Reconciling modern machine-learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019

    Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine-learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019

  5. [5]

    Equivariant subgraph aggregation networks.arXiv preprint arXiv:2110.02910, 2021

    Beatrice Bevilacqua, Fabrizio Frasca, Derek Lim, Balasubramaniam Srinivasan, Chen Cai, Gopinath Balamu- rugan, Michael M Bronstein, and Haggai Maron. Equivariant subgraph aggregation networks.arXiv preprint arXiv:2110.02910, 2021

  6. [6]

    Weisfeiler and lehman go cellular: Cw networks.Advances in neural information processing systems, 34:2625–2640, 2021

    Cristian Bodnar, Fabrizio Frasca, Nina Otter, Yuguang Wang, Pietro Lio, Guido F Montufar, and Michael Bronstein. Weisfeiler and lehman go cellular: Cw networks.Advances in neural information processing systems, 34:2625–2640, 2021

  7. [7]

    Improving graph neural network expressivity via subgraph isomorphism counting.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):657–668, 2022

    Giorgos Bouritsas, Fabrizio Frasca, Stefanos Zafeiriou, and Michael M Bronstein. Improving graph neural network expressivity via subgraph isomorphism counting.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):657–668, 2022

  8. [8]

    Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇckovi´c

    Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇckovi´c. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges.arXiv preprint arXiv:2104.13478, 2021

  9. [9]

    Reconstruction for powerful graph representations

    Leonardo Cotta, Christopher Morris, and Bruno Ribeiro. Reconstruction for powerful graph representations. Advances in Neural Information Processing Systems, 34:1713–1726, 2021

  10. [10]

    Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314, 1989

    George Cybenko. Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314, 1989

  11. [11]

    Laplace redux-effortless bayesian deep learning.Advances in neural information processing systems, 34:20089– 20103, 2021

    Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig. Laplace redux-effortless bayesian deep learning.Advances in neural information processing systems, 34:20089– 20103, 2021

  12. [12]

    Deep learning on implicit neural representations of shapes.arXiv preprint arXiv:2302.05438, 2023

    Luca De Luigi, Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, and Luigi Di Stefano. Deep learning on implicit neural representations of shapes.arXiv preprint arXiv:2302.05438, 2023

  13. [13]

    Gradient descent finds global minima of deep neural networks

    Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. Gradient descent finds global minima of deep neural networks. InInternational conference on machine learning, pages 1675–1685. PMLR, 2019

  14. [14]

    From data to functa: Your data point is a function and you can treat it like one.arXiv preprint arXiv:2201.12204, 2022

    Emilien Dupont, Hyunjik Kim, SM Eslami, Danilo Rezende, and Dan Rosenbaum. From data to functa: Your data point is a function and you can treat it like one.arXiv preprint arXiv:2201.12204, 2022

  15. [15]

    Benchmarking graph neural networks.Journal of Machine Learning Research, 24(43):1–48, 2023

    Vijay Prakash Dwivedi, Chaitanya K Joshi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Benchmarking graph neural networks.Journal of Machine Learning Research, 24(43):1–48, 2023

  16. [16]

    On the universality of rotation equivariant point cloud networks.arXiv preprint arXiv:2010.02449, 2020

    Nadav Dym and Haggai Maron. On the universality of rotation equivariant point cloud networks.arXiv preprint arXiv:2010.02449, 2020

  17. [17]

    Classifying the classifier: dissecting the weight space of neural networks.arXiv preprint arXiv:2002.05688, 2020

    Gabriel Eilertsen, Daniel J¨onsson, Timo Ropinski, Jonas Unger, and Anders Ynnerman. Classifying the classifier: dissecting the weight space of neural networks.arXiv preprint arXiv:2002.05688, 2020. 11

  18. [18]

    Topological blindspots: Understanding and extending topological deep learning through the lens of expressivity.arXiv preprint arXiv:2408.05486, 2024

    Yam Eitan, Yoav Gelberg, Guy Bar-Shalom, Fabrizio Frasca, Michael Bronstein, and Haggai Maron. Topological blindspots: Understanding and extending topological deep learning through the lens of expressivity.arXiv preprint arXiv:2408.05486, 2024

  19. [19]

    On the expressive power of gnn derivatives.arXiv preprint arXiv:2510.02565, 2025

    Yam Eitan, Moshe Eliasof, Yoav Gelberg, Fabrizio Frasca, Guy Bar-Shalom, and Haggai Maron. On the expressive power of gnn derivatives.arXiv preprint arXiv:2510.02565, 2025

  20. [20]

    Fs-kan: Permutation equivariant kolmogorov-arnold networks via function sharing.arXiv preprint arXiv:2509.24472, 2025

    Ran Elbaz, Guy Bar-Shalom, Yam Eitan, Fabrizio Frasca, and Haggai Maron. Fs-kan: Permutation equivariant kolmogorov-arnold networks via function sharing.arXiv preprint arXiv:2509.24472, 2025

  21. [21]

    Equivariance everywhere all at once: A recipe for graph foundation models.arXiv preprint arXiv:2506.14291, 2025

    Ben Finkelshtein, ˙Ismail ˙Ilkan Ceylan, Michael Bronstein, and Ron Levie. Equivariance everywhere all at once: A recipe for graph foundation models.arXiv preprint arXiv:2506.14291, 2025

  22. [22]

    Understanding and extending subgraph gnns by rethinking their symmetries.Advances in Neural Information Processing Systems, 35:31376– 31390, 2022

    Fabrizio Frasca, Beatrice Bevilacqua, Michael Bronstein, and Haggai Maron. Understanding and extending subgraph gnns by rethinking their symmetries.Advances in Neural Information Processing Systems, 35:31376– 31390, 2022

  23. [23]

    Gradmetanet: An equivariant architecture for learning on gradients.arXiv preprint arXiv:2507.01649, 2025

    Yoav Gelberg, Yam Eitan, Aviv Navon, Aviv Shamsian, Michael Bronstein, Haggai Maron, et al. Gradmetanet: An equivariant architecture for learning on gradients.arXiv preprint arXiv:2507.01649, 2025

  24. [24]

    Descriptive complexity, canonisation, and definable graph structure theory, volume 47 of lecture notes in logic

    Martin Grohe. Descriptive complexity, canonisation, and definable graph structure theory, volume 47 of lecture notes in logic. association for symbolic logic, 2017

  25. [25]

    Studying large language model generalization with influence functions.arXiv preprint arXiv:2308.03296, 2023

    Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al. Studying large language model generalization with influence functions.arXiv preprint arXiv:2308.03296, 2023

  26. [26]

    Learning both weights and connections for efficient neural network.Advances in neural information processing systems, 28, 2015

    Song Han, Jeff Pool, John Tran, and William Dally. Learning both weights and connections for efficient neural network.Advances in neural information processing systems, 28, 2015

  27. [27]

    Optimal brain surgeon and general network pruning

    Babak Hassibi, David G Stork, and Gregory J Wolff. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, pages 293–299. IEEE, 1993

  28. [28]

    Sparsified model zoo twins: Investigating popula- tions of sparsified neural network models.arXiv preprint arXiv:2304.13718, 2023

    Dominik Honegger, Konstantin Sch¨urholt, and Damian Borth. Sparsified model zoo twins: Investigating popula- tions of sparsified neural network models.arXiv preprint arXiv:2304.13718, 2023

  29. [29]

    Eurosat model zoo: A dataset and benchmark on populations of neural networks and its sparsified model twins

    Dominik Honegger, Konstantin Sch¨urholt, Linus Scheibenreif, and Damian Borth. Eurosat model zoo: A dataset and benchmark on populations of neural networks and its sparsified model twins. InIGARSS 2023-2023 IEEE International Geoscience and Remote Sensing Symposium, pages 888–891. IEEE, 2023

  30. [30]

    Weisfeiler leman for euclidean equivariant machine learning.arXiv preprint arXiv:2402.02484, 2024

    Snir Hordan, Tal Amir, and Nadav Dym. Weisfeiler leman for euclidean equivariant machine learning.arXiv preprint arXiv:2402.02484, 2024

  31. [31]

    Complete neural networks for complete euclidean graphs

    Snir Hordan, Tal Amir, Steven J Gortler, and Nadav Dym. Complete neural networks for complete euclidean graphs. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 12482–12490, 2024

  32. [32]

    Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257, 1991

    Kurt Hornik. Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257, 1991

  33. [33]

    Scalable marginal likelihood estimation for model selection in deep learning

    Alexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar R¨atsch, and Khan Mohammad Emtiyaz. Scalable marginal likelihood estimation for model selection in deep learning. InInternational Conference on Machine Learning, pages 4563–4573. PMLR, 2021

  34. [34]

    Scale equivariant graph metanetworks

    Ioannis Kalogeropoulos, Giorgos Bouritsas, and Yannis Panagakis. Scale equivariant graph metanetworks. Advances in neural information processing systems, 37:106800–106840, 2024

  35. [35]

    Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020. 12

  36. [36]

    Universal invariant and equivariant graph neural networks.Advances in neural information processing systems, 32, 2019

    Nicolas Keriven and Gabriel Peyr ´e. Universal invariant and equivariant graph neural networks.Advances in neural information processing systems, 32, 2019

  37. [37]

    Graph neural networks for learning equivariant representations of neural networks

    Miltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen, Gertjan J Burghouts, Efstratios Gavves, Cees GM Snoek, and David W Zhang. Graph neural networks for learning equivariant representations of neural networks. arXiv preprint arXiv:2403.12143, 2024

  38. [38]

    Understanding black-box predictions via influence functions

    Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. InInternational conference on machine learning, pages 1885–1894. PMLR, 2017

  39. [39]

    Optimal brain damage.Advances in neural information processing systems, 2, 1989

    Yann LeCun, John Denker, and Sara Solla. Optimal brain damage.Advances in neural information processing systems, 2, 1989

  40. [40]

    Sign and basis invariant networks for spectral graph representation learning.arXiv preprint arXiv:2202.13013, 2022

    Derek Lim, Joshua Robinson, Lingxiao Zhao, Tess Smidt, Suvrit Sra, Haggai Maron, and Stefanie Jegelka. Sign and basis invariant networks for spectral graph representation learning.arXiv preprint arXiv:2202.13013, 2022

  41. [41]

    Graph metanetworks for processing diverse neural architectures.arXiv preprint arXiv:2312.04501, 2023

    Derek Lim, Haggai Maron, Marc T Law, Jonathan Lorraine, and James Lucas. Graph metanetworks for processing diverse neural architectures.arXiv preprint arXiv:2312.04501, 2023

  42. [42]

    Provably powerful graph networks

    Haggai Maron, Heli Ben-Hamu, Hadar Serviansky, and Yaron Lipman. Provably powerful graph networks. Advances in neural information processing systems, 32, 2019

  43. [43]

    On the universality of invariant networks

    Haggai Maron, Ethan Fetaya, Nimrod Segol, and Yaron Lipman. On the universality of invariant networks. In International conference on machine learning, pages 4363–4371. PMLR, 2019

  44. [44]

    On learning sets of symmetric elements

    Haggai Maron, Or Litany, Gal Chechik, and Ethan Fetaya. On learning sets of symmetric elements. InInternational conference on machine learning, pages 6734–6744. PMLR, 2020

  45. [45]

    Nerf: Representing scenes as neural radiance fields for view synthesis.Communications of the ACM, 65(1):99–106, 2021

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis.Communications of the ACM, 65(1):99–106, 2021

  46. [46]

    On the number of linear regions of deep neural networks.Advances in neural information processing systems, 27, 2014

    Guido Mont´ufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio. On the number of linear regions of deep neural networks.Advances in neural information processing systems, 27, 2014

  47. [47]

    Weisfeiler and leman go neural: Higher-order graph neural networks

    Christopher Morris, Martin Ritzert, Matthias Fey, William L Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and leman go neural: Higher-order graph neural networks. InProceedings of the AAAI conference on artificial intelligence, volume 33, pages 4602–4609, 2019

  48. [48]

    Weisfeiler and leman go machine learning: The story so far.Journal of Machine Learning Research, 24(333):1–59, 2023

    Christopher Morris, Yaron Lipman, Haggai Maron, Bastian Rieck, Nils M Kriege, Martin Grohe, Matthias Fey, and Karsten Borgwardt. Weisfeiler and leman go machine learning: The story so far.Journal of Machine Learning Research, 24(333):1–59, 2023

  49. [49]

    Equivariant architectures for learning in deep weight spaces

    Aviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya, Gal Chechik, and Haggai Maron. Equivariant architectures for learning in deep weight spaces. InInternational Conference on Machine Learning, pages 25790–25816. PMLR, 2023

  50. [50]

    On universality of deep equivariant networks.arXiv preprint arXiv:2510.15814, 2025

    Marco Pacini, Mircea Petrache, Bruno Lepri, Shubhendu Trivedi, and Robin Walters. On universality of deep equivariant networks.arXiv preprint arXiv:2510.15814, 2025

  51. [51]

    On universality classes of equivariant networks.arXiv preprint arXiv:2506.02293, 2025

    Marco Pacini, Gabriele Santin, Bruno Lepri, and Shubhendu Trivedi. On universality classes of equivariant networks.arXiv preprint arXiv:2506.02293, 2025

  52. [52]

    Learning on loras: Gl-equivariant processing of low-rank weight spaces for large finetuned models.arXiv preprint arXiv:2410.04207, 2024

    Theo Putterman, Derek Lim, Yoav Gelberg, Stefanie Jegelka, and Haggai Maron. Learning on loras: Gl-equivariant processing of low-rank weight spaces for large finetuned models.arXiv preprint arXiv:2410.04207, 2024

  53. [53]

    Pointnet: Deep learning on point sets for 3d classification and segmentation

    Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660, 2017. 13

  54. [54]

    On the expressive power of deep neural networks

    Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein. On the expressive power of deep neural networks. Ininternational conference on machine learning, pages 2847–2854. PMLR, 2017

  55. [55]

    Universal equivariant multilayer perceptrons

    Siamak Ravanbakhsh. Universal equivariant multilayer perceptrons. InInternational Conference on Machine Learning, pages 7996–8006. PMLR, 2020

  56. [56]

    A persistent weisfeiler-lehman procedure for graph classification

    Bastian Rieck, Christian Bock, and Karsten Borgwardt. A persistent weisfeiler-lehman procedure for graph classification. InInternational Conference on Machine Learning, pages 5448–5458. PMLR, 2019

  57. [57]

    Self-supervised representation learning on neural network weights for model characteristic prediction.Advances in Neural Information Processing Systems, 34: 16481–16493, 2021

    Konstantin Sch¨urholt, Dimche Kostadinov, and Damian Borth. Self-supervised representation learning on neural network weights for model characteristic prediction.Advances in Neural Information Processing Systems, 34: 16481–16493, 2021

  58. [58]

    A model zoo on phase transitions in neural networks.arXiv preprint arXiv:2504.18072, 2025

    Konstantin Sch¨urholt, L´eo Meynent, Yefan Zhou, Haiquan Lu, Yaoqing Yang, and Damian Borth. A model zoo on phase transitions in neural networks.arXiv preprint arXiv:2504.18072, 2025

  59. [59]

    On universal equivariant set networks.arXiv preprint arXiv:1910.02421, 2019

    Nimrod Segol and Yaron Lipman. On universal equivariant set networks.arXiv preprint arXiv:1910.02421, 2019

  60. [60]

    Bounding and counting linear regions of deep neural networks

    Thiago Serra, Christian Tjandraatmadja, and Srikumar Ramalingam. Bounding and counting linear regions of deep neural networks. InInternational conference on machine learning, pages 4558–4566. PMLR, 2018

  61. [61]

    Improved generalization of weight space networks via augmentations.arXiv preprint arXiv:2402.04081, 2024

    Aviv Shamsian, Aviv Navon, David W Zhang, Yan Zhang, Ethan Fetaya, Gal Chechik, and Haggai Maron. Improved generalization of weight space networks via augmentations.arXiv preprint arXiv:2402.04081, 2024

  62. [62]

    Implicit neural representations with periodic activation functions.Advances in neural information processing systems, 33: 7462–7473, 2020

    Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions.Advances in neural information processing systems, 33: 7462–7473, 2020

  63. [63]

    Balancing efficiency and expressiveness: Subgraph gnns with walk-based centrality.arXiv preprint arXiv:2501.03113, 2025

    Joshua Southern, Yam Eitan, Guy Bar-Shalom, Michael Bronstein, Haggai Maron, and Fabrizio Frasca. Balancing efficiency and expressiveness: Subgraph gnns with walk-based centrality.arXiv preprint arXiv:2501.03113, 2025

  64. [64]

    Monomial matrix group equivariant neural functional networks.Advances in Neural Information Processing Systems, 37:48628–48665, 2024

    Hoang Tran, Thieu V o, Tho Huu, Tan Nguyen, et al. Monomial matrix group equivariant neural functional networks.Advances in Neural Information Processing Systems, 37:48628–48665, 2024

  65. [65]

    Predicting neural network accuracy from weights.arXiv preprint arXiv:2002.11448, 2020

    Thomas Unterthiner, Daniel Keysers, Sylvain Gelly, Olivier Bousquet, and Ilya Tolstikhin. Predicting neural network accuracy from weights.arXiv preprint arXiv:2002.11448, 2020

  66. [66]

    Equivariant polynomial functional networks.arXiv preprint arXiv:2410.04213, 2024

    Thieu N V o, Viet-Hoang Tran, Tho Tran Huu, An Nguyen The, Thanh Tran, Minh-Khoi Nguyen-Nhat, Duy-Tung Pham, and Tan Minh Nguyen. Equivariant polynomial functional networks.arXiv preprint arXiv:2410.04213, 2024

  67. [67]

    Recurrent diffusion for large-scale parameter generation.arXiv preprint arXiv:2501.11587, 2025

    Kai Wang, Dongwen Tang, Wangbo Zhao, Konstantin Sch¨urholt, Zhangyang Wang, and Yang You. Recurrent diffusion for large-scale parameter generation.arXiv preprint arXiv:2501.11587, 2025

  68. [68]

    How powerful are graph neural networks?arXiv preprint arXiv:1810.00826, 2018

    Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks?arXiv preprint arXiv:1810.00826, 2018

  69. [69]

    Deep sets.Advances in neural information processing systems, 30, 2017

    Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola. Deep sets.Advances in neural information processing systems, 30, 2017

  70. [70]

    Nested graph neural networks.Advances in Neural Information Processing Systems, 34:15734–15747, 2021

    Muhan Zhang and Pan Li. Nested graph neural networks.Advances in Neural Information Processing Systems, 34:15734–15747, 2021

  71. [71]

    Permutation equivariant neural functionals.Advances in neural information processing systems, 36: 24966–24992, 2023

    Allan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace, Yiding Jiang, Samuel Sokota, J Zico Kolter, and Chelsea Finn. Permutation equivariant neural functionals.Advances in neural information processing systems, 36: 24966–24992, 2023

  72. [72]

    Neural functional transformers.Advances in neural information processing systems, 36:77485–77502, 2023

    Allan Zhou, Kaien Yang, Yiding Jiang, Kaylee Burns, Winnie Xu, Samuel Sokota, J Zico Kolter, and Chelsea Finn. Neural functional transformers.Advances in neural information processing systems, 36:77485–77502, 2023

  73. [73]

    Universal neural functionals.Advances in neural information processing systems, 37:104754–104775, 2024

    Allan Zhou, Chelsea Finn, and James Harrison. Universal neural functionals.Advances in neural information processing systems, 37:104754–104775, 2024. 14 Contents A. Extended previous work. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 B. Extended pre...

  74. [74]

    preserved

    +c in is: hν(ℓ) i :=    nlayer(ν(0) i ),ntype(ν (0) i ),0 cin ifℓ= 0, nlayer(ν(ℓ) i ),ntype(ν (ℓ) i ),(b ℓ)i,: if0< ℓ≤L. 24 MPNN layer update (no global features).AnNG layeris a tuple of MLPs L= (ϕ m, ϕh, ϕe) that updates the state (h,e)7→(h ′,e ′)via: mj→i =ϕ m hi,h j,e ij ,(31) h′ i =ϕ h hi, X j∈N(i) mj→i ,(32) e′ ij =ϕ e hi,h j,e ij ,(33) whereN(i)d...

  75. [75]

    There exists a compact setK ′ ⊆ Vwhich is aδ-approximator ofK(see Definition B.4)

  76. [76]

    Before proving Theorem F.9, we introduce a convenient class of functions that captures the geometric complexity of ReLU networks with a fixed architecture

    For any such setK ′ there exists a DWS networkΦ :V A → VA such that, for everyv∈K ′, ∥f Φ(v) −Ψ(f v)∥ ∞ < ϵ.(232) F.2.1 Proof of Theorem F.9 For the remainder of this subsection, we fix X= [0,1] n. Before proving Theorem F.9, we introduce a convenient class of functions that captures the geometric complexity of ReLU networks with a fixed architecture. Def...

  77. [77]

    ,vj,rj }(233) for somev j,ℓ ∈Xwith1≤r j ≤R

    for eachj, the polytopeP j can be written as Pj = conv{vj,1, . . . ,vj,rj }(233) for somev j,ℓ ∈Xwith1≤r j ≤R

  78. [78]

    the interiors are pairwise disjoint and coverX, X= m[ j=1 Pj,int(P i)∩int(P j) =∅fori̸=j;(234) 56

  79. [79]

    f(x) =A j(x)∀x∈P j.(235) We refer toP 1,

    on each polytopeP j the functionfis affine, i.e. f(x) =A j(x)∀x∈P j.(235) We refer toP 1, . . . , Pm as thelinearity regionsoff. The next lemma shows that, for a fixed ReLU architecture, there are uniform bounds on both the number of linearity regions induced by each network and the number of extreme points of each such region. The arguments used in this ...

  80. [80]

    an original vertex ofP, or

Showing first 80 references.