REVIEW 10 cited by
Neural Network Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models have achieved remarkable success in image and video generation. In this work, we demonstrate that diffusion models can also \textit{generate high-performing neural network parameters}. Our approach is simple, utilizing an autoencoder and a diffusion model. The autoencoder extracts latent representations of a subset of the trained neural network parameters. Next, a diffusion model is trained to synthesize these latent representations from random noise. This model then generates new representations, which are passed through the autoencoder's decoder to produce new subsets of high-performing network parameters. Across various architectures and datasets, our approach consistently generates models with comparable or improved performance over trained networks, with minimal additional cost. Notably, we empirically find that the generated models are not memorizing the trained ones. Our results encourage more exploration into the versatile use of diffusion models. Our code is available \href{https://github.com/NUS-HPC-AI-Lab/Neural-Network-Diffusion}{here}.
Forward citations
Cited by 10 Pith papers
-
ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization
A calibration-free rounding scheme uses a diffusion model trained on a model's own weights to resolve ambiguous midpoint rounding, improving 3- and 4-bit weight-only quantization on small LLMs.
-
Model Diffusion for Certifiable Few-shot Transfer Learning
STEEL samples a finite set of PEFT adapters from a diffusion model and selects the best on the downstream support set, producing non-vacuous PAC-Bayes generalization certificates for low-shot LLM and vision transfer learning.
-
Architecture Generalization with MetaNCA
A learned local rule (Weight Transformer) iteratively self-organizes task-network weights from local graph neighborhoods and generalizes across unseen MLP, CNN, and ResNet architectures up to ~2M parameters.
-
WeightCLIP: Aligning Datasets and Models for Weight Space Learning
Contrastive dataset–weight alignment reshapes weight-space latents so dataset prompts retrieve, generate, and refine neural nets better than prior weight-space methods.
-
Weight Space Representation Learning via Neural Field Adaptation
Multiplicative LoRA weights of pre-trained neural fields form structured, semantically meaningful representations that outperform prior weight-space methods for generation and classification.
-
Delta Activations: A Representation for Finetuned Large Language Models
Delta Activations embed finetuned LLMs as the average difference in hidden states between the finetuned model and its base model on a small set of generic prompts, yielding domain clusters and approximate additive com...
-
Recurrent Diffusion for Large-Scale Parameter Generation
RPG generates full weights for models up to 200M parameters, including ConvNeXt-L and LLaMA LoRA adapters, at accuracy comparable to trained checkpoints, using recurrent token prototypes to condition a 1D diffusion model.
-
Learning with springs and sticks
A damped spring-and-stick lattice performs regression by energy relaxation, and a reported 'thermodynamic learning barrier' sets the minimum free energy needed for learning.
-
DiffPINN: Generative diffusion-initialized physics-informed neural networks for accelerating seismic wavefield representation
A conditional latent diffusion model generates physics-informed neural network initializations that speed up seismic wavefield PINN training and improve accuracy.
-
Few-shot Implicit Function Generation via Equivariance
EquiGen generates diverse and functionally similar implicit neural network weights from only a few examples by exploiting permutation equivariance in the weight space.
Discussion (0). Continue with ORCID to comment.