Pith. sign in

REVIEW 10 cited by

Neural Network Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13144 v3 pith:A42OAK6F submitted 2024-02-20 cs.LG cs.CV

classification cs.LGcs.CV
keywords diffusionmodelsnetworktrainedautoencodermodelneuralparameters
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models have achieved remarkable success in image and video generation. In this work, we demonstrate that diffusion models can also \textit{generate high-performing neural network parameters}. Our approach is simple, utilizing an autoencoder and a diffusion model. The autoencoder extracts latent representations of a subset of the trained neural network parameters. Next, a diffusion model is trained to synthesize these latent representations from random noise. This model then generates new representations, which are passed through the autoencoder's decoder to produce new subsets of high-performing network parameters. Across various architectures and datasets, our approach consistently generates models with comparable or improved performance over trained networks, with minimal additional cost. Notably, we empirically find that the generated models are not memorizing the trained ones. Our results encourage more exploration into the versatile use of diffusion models. Our code is available \href{https://github.com/NUS-HPC-AI-Lab/Neural-Network-Diffusion}{here}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A calibration-free rounding scheme uses a diffusion model trained on a model's own weights to resolve ambiguous midpoint rounding, improving 3- and 4-bit weight-only quantization on small LLMs.

  2. Model Diffusion for Certifiable Few-shot Transfer Learning

    cs.LG 2025-02 conditional novelty 7.0 of 10

    STEEL samples a finite set of PEFT adapters from a diffusion model and selects the best on the downstream support set, producing non-vacuous PAC-Bayes generalization certificates for low-shot LLM and vision transfer learning.

  3. Architecture Generalization with MetaNCA

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A learned local rule (Weight Transformer) iteratively self-organizes task-network weights from local graph neighborhoods and generalizes across unseen MLP, CNN, and ResNet architectures up to ~2M parameters.

  4. WeightCLIP: Aligning Datasets and Models for Weight Space Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Contrastive dataset–weight alignment reshapes weight-space latents so dataset prompts retrieve, generate, and refine neural nets better than prior weight-space methods.

  5. Weight Space Representation Learning via Neural Field Adaptation

    cs.LG 2025-12 conditional novelty 6.0 of 10

    Multiplicative LoRA weights of pre-trained neural fields form structured, semantically meaningful representations that outperform prior weight-space methods for generation and classification.

  6. Delta Activations: A Representation for Finetuned Large Language Models

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Delta Activations embed finetuned LLMs as the average difference in hidden states between the finetuned model and its base model on a small set of generic prompts, yielding domain clusters and approximate additive com...

  7. Recurrent Diffusion for Large-Scale Parameter Generation

    cs.LG 2025-01 conditional novelty 6.0 of 10

    RPG generates full weights for models up to 200M parameters, including ConvNeXt-L and LLaMA LoRA adapters, at accuracy comparable to trained checkpoints, using recurrent token prototypes to condition a 1D diffusion model.

  8. Learning with springs and sticks

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A damped spring-and-stick lattice performs regression by energy relaxation, and a reported 'thermodynamic learning barrier' sets the minimum free energy needed for learning.

  9. DiffPINN: Generative diffusion-initialized physics-informed neural networks for accelerating seismic wavefield representation

    physics.geo-ph 2025-05 conditional novelty 5.0 of 10

    A conditional latent diffusion model generates physics-informed neural network initializations that speed up seismic wavefield PINN training and improve accuracy.

  10. Few-shot Implicit Function Generation via Equivariance

    cs.CV 2025-01 conditional novelty 5.0 of 10

    EquiGen generates diverse and functionally similar implicit neural network weights from only a few examples by exploiting permutation equivariance in the weight space.

Pith tools