Pith. sign in

REVIEW 16 cited by

Bayesian Flow Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.07037 v6 pith:BC5TGPQW submitted 2023-08-14 cs.LG cs.AI

classification cs.LGcs.AI
keywords discretedatabayesianmodellingnetworkbfnsdiffusiondistributions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces Bayesian Flow Networks (BFNs), a new class of generative model in which the parameters of a set of independent distributions are modified with Bayesian inference in the light of noisy data samples, then passed as input to a neural network that outputs a second, interdependent distribution. Starting from a simple prior and iteratively updating the two distributions yields a generative procedure similar to the reverse process of diffusion models; however it is conceptually simpler in that no forward process is required. Discrete and continuous-time loss functions are derived for continuous, discretised and discrete data, along with sample generation procedures. Notably, the network inputs for discrete data lie on the probability simplex, and are therefore natively differentiable, paving the way for gradient-based sample guidance and few-step generation in discrete domains such as language modelling. The loss function directly optimises data compression and places no restrictions on the network architecture. In our experiments BFNs achieve competitive log-likelihoods for image modelling on dynamically binarized MNIST and CIFAR-10, and outperform all known discrete diffusion models on the text8 character-level language modelling task.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Large Language Diffusion Models

    cs.CL 2025-02 unverdicted novelty 8.0 of 10

    LLaDA is a scalable diffusion-based language model that matches autoregressive LLMs like LLaMA3 8B on tasks and surpasses GPT-4o on reversal poem completion.

  2. Observation-Aligned Mask Priors for Learning Physical Dynamics from Authentic Occlusions

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    A framework pretrained on authentic binary occlusion masks uses guided sampling and intersection-based partitioning to train diffusion models on incomplete physical observations without zero-query regions.

  3. Demystifying Multimodal Biomolecular Co-design With Intrinsic Geodesic Coupling

    q-bio.BM 2026-06 unverdicted novelty 6.0 of 10

    GeoCoupling optimizes temporal couplings between modalities in biomolecular generative models and outperforms synchronous baselines on drug design and protein design tasks.

  4. A renormalization-group inspired lattice-based framework for piecewise generalized linear models

    stat.ME 2026-05 unverdicted novelty 6.0 of 10

    RG-inspired lattice models for piecewise GLMs provide explicit interpretable partitions and a replica-analysis-derived scaling law for regularization that allows increasing complexity without expected rise in generali...

  5. Discrete Bayesian Sample Inference for Graph Generation

    cs.LG 2025-11 unverdicted novelty 6.0 of 10

    GraphBSI uses Bayesian Sample Inference as noise-controlled SDEs to generate discrete graphs in one shot, achieving state-of-the-art results on molecular benchmarks Moses and GuacaMol.

  6. Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    TraceRL aligns the RL objective with the DLM's actual step-by-step decoding, producing TraDo-4B/8B models that beat autoregressive baselines on math reasoning.

  7. Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Seed Diffusion Preview is a discrete diffusion language model that reaches 2146 tokens per second inference on H20 GPUs with competitive code benchmark performance, establishing a new speed-quality Pareto frontier.

  8. Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TarFlowLM models language in a continuous latent space with transformer-based autoregressive normalizing flows, using mixture-CDF and Rosenblatt couplings, and reports competitive NELBO on TEXT8 and OpenWebText.

  9. LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLaDA-V is a diffusion-based multimodal large language model that reaches competitive or state-of-the-art results on visual instruction tasks while using a non-autoregressive architecture.

  10. Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Parallel Trajectory Tempering with reservoirs and adaptive steps makes equilibrium maximum-likelihood training of EBMs practical and often better than PCD and deep generators on multimodal scientific data.

  11. Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    BA-Att introduces pre-downsampled block selection with norm-sorting and diagonal covariance correction to approximate sparse attention, yielding up to 6.95x speedup at 50% sparsity across language, multimodal, and vid...

  12. Controllable 3D Molecular Generation for Structure-Based Drug Design Through Bayesian Flow Networks and Gradient Integration

    cs.LG 2025-08 reject novelty 5.0 of 10

    Gradient guidance inside Bayesian Flow Network updates generates 3D drug candidates with stronger predicted docking scores, better retrosynthesis feasibility, and improved kinase selectivity than diffusion baselines.

  13. AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

    q-bio.BM 2025-07 conditional novelty 5.0 of 10

    AMix-1, a 1.7B-parameter Bayesian Flow Network protein model conditioned on MSA profiles and refined by an evolutionary test-time scaling loop, produced AmeR variants with up to 50x wild-type activity in wet-lab tests.

  14. A Unified Measure-Theoretic View of Diffusion, Score-Based, and Flow Matching Generative Models

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    Diffusion, score-based, and flow matching models are unified as instances of learning time-dependent vector fields inducing marginal distributions governed by continuity and Fokker-Planck equations.

  15. MolPIF: A Parameter Interpolation Flow Model for Molecule Generation

    cs.LG 2025-07 conditional novelty 4.0 of 10

    MolPIF generates 3D ligands by interpolating the parameters of Gaussian coordinate and Dirichlet atom-type distributions, reporting stronger docking scores and geometric fidelity than prior flow and diffusion models o...

  16. A Survey on Latent Reasoning

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey that organizes latent reasoning methods into vertical recurrence, horizontal recurrence, and infinite-depth diffusion, arguing that silent reasoning can beat explicit chain-of-thought.

Pith tools