Pith. sign in

GFlowNet-EM for learning compositional latent variable models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Latent variable models (LVMs) with discrete compositional latents are an important but challenging setting due to a combinatorially large number of possible configurations of the latents. A key tradeoff in modeling the posteriors over latents is between expressivity and tractable optimization. For algorithms based on expectation-maximization (EM), the E-step is often intractable without restrictive approximations to the posterior. We propose the use of GFlowNets, algorithms for sampling from an unnormalized density by learning a stochastic policy for sequential construction of samples, for this intractable E-step. By training GFlowNets to sample from the posterior over latents, we take advantage of their strengths as amortized variational inference algorithms for complex distributions over discrete structures. Our approach, GFlowNet-EM, enables the training of expressive LVMs with discrete compositional latents, as shown by experiments on non-context-free grammar induction and on images using discrete variational autoencoders (VAEs) without conditional independence enforced in the encoder.

fields

cs.LG 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Effective Reward Specification in Deep Reinforcement Learning

cs.LG · 2024-12-10 · conditional · novelty 4.0

A thesis presenting four methods (ASAF, TeamReg, CoachReg, constrained RL, goal-conditioned GFlowNets) that improve reward specification for deep RL through demonstrations, policy regularization, behavior constraints, and multi-objective conditioning.

citing papers explorer

Showing 1 of 1 citing paper.

  • Effective Reward Specification in Deep Reinforcement Learning cs.LG · 2024-12-10 · conditional · none · ref 134 · internal anchor

    A thesis presenting four methods (ASAF, TeamReg, CoachReg, constrained RL, goal-conditioned GFlowNets) that improve reward specification for deep RL through demonstrations, policy regularization, behavior constraints, and multi-objective conditioning.