Pith. sign in

REVIEW 1 cited by

Trajectory balance: Improved credit assignment in GFlowNets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.13259 v3 pith:64FUE5SJ submitted 2022-01-31 cs.LG stat.ML

classification cs.LGstat.ML
keywords balanceactiongflownetslearningsequencestrajectoryobjectivecredit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative flow networks (GFlowNets) are a method for learning a stochastic policy for generating compositional objects, such as graphs or strings, from a given unnormalized density by sequences of actions, where many possible action sequences may lead to the same object. We find previously proposed learning objectives for GFlowNets, flow matching and detailed balance, which are analogous to temporal difference learning, to be prone to inefficient credit propagation across long action sequences. We thus propose a new learning objective for GFlowNets, trajectory balance, as a more efficient alternative to previously used objectives. We prove that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution. In experiments on four distinct domains, we empirically demonstrate the benefits of the trajectory balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequences and large action spaces.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Machine learning for sample-based quantum diagonalization: generative configuration recovery and the classical-simulability frontier

    quant-ph 2026-08 conditional novelty 6.0 of 10

    A critical review plus small exact-FCI experiments concludes that sample-based quantum diagonalization has not beaten classical selected CI and maps where, if anywhere, a quantum or generative advantage could survive.

Pith tools