REVIEW 2 major objections 6 minor 300 references
Path-dependent Discrete Amortized Inference
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A deterministic latent memory added to a GFlowNet's state space lets the policy depend on the whole past trajectory, and the paper proves the standard balance losses still yield the correct terminal distribution.
desk verdict Solid central contribution (balance losses on a lifted DAG for path-dependent GFlowNets) undermined by a flawed expressivity proof in Proposition 3.1 that should be fixed or downgraded before peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the lifted pointed DAG $S^\star = S \times \Omega$ with deterministic latent dynamics $W_t = W_{t-1} + \phi(s_t, W_{t-1})$: it expands the state graph so that each distinct history receives its own latent context, while keeping the lifted space finite. The paper implements $\phi$ with a modified self-referential weight matrix (a fast-weight update rule that rotates the weight matrix by a fixed ergodic rotation and self-adjusts its update strength), so the memory both encodes the past and actively reduces state aliasing. The subtrajectory balance, trajectory balance, and contrastive balance conditions are then evaluated on this lifted graph, and the balance identity is what transfers the correctness guarantee from the Markovian setting to the path-dependent one.
What would settle it
Train the LINES environment from Section 3 with the same recurrent policy but a stochastic latent update $W_t = W_{t-1} + \phi(s_t, W_{t-1}) + \varepsilon_t$ and minimize the TB loss; if the terminal marginal's distance to $R$ does not approach the deterministic version's level while all else is held equal, the finiteness assumption behind Proposition 4.5 is the reason.
Extended reading notes
Core claim
On a lifted pointed DAG $S^\star = S \times \Omega$, where each state carries a latent memory updated by a deterministic, state-conditional rule, the central result (Proposition 4.5) is that if the subtrajectory balance condition $F(s^\star_i)p_F(s^\star_{i+1:j} \mid s^\star_i) = p_B(s^\star_{i:j-1} \mid s^\star_j)F(s^\star_j)$ holds for every slice of every trajectory, then the marginal of the forward policy over terminal states is proportional to the flow $F$, and therefore to the target $R$ on terminal objects. Trajectory balance and contrastive balance follow as corollaries. Because the latent update is deterministic, each discrete path determines a unique latent sequence, keeping the number of paths through each lifted state finite, which is what carries the proof. The same lifting is shown to strictly enlarge the representable class of distributions for linear policies and for GNN-based policies, and a two-terminal toy graph exhibits a uniform Markovian output where the path-dependent model can fit any target. A companion result shows that a Markovian backward policy forces the optimal forward policy to be Markovian as well, so the lift buys expressivity only when the backward model is also allowed to be path-dependent.
Load-bearing premise
The load-bearing premise is that the latent update is deterministic, $W_t = W_{t-1} + \phi(s_t, W_{t-1})$, so each discrete path determines a unique latent trajectory and the lifted graph stays finite; a stochastic latent update breaks that finiteness and the paper leaves that case open.
Editorial extensions
If this is right
- Existing GFlowNet training pipelines can swap a Markovian policy head for a recurrent policy and keep the same balance-loss guarantees that the terminal marginal is proportional to the target.
- Path-dependent parameterizations can represent distributions that Markovian parameterizations of the same architecture cannot, so state aliasing is an expressivity ceiling, not just an optimization difficulty.
- A Markovian backward policy drives the optimal lifted forward policy back to being Markovian; to benefit from path-dependence, the backward model must also be path-dependent, which the paper's experimental designs do.
- The vector field $\phi$ and the policy are learned jointly from a single loss, so the method adds memory without introducing a separate training objective.
- Distributed and streaming amortized inference schemes transfer to the lifted setting, as shown in the supplement, because the balance conditions remain valid there.
Reading between the lines
- A natural testable extension is stochastic latent dynamics: the paper sketches a measure-theoretic formulation but leaves the learning guarantees open, so an open problem is whether any SubTB-type condition can be proven when each history branches over infinitely many latent continuations.
- The expressivity results suggest that symmetry-rich generation tasks, where GNN-based GFlowNets are known to collapse to uniform or near-uniform outputs, are prime candidates for this memory-based lifting.
- Because the lifted policy is effectively a recurrent policy, techniques from partially observable MDPs and recurrent reinforcement learning---belief-state compression, memory regularization, action abstraction---could be imported directly into discrete amortized sampling.
- The ablation study singles out the rotation matrix as the component that most accelerates convergence, which suggests a cheap improvement test for other recurrent samplers: swap the learned memory update for a rotation-based fast-weight rule before scaling the network.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes to lift the pointed DAG of a GFlowNet by attaching a deterministic, learnable latent dynamical system W_t = W_{t-1} + phi(s_t, W_{t-1}), so that the forward policy can condition on the whole trajectory through W_t. It argues that Markovian policies suffer from state aliasing and limited expressivity, illustrates this with linear LINES and GNN-based graph examples, and proves (Propositions 4.5-4.7) that trajectory balance, subtrajectory balance, and contrastive balance remain valid on the lifted DAG. It also proves (Proposition 4.8) that a Markovian backward policy forces the optimal forward policy to be Markovian. Experiments on set generation, sequence design, grid worlds, preference learning, bit sequences, lazy random walks, and Ising models report improved goodness-of-fit and faster convergence relative to Markovian baselines with comparable parameter counts.
Significance. If the main theoretical claims held as stated, the paper would make a useful contribution: a principled recipe for recurrent/path-dependent GFlowNet policies with the same training objectives as standard GFlowNets, plus a formal demonstration that lifting can strictly increase the representable family in some settings. The proof strategy for Propositions 4.5-4.7 is sound: the lifted graph is a finite pointed DAG for fixed deterministic phi, and the standard balance-condition argument applies. The experimental section is broad, code is promised in the supplement, and the ablations (rotation matrix, link function, RNN baselines) support the claimed mechanism. The main reservation is that one of the headline expressivity results (Proposition 3.1) is not established by its proof, and the appendix proof of Proposition 4.8 contains a nontrivial omission. The central balance-loss extension is nonetheless credible and worth publishing after revision.
major comments (2)
- [Section 3.1, Proposition 3.1; Appendix A.1] The proof equates the expressive power of the Markovian sampler with the dimension of the set of achievable policy log-odds, writing o = Psi k (Eq. 13) and concluding dim W_M = rank Psi. This is not valid: the map from log-odds o to the terminal distribution r is nonlinear, since for M = 1, r_i = (1 - sigma(o_i)) prod_{j < i} sigma(o_j). For N = 2, d = 1 and psi(p_i) = i, the three achievable normalized terminal vectors r(0) = (1/2, 1/4, 1/4), r(-ln 2) = (2/3, 4/15, 1/15), and r(ln 2) = (1/3, 2/15, 8/15) are linearly independent (determinant -1/60), while rank Psi = 1. Hence the linear span of the achievable terminal distributions can have dimension 3 even though the log-odds lie on a line; the claimed inequality dim span(W_M) <= dim span(W_P) and the conclusion dim W_M <= d are not established. The proposition should be reformulated (e.g., in terms of the intrinsic dimension of the achievable distribution manifold) and proved for that statement.
- [Appendix A.4, proof of Proposition 4.8] In the marginalization step, the proof uses Eq. (16) to replace products of pF with products of pB-tilde, but Eq. (16) contains the factor R(s_T) on the right-hand side. When computing pS(s_l | xi*) as a ratio of sums over continuations tau: s_l -> X and tau': s_{l-1} -> X, the factors R(x_tau) and R(x_tau') appear inside the sums and do not cancel with the displayed pB-tilde factors. They are dropped in the manuscript's derivation. The conclusion is recoverable: define A(s_l) = sum_{tau: s_l -> X} R(x_tau) pB-tilde(tau | s_l) and B(s_{l-1}) = sum_{tau': s_{l-1} -> X} R(x_tau') pB-tilde(tau' | s_{l-1}); then the conditional is pB-tilde(s_{l-1} | s_l) A(s_l) / B(s_{l-1}), which still depends only on s_{l-1} and s_l. The proof should be rewritten accordingly, and the notation pB(xi* | s_T) for an incomplete prefix xi* should be clarified or replaced.
minor comments (6)
- [Appendix A.2, proof of Proposition 3.2] The first inclusion is written R_P(S) subset of R_M(S); it should be R_M(S) subset of R_P(S). The same reversal appears in the sentence beginning 'Consequently, R_P(S) subset of R_M(S)'.
- [Appendix A.3] The sentence 'Corollaries 4.6, 4.7, B.1 follow directly from Proposition 3.2' should refer to Proposition 4.5, not Proposition 3.2.
- [Appendix A.1, Eq. (14)] The definition c_i = sum_{0 <= j <= i} psi(p_i) uses the wrong index; it should be psi(p_j).
- [Appendix A.3, proof of Proposition 4.5] In the line 'P_{tau: s_o* -> s*} pF(tau | s*) = 1', the sum should be over the backward policy pB(tau | s*), not pF.
- [Section 5, Table 1 and surrounding text] The text says K is in {16, 32}, but Table 1 reports K = 16 and K = 24; the cited cardinalities 10^14 and 10^18 correspond to K = 16 and K = 32, respectively. Please align the text and table.
- [Section D and Conclusions] The main text and Section D appear to claim that a stochastic latent dynamical system can be developed, while the Conclusions state that the feasibility of a stochastic dynamics remains open; please clarify which parts of Section D are formal results and which are preliminary.
Circularity Check
No significant circularity; the core balance-correctness theorem is proved from stated assumptions and the self-citations are not load-bearing.
full rationale
The paper's central formal claim, Proposition 4.5, is a self-contained proof: assuming the SubTB condition (Eq. 10) on the lifted DAG, it derives p^T(s*) proportional to F(s*) by summing the backward-policy decomposition and using finiteness of trajectories, which follows from the deterministic latent update. No fitted parameter is renamed as a prediction, and no theorem is assumed into its own conclusion. The expressivity propositions are stated as parameter-free comparisons of representable sets; although the proof of Proposition 3.1 contains a gap (it identifies the dimension of policy log-odds with the dimension of the terminal-distribution span), that is a correctness issue rather than a circular reduction. Several cited prior works are by overlapping authors (Silva et al. 2025a for GNN expressivity, da Silva et al. 2024a for CB-implies-TB), but those citations are supporting steps, not the definition of the target result, and the manuscript also gives its own construction in Appendix A.2. The FCS metric is cited from the same authors but is explicitly defined in Eq. (18) and used only for evaluation. Thus the derivation chain is not circular; the modest score reflects only the presence of minor non-load-bearing self-citations.
Assumptions & free parameters
free parameters (3)
- Latent dimension d =
d=32 in experiments
- Rotation matrix R =
fixed ergodic rotation matrix
- Learning rates =
1e-3 for policy, 1e-1 for logZ
assumptions (5)
- domain assumption The lifted DAG SG* (Definition 4.1) with deterministic phi and fixed W0 is a pointed DAG with finitely many trajectories per terminal state.
- standard math The balance conditions (TB, CB, SubTB) are sufficient and their unrestricted optima are achieved by the corresponding objectives.
- domain assumption In Proposition 3.1, the state embeddings are fixed (e.g., psi(p_i)=i for the natural embedding), and the comparison holds within the restricted linear policy class.
- standard math In Proposition 3.2, GNN expressivity is bounded by the Weisfeiler-Lehman hierarchy.
- domain assumption The target distribution is positive: R(x)>0 for all terminal x.
invented entities (2)
-
Latent dynamical system {W_t}
-
Ergodic rotation matrix R inside the SRWM
Cite this review
Pith. "Pith review of Path-dependent Discrete Amortized Inference." pith.science (2026). https://pith.science/paper/CH2YKO6U
@misc{pith2026260808644,
author = {Pith},
title = {Pith review of: Path-dependent Discrete Amortized Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/CH2YKO6U}},
note = {Machine review of arXiv:2608.08644}
}
read the original abstract
We consider the problem of sampling compositional and discrete objects from a given unnormalized posterior distribution. Notably, recent studies have shown that this problem can be efficiently solved by learning a deterministic Markov Decision Process (MDP) that progressively builds each object in proportion to the posterior. In this work, however, we demonstrate that the Markovian assumption can both hamper signal propagation during training and catastrophically reduce the learned sampler's expressivity due to state aliasing. To address these issues, we propose lifting the MDP with a learnable latent dynamical system that allows the underlying policy to depend on the entire past trajectory---and not only on the current state. In view of this, we refer to the resulting method as path-dependent discrete amortized inference. Importantly, we provably extend existing learning algorithms for discrete amortized samplers to our setting. In experiments on standard benchmark problems, we also show that our approach often leads to faster learning convergence and improved state space exploration relatively to prior techniques.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Bengio, Yoshua and LeCun, Yann , booktitle =
-
[2]
and Osindero, Simon and Teh, Yee Whye , journal =
Hinton, Geoffrey E. and Osindero, Simon and Teh, Yee Whye , journal =
-
[3]
2016 , publisher=
Deep Learning , author=. 2016 , publisher=
2016
-
[4]
Pat Langley , booktitle =
-
[5]
Sitan Chen and Sinho Chewi and Jerry Li and Yuanzhi Li and Adil Salim and Anru Zhang , booktitle=
-
[6]
Transactions on Machine Learning Research , note=
Ricky Mol. Transactions on Machine Learning Research , note=
-
[7]
Matsen IV , booktitle=
Cheng Zhang and Frederick A. Matsen IV , booktitle=
-
[8]
Mnih, Andriy and Rezende, Danilo , booktitle=
Show all 300 references
-
[9]
Xu, Pan and Gao, Felicia and Gu, Quanquan , booktitle =
-
[10]
Papini, Matteo and Binaghi, Damiano and Canonaco, Giuseppe and Pirotta, Matteo and Restelli, Marcello , booktitle =
-
[11]
Shakir Mohamed and Mihaela Rosca and Michael Figurnov and Andriy Mnih , journal =
-
[12]
Ming Yang Zhou and Zichao Yan and Elliot Layne and Nikolay Malkin and Dinghuai Zhang and Moksh Jain and Mathieu Blanchette and Yoshua Bengio , booktitle=
-
[13]
Information theory: coding theorems for discrete memoryless systems , author=
-
[14]
Cole, T. J. , year =. Applied Statistics , publisher =. doi:10.2307/2347432 , number =
-
[15]
Elaine Lau and Nikhil Murali Vemgal and Doina Precup and Emmanuel Bengio , booktitle=
-
[16]
Broderick, Tamara and Boyd, Nicholas and Wibisono, Andre and Wilson, Ashia C and Jordan, Michael I , booktitle =
-
[17]
Bui, Thang D and Nguyen, Cuong and Turner, Richard E , journal=
-
[18]
, booktitle =
Fang, Shikai and Wang, Zheng and Pan, Zhimeng and et al. , booktitle =
-
[19]
Schaeffer, Rylan and Du, Yilun and Liu, Gabrielle K and Fiete, Ila , booktitle =
-
[20]
Jaakkola , booktitle=
Timur Garipov and Sebastiaan De Peuter and Ge Yang and Vikas Garg and Samuel Kaski and Tommi S. Jaakkola , booktitle=
-
[21]
Bayesian statistics: A review , author=
-
[22]
Systematic Biology , publisher =
Dinh, Vu and Darling, Aaron E and Matsen IV, Frederick A , editor =. Systematic Biology , publisher =. doi:10.1093/sysbio/syx087 , number =
-
[23]
Kingma, Diederik P and Ba, Jimmy , booktitle=
-
[24]
Sida Li and Ioana Marinescu and Sebastian Musslick , eprint=
-
[25]
, editor =
Bouckaert, Remco et al. , editor =. PLOS Computational Biology , publisher =. doi:10.1371/journal.pcbi.1006650 , number =
-
[26]
Mammalian Protein Metabolism , author =
-
[27]
and Mesquita, Diego and Kaski, Samuel and Acerbi, Luigi , booktitle =
De Souza, Daniel A. and Mesquita, Diego and Kaski, Samuel and Acerbi, Luigi , booktitle =
-
[28]
Ritter, Hippolyt and Botev, Aleksandar and Barber, David , booktitle =
-
[29]
Arindam RoyChoudhury and Amy Willis and John Bunge , journal =
-
[30]
Chen, Pei-Hung and Wei, Wei and Hsieh, Cho-Jui and Dai, Bo , booktitle =
-
[31]
International Conference on Machine Learning , organization=
Gonz. International Conference on Machine Learning , organization=
-
[32]
and Habraken, Hilde and Bloch, Daniel A
Hornberger, John C. and Habraken, Hilde and Bloch, Daniel A. , year =. Medical Care , publisher =. doi:10.1097/00005650-199503000-00008 , number =
-
[33]
Molecular Evolution: A Statistical Approach , ISBN =
Yang, Ziheng , year =. Molecular Evolution: A Statistical Approach , ISBN =. doi:10.1093/acprof:oso/9780199602605.001.0001 , publisher =
-
[34]
T. M. Mitchell , institution =
-
[35]
M. J. Kearns , school =
-
[36]
Machine Learning: An Artificial Intelligence Approach, Vol. I. 1983
1983
-
[37]
R. O. Duda and P. E. Hart and D. G. Stork. Pattern Classification. 2000
2000
-
[38]
Suppressed for Anonymity , author=
-
[39]
Newell and P
A. Newell and P. S. Rosenbloom. Mechanisms of Skill Acquisition and the Law of Practice. Cognitive Skills and Their Acquisition. 1981
1981
-
[40]
A. L. Samuel. Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development. 1959
1959
-
[41]
Zhiqing Sun and Yiming Yang , eprint=
-
[42]
doi:10.1609/aaai.v37i10.26430 , year =
Kai Wang and Zhao Song and Georgios Theocharous and Sridhar Mahadevan , journal =. doi:10.1609/aaai.v37i10.26430 , year =
-
[43]
and et al
Blei, David M. and et al. , journal =
-
[44]
Yutian Chen and Aja Huang and Ziyun Wang and Ioannis Antonoglou and Julian Schrittwieser and David Silver and Nando de Freitas , howpublished=
-
[45]
Larochelle and Ryan P
Jasper Snoek and H. Larochelle and Ryan P. Adams , booktitle=
-
[46]
Javier I. Gonz. International Conference on Machine Learning , year=
-
[47]
Wei Chu and Zoubin Ghahramani , journal=
-
[48]
Nicholas Kramer and Jonathan Schmidt and Philipp Hennig , howpublished=
-
[49]
International Conference on Artificial Intelligence and Statistics , pages=
Kr. International Conference on Artificial Intelligence and Statistics , pages=
-
[50]
Advances in Neural Information Processing Systems , volume=
Kr. Advances in Neural Information Processing Systems , volume=
-
[51]
doi:10.1017/9781316681411 , year =
Probabilistic Numerics , author =. doi:10.1017/9781316681411 , year =
-
[52]
1983 , publisher=
Kirkpatrick, Scott and Gelatt Jr, C Daniel and Vecchi, Mario P , journal=. 1983 , publisher=
1983
-
[53]
Polato, Mirko , booktitle=
-
[54]
and Zhou, Jiayu , year =
Hong, Junyuan and Zhu, Zhuangdi and Yu, Shuyang and Wang, Zhangyang and Dodge, Hiroko H. and Zhou, Jiayu , year =
-
[55]
Chang and H
Q. Chang and H. Qu and Y. Zhang and M. Sabuncu and C. Chen and T. Zhang and D. N. Metaxas , booktitle =
-
[56]
Learn Distributed GAN with Temporary Discriminators
Qu, Hui and Zhang, Yikai and Chang, Qi and Yan, Zhennan and Chen, Chao and Metaxas, Dimitris. Learn Distributed GAN with Temporary Discriminators. European Conference on Computer Vision (ECCV). 2020
2020
-
[57]
and Dunson, David B
Wang, Xiangyu and Guo, Fangjian and Heller, Katherine A. and Dunson, David B. , booktitle =
-
[58]
Vono, Maxime and Plassier, Vincent and Durmus, Alain and Dieuleveut, Aymeric and Moulines, Eric , booktitle =
-
[59]
Scott and Alexander W
Steven L. Scott and Alexander W. Blocker and Fernando V. Bonassi and Hugh A. Chipman and Edward I. George and Robert E. McCulloch , year =. International Journal of Management Science and Engineering Management , volume =
-
[60]
Joseph Felsenstein , journal =
-
[61]
de Souza and D
D. de Souza and D. Mesquita and S. Kaski and L. Acerbi , booktitle =
-
[62]
Neiswanger and C
W. Neiswanger and C. Wang and E. P. Xing , booktitle =
-
[63]
Mesquita and P
D. Mesquita and P. Blomstedt and S. Kaski , booktitle =
-
[64]
Bayesian Analysis
Merging MCMC Subposteriors through G aussian-Process Approximations. Bayesian Analysis. 2018
2018
-
[65]
, journal =
Hinton, Geoffrey E. , journal =. Training Products of Experts by Minimizing Contrastive Divergence. 2002 , month =
2002
-
[66]
Malkin, Nikolay and Lahlou, Salem and Deleu, Tristan and Ji, Xu and Hu, Edward and Everett, Katie and Zhang, Dinghuai and Bengio, Yoshua , booktitle =
-
[67]
Nikolay Malkin and Moksh Jain and Emmanuel Bengio and Chen Sun and Yoshua Bengio , booktitle=
-
[68]
Pan, Ling and Malkin, Nikolay and Zhang, Dinghuai and Bengio, Yoshua , booktitle =
-
[69]
Ling Pan and Dinghuai Zhang and Aaron Courville and Longbo Huang and Yoshua Bengio , booktitle=
-
[70]
Hu and Mo Tiwari and Emmanuel Bengio , journal =
Yoshua Bengio and Salem Lahlou and Tristan Deleu and Edward J. Hu and Mo Tiwari and Emmanuel Bengio , journal =
-
[71]
Bengio, Emmanuel and Jain, Moksh and Korablyov, Maksym and Precup, Doina and Bengio, Yoshua , booktitle =
-
[72]
Tiago da Silva and Eliezer Silva and Adèle Ribeiro and António Góis and Dominik Heider and Samuel Kaski and Diego Mesquita , howpublished=
-
[73]
Uncertainty in Artificial Intelligence (UAI) , year=
Deleu, Tristan and G. Uncertainty in Artificial Intelligence (UAI) , year=
-
[74]
Deleu, Tristan and Nishikawa-Toomey, Mizu and Subramanian, Jithendaraa and Malkin, Nikolay and Charlin, Laurent and Bengio, Yoshua , booktitle=
-
[75]
Dinghuai Zhang and Hanjun Dai and Nikolay Malkin and Aaron Courville and Yoshua Bengio and Ling Pan , booktitle =
-
[76]
Jain, Moksh and Bengio, Emmanuel and Hernandez-Garcia, Alex and Rector-Brooks, Jarrid and Dossou, Bonaventure F. P. and Ekbote, Chanakya Ajit and Fu, Jie and Zhang, Tianyu and Kilgour, Michael and Zhang, Dinghuai and Simine, Lena and Das, Payel and Bengio, Yoshua , booktitle =
-
[77]
Jain, Moksh and Raparthy, Sharath Chandra and Hernandez-Garcia, Alex and Rector-Brooks, Jarrid and Bengio, Yoshua and Miret, Santiago and Bengio, Emmanuel , booktitle =
-
[78]
Linear Programming and Network Flows , author =
-
[79]
Brendan McMahan and E
H. Brendan McMahan and E. Moore and D. Ramage and S. Hampson and B. Agüera y Arcas , booktitle=
-
[80]
El Mekkaoui, Khaoula and Mesquita, Diego and Blomstedt, Paul and Kaski, Samuel , booktitle=
-
[81]
Daulton and M
S. Daulton and M. Balandat and E. Bakshy , booktitle =
-
[82]
Zhang, Cheng and Matsen IV, Frederick A , booktitle=
-
[83]
Zhang, David W and Rainone, Corrado and Peschl, Markus and Bondesan, Roberto , booktitle=
-
[84]
2306.11715 , archivePrefix=
Alex Hernandez-Garcia and Nikita Saxena and Moksh Jain and Cheng-Hao Liu and Yoshua Bengio , year=. 2306.11715 , archivePrefix=
-
[85]
Zhang, Dinghuai and Chen, Ricky TQ and Malkin, Nikolay and Bengio, Yoshua , journal=
-
[86]
Hu, Edward J and Malkin, Nikolay and Jain, Moksh and Everett, Katie E and Graikos, Alexandros and Bengio, Yoshua , booktitle=
-
[87]
Zhang, Dinghuai and Malkin, Nikolay and Liu, Zhen and Volokhova, Alexandra and Courville, Aaron and Bengio, Yoshua , booktitle=
-
[88]
2023 , publisher =
Liu et al., Dianbo , booktitle =. 2023 , publisher =
2023
-
[89]
and Bengio, Emmanuel and Hajiramezanali, Ehsan and Loukas, Andreas and Cho, Kyunghyun and Biancalani, Tommaso , booktitle =
Shen, Max W. and Bengio, Emmanuel and Hajiramezanali, Ehsan and Loukas, Andreas and Cho, Kyunghyun and Biancalani, Tommaso , booktitle =
-
[90]
Maddison and Andriy Mnih and Yee Whye Teh , booktitle=
Chris J. Maddison and Andriy Mnih and Yee Whye Teh , booktitle=
-
[91]
Han, Jun and Ding, Fan and Liu, Xianglong and Torresani, Lorenzo and Peng, Jian and Liu, Qiang , booktitle=
-
[92]
2022 , organization=
Wojnowicz, Michael T and Aeron, Shuchin and Miller, Eric L and Hughes, Michael , booktitle=. 2022 , organization=
2022
-
[93]
On the asymptotic behaviour of posterior distributions
Walker, A M. On the asymptotic behaviour of posterior distributions. J. R. Stat. Soc
-
[94]
IEEE Transactions on Pattern Analysis and Machine Intelligence , author =
-
[95]
Eric Jang and Shixiang Gu and Ben Poole , booktitle=
-
[96]
Hu and Moksh Jain and Eric Elmoznino and Younesse Kaddar and et al
Edward J. Hu and Moksh Jain and Eric Elmoznino and Younesse Kaddar and et al. , booktitle=
-
[97]
Li, Yinchuan and Luo, Shuang and Wang, Haozhi and Hao, Jianye , journal=
-
[98]
An Excursion through Elementary Mathematics, Volume I , ISBN =
Caminha Muniz Neto, Antonio , year =. An Excursion through Elementary Mathematics, Volume I , ISBN =. doi:10.1007/978-3-319-53871-6 , journal =
-
[99]
Lahlou at al., Salem , booktitle=
-
[100]
International Conference on Machine Learning , pages=
A theory of continuous generative flow networks , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[101]
Pan, Ling and Zhang, Dinghuai and Jain, Moksh and Huang, Longbo and Bengio, Yoshua , booktitle=
-
[102]
Zhang, Dinghuai and Pan, Ling and Chen, Ricky TQ and Courville, Aaron and Bengio, Yoshua , journal=
-
[103]
2023 , eprint=
Tristan Deleu and Yoshua Bengio , howpublished=. 2023 , eprint=
2023
-
[104]
Shen, Max W and Bengio, Emmanuel and Hajiramezanali, Ehsan and Loukas, Andreas and Cho, Kyunghyun and Biancalani, Tommaso , booktitle=
-
[105]
Lee and Bo Wang and Yoshua Bengio , booktitle=
Lazar Atanackovic and Alexander Tong and Jason Hartford and Leo J. Lee and Bo Wang and Yoshua Bengio , booktitle=
-
[106]
Ilya Loshchilov and Frank Hutter , booktitle=
-
[107]
Maas, Andrew L and Hannun, Awni Y and Ng, Andrew Y and others , booktitle=
-
[108]
, booktitle=
Fey, Matthias and Lenssen, Jan E. , booktitle=
-
[109]
Paszke, Adam and Gross, Sam and Massa, Francisco and Lerer, Adam and Bradbury, James and Chanan, Gregory and Killeen, Trevor and Lin, Zeming and Gimelshein, Natalia and Antiga, Luca and Desmaison, Alban and Kopf, Andreas and Yang, Edward and DeVito, Zachary and Raison, Martin ...
-
[110]
Xu, Keyulu and Hu, Weihua and Leskovec, Jure and Jegelka, Stefanie , journal=
-
[111]
2310.02423 , archivePrefix=
Jean-Pierre Falet and Hae Beom Lee and Nikolay Malkin and Chen Sun and Dragos Secrieru and Dinghuai Zhang and Guillaume Lajoie and Yoshua Bengio , year=. 2310.02423 , archivePrefix=
-
[112]
Marco Jiralerspong and Bilun Sun and Danilo Vucetic and Tianyu Zhang and Yoshua Bengio and Gauthier Gidel and Nikolay Malkin , booktitle=
-
[113]
1953 , publisher=
Metropolis, Nicholas and Rosenbluth, Arianna W and Rosenbluth, Marshall N and Teller, Augusta H and Teller, Edward , journal=. 1953 , publisher=
1953
-
[114]
Dinghuai Zhang and Ricky Tian Qi Chen and Cheng-Hao Liu and Aaron Courville and Yoshua Bengio , booktitle=
-
[115]
Kanika Madan and Jarrid Rector-Brooks and Maksym Korablyov and Emmanuel Bengio and Moksh Jain and Andrei Cristian Nica and Tom Bosc and Yoshua Bengio and Nikolay Malkin , booktitle=
-
[116]
2026 , eprint=
Boosted GFlowNets: Improving Exploration via Sequential Learning , author=. 2026 , eprint=
2026
-
[117]
2024 , eprint=
Generative Marginalization Models , author=. 2024 , eprint=
2024
-
[118]
2022 , eprint=
Generative Flow Networks for Discrete Probabilistic Modeling , author=. 2022 , eprint=
2022
-
[119]
Nica, Andrei Cristian and Jain, Moksh and Bengio, Emmanuel and Liu, Cheng-Hao and Korablyov, Maksym and Bronstein, Michael M and Bengio, Yoshua , journal=
-
[120]
Ihler and John W
Qiang Liu and Jian Peng and Alexander T. Ihler and John W. Fisher III , booktitle=
-
[121]
Jang, Hyosoon and Kim, Minsu and Ahn, Sungsoo , booktitle=
-
[122]
Zhou, Mingyang and Yan, Zichao and Layne, Elliot and Malkin, Nikolay and Zhang, Dinghuai and Jain, Moksh and Blanchette, Mathieu and Bengio, Yoshua , booktitle=
-
[123]
Silva, Tiago and Souza, Amauri H and Carvalho, Luiz Max and Kaski, Samuel and Mesquita, Diego , howpublished=
-
[124]
International Conference on Machine Learning (ICML) , publisher =
Sohl-Dickstein, Jascha and Weiss, Eric and Maheswaranathan, Niru and Ganguli, Surya , year =. International Conference on Machine Learning (ICML) , publisher =
-
[125]
2020 , organization=
Buesing, Lars and Heess, Nicolas and Weber, Theophane , booktitle=. 2020 , organization=
2020
-
[126]
Zhang and Shaoqing Ren and Jian Sun , journal=
Kaiming He and X. Zhang and Shaoqing Ren and Jian Sun , journal=
-
[127]
doi:10.1109/tpami.1984.4767596 , year =
Stuart Geman and Donald Geman , journal =. doi:10.1109/tpami.1984.4767596 , year =
1984
-
[128]
International Conference on Machine Learning , year=
Yoshua Bengio and J. International Conference on Machine Learning , year=
-
[129]
Bissiri, P. G. and Holmes, C. C. and Walker, S. G. , title = ". Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =. 2016 , month =
2016
-
[130]
2009 , isbn =
Analytic Combinatorics , author =. 2009 , isbn =
2009
-
[131]
Goodfellow and Mehdi Mirza and Xia Da and Aaron C
Ian J. Goodfellow and Mehdi Mirza and Xia Da and Aaron C. Courville and Yoshua Bengio , booktitle =
-
[132]
Cohen , editor =
Michael McCloskey and Neal J. Cohen , editor =. Psychology of Learning and Motivation , publisher =. doi:https://doi.org/10.1016/S0079-7421(08)60536-8 , year =
-
[133]
2022 , volume =
Jeremias Knoblauch and Jack Jewson and Theodoros Damoulas , journal =. 2022 , volume =
2022
-
[134]
2024 , eprint=
Lazar Atanackovic and Emmanuel Bengio , journal=. 2024 , eprint=
2024
-
[135]
2407.03105 , archivePrefix=
Anas Krichel and Nikolay Malkin and Salem Lahlou and Yoshua Bengio , year=. 2407.03105 , archivePrefix=
-
[136]
Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2 , pages=
Seldin, Yevgeny and Cesa-Bianchi, Nicol. Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2 , pages=. 2012 , organization=
2012
-
[137]
Sanae Lotfi and Yilun Kuang and Brandon Amos and Micah Goldblum and Marc Finzi and Andrew Gordon Wilson , booktitle=
-
[138]
Sanae Lotfi and Marc Finzi and Yilun Kuang and Tim G. J. Rudner and Micah Goldblum and Andrew Gordon Wilson , booktitle=. 2024 , organization=
2024
-
[139]
2012 , publisher=
Yevgeny Seldin and François Laviolette and Nicolò Cesa-Bianchi and John Shawe-Taylor and Peter Auer , journal=. 2012 , publisher=
2012
-
[140]
Haddouche, Maxime and Guedj, Benjamin , journal=
-
[141]
Schapire , booktitle=
Alina Beygelzimer and John Langford and Lihong Li and Lev Reyzin and Robert E. Schapire , booktitle=. 2011 , organization=
2011
-
[142]
, booktitle=
Dziugaite, Gintare Karolina and Hsu, Kyle and Gharbieh, Waseem and Arpino, Gabriel and Roy, Daniel M. , booktitle=. 2021 , organization=
2021
-
[143]
Dziugaite, Gintare Karolina and Roy, Daniel M , booktitle=
-
[144]
Entropy , publisher=
Haddouche, Maxime and Guedj, Benjamin and Rivasplata, Omar and Shawe-Taylor, John , year=. Entropy , publisher=. doi:10.3390/e23101330 , number=
-
[145]
Advances in Neural Information Processing Systems , year=
Casado, Ioar and Ortega, Luis A and P. Advances in Neural Information Processing Systems , year=
-
[146]
Machine Learning , publisher=
Alquier, Pierre and Guedj, Benjamin , year=. Machine Learning , publisher=. doi:10.1007/s10994-017-5690-0 , number=
-
[147]
Journal of Machine Learning Research , volume=
Rodr. Journal of Machine Learning Research , volume=
-
[148]
2024 IEEE International Symposium on Information Theory (ISIT) , pages=
Rodr. 2024 IEEE International Symposium on Information Theory (ISIT) , pages=. 2024 , organization=
2024
-
[149]
Rivasplata, Omar and Tankasali, Vikram M and Szepesvari, Csaba , howpublished=
-
[150]
2013 , eprint=
David McAllester , howpublished=. 2013 , eprint=
2013
-
[151]
2004 , eprint=
Andreas Maurer , howpublished=. 2004 , eprint=
2004
-
[152]
2013 , publisher=
Concentration inequalities: A nonasymptotic theory of independence , author=. 2013 , publisher=
2013
-
[153]
2402.05098 , archivePrefix=
Marcin Sendera and Minsu Kim and Sarthak Mittal and Pablo Lemos and Luca Scimeca and Jarrid Rector-Brooks and Alexandre Adam and Yoshua Bengio and Nikolay Malkin , year=. 2402.05098 , archivePrefix=
-
[154]
Siddarth Venkatraman and Moksh Jain and Luca Scimeca and Minsu Kim and Marcin Sendera and Mohsin Hasan and Luke Rowe and Sarthak Mittal and Pablo Lemos and Emmanuel Bengio and Alexandre Adam and Jarrid Rector-Brooks and Yoshua Bengio and Glen Berseth and Nikolay Malkin , booktitle=
-
[155]
2023 , url=
Mbacke, Sokhna Diarra and Clerc, Florence and Germain, Pascal , booktitle=. 2023 , url=
2023
-
[156]
2309.04381 , archivePrefix=
Fredrik Hellström and Giuseppe Durisi and Benjamin Guedj and Maxim Raginsky , year=. 2309.04381 , archivePrefix=
-
[157]
Vincent Poor , booktitle=
Semih Yagli and Alex Dytso and H. Vincent Poor , booktitle=. 2020 , url=
2020
-
[158]
Entropy , publisher=
Barnes, Leighton Pate and Dytso, Alex and Poor, Harold Vincent , year=. Entropy , publisher=. doi:10.3390/e24091178 , number=
-
[159]
Milad Sefidgaran and Romain Chor and Abdellatif Zaidi , booktitle=
-
[160]
2024 , eprint=
Pierre Alquier , journal=. 2024 , eprint=
2024
-
[161]
2019 , url=
Guedj, Benjamin , booktitle=. 2019 , url=
2019
-
[162]
2014 , editor =
Honorio, Jean and Jaakkola, Tommi , booktitle =. 2014 , editor =
2014
-
[163]
1910.10367 , archivePrefix=
Sanjay Thakur and Herke Van Hoof and Gunshi Gupta and David Meger , year=. 1910.10367 , archivePrefix=
1910 arXiv
-
[164]
2202.11455 , archivePrefix=
Badr-Eddine Chérief-Abdellatif and Yuyang Shi and Arnaud Doucet and Benjamin Guedj , year=. 2202.11455 , archivePrefix=
-
[165]
Pascal Germain and Francis Bach and Alexandre Lacoste and Simon Lacoste-Julien , booktitle=
-
[166]
Symposium on Advances in Approximate Bayesian Inference , pages=
Tasdighi, Bahareh and Akg. Symposium on Advances in Approximate Bayesian Inference , pages=. 2024 , url=
2024
-
[167]
and Pineau, Joelle , booktitle =
Fard, M. and Pineau, Joelle , booktitle =
-
[168]
Catoni, Olivier , journal=
-
[169]
McAllester, David A , booktitle=
-
[170]
Probabilistic Methods for Algorithmic Discrete Mathematics , pages=
Concentration , author=. Probabilistic Methods for Algorithmic Discrete Mathematics , pages=. 1998 , publisher=
1998
-
[171]
Zhang , year=
Haotian Ju and Dongyue Li and Aneesh Sharma and Hongyang R. Zhang , year=. 2302.04451 , archivePrefix=
-
[172]
2310.02710 , archivePrefix=
Minsu Kim and Taeyoung Yun and Emmanuel Bengio and Dinghuai Zhang and Yoshua Bengio and Sungsoo Ahn and Jinkyoo Park , year=. 2310.02710 , archivePrefix=
-
[173]
2306.17693 , archivePrefix=
Jarrid Rector-Brooks and Kanika Madan and Moksh Jain and Maksym Korablyov and Cheng-Hao Liu and Sarath Chandar and Nikolay Malkin and Yoshua Bengio , year=. 2306.17693 , archivePrefix=
-
[174]
2023 , eprint=
Nikhil Vemgal and Elaine Lau and Doina Precup , journal=. 2023 , eprint=
2023
-
[175]
Statistical Learning Theory , author=
-
[176]
, year =
Vapnik, Vladimir N. , year =. The Nature of Statistical Learning Theory , url =
-
[177]
Communications of the ACM , volume=
A theory of the learnable , author=. Communications of the ACM , volume=. 1984 , publisher=
1984
-
[178]
2024 , url=
Hyosoon Jang and Minsu Kim and Sungsoo Ahn , booktitle=. 2024 , url=
2024
-
[179]
Shawe-Taylor, John and Bartlett, Peter L and Williamson, Robert C and Anthony, Martin , booktitle=
-
[180]
Shawe-Taylor, John and Williamson, Robert C , booktitle=
-
[181]
2014 , publisher=
Understanding Machine Learning: From Theory to Algorithms , author=. 2014 , publisher=
2014
-
[182]
2024 , url=
Tiago Silva and Eliezer de Souza da Silva and Rodrigo Barreto Alves and Luiz Max Carvalho and Amauri H Souza and Samuel Kaski and Vikas Garg and Diego Mesquita , booktitle=. 2024 , url=
2024
-
[183]
2406.03288 , archivePrefix=
Tiago da Silva and Luiz Max Carvalho and Amauri Souza and Samuel Kaski and Diego Mesquita , year=. 2406.03288 , archivePrefix=
-
[184]
2013 , editor =
Ma, Jianzhu and Peng, Jian and Wang, Sheng and Xu, Jinbo , booktitle =. 2013 , editor =
2013
-
[185]
1967 , publisher=
Azuma, Kazuoki , journal=. 1967 , publisher=
1967
-
[186]
Third Symposium on Advances in Approximate Bayesian Inference , year=
Emiel Hoogeboom and Didrik Nielsen and Priyank Jaini and Patrick Forr. Third Symposium on Advances in Approximate Bayesian Inference , year=
-
[187]
2019 , url=
Dustin Tran and Keyon Vafa and Kumar Krishna Agrawal and Laurent Dinh and Ben Poole , booktitle=. 2019 , url=. 1905.10347 , archivePrefix=
2019 arXiv
-
[188]
2024 , eprint=
Daniil Tiapkin and Nikita Morozov and Alexey Naumov and Dmitry Vetrov , booktitle=. 2024 , eprint=
2024
-
[189]
2310.02823 , archivePrefix=
Minsu Kim and Joohwan Ko and Taeyoung Yun and Dinghuai Zhang and Ling Pan and Woochang Kim and Jinkyoo Park and Emmanuel Bengio and Yoshua Bengio , year=. 2310.02823 , archivePrefix=
-
[190]
Statistical tables for biological, agricultural and medical research , author=
-
[191]
Yihang Chen and Lukas Mauch , booktitle=
-
[192]
2022 , organization=
Brandon Trabucco and Xinyang Geng and Aviral Kumar and Sergey Levine , booktitle=. 2022 , organization=. 2202.08450 , archivePrefix=
2022 arXiv
-
[193]
International Conference on Machine Learning , pages=
Design-bench: Benchmarks for data-driven offline model-based optimization , author=. International Conference on Machine Learning , pages=. 2022 , organization=
2022
-
[194]
2016 , publisher=
Barrera, Luis A and Vedenko, Anastasia and Kurland, Jesse V and Rogers, Julia M and Gisselbrecht, Stephen S and Rossin, Elizabeth J and Woodard, Jaie and Mariani, Luca and Kock, Kian Hong and Inukai, Sachi and others , journal=. 2016 , publisher=
2016
-
[195]
Li, Puheng and Li, Zhong and Zhang, Huishuai and Bian, Jiang , booktitle=
-
[196]
2023 , organization=
Ju, Haotian and Li, Dongyue and Sharma, Aneesh and Zhang, Hongyang R , booktitle=. 2023 , organization=
2023
-
[197]
2023 , organization=
Tang, Huayi and Liu, Yong , booktitle=. 2023 , organization=
2023
-
[198]
2015 , publisher=
Vapnik, Vladimir N and Chervonenkis, A Ya , booktitle=. 2015 , publisher=
2015
-
[199]
2402.05234 , archivePrefix=
Elaine Lau and Stephen Zhewen Lu and Ling Pan and Doina Precup and Emmanuel Bengio , year=. 2402.05234 , archivePrefix=
-
[200]
Mankowitz and Timothy A
Daniel J. Mankowitz and Timothy A. Mann and Shie Mannor , booktitle=. 2016 , eprint=
2016
-
[201]
Advances in neural information processing systems , volume=
Adaptive skills adaptive partitions (ASAP) , author=. Advances in neural information processing systems , volume=
-
[202]
Malik and Salem Lahlou and Andrew Jesson and Moksh Jain and Nikolay Malkin and Tristan Deleu and Yoshua Bengio and Yarin Gal , year=
Shreshth A. Malik and Salem Lahlou and Andrew Jesson and Moksh Jain and Nikolay Malkin and Tristan Deleu and Yoshua Bengio and Yarin Gal , year=. 2306.15058 , archivePrefix=
-
[203]
1994 , url=
Cohn, David and Atlas, Les and Ladner, Richard , journal=. 1994 , url=
1994
-
[204]
2017 , organization=
Gal, Yarin and Islam, Riashat and Ghahramani, Zoubin , booktitle=. 2017 , organization=
2017
-
[205]
2402.10309 , archivePrefix=
Tristan Deleu and Padideh Nouri and Nikolay Malkin and Doina Precup and Yoshua Bengio , year=. 2402.10309 , archivePrefix=
-
[206]
2014 , editor =
London, Ben and Huang, Bert and Taskar, Ben and Getoor, Lise , booktitle =. 2014 , editor =
2014
-
[207]
Wu, Yi-Shan and Masegosa, Andres and Lorenzen, Stephan and Igel, Christian and Seldin, Yevgeny , booktitle =
-
[208]
2019 , eprint=
Kohei Miyaguchi , howpublished=. 2019 , eprint=
2019
-
[209]
Holland, Matthew , booktitle =
-
[210]
2015 , eprint=
Akshay Balsubramani , howpublished=. 2015 , eprint=
2015
-
[211]
2022 , url=
Maxime Haddouche and Benjamin Guedj , booktitle=. 2022 , url=
2022
-
[212]
1211.1847 , archivePrefix=
Pierre Alquier and Xiaoyin Li and Olivier Wintenberger , year=. 1211.1847 , archivePrefix=
-
[213]
2020 , eprint=
Omar Rivasplata and Ilja Kuzborskij and Csaba Szepesvari and John Shawe-Taylor , booktitle=. 2020 , eprint=
2020
-
[214]
Tim Dettmers and Artidoro Pagnoni and Ari Holtzman and Luke Zettlemoyer , booktitle=
-
[215]
Sharma, Apoorva and Veer, Sushant and Hancock, Asher and Yang, Heng and Pavone, Marco and Majumdar, Anirudha , journal=
-
[216]
2023 , organization=
Sakhi, Otmane and Alquier, Pierre and Chopin, Nicolas , booktitle=. 2023 , organization=
2023
-
[217]
2023 , organization=
Biggs, Felix and Guedj, Benjamin , booktitle=. 2023 , organization=
2023
-
[218]
2022 , organization=
Biggs, Felix and Guedj, Benjamin , booktitle=. 2022 , organization=
2022
-
[219]
Journal of Machine Learning Research , volume=
P. Journal of Machine Learning Research , volume=
-
[220]
Bengio, Yoshua and Malkin, Nikolay , journal=
-
[221]
2001 , publisher=
Myrvold, Wendy and Ruskey, Frank , journal=. 2001 , publisher=
2001
-
[222]
1997 , publisher=
Liebehenschel, Jens , journal=. 1997 , publisher=
1997
-
[223]
2024 , organization=
Malach, Eran , booktitle=. 2024 , organization=
2024
-
[224]
1609.04836 , archivePrefix=
Nitish Shirish Keskar and Dheevatsa Mudigere and Jorge Nocedal and Mikhail Smelyanskiy and Ping Tak Peter Tang , year=. 1609.04836 , archivePrefix=
-
[225]
2024 , eprint=
Maxime Haddouche and Paul Viallard and Umut Simsekli and Benjamin Guedj , booktitle=. 2024 , eprint=
2024
-
[226]
Pan Zhou and Jiashi Feng and Chao Ma and Caiming Xiong and Steven Hoi and Weinan E , booktitle=
-
[227]
Neural computation , volume=
Hochreiter, Sepp and Schmidhuber, J. Neural computation , volume=. 1997 , publisher=
1997
-
[228]
Pandey, Mohit and Subbaraj, Gopeshh and Bengio, Emmanuel , journal=
-
[229]
Roy, Julien and Bacon, Pierre-Luc and Pal, Christopher and Bengio, Emmanuel , journal=
-
[230]
2402.01103 , archivePrefix=
Yilun Du and Leslie Kaelbling , year=. 2402.01103 , archivePrefix=
-
[231]
International conference on machine learning , pages=
Deep linear networks with arbitrary loss: All local minima are global , author=. International conference on machine learning , pages=. 2018 , organization=
2018
-
[232]
arXiv preprint arXiv:1810.02032 , year=
Gradient descent aligns the layers of deep linear networks , author=. arXiv preprint arXiv:1810.02032 , year=
-
[233]
Neural Networks , volume=
Complex-valued autoencoders , author=. Neural Networks , volume=. 2012 , publisher=
2012
-
[234]
Neural networks , volume=
Neural networks and principal component analysis: Learning from examples without local minima , author=. Neural networks , volume=. 1989 , publisher=
1989
-
[235]
Generalization and Distributed Learning of
Tiago Silva and Amauri H Souza and Omar Rivasplata and Vikas Garg and Samuel Kaski and Diego Mesquita , booktitle=. Generalization and Distributed Learning of. 2025 , url=
2025
-
[236]
2024 , eprint=
Streaming Bayes GFlowNets , author=. 2024 , eprint=
2024
-
[237]
Tiago Silva and Rodrigo Barreto Alves and Eliezer de Souza da Silva and Amauri H Souza and Vikas Garg and Samuel Kaski and Diego Mesquita , booktitle=. When do. 2025 , url=
2025
-
[238]
2022 , eprint=
A Modern Self-Referential Weight Matrix That Learns to Modify Itself , author=. 2022 , eprint=
2022
-
[239]
2021 , eprint=
Linear Transformers Are Secretly Fast Weight Programmers , author=. 2021 , eprint=
2021
-
[240]
Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks , volume =
Schmidhuber, J\". Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks , volume =. Neural Computation , publisher =. 1992 , month = jan, pages =. doi:10.1162/neco.1992.4.1.131 , number =
1992 doi
-
[241]
2024 , eprint=
Action abstractions for amortized sampling , author=. 2024 , eprint=
2024
-
[242]
1991 , publisher=
Markov chain Monte Carlo maximum likelihood , author=. 1991 , publisher=
1991
-
[243]
1999 , publisher=
Monte Carlo statistical methods , author=. 1999 , publisher=
1999
-
[244]
Handbook of markov chain monte carlo , volume=
MCMC using Hamiltonian dynamics , author=. Handbook of markov chain monte carlo , volume=. 2011 , publisher=
2011
-
[245]
1833 , publisher=
On a General Method of Expressing the Paths of Light, & of the Planets, by the Coefficients of a Characteristic Function , author=. 1833 , publisher=
-
[246]
arXiv preprint arXiv:1701.02434 , year=
A conceptual introduction to Hamiltonian Monte Carlo , author=. arXiv preprint arXiv:1701.02434 , year=
-
[247]
Journal of statistical software , volume=
Stan: A probabilistic programming language , author=. Journal of statistical software , volume=
-
[248]
2025 , eprint=
Secrets of GFlowNets' Learning Behavior: A Theoretical Study , author=. 2025 , eprint=
2025
-
[249]
Neural networks , volume=
Multilayer feedforward networks are universal approximators , author=. Neural networks , volume=. 1989 , publisher=
1989
-
[250]
2020 , eprint=
Stein Variational Inference for Discrete Distributions , author=. 2020 , eprint=
2020
-
[251]
1998 , publisher=
Reinforcement learning: An introduction , author=. 1998 , publisher=
1998
-
[252]
2018 , publisher=
Foundations of machine learning , author=. 2018 , publisher=
2018
-
[253]
2025 , eprint=
Symmetry-Aware GFlowNets , author=. 2025 , eprint=
2025
-
[254]
2019 , eprint=
Neural Ordinary Differential Equations , author=. 2019 , eprint=
2019
-
[255]
Neural computation , volume=
Long short-term memory , author=. Neural computation , volume=. 1997 , publisher=
1997
-
[256]
The annals of mathematical statistics , volume=
Statistical inference for probabilistic functions of finite state Markov chains , author=. The annals of mathematical statistics , volume=. 1966 , publisher=
1966
-
[257]
Advances in neural information processing systems , volume=
Amortizing intractable inference in diffusion models for vision, language, and control , author=. Advances in neural information processing systems , volume=
-
[258]
2025 , eprint=
Fast weight programming and linear transformers: from machine learning to neurobiology , author=. 2025 , eprint=
2025
-
[259]
2020 , eprint=
Temporal Graph Networks for Deep Learning on Dynamic Graphs , author=. 2020 , eprint=
2020
-
[260]
2022 , eprint=
Provably expressive temporal graph networks , author=. 2022 , eprint=
2022
-
[261]
Towards Improving Exploration through Sibling Augmented
Kanika Madan and Alex Lamb and Emmanuel Bengio and Glen Berseth and Yoshua Bengio , booktitle=. Towards Improving Exploration through Sibling Augmented. 2025 , url=
2025
-
[262]
Mathematics of control, signals and systems , volume=
Approximation by superpositions of a sigmoidal function , author=. Mathematics of control, signals and systems , volume=. 1989 , publisher=
1989
-
[263]
Advances in neural information processing systems , volume=
The expressive power of neural networks: A view from the width , author=. Advances in neural information processing systems , volume=
-
[264]
2018 , eprint=
Approximating Continuous Functions by ReLU Nets of Minimal Width , author=. 2018 , eprint=
2018
-
[265]
2014 , eprint=
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks , author=. 2014 , eprint=
2014
-
[266]
2020 , eprint=
A Survey of Inverse Reinforcement Learning: Challenges, Methods and Progress , author=. 2020 , eprint=
2020
-
[267]
2023 , eprint=
Towards Theoretical Understanding of Inverse Reinforcement Learning , author=. 2023 , eprint=
2023
-
[268]
2017 , eprint=
Semi-Supervised Classification with Graph Convolutional Networks , author=. 2017 , eprint=
2017
-
[269]
2020 , publisher=
Graph representation learning , author=. 2020 , publisher=
2020
-
[270]
2021 , eprint=
Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges , author=. 2021 , eprint=
2021
-
[271]
Nature Reviews Methods Primers , volume=
Graph neural networks , author=. Nature Reviews Methods Primers , volume=. 2024 , publisher=
2024
-
[272]
2020 , eprint=
Inductive Representation Learning on Temporal Graphs , author=. 2020 , eprint=
2020
-
[273]
2022 , eprint=
Inductive Representation Learning in Temporal Networks via Causal Anonymous Walks , author=. 2022 , eprint=
2022
-
[274]
2025 , eprint=
torchgfn: A PyTorch GFlowNet library , author=. 2025 , eprint=
2025
-
[275]
2013 , publisher=
Stochastic differential equations: an introduction with applications , author=. 2013 , publisher=
2013
-
[276]
Applied Stochastic Differential Equations , ISBN =
S\". Applied Stochastic Differential Equations , ISBN =. doi:10.1017/9781108186735 , publisher =
-
[277]
The Thirteenth International Conference on Learning Representations , year=
Adaptive teachers for amortized samplers , author=. The Thirteenth International Conference on Learning Representations , year=
-
[278]
2017 , eprint=
QMDP-Net: Deep Learning for Planning under Partial Observability , author=. 2017 , eprint=
2017
-
[279]
and Hinton, Geoffrey E
Rumelhart, David E. and Hinton, Geoffrey E. and Williams, Ronald J. , year =. Learning representations by back-propagating errors , volume =. Nature , publisher =. doi:10.1038/323533a0 , number =
-
[280]
2014 , eprint=
Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling , author=. 2014 , eprint=
2014
-
[281]
2022 , eprint=
Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs , author=. 2022 , eprint=
2022
-
[282]
International conference on artificial neural networks , pages=
A ‘self-referential’weight matrix , author=. International conference on artificial neural networks , pages=. 1993 , organization=
1993
-
[283]
1993 third international conference on artificial neural networks , pages=
An'introspective'network that can learn to run its own weight change algorithm , author=. 1993 third international conference on artificial neural networks , pages=. 1993 , organization=
1993
-
[284]
Communications of the ACM , volume=
The hardware lottery , author=. Communications of the ACM , volume=. 2021 , publisher=
2021
-
[285]
2025 , eprint=
An efficient probabilistic hardware architecture for diffusion-like models , author=. 2025 , eprint=
2025
-
[286]
2024 , eprint=
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks , author=. 2024 , eprint=
2024
-
[287]
2020 , eprint=
VarGrad: A Low-Variance Gradient Estimator for Variational Inference , author=. 2020 , eprint=
2020
-
[288]
A Theory of Non-acyclic Generative Flow Networks , volume=
Brunswic, Leo and Li, Yinchuan and Xu, Yushun and Feng, Yijun and Jui, Shangling and Ma, Lizhuang , year=. A Theory of Non-acyclic Generative Flow Networks , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , publisher=. doi:10.1609/aaai.v38i10.28989 , number=
-
[289]
2025 , eprint=
Revisiting Non-Acyclic GFlowNets in Discrete Environments , author=. 2025 , eprint=
2025
-
[290]
James Bradbury and Roy Frostig and Peter Hawkins and Matthew James Johnson and Chris Leary and Dougal Maclaurin and George Necula and Adam Paszke and Jake Vander
-
[291]
2018 , eprint=
Underdamped Langevin MCMC: A non-asymptotic analysis , author=. 2018 , eprint=
2018
-
[292]
Sur quelques propri
Cayley, Arthur , year=. Sur quelques propri
-
[293]
2022 , eprint=
Time Limits in Reinforcement Learning , author=. 2022 , eprint=
2022
-
[294]
and Ballard, Dana H
Whitehead, Steven D. and Ballard, Dana H. , year =. Learning to Perceive and Act by Trial and Error , volume =. Machine Learning , publisher =. doi:10.1023/a:1022619109594 , number =
-
[295]
2022 , eprint=
The Logic of Graph Neural Networks , author=. 2022 , eprint=
2022
-
[296]
Weisfeiler, Boris and Lehman, A. A. , biburl =. Nauchno-Technicheskaya Informatsia , keywords =
-
[297]
2019 , eprint=
Lipschitz regularity of deep neural networks: analysis and efficient estimation , author=. 2019 , eprint=
2019
-
[298]
The annals of mathematical statistics , volume=
On information and sufficiency , author=. The annals of mathematical statistics , volume=. 1951 , publisher=
1951
-
[299]
Journal of the American statistical association , volume=
Sampling-based approaches to calculating marginal densities , author=. Journal of the American statistical association , volume=. 1990 , publisher=
1990
-
[300]
Statistical Science , pages=
Honest exploration of intractable probability distributions via Markov chain Monte Carlo , author=. Statistical Science , pages=. 2001 , publisher=
2001
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.