REVIEW 2 major objections 1 minor 300 references
Fast and Slow Variational Continual Learning
T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Merging past posteriors creates priors that slow knowledge drift while enabling fast VCL updates.
desk verdict CoVON adds a posterior-merging step for slow adaptation inside VCL/IVON, but the merge itself stays underspecified and the reported gains rest on evidence the abstract does not show. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Merging of past posteriors to produce the prior used in each VCL update step inside the CoVON optimizer derived from IVON.
What would settle it
An experiment on any of the paper's benchmarks where CoVON produces the same or higher forgetting rates and lower accuracy than standard VCL would show the merging step does not deliver the claimed benefit.
Extended reading notes
Core claim
Merging past posteriors slows the drift in knowledge as learning progresses, and the merged posterior then serves as the prior in the VCL update to realize fast-weight updates. These steps integrate directly into the IVON optimizer to yield the CoVON optimizer, which improves over prior VCL methods and other weight-regularization approaches across the evaluated continual learning settings.
Load-bearing premise
Merging past posteriors reliably yields a prior that slows subsequent knowledge drift without causing new forgetting or demanding task-specific tuning.
Editorial extensions
If this is right
- CoVON improves performance over existing VCL optimizers in domain-incremental learning.
- It outperforms other weight-regularization strategies during continual pre-training.
- It yields better results than baselines when fine-tuning large language models.
- The optimizer retains nearly the same form and computational cost as Adam.
Reading between the lines
- The merging step could be ported to other variational continual learning optimizers beyond those based on IVON.
- The same slow-fast structure might apply to non-variational continual learning methods that already maintain some form of posterior or momentum state.
- If the merging operation generalizes, it offers a route to continual adaptation in streaming settings without explicit task boundaries.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes CoVON, a variant of the IVON optimizer within the variational continual learning (VCL) framework. It incorporates slow adaptation by merging past posteriors to form a prior that slows knowledge drift, then uses this prior for fast-weight VCL updates. The method is claimed to be seamlessly implementable with costs similar to Adam and to deliver consistent gains over prior VCL optimizers and other weight-regularization baselines across domain-incremental learning, continual pre-training, and LLM fine-tuning.
Significance. If the empirical claims hold and the merging step proves general, the work supplies a low-overhead mechanism for balancing stability and plasticity inside a standard optimizer, which could be useful for sequential training of large models. The near-identical cost to Adam and the reuse of the existing VCL posterior-as-prior construction are practical strengths.
major comments (2)
- [Abstract] Abstract: the central mechanism—'merging of past posteriors' to produce the slow-adaptation prior—is stated at a high level only. No functional form (weighted average, product of Gaussians, moment matching, etc.), no derivation showing the result remains a valid regularizer under domain shift, and no analysis of whether the merge introduces order-dependent or sequence-specific hyperparameters are supplied. Because this operation is load-bearing for both the 'slows knowledge drift' claim and the 'no task-specific tuning' assertion, its underspecification prevents verification that performance differences arise from the fast-slow principle rather than from the choice of merge.
- [Abstract] The VCL posterior-as-prior construction is inherited without additional justification that the merged prior reliably slows drift without new forgetting; the abstract supplies no quantitative results, ablation studies, or error bars to support the 'consistent improvements' claim, leaving the soundness of the central empirical assertion unassessable from the provided text.
minor comments (1)
- [Abstract] Abstract: the phrase 'seamlessly implemented in the IVON optimizer' would benefit from a one-sentence clarification of the exact code-level change required.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address the two major comments on the abstract point by point below, clarifying the manuscript content and indicating revisions where appropriate.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central mechanism—'merging of past posteriors' to produce the slow-adaptation prior—is stated at a high level only. No functional form (weighted average, product of Gaussians, moment matching, etc.), no derivation showing the result remains a valid regularizer under domain shift, and no analysis of whether the merge introduces order-dependent or sequence-specific hyperparameters are supplied. Because this operation is load-bearing for both the 'slows knowledge drift' claim and the 'no task-specific tuning' assertion, its underspecification prevents verification that performance differences arise from the fast-slow principle rather than from the choice of merge.
Authors: The abstract provides a high-level summary consistent with its length constraints. The functional form is moment matching of the Gaussian posteriors, the derivation that the result remains a valid regularizer is given in Section 3.2, and the analysis confirming no new order-dependent hyperparameters is in Section 3.3. We will revise the abstract to include one sentence specifying the moment-matching merge and referencing the section for the derivation. revision: yes
-
Referee: [Abstract] The VCL posterior-as-prior construction is inherited without additional justification that the merged prior reliably slows drift without new forgetting; the abstract supplies no quantitative results, ablation studies, or error bars to support the 'consistent improvements' claim, leaving the soundness of the central empirical assertion unassessable from the provided text.
Authors: Abstracts conventionally omit detailed quantitative results, ablations, and error bars; these appear in Sections 4–6 with multiple runs, error bars, and statistical tests demonstrating reduced forgetting and consistent gains. The merged prior's effect on drift is justified both by the VCL construction (Section 2) and by the reported experiments. We will add a short clause to the abstract noting that the stability-plasticity benefits are empirically validated in the main text. revision: partial
Circularity Check
No significant circularity; method is an empirical extension of VCL
full rationale
The paper proposes CoVON as a practical modification to the existing VCL framework by adding a merging step for past posteriors to implement slow adaptation. This is presented as a design choice implemented in the IVON optimizer, with performance claims resting on empirical comparisons across domain-incremental, continual pre-training, and LLM fine-tuning tasks rather than any first-principles derivation or prediction that reduces to the inputs by construction. No equations or uniqueness theorems are invoked that would trigger self-definitional, fitted-input, or self-citation load-bearing patterns. The merging operation is introduced as the novel mechanism, not presupposed as its own output.
Assumptions & free parameters
assumptions (1)
- domain assumption Merging of past posteriors produces a usable prior that slows parameter drift while preserving the fast-update properties of IVON
Cite this review
Pith. "Pith review of Fast and Slow Variational Continual Learning." pith.science (2026). https://pith.science/paper/7L2GFRHV
@misc{pith2026260624007,
author = {Pith},
title = {Pith review of: Fast and Slow Variational Continual Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/7L2GFRHV}},
note = {Machine review of arXiv:2606.24007}
}
read the original abstract
Continual learning remains a major challenge for modern deep networks, partly because commonly used optimizers lack inherent mechanisms for continual adaptation. One such natural mechanism is fast and slow adaptation to balance stability and plasticity. This mechanism has deep roots in neuroscience and biology, but there is no consensus on how to best incorporate it in commonly used optimizers. Here, we show that this can be easily done via the VCL framework, where past posteriors are used as priors in the future. Our key idea is to incorporate slow adaptation via merging of past posteriors to slow down the drift in the knowledge as learning progresses. The merged posterior is then used as the prior in the VCL update to implement the fast-weight updates. These steps can be seamlessly implemented in the IVON optimizer, whose form and costs are nearly identical to that of Adam. We call this new optimizer the Continual IVON (CoVON) optimizer and show that it not only consistently improves over existing VCL optimizers, but also performs better than other weight-regularization strategies across domain-incremental learning, continual pre-training, and fine-tuning of large language models.
Figures
Reference graph
Works this paper leans on
-
[1]
Constrained optimization and
Bertsekas, Dimitri P , year =. Constrained optimization and
-
[2]
and Bodard, Alexander and Laude, Emanuel and Patrinos, Panagiotis , year =
Oikonomidis, Konstantinos A. and Bodard, Alexander and Laude, Emanuel and Patrinos, Panagiotis , year =. Global convergence analysis of the power proximal point and augmented Lagrangian method , volume =. Computational Optimization and Applications , publisher =
-
[3]
icml, , year =
Mishchenko, Konstantin and Malinovsky, Grigory and Stich, Sebastian and Richt. icml, , year =
-
[4]
Federated Learning Via Inexact
Zhou, Shenglong and Li, Geoffrey Ye , journal =. Federated Learning Via Inexact. 2023 , volume =
2023
-
[5]
2021 , volume =
Zhang, Xinwei and Hong, Mingyi and Dhople, Sairaj and Yin, Wotao and Liu, Yang , journal =. 2021 , volume =
2021
-
[6]
, booktitle =
Gong, Yonghai and Li, Yichuan and Freris, Nikolaos M. , booktitle =
-
[7]
Mutambara, Arthur G. O. , title =. 1998 , isbn =
1998
-
[8]
aistats, , year =
Communication-Efficient Learning of Deep Networks from Decentralized Data , author =. aistats, , year =
Show all 300 references
-
[9]
Federated Optimization in Heterogeneous Networks , year =
Li, Tian and Sahu, Anit Kumar and Zaheer, Manzil and Sanjabi, Maziar and Talwalkar, Ameet and Smith, Virginia , booktitle =. Federated Optimization in Heterogeneous Networks , year =
-
[10]
iclr, , year =
Federated Learning as Variational Inference: A Scalable Expectation Propagation Approach , author =. iclr, , year =
-
[11]
Han Wang and Siddartha Marella and James Anderson , journal =. Fed
-
[12]
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization , booktitle =
Hendrycks, Dan and Basart, Steven and Mu, Norman and Kadavath, Saurav and Wang, Frank and Dorundo, Evan and Desai, Rahul and Zhu, Tyler and Parajuli, Samyak and Guo, Mike and Song, Dawn and Steinhardt, Jacob and Gilmer, Justin , year =. The Many Faces of Robustness: A Critical...
-
[13]
International Conference on Information Fusion , year =
Data fusion in decentralised sensing networks , author =. International Conference on Information Fusion , year =
-
[14]
Babagholami-Mohamadabadi, Behnam and Yoon, Sejong and Pavlovic, Vladimir , journal=
-
[15]
Federated Learning via Posterior Averaging: A New Perspective and Practical Algorithms , author=
-
[16]
Federated Learning Based on Dynamic Regularization , author=
-
[17]
Variational learning is effective for large deep networks , author=
-
[18]
Caldarola, Debora and Caputo, Barbara and Ciccone, Marco , booktitle = eccv, title =
-
[19]
Qu, Zhe and Li, Xingyu and Duan, Rui and Liu, Yao and Tang, Bo and Lu, Zhuo , booktitle = icml, title =
-
[20]
Fan, Ziqing and Hu, Shengchao and Yao, Jiangchao and Niu, Gang and Zhang, Ya and Sugiyama, Masashi and Wang, Yanfeng , booktitle = icml, title =
-
[21]
Nonlinear proximal point algorithms using
Eckstein, Jonathan , journal =. Nonlinear proximal point algorithms using
-
[22]
Wang, Huahua and Banerjee, Arindam , journal = nips, title =
-
[23]
Applications of a splitting algorithm to decomposition in convex programming and variational inequalities , volume =
Tseng, Paul , journal = siconopt, number =. Applications of a splitting algorithm to decomposition in convex programming and variational inequalities , volume =
-
[24]
A dual algorithm for the solution of nonlinear variational problems via finite element approximation , volume =
Gabay, Daniel and Mercier, Bertrand , journal =. A dual algorithm for the solution of nonlinear variational problems via finite element approximation , volume =
-
[25]
Sur l'approximation, par
Glowinski, Roland and Marroco, Americo , journal =. Sur l'approximation, par
-
[26]
Zhu, Jia-Jie and Mielke, Alexander , title =
-
[27]
Gradient flows for sampling: mean-field models,
Chen, Yifan and Huang, Daniel Zhengyu and Huang, Jiaoyang and Reich, Sebastian and Stuart, Andrew M , journal =. Gradient flows for sampling: mean-field models,
-
[28]
Learning without forgetting , author=
-
[29]
Zifeng Wang and Zizhao Zhang and Sayna Ebrahimi and Ruoxi Sun and Han Zhang and Chen-Yu Lee and Xiaoqi Ren and Guolong Su and Vincent Perot and Jennifer Dy and Tomas Pfister , year=
-
[30]
Learning to Prompt for Continual Learning , author=
-
[31]
Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning , author=
-
[32]
Non-exemplar domain incremental learning via cross-domain concept integration , author=
-
[33]
S-prompts learning with pre-trained transformers: An
Wang, Yabin and Huang, Zhiwu and Hong, Xiaopeng , booktitle=nips, year=. S-prompts learning with pre-trained transformers: An
-
[34]
Conference on Robot Learning (CoRL) , year=
Core50: a new dataset and benchmark for continuous object recognition , author=. Conference on Robot Learning (CoRL) , year=
-
[35]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , year=
A continual deepfake detection benchmark: Dataset, methods, and essentials , author=. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , year=
-
[36]
Moment matching for multi-source domain adaptation , author=
-
[37]
Stephenson, Will and Frangella, Zachary and Udell, Madeleine and Broderick, Tamara , journal = nips, title =
-
[38]
Introduction to stochastic search and optimization: estimation, simulation, and control , year =
Spall, James C , publisher =. Introduction to stochastic search and optimization: estimation, simulation, and control , year =
-
[39]
Kiral, Eren Mehmet and M. The
-
[40]
Mohamed, Shakir and Rosca, Mihaela and Figurnov, Michael and Mnih, Andriy , journal = jmlr, number =. Monte
-
[41]
Figurnov, Mikhail and Mohamed, Shakir and Mnih, Andriy , booktitle = nips, title =
-
[42]
On the Convergence of IRLS and Its Variants in Outlier-Robust Estimation , year =
Peng, Liangzu and K. On the Convergence of IRLS and Its Variants in Outlier-Robust Estimation , year =
-
[43]
Yang, Rubing and Mao, Jialin and Chaudhari, Pratik , booktitle = icml, title =
-
[44]
Control systems and reinforcement learning , year =
Meyn, Sean , publisher =. Control systems and reinforcement learning , year =
-
[45]
Statistical learning theory and stochastic optimization , year =
Catoni, Olivier , publisher =. Statistical learning theory and stochastic optimization , year =
-
[46]
Generalization Guarantees via Algorithm-dependent
Sachs, Sarah and van Erven, Tim and Hodgkinson, Liam and Khanna, Rajiv and. Generalization Guarantees via Algorithm-dependent
-
[47]
Sefidgaran, Milad and Gohari, Amin and Richard, Gael and Simsekli, Umut , booktitle = colt, title =
-
[48]
Arora, Sanjeev and Ge, Rong and Neyshabur, Behnam and Zhang, Yi , booktitle = icml, title =
-
[49]
Suzuki, Taiji and Abe, Hiroshi and Nishimura, Tomoaki , booktitle = iclr, title =
-
[50]
Spectral pruning: Compressing deep neural networks via spectral analysis and its generalization error , year =
Suzuki, Taiji and Abe, Hiroshi and Murata, Tomoya and Horiuchi, Shingo and Ito, Kotaro and Wachi, Tokuma and Hirai, So and Yukishima, Masatoshi and Nishimura, Tomoaki , booktitle =. Spectral pruning: Compressing deep neural networks via spectral analysis and its generalization...
-
[51]
Barsbey, Melih and Sefidgaran, Milad and Erdogdu, Murat A and Richard, Gael and Simsekli, Umut , journal = nips, title =
-
[52]
Burke, James V and Ferris, Michael C , journal = mp, number =. A
-
[53]
Harmonic exponential families on homogeneous spaces , volume =
Tojo, Koichi and Yoshino, Taro , journal =. Harmonic exponential families on homogeneous spaces , volume =
-
[54]
Duality for Neural Networks through Reproducing Kernel
Spek, Len and Heeringa, Tjeerd Jan and Brune, Christoph , journal =. Duality for Neural Networks through Reproducing Kernel
-
[55]
Approximation accuracy, gradient methods, and error bound for structured convex optimization , volume =
Tseng, Paul , journal = mp, number =. Approximation accuracy, gradient methods, and error bound for structured convex optimization , volume =
-
[56]
Welling, Max and Teh, Yee W , booktitle = icml, title =
-
[57]
SAE: Sequential Anchored Ensembles , year =
Delaunoy, Arnaud and Louppe, Gilles , journal =. SAE: Sequential Anchored Ensembles , year =
-
[58]
Decoupled Weight Decay Regularization , author=
-
[59]
Gupta, Vineet and Koren, Tomer and Singer, Yoram , booktitle = icml, title =
-
[60]
A stochastic approximation method , year =
Robbins, Herbert and Monro, Sutton , journal =. A stochastic approximation method , year =
-
[61]
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, Lukasz and Polosukhin, Illia , journal = nips, title =
-
[62]
arXiv:2205.15902 , title =
Lambert, Marc and Chewi, Sinho and Bach, Francis and Bonnabel, Silv. arXiv:2205.15902 , title =
-
[63]
Minimax in Geodesic Metric Spaces: Sion's Theorem and Algorithms , year =
Zhang, Peiyuan and Zhang, Jingzhao and Sra, Suvrit , journal =. Minimax in Geodesic Metric Spaces: Sion's Theorem and Algorithms , year =
-
[64]
Linear convergence of gradient and proximal-gradient methods under the
Karimi, Hamed and Nutini, Julie and Schmidt, Mark , booktitle =. Linear convergence of gradient and proximal-gradient methods under the
-
[65]
Lifting the convex conjugate in
Bauermeister, Hartmut and Laude, Emanuel and Mollenhoff, Thomas and Moeller, Michael and Cremers, Daniel , journal = siims, number =. Lifting the convex conjugate in
-
[66]
Bayesian neural network priors revisited , year =
Fortuin, Vincent and Garriga-Alonso, Adri. Bayesian neural network priors revisited , year =
-
[67]
Coker, Beau and Bruinsma, Wessel P and Burt, David R and Pan, Weiwei and Doshi-Velez, Finale , booktitle = aistats, title =
-
[68]
On the expressiveness of approximate inference in
Foong, Andrew and Burt, David and Li, Yingzhen and Turner, Richard , journal =. On the expressiveness of approximate inference in
-
[69]
Priors in
Fortuin, Vincent , journal =. Priors in
-
[70]
Zhang, Tong , booktitle = colt, title =
-
[71]
Adversarial Interpretation of
Husain, Hisham and Knoblauch, Jeremias , booktitle =. Adversarial Interpretation of
-
[72]
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, Shai and Zhang, Tong , journal = jmlr, number =. Stochastic dual coordinate ascent methods for regularized loss minimization. , volume =
-
[73]
Deep learning in neural networks: An overview , volume =
Schmidhuber, J. Deep learning in neural networks: An overview , volume =. Neural networks , pages =
-
[74]
Deep learning , volume =
LeCun, Yann and Bengio, Yoshua and Hinton, Geoffrey , journal =. Deep learning , volume =
-
[75]
Evaluating Approximate Inference in
Wilson, Andrew Gordon and Izmailov, Pavel and Hoffman, Matthew D and Gal, Yarin and Li, Yingzhen and Pradier, Melanie F and Vikram, Sharad and Foong, Andrew and Lotfi, Sanae and Farquhar, Sebastian , booktitle =. Evaluating Approximate Inference in
-
[76]
Amid, Ehsan and Anil, Rohan and Warmuth, Manfred , booktitle = aistats, title =
-
[77]
Stochastic gradient descent as approximate
Mandt, Stephan and Hoffman, Matthew D and Blei, David M , journal = jmlr, pages =. Stochastic gradient descent as approximate
-
[78]
Khan, Mohammad Emtiyaz and Mohamed, Shakir and Marlin, Benjamin and Murphy, Kevin , booktitle = aistats, title =
-
[79]
Questions for Flat-Minima Optimization of Modern Neural Networks , year =
Kaddour, Jean and Liu, Linqing and Silva, Ricardo and Kusner, Matt J , journal =. Questions for Flat-Minima Optimization of Modern Neural Networks , year =
-
[80]
Guo, Chuan and Pleiss, Geoff and Sun, Yu and Weinberger, Kilian Q , booktitle = icml, title =
-
[81]
Variational bounds for mixed-data factor analysis , year =
Khan, Mohammad Emtiyaz , school =. Variational bounds for mixed-data factor analysis , year =
-
[82]
Khan, Mohammad Emtiyaz and Bouchard, Guillaume and Murphy, Kevin P and Marlin, Benjamin M , journal = nips, title =
-
[83]
Jaakkola, Tommi S and Jordan, Michael I , booktitle = aistats, title =
-
[84]
Efficient bounds for the softmax function, applications to inference in hybrid models , year =
Bouchard, Guillaume , booktitle =. Efficient bounds for the softmax function, applications to inference in hybrid models , year =
-
[85]
Kwon, Jungmin and Kim, Jeongseop and Park, Hyunseo and Choi, In Kwon , booktitle = icml, title =
-
[86]
Explicit Regularization in Overparametrized Models via Noise Injection , year =
Orvieto, Antonio and Raj, Anant and Kersting, Hans and Bach, Francis , journal =. Explicit Regularization in Overparametrized Models via Noise Injection , year =
-
[87]
Jin, Chi and Netrapalli, Praneeth and Jordan, Michael , booktitle = icml, title =
-
[88]
Scalable marginal likelihood estimation for model selection in deep learning , year =
Immer, Alexander and Bauer, Matthias and Fortuin, Vincent and R. Scalable marginal likelihood estimation for model selection in deep learning , year =
-
[89]
Anticorrelated Noise Injection for Improved Generalization , year =
Orvieto, Antonio and Kersting, Hans and Proske, Frank and Bach, Francis and Lucchi, Aurelien , journal =. Anticorrelated Noise Injection for Improved Generalization , year =
-
[90]
Bisla, Devansh and Wang, Jing and Choromanska, Anna , journal = aistats, title =
-
[91]
Ockham's razor and
Jefferys, William H and Berger, James O , journal =. Ockham's razor and
-
[92]
The minimum description length principle , year =
Gr. The minimum description length principle , year =
-
[93]
A universal prior for integers and estimation by minimum description length , volume =
Rissanen, Jorma , journal =. A universal prior for integers and estimation by minimum description length , volume =
-
[94]
Smith and Quoc V
Samuel L. Smith and Quoc V. Le , booktitle = iclr, title =
-
[95]
Sharpness-Aware Minimization Improves Language Model Generalization , year =
Bahri, Dara and Mobahi, Hossein and Tay, Yi , journal =. Sharpness-Aware Minimization Improves Language Model Generalization , year =
-
[96]
Learning with submodular functions: A convex optimization perspective , volume =
Bach, Francis and others , journal =. Learning with submodular functions: A convex optimization perspective , volume =
-
[97]
Online model selection based on the variational
Sato, Masa-Aki , journal =. Online model selection based on the variational
-
[98]
Conjugate dualities for relative smoothness and strong convexity under the light of generalized convexity , year =
Laude, Emanuel and Themelis, Andreas and Patrinos, Panagiotis , journal =. Conjugate dualities for relative smoothness and strong convexity under the light of generalized convexity , year =
-
[99]
Natural gradient works efficiently in learning , volume =
Amari, Shun-Ichi , journal =. Natural gradient works efficiently in learning , volume =
-
[100]
User-friendly introduction to
Alquier, Pierre , journal =. User-friendly introduction to
-
[101]
Tomczak , publisher =
Jakub M. Tomczak , publisher =. Deep Generative Modeling , year =
-
[102]
Variational Methods for Machine Learning with Applications to Deep Networks , year =
Cinelli, Lucas Pinheiro and Marins, Matheus Ara. Variational Methods for Machine Learning with Applications to Deep Networks , year =
-
[103]
Probabilistic deep learning: With python, keras and tensorflow probability , year =
D. Probabilistic deep learning: With python, keras and tensorflow probability , year =
-
[104]
Catoni, Olivier , number =
-
[105]
Deep learning , year =
Goodfellow, Ian and Bengio, Yoshua and Courville, Aaron , publisher =. Deep learning , year =
-
[106]
Information theory, inference and learning algorithms , year =
MacKay, David JC , publisher =. Information theory, inference and learning algorithms , year =
-
[107]
Bayesian reasoning and machine learning , year =
Barber, David , publisher =. Bayesian reasoning and machine learning , year =
-
[108]
Probabilistic machine learning: advanced topics , year =
Murphy, Kevin Patrick , publisher =. Probabilistic machine learning: advanced topics , year =
-
[109]
Probabilistic machine learning: an introduction , year =
Murphy, Kevin Patrick , publisher =. Probabilistic machine learning: an introduction , year =
-
[110]
Machine learning: a probabilistic perspective , year =
Murphy, Kevin Patrick , publisher =. Machine learning: a probabilistic perspective , year =
-
[111]
Bishop , publisher =
Christopher M. Bishop , publisher =. Pattern recognition and machine learning , year =
-
[112]
An inertial
Castera, Camille and Bolte, J. An inertial
-
[113]
A second-order gradient-like dissipative dynamical system with
Alvarez, Felipe and Attouch, Hedy and Bolte, J. A second-order gradient-like dissipative dynamical system with. Journal de math
-
[114]
Hoffman and Andrew Gordon Wilson , booktitle = icml, title =
Pavel Izmailov and Sharad Vikram and Matthew D. Hoffman and Andrew Gordon Wilson , booktitle = icml, title =
-
[115]
DAGM German Conference on Pattern Recognition (GCPR) , title =
Ye, Zhenzhang and Haefner, Bjoern and Qu. DAGM German Conference on Pattern Recognition (GCPR) , title =
-
[116]
Fixed-form variational posterior approximation through stochastic linear regression , volume =
Salimans, Tim and Knowles, David A , journal =. Fixed-form variational posterior approximation through stochastic linear regression , volume =
-
[117]
Human-level concept learning through probabilistic program induction , volume =
Lake, Brenden M and Salakhutdinov, Ruslan and Tenenbaum, Joshua B , journal =. Human-level concept learning through probabilistic program induction , volume =
-
[118]
arXiv:2106.05237 , title =
Beyer, Lucas and Zhai, Xiaohua and Royer, Am. arXiv:2106.05237 , title =
-
[119]
James Martens , booktitle = icml, title =
-
[120]
Donald Goldfarb and Yi Ren and Achraf Bahamou , booktitle = nips, title =
-
[121]
Grosse , booktitle = icml, title =
James Martens and Roger B. Grosse , booktitle = icml, title =
-
[122]
Revisiting
Bello, Irwan and Fedus, William and Du, Xianzhi and Cubuk, Ekin D and Srinivas, Aravind and Lin, Tsung-Yi and Shlens, Jonathon and Zoph, Barret , journal =. Revisiting
-
[123]
Scalable second order optimization for deep learning , year =
Anil, Rohan and Gupta, Vineet and Koren, Tomer and Regan, Kevin and Singer, Yoram , journal =. Scalable second order optimization for deep learning , year =
-
[124]
Mahoney , booktitle = aaai, title =
Zhewei Yao and Amir Gholami and Sheng Shen and Mustafa Mustafa and Kurt Keutzer and Michael W. Mahoney , booktitle = aaai, title =
-
[125]
Convergence of Entropic Schemes for Optimal Transport and Gradient Flows , volume =
Guillaume Carlier and Vincent Duval and Gabriel Peyr. Convergence of Entropic Schemes for Optimal Transport and Gradient Flows , volume =
-
[126]
A weak convergence approach to the theory of large deviations , volume =
Dupuis, Paul and Ellis, Richard S , publisher =. A weak convergence approach to the theory of large deviations , volume =
-
[127]
Large Deviation Principles , year =
Scott Robertson , howpublished =. Large Deviation Principles , year =
-
[128]
Diffusions for global optimization , volume =
Geman, Stuart and Hwang, Chii-Ruey , journal = siconopt, number =. Diffusions for global optimization , volume =
-
[129]
Maximum d'entropie et probl
Dacunha-Castelle, Didier and Gamboa, Fabrice , booktitle =. Maximum d'entropie et probl
-
[130]
Rioux, Gabriel Erik , title =
-
[131]
The maximum entropy on the mean method, noise and sensitivity , year =
Bercher, Jean-Francois and Le Besnerais, Guy and Demoment, Guy , booktitle =. The maximum entropy on the mean method, noise and sensitivity , year =
-
[132]
A Generalized Representer Theorem , volume =
Bernhard Sch. A Generalized Representer Theorem , volume =
-
[133]
Banach Space Representer Theorems for Neural Networks and Ridge Splines
Parhi, Rahul and Nowak, Robert D , journal = jmlr, pages =. Banach Space Representer Theorems for Neural Networks and Ridge Splines. , volume =
-
[134]
Spline solutions to
Fisher, SD and Jerome, Joseph W , journal =. Spline solutions to
-
[135]
Chang and Mohammad Emtiyaz Khan and Arno Solin , booktitle = nips, title =
Vincent Adam and Paul E. Chang and Mohammad Emtiyaz Khan and Arno Solin , booktitle = nips, title =
-
[136]
An extended variational principle , year =
Jordan, Richard and Kinderlehrer, David , booktitle =. An extended variational principle , year =
-
[137]
Estimation of Non-Normalized Statistical Models by Score Matching , volume =
Aapo Hyv. Estimation of Non-Normalized Statistical Models by Score Matching , volume =
-
[138]
Maddox and Samuel Stanton and Andrew Gordon Wilson , journal = nips, title =
Wesley J. Maddox and Samuel Stanton and Andrew Gordon Wilson , journal = nips, title =
-
[139]
Bach , journal =
Francis R. Bach , journal =. Duality Between Subgradient and Conditional Gradient Methods , volume =
-
[140]
Sparse On-Line
Lehel Csat. Sparse On-Line. Neural Comput. , number =
-
[141]
Haowei He and Gao Huang and Yang Yuan , booktitle = nips, title =
-
[142]
Differential geometry and statistics , year =
Murray, Michael K and Rice, John W , publisher =. Differential geometry and statistics , year =
-
[143]
Goodfellow and Jonathon Shlens and Christian Szegedy , booktitle = iclr, title =
Ian J. Goodfellow and Jonathon Shlens and Christian Szegedy , booktitle = iclr, title =
-
[144]
Goodfellow and Rob Fergus , booktitle = iclr, title =
Christian Szegedy and Wojciech Zaremba and Ilya Sutskever and Joan Bruna and Dumitru Erhan and Ian J. Goodfellow and Rob Fergus , booktitle = iclr, title =
-
[145]
Pierre Alquier and The Tien Mai and Massimiliano Pontil , booktitle = aistats, title =
-
[146]
Ron Amit and Ron Meir , booktitle = icml, title =
-
[147]
Stochastic Training is Not Necessary for Generalization , year =
Geiping, Jonas and Goldblum, Micah and Pope, Phillip E and Moeller, Michael and Goldstein, Tom , journal =. Stochastic Training is Not Necessary for Generalization , year =
-
[148]
Logistic regression diagnostics , volume =
Pregibon, Daryl , journal =. Logistic regression diagnostics , volume =
-
[149]
Concentration inequalities: A nonasymptotic theory of independence , year =
Boucheron, St. Concentration inequalities: A nonasymptotic theory of independence , year =
-
[150]
Introduction to Riemannian manifolds , year =
Lee, John M , publisher =. Introduction to Riemannian manifolds , year =
-
[151]
Mert Pilanci and Laurent El Ghaoui and Venkat Chandrasekaran , booktitle = nips, title =
-
[152]
The variational formulation of the
Jordan, Richard and Kinderlehrer, David and Otto, Felix , journal =. The variational formulation of the
-
[153]
Andrew Holbrook and Shiwei Lan and Jeffrey Streets and Babak Shahbaba , booktitle = uai, title =
-
[154]
Amari, Shun-Ichi and Ba, Jimmy and Grosse, Roger and Li, Xuechen and Nitanda, Atsushi and Suzuki, Taiji and Wu, Denny and Xu, Ji , booktitle = iclr, title =
-
[155]
Finding Mixed Nash Equilibria of Generative Adversarial Networks , year =
Ya. Finding Mixed Nash Equilibria of Generative Adversarial Networks , year =
-
[156]
Robert J. N. Baldock and Hartmut Maennel and Behnam Neyshabur , booktitle = nips, title =
-
[157]
Warmuth , booktitle = nips, title =
Michal Derezinski and Manfred K. Warmuth , booktitle = nips, title =
-
[158]
On the properties of variational approximations of
Pierre Alquier and James Ridgway and Nicolas Chopin , journal = jmlr, pages =. On the properties of variational approximations of
-
[159]
Behnam Neyshabur and Srinadh Bhojanapalli and David McAllester and Nati Srebro , booktitle = nips, title =
-
[160]
Sixin Zhang and Anna Choromanska and Yann LeCun , booktitle = nips, title =
-
[161]
McAllester , booktitle = colt, title =
David A. McAllester , booktitle = colt, title =
-
[162]
Nitish Shirish Keskar and Dheevatsa Mudigere and Jorge Nocedal and Mikhail Smelyanskiy and Ping Tak Peter Tang , booktitle = iclr, title =
-
[163]
Backpropagation and stochastic gradient descent method , volume =
Shun. Backpropagation and stochastic gradient descent method , volume =. Neurocomputing , number =
-
[164]
I-divergence geometry of probability distributions and minimization problems , year =
Csisz. I-divergence geometry of probability distributions and minimization problems , year =. Annals of Probability , pages =
-
[165]
Asymptotic evaluation of certain
Donsker, Monroe D and Varadhan, SR Srinivasa , journal =. Asymptotic evaluation of certain
-
[166]
Bayesian data analysis , year =
Gelman, Andrew and Carlin, John B and Stern, Hal S and Rubin, Donald B , publisher =. Bayesian data analysis , year =
-
[167]
Estimation of exponential-polynomial distribution by holonomic gradient descent , volume =
Hayakawa, Jumpei and Takemura, Akimichi , journal =. Estimation of exponential-polynomial distribution by holonomic gradient descent , volume =
-
[168]
Mode-Finding for Mixtures of
Miguel. Mode-Finding for Mixtures of
-
[169]
Mathematical foundations of infinite-dimensional statistical models , year =
Gin. Mathematical foundations of infinite-dimensional statistical models , year =
-
[170]
A generalization of the maximum entropy principle for curved statistical manifolds , year =
Morales, Pablo A and Rosas, Fernando E , journal =. A generalization of the maximum entropy principle for curved statistical manifolds , year =
-
[171]
Introduction to the theory of regular exponential families , volume =
Johansen, S. Introduction to the theory of regular exponential families , volume =
-
[172]
Mahoney , howpublished =
Ahmed El Alaoui and Michael W. Mahoney , howpublished =. Fast Randomized Kernel Methods With Statistical Guarantees , year =
-
[173]
Variational Bayesian Reinforcement Learning with Regret Bounds , year =
Brendan O'Donoghue , howpublished =. Variational Bayesian Reinforcement Learning with Regret Bounds , year =
-
[174]
Understanding priors in
Vladimirova, Mariia and Verbeek, Jakob and Mesejo, Pablo and Arbel, Julyan , booktitle = icml, pages =. Understanding priors in
-
[175]
Lin, Wu and Khan, Mohammad Emtiyaz and Schmidt, Mark , booktitle = icml, title =
-
[176]
Automatic smoothing of regression functions in generalized linear models , volume =
O'Sullivan, Finbarr and Yandell, Brian S and Raynor Jr, William J , journal =. Automatic smoothing of regression functions in generalized linear models , volume =
-
[177]
Computing Statistical Divergences with Sigma Points , year =
Nielsen, Frank and Nock, Richard , booktitle =. Computing Statistical Divergences with Sigma Points , year =
-
[178]
Lifting the Convex Conjugate in Lagrangian Relaxations: A Tractable Approach for Continuous
Bauermeister, Hartmut and Laude, Emanuel and M. Lifting the Convex Conjugate in Lagrangian Relaxations: A Tractable Approach for Continuous
-
[179]
Stochastic reformulations of linear systems: algorithms and convergence theory , volume =
Richt. Stochastic reformulations of linear systems: algorithms and convergence theory , volume =. SIAM Journal on Matrix Analysis and Applications , number =
-
[180]
Jonas Rothfuss and Vincent Fortuin and Martin Josifoski and Andreas Krause , booktitle =
-
[181]
Justin Domke , booktitle = icml, title =
-
[182]
Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun , booktitle = cvpr, title =
-
[183]
New Insights and Perspectives on the Natural Gradient Method , volume =
James Martens , journal = jmlr, pages =. New Insights and Perspectives on the Natural Gradient Method , volume =
-
[184]
The equivalence between
Casey Chu and Kentaro Minami and Kenji Fukumizu , howpublished =. The equivalence between
-
[185]
Roy , booktitle = colt, title =
Gergely Neu and Gintare Karolina Dziugaite and Mahdi Haghifam and Daniel M. Roy , booktitle = colt, title =
-
[186]
Guy Blanc and Neha Gupta and Gregory Valiant and Paul Valiant , booktitle = colt, title =
-
[187]
Label Noise
Damian, Alex and Ma, Tengyu and Lee, Jason , howpublished =. Label Noise
-
[188]
Connecting information geometry and geometric mechanics , volume =
Leok, Melvin and Zhang, Jun , journal =. Connecting information geometry and geometric mechanics , volume =
-
[189]
Symplectic structures on statistical manifolds , volume =
Noda, Tomonori , journal =. Symplectic structures on statistical manifolds , volume =
-
[190]
Jerry Chee and Panos Toulis , booktitle = aistats, title =
-
[191]
Denker and Sara A
Yann LeCun and John S. Denker and Sara A. Solla , booktitle = nips, title =
-
[192]
Osborne and Yarin Gal , booktitle = aistats, title =
Sebastian Farquhar and Michael A. Osborne and Yarin Gal , booktitle = aistats, title =
-
[193]
Jiancheng Yang and Rui Shi and Bingbing Ni , booktitle =. Med
-
[194]
Satoshi Hara and Atsushi Nitanda and Takanori Maehara , booktitle = nips, title =
-
[195]
Differential geometry of curved exponential families-curvatures and information loss , year =
Amari, Shun-Ichi , journal =. Differential geometry of curved exponential families-curvatures and information loss , year =
-
[196]
Nguyen and Yingzhen Li and Thang D
Cuong V. Nguyen and Yingzhen Li and Thang D. Bui and Richard E. Turner , booktitle = iclr, title =
-
[197]
Robert Kleinberg and Yuanzhi Li and Yang Yuan , booktitle = icml, title =
-
[198]
Sculley and Joshua V
Jasper Snoek and Yaniv Ovadia and Emily Fertig and Balaji Lakshminarayanan and Sebastian Nowozin and D. Sculley and Joshua V. Dillon and Jie Ren and Zachary Nado , booktitle = nips, title =
-
[199]
Weinberger , booktitle = icml, title =
Chuan Guo and Geoff Pleiss and Yu Sun and Kilian Q. Weinberger , booktitle = icml, title =
-
[200]
David J. C. MacKay , journal =. A Practical Bayesian Framework for Backpropagation Networks , volume =
-
[201]
Denker and Yann LeCun , booktitle = nips, title =
John S. Denker and Yann LeCun , booktitle = nips, title =
-
[202]
Andrew Hicks and Ronald K
R. Andrew Hicks and Ronald K. Perline , booktitle = cvpr, title =
-
[203]
Vitaly Feldman and Chiyuan Zhang , booktitle = nips, title =
-
[204]
Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory , volume =
Gr. Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory , volume =. Annals of Statistics , number =
-
[205]
Sensitivity analysis in linear regression , volume =
Chatterjee, Samprit and Hadi, Ali S , publisher =. Sensitivity analysis in linear regression , volume =
-
[206]
Influence sketching
Michael Wojnowicz and Ben Cruz and Xuan Zhao and Brian Wallace and Matt Wolff and Jay Luan and Caleb Crable , booktitle =. "Influence sketching": Finding influential samples in large-scale regressions , year =
-
[207]
Generalized linear models , year =
McCullagh, Peter and Nelder, John A , editor =. Generalized linear models , year =
-
[208]
Regression diagnostics: Identifying influential data and sources of collinearity , year =
Belsley, David A and Kuh, Edwin and Welsch, Roy E , publisher =. Regression diagnostics: Identifying influential data and sources of collinearity , year =
-
[209]
Moreno and Andrew Gordon Wilson and Andreas Damianou , booktitle = aistats, title =
Wesley Maddox and Shuai Tang and Pablo G. Moreno and Andrew Gordon Wilson and Andreas Damianou , booktitle = aistats, title =
-
[210]
Frederik Kunstner and Philipp Hennig and Lukas Balles , booktitle = nips, title =
-
[211]
Mohammad Emtiyaz Khan and Siddharth Swaroop , howpublished = nips, title =
-
[212]
Sampling with Mirrored Stein Operators , year =
Jiaxin Shi and Chang Liu and Lester Mackey , howpublished =. Sampling with Mirrored Stein Operators , year =
-
[213]
Convex analysis in general vector spaces , year =
Zalinescu, Constantin , publisher =. Convex analysis in general vector spaces , year =
-
[214]
Support vector machines, reproducing kernel
Wahba, Grace , journal =. Support vector machines, reproducing kernel
-
[215]
The relation of
Jaynes, Edwin Thompson , booktitle =. The relation of
-
[216]
How does the brain do plausible reasoning? , year =
Jaynes, Edwin Thompson , booktitle =. How does the brain do plausible reasoning? , year =
-
[217]
A theoretical framework for back-propagation , volume =
LeCun, Yann , booktitle =. A theoretical framework for back-propagation , volume =
-
[218]
Statistical inference on time series by
Parzen, Emanuel , institution =. Statistical inference on time series by
-
[219]
Some results on
Kimeldorf, George S and Wahba, Grace , journal =. Some results on
-
[220]
A correspondence between
Kimeldorf, George S and Wahba, Grace , journal =. A correspondence between
-
[221]
Tilt stability of a local minimum , volume =
Poliquin, Ren. Tilt stability of a local minimum , volume =
-
[222]
Continual deep learning by functional regularisation of memorable past , year =
Pan, Pingbo and Swaroop, Siddharth and Immer, Alexander and Eschenhagen, Runa and Turner, Richard E and Khan, Mohammad Emtiyaz , journal =. Continual deep learning by functional regularisation of memorable past , year =
-
[223]
Kakade and Aaron Sidford , booktitle = icml, title =
Roy Frostig and Rong Ge and Sham M. Kakade and Aaron Sidford , booktitle = icml, title =
-
[224]
The geometry of exponential families , volume =
Efron, Bradley , journal =. The geometry of exponential families , volume =
-
[225]
Azoury and Manfred K
Katy S. Azoury and Manfred K. Warmuth , journal =. Relative Loss Bounds for On-Line Density Estimation with the Exponential Family of Distributions , volume =
-
[226]
Anant Raj and Cameron Musco and Lester Mackey , booktitle = aistats, title =
-
[227]
Efficient Sequential Learning in Structured and Constrained Environments , year =
Calandriello, Daniele , school =. Efficient Sequential Learning in Structured and Constrained Environments , year =
-
[228]
A method to construct exponential families by representation theory , year =
Tojo, Koichi and Yoshino, Taro , howpublished =. A method to construct exponential families by representation theory , year =
-
[229]
A Bregman Learning Framework for Sparse Neural Networks , year =
Leon Bungert and Tim Roith and Daniel Tenbrinck and Martin Burger , howpublished =. A Bregman Learning Framework for Sparse Neural Networks , year =
-
[230]
Hidden convexity in a problem of nonlinear elasticity , volume =
Ghoussoub, Nassif and Kim, Young-Heon and Lavenant, Hugo and Palmer, Aaron Zeff , journal =. Hidden convexity in a problem of nonlinear elasticity , volume =
-
[231]
Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks , year =
Jos. Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks , year =
-
[232]
Behnam Neyshabur , booktitle = nips, title =
-
[233]
Yarin Gal and Zoubin Ghahramani , booktitle = icml, title =
-
[234]
Alex Graves , booktitle = nips, title =
-
[235]
Gradient Regularisation as Approximate Variational Inference , year =
Ali Unlu and Laurence Aitchison , howpublished =. Gradient Regularisation as Approximate Variational Inference , year =
-
[236]
Mathematical statistics: basic ideas and selected topics , volume =
Bickel, Peter J and Doksum, Kjell A , publisher =. Mathematical statistics: basic ideas and selected topics , volume =
-
[237]
Blei , journal =
Chong Wang and David M. Blei , journal =. Variational inference in nonconjugate models , volume =
-
[238]
Variational
Ali Unlu and Laurence Aitchison , howpublished =. Variational
-
[239]
G. M. Korpelevich , journal =. An Extragradient Method for Finding Saddle Points and for Other Problems , volume =
-
[240]
Giacomo Meanti and Luigi Carratino and Lorenzo Rosasco and Alessandro Rudi , booktitle = nips, title =
-
[241]
Alessandro Rudi and Luigi Carratino and Lorenzo Rosasco , booktitle = nips, title =
-
[242]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale , year =
Alexey Dosovitskiy and Lucas Beyer and Alexander Kolesnikov and Dirk Weissenborn and Xiaohua Zhai and Thomas Unterthiner and Mostafa Dehghani and Matthias Minderer and Georg Heigold and Sylvain Gelly and Jakob Uszkoreit and Neil Houlsby , howpublished =. An Image is Worth 16x1...
-
[243]
When Vision Transformers Outperform
Xiangning Chen and Cho-Jui Hsieh and Boqing Gong , howpublished =. When Vision Transformers Outperform
-
[244]
Hugo Touvron and Piotr Bojanowski and Mathilde Caron and Matthieu Cord and Alaaeldin El. Res
-
[245]
Linear and integer programming vs linear integration and counting: a duality viewpoint , year =
Lasserre, Jean-Bernard , publisher =. Linear and integer programming vs linear integration and counting: a duality viewpoint , year =
-
[246]
Duality between probability and optimization , volume =
Akian, Marianne and Quadrat, Jean-Pierre and Viot, Michel , journal =. Duality between probability and optimization , volume =
-
[247]
On the marginal likelihood and cross-validation , volume =
Fong, Edwin and Holmes, CC , journal =. On the marginal likelihood and cross-validation , volume =
-
[248]
An introduction to optimization on smooth manifolds , year =
Boumal, Nicolas , publisher =. An introduction to optimization on smooth manifolds , year =
-
[249]
Knowledge Distillation:
Jianping Gou and Baosheng Yu and Stephen John Maybank and Dacheng Tao , howpublished =. Knowledge Distillation:
-
[250]
Hinton and Oriol Vinyals and Jeffrey Dean , howpublished =
Geoffrey E. Hinton and Oriol Vinyals and Jeffrey Dean , howpublished =. Distilling the Knowledge in a Neural Network , year =
-
[251]
Murphy and Max Welling , booktitle = nips, title =
Anoop Korattikara Balan and Vivek Rathod and Kevin P. Murphy and Max Welling , booktitle = nips, title =
-
[252]
Maddox and Pavel Izmailov and Timur Garipov and Dmitry P
Wesley J. Maddox and Pavel Izmailov and Timur Garipov and Dmitry P. Vetrov and Andrew Gordon Wilson , booktitle = nips, title =
-
[253]
Benton and Wesley J
Gregory W. Benton and Wesley J. Maddox and Sanae Lotfi and Andrew Gordon Wilson , howpublished =. Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling , year =
-
[254]
Simon Baker and Daniel Scharstein and J. P. Lewis and Stefan Roth and Michael J. Black and Richard Szeliski , journal = ijcv, number =. A Database and Evaluation Methodology for Optical Flow , volume =
-
[255]
Numerical treatment of a class of semi-infinite programming problems , volume =
Gustafson, S. Numerical treatment of a class of semi-infinite programming problems , volume =. Naval Research Logistics Quarterly , number =
-
[256]
On a class of multidimensional optimal transportation problems , volume =
Carlier, Guillaume , journal =. On a class of multidimensional optimal transportation problems , volume =
-
[257]
Boyd and Neal Parikh and Eric Chu and Borja Peleato and Jonathan Eckstein , journal =
Stephen P. Boyd and Neal Parikh and Eric Chu and Borja Peleato and Jonathan Eckstein , journal =. Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers , volume =
-
[258]
Nonparametric regression between general
Steinke, Florian and Hein, Matthias and Sch. Nonparametric regression between general
-
[259]
Introduction to Optimization , year =
Polyak, Boris Teodorovich , publisher =. Introduction to Optimization , year =
-
[260]
Pratik Chaudhari and Stefano Soatto , booktitle = iclr, title =
-
[261]
Smith and Benoit Dherin and David G
Samuel L. Smith and Benoit Dherin and David G. T. Barrett and Soham De , booktitle = iclr, title =
-
[262]
David G. T. Barrett and Benoit Dherin , booktitle = iclr, title =
-
[263]
Penalty function theory for general convex programming problems , volume =
Mine, Hisashi and Fukushima, Masao , journal =. Penalty function theory for general convex programming problems , volume =
-
[264]
A generalized proximal point algorithm for certain non-convex minimization problems , volume =
Fukushima, Masao and Mine, Hisashi , journal =. A generalized proximal point algorithm for certain non-convex minimization problems , volume =
-
[265]
A minimization method for the sum of a convex function and a continuously differentiable function , volume =
Mine, Hisashi and Fukushima, Masao , journal =. A minimization method for the sum of a convex function and a continuously differentiable function , volume =
-
[266]
Wainwright , journal = jmlr, pages =
Pradeep Ravikumar and Alekh Agarwal and Martin J. Wainwright , journal = jmlr, pages =. Message-passing for Graph-structured Linear Programs: Proximal Methods and Rounding Schemes , volume =
-
[267]
Entropic proximal mappings with applications to nonlinear programming , volume =
Teboulle, Marc , journal =. Entropic proximal mappings with applications to nonlinear programming , volume =
-
[268]
Generalized semi-infinite programming: a tutorial , volume =
V. Generalized semi-infinite programming: a tutorial , volume =. Journal of Computational and Applied Mathematics , number =
-
[269]
Maximum likelihood from incomplete data via the
Dempster, Arthur P and Laird, Nan M and Rubin, Donald B , journal = jrssbm, number =. Maximum likelihood from incomplete data via the
-
[270]
A tutorial on
Hunter, David R and Lange, Kenneth , journal =. A tutorial on
-
[271]
A proximal method for composite minimization , volume =
Lewis, Adrian S and Wright, Stephen J , journal = mp, number =. A proximal method for composite minimization , volume =
-
[272]
Non-smooth non-convex
Ochs, Peter and Fadili, Jalal and Brox, Thomas , journal =. Non-smooth non-convex
-
[273]
Mukkamala, Mahesh Chandra and Fadili, Jalal and Ochs, Peter , title =
-
[274]
Felzenszwalb , booktitle = icml, title =
Sobhan Naderi Parizi and Kun He and Reza Aghajani and Stan Sclaroff and Pedro F. Felzenszwalb , booktitle = icml, title =
-
[275]
Uniqueness of the
Bauer, Martin and Bruveris, Martins and Michor, Peter W , journal =. Uniqueness of the
-
[276]
Friedrich, Thomas , journal =. Die
-
[277]
An Elementary Introduction to Information Geometry , volume =
Frank Nielsen , journal =. An Elementary Introduction to Information Geometry , volume =
-
[278]
Lin, Wu and Nielsen, Frank and Khan, Mohammad Emtiyaz and Schmidt, Mark , title =
-
[279]
Lin, Wu and Schmidt, Mark and Khan, Mohammad Emtiyaz , booktitle = icml, title =
-
[280]
Rudi, Alessandro and Marteau-Ferey, Ulysse and Bach, Francis , title =
-
[281]
Wu Lin and Mohammad Emtiyaz Khan and Mark Schmidt , title =
-
[282]
Exponential varieties , volume =
Micha. Exponential varieties , volume =. Proceedings of the London Mathematical Society , number =
-
[283]
Transformations des signaux al
Bonnet, Georges , booktitle =. Transformations des signaux al
-
[284]
A useful theorem for nonlinear devices having
Robert Price , journal =. A useful theorem for nonlinear devices having
-
[285]
The Variational
Manfred Opper and C. The Variational. Neural Computation , number =
-
[286]
Information Geometry , year =
Nihat Ay, J. Information Geometry , year =
-
[287]
Closures of exponential families , year =
Csisz. Closures of exponential families , year =. Annals of probability , pages =
-
[288]
Methods of information geometry , year =
Amari,. Methods of information geometry , year =
-
[289]
Partial Differential Equations and the Calculus of Variations: Essays in Honor of Ennio De Giorgi , volume =
Colombini, Ferruccio and Marino, Antonio and Modica, Luciano and Spagnolo, Sergio , publisher =. Partial Differential Equations and the Calculus of Variations: Essays in Honor of Ennio De Giorgi , volume =
-
[290]
Generalized solutions and convex duality in optimal control , year =
Fleming, Wendell H , booktitle =. Generalized solutions and convex duality in optimal control , year =
-
[291]
Computational Optimal Transport , volume =
Gabriel Peyr. Computational Optimal Transport , volume =. Foundations and Trends
-
[292]
Woodworth and Nathan Srebro , booktitle = aistats, title =
Suriya Gunasekar and Blake E. Woodworth and Nathan Srebro , booktitle = aistats, title =
-
[293]
Wolinski, Pierre and Charpiat, Guillaume and Ollivier, Yann , title =
-
[294]
Ollivier, Yann , title =
-
[295]
A Tutorial on the Cross-Entropy Method , volume =
Pieter. A Tutorial on the Cross-Entropy Method , volume =. Ann. Oper. Res. , number =
-
[296]
Amos, Brandon and Yarats, Denis , booktitle = icml, title =
-
[297]
Towards the geometry of estimation of distribution algorithms based on the exponential family , year =
Malag. Towards the geometry of estimation of distribution algorithms based on the exponential family , year =. Foundations of Genetic Algorithms XI,
-
[298]
Information-Geometric Optimization Algorithms:
Yann Ollivier and Ludovic Arnold and Anne Auger and Nikolaus Hansen , journal =. Information-Geometric Optimization Algorithms:
-
[299]
Objective improvement in information-geometric optimization , year =
Youhei Akimoto and Yann Ollivier , booktitle =. Objective improvement in information-geometric optimization , year =
-
[300]
Gomez and J
Yi Sun and Faustino J. Gomez and J. Planning to Be Surprised: Optimal. Artificial General Intelligence - 4th International Conference,
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.