REVIEW 3 major objections 4 minor 103 references
Parametric Majorization for Data-Driven Energy Minimization Methods
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Bi-level training of parametric energy models can be replaced by a single-level majorizing surrogate, and the paper proves a hierarchy from exact Bregman distances down to a cheap gradient penalty.
desk verdict A solid, transparent framework for majorizer-based training of energy models, but the headline 'without collapse' claim outruns the theory for the iterative scheme used in the experiments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the parametric majorizer $S(x,y,\theta)$, a function satisfying $l(x,x(\theta)) \le S(x,y,\theta)$ for every $\theta$ and vanishing exactly when the loss vanishes. The paper's canonical example is the Bregman distance of the lower-level energy, $D^0_{E_\theta}(x^*, x(\theta)) = E(x^*) - E(x(\theta))$, a convex-analysis measure of deviation from a linear lower bound. Bregman duality rewrites it as $D^{x^*}_{E^*_\theta}(0,q)$, a Bregman distance in the convex conjugate, and decomposing $E=E_1+E_2$ yields a partial surrogate $\min_{z\in\partial E_2(x^*)} W_{E_1}(-z,x^*)$ plus, under strong convexity, the final gradient penalty $\frac{1}{m(\theta,y)}\|q\|^2$. The iterative variant uses the Bregman three-point identity to linearize the loss around the current solution $x(\theta_k)$, producing a descent lemma for the exact majorizer and an implementable over-approximation.
What would settle it
On the CT setup with a rank-deficient angular sampling operator $A$, evaluate $l(x^*_i, x_i(\theta))$ and the gradient-penalty surrogate for a dense random grid of $\theta$ values; if any sample has surrogate value below the true loss, the majorization property that the practical claim relies on is violated, and the CT result would rest on an unproven inequality rather than the paper's theorem.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that the bi-level problem $\min_\theta \sum_i l(x^*_i, x_i(\theta))$ with $x_i(\theta) = \arg\min_x E(x,y_i,\theta)$ can be solved by minimizing a single-level surrogate provided $l(x,z) \le D_{E_\theta}(x,z)$ for all $x,z$ and the surrogate vanishes when the loss vanishes. The Bregman surrogate $D^0_{E_\theta}(x^*_i,x_i(\theta)) = E(x^*_i) - E(x_i(\theta))$ is the tightest member of the family; by Bregman duality it equals $D^{x^*_i}_{E^*_\theta}(0,q_i)$ with $q_i \in \partial E(x^*_i)$, and when $E$ is $m(\theta,y)$-strongly convex this is bounded by the gradient penalty $\frac{1}{m(\theta,y)}\|q_i\|^2$. Proposition 2 records the whole chain $l(x^*_i,x_i(\theta)) \le D^0_{E_\theta}(x^*_i,x_i(\theta)) \le \min_{z\in\partial E_2(x^*_i)} W_{E_1}(-z,x^*_i) \le \frac{1}{m(\theta,y)}\|q_i\|^2$, so the practitioner may choose the cheapest bound that still majorizes the loss. The iterative variant re-linearizes around the current solution and, for the exact majorizer, yields a descent property; the implemented version is an acknowledged over-approximation of that ideal. Experiments on learned CT corrections, TV-entropy segmentation, and analysis-operator denoising show that the surrogates are trainable with standard first-order tools and, in the denoising case, cut training time by an order of magnitude relative to implicit-differentiation baselines while matching their quality.
Load-bearing premise
The load-bearing premise is Equation (8), that the chosen loss is everywhere at most the Bregman distance of the energy; when that fails, as the authors acknowledge it formally does for rank-deficient measurement operators in their CT experiment, the surrogates are no longer guaranteed to bound the true loss.
Editorial extensions
If this is right
- Training an energy-based model can be done with ordinary first-order optimization on the surrogate, avoiding second-order or implicit differentiation of the argmin.
- A zero surrogate value certifies a perfect match, because $S(x,y,\theta)=0$ implies $l(x,x(\theta))=0$, so the training signal is aligned with the model's actual minimizer.
- The ordering in Proposition 2 gives a principled trade-off between surrogate tightness and computational cost, with the gradient penalty usable exactly when the energy is strongly convex.
- The iterative surrogate offers a way to correct for non-separable problems, and the experiments indicate that iterating is needed to reach competitive denoising performance.
- The approach applies to non-smooth convex energies, as demonstrated by the TV-plus-entropy segmentation model, so it is not restricted to differentiable variational models.
Reading between the lines
- One can read the Bregman surrogate as a continuous analogue of margin-based structured prediction: instead of enforcing a fixed margin, the energy is required to grow at least as fast as the loss away from the optimum, which suggests a direct bridge between energy-based learning and large-margin classifiers.
- If the gradient penalty is a valid majorizer, the training objective becomes a regression on the first-order optimality residuals of the energy; this could let practitioners attach a learned correction to a known physical model and train purely on violations of the optimality conditions, a testable recipe for inverse problems beyond CT.
- Because the exact descent guarantee applies only to the unimplemented iterative majorizer, a natural empirical check is to monitor the true bilevel loss and the surrogate per iteration; the paper's step-size-reduction heuristic already hints that the over-approximation can break monotonicity.
- Since Proposition 2 gives a nested sequence of bounds, the gap between the Bregman surrogate and the gradient penalty could serve as a diagnostic for how much accuracy is being traded for speed, guiding adaptive selection of the surrogate during training.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses the bilevel training problem for parametric energy minimization models, in which parameters θ are learned so that the minimizer x(θ)=argmin_x E(x,y,θ) matches a supervised target x* according to a loss l. Because direct differentiation through the argmin is costly and problematic for nonsmooth energies, the authors propose to replace the bilevel objective by a single-level surrogate S(x,y,θ) that majorizes the loss l(x,x(θ)) and satisfies a zero-equivalence condition, then minimize Σ_i S(x_i^*, y_i, θ). They construct a chain of surrogates: the Bregman distance of the lower-level energy, its dual Bregman-conjugate form, partial surrogates via the W-function, and a gradient penalty under strong convexity, and they prove an ordering result in Proposition 2. They also introduce an iterative majorizer based on a linearized bound, prove a descent lemma for the exact iteration (Proposition 4), and propose an over-approximated iterative surrogate (Eqs. 22-23) that is used in the experiments. Applications are presented for computed tomography, variational segmentation, and analysis-operator denoising.
Significance. If the claims held, the framework would be a useful bridge between energy-based modeling and deep learning: it gives a single-level training objective, avoids differentiating through an argmin, is applicable to convex nonsmooth energies, and empirically trains denoising models substantially faster than the implicit-differentiation baseline of [26]. The ordering of surrogates in Proposition 2 is elegant, the toy example in Section 3.3 is consistent, and the authors provide a public implementation. The main caveat is that the implemented iterative scheme used for the headline denoising results is not a parametric majorizer in the paper's own sense, and the CT experiment uses the gradient penalty outside the regime in which it is proven to be a majorizer. Thus the central 'without collapse' guarantee is established only for the non-iterative surrogates, not for all algorithms used in the experiments.
major comments (3)
- [§3.4, Eq. (22) and Appendix A.1.5] The iterative surrogate in Eq. (22) is not a parametric majorizer under Definition 1. For a parameter θ with x(θ)=x* and a previous iterate xbar≠x*, the true loss is l(x*,x(θ))=0, but Eq. (22) evaluates to l(x*,xbar)-⟨∇l(x*,xbar),xbar⟩+E(xbar)+E*(∇l(x*,xbar)) = l(x*,xbar)+W_E(∇l(x*,xbar),xbar), which is strictly positive for the squared-Euclidean and cross-entropy losses used in Sections 4.2 and 4.3 because ∇l(x*,xbar) is not a subgradient of E at xbar. Hence the objective actually minimized in Eq. (23) fails the zero-equivalence clause of Definition 1. Appendix A.1.5 already concedes that Proposition 4 holds only approximately for Eq. (22); the zero-equivalence failure is a separate, sharper reason that the descent and no-collapse guarantees do not transfer to the implemented iterative scheme. Since Table 1 reports denoising results from exactly this scheme, the paper's central claim of 'without collapse' is not established for the algorithm used in its main experiment. The revision should either prove a descent/no-collapse property for the over-approximation under explicit conditions or clearly label Eq. (23) as a heuristic with no monotonicity guarantee.
- [§4.1, Eq. (17)] The CT experiment uses the gradient penalty (17) for E(x)=1/2||Ax-y||²+βR(x)+⟨x,N(θ,y)⟩. The text says this is a parametric majorizer for the Euclidean loss 'if A has full rank (and practically even works beyond this setting...)'. When A is rank-deficient, the energy is not strongly convex and the modulus m(θ,y) in Eq. (17) is zero (or undefined), so the last inequality of Proposition 2 does not apply. The parenthetical 'practically works' is an empirical claim, not a proof; in particular, the data term controls only the range of Aᵀ, not the nullspace component of x*-x(θ). The revision should state the full-rank condition as an explicit assumption for the CT experiment or provide a separate bound that handles rank-deficient A.
- [§3.4, Eq. (23)] Proposition 4 is proved for the exact iterative procedure that minimizes the right-hand side of Eq. (20b), but the implemented iteration (23) minimizes a different objective obtained by an additional Fenchel upper bound and by dropping the constant C. The paper should state explicitly that Eq. (23) is not the algorithm covered by Proposition 4, and should report which loss values are monitored during the step-size reduction heuristic described in Appendix A.2.3. As written, the connection between the theoretical descent lemma and the experimental iterative scheme is too loose to support the claims made in Section 3.4.
minor comments (4)
- [§3.4 / §4.2] The argument order of the Bregman distance is inconsistent: Proposition 3 writes l(x,y)=D_w(y,x), while Eq. (26) writes the log loss as D_h(x_i^*, x_i(θ)). For nonsymmetric distances these two expressions are not equal. Please use one convention throughout and adjust the proof of Proposition 3 accordingly.
- [Figure 1] The caption says the orange curve is the partial surrogate obtained from Eq. (15) by inserting z=∇E1(x*), while Eq. (16) defines the partial surrogate by choosing z∈∂E2(x*). The correspondence between the plotted curves and the displayed equations should be clarified.
- [Appendix A.1.3] The statement of Proposition 2 upper-bounds by 1/(m(θ,y))||q_i||², but the proof arrives at 1/(2m(θ,y))||q_i||². Since the former is the larger bound, the proof is not wrong, but the constants should be reconciled for clarity.
- [§3.2, Eq. (9)] Eq. (9) writes D_Eθ without the superscript 0 that is introduced in Eq. (10); please use D^0_Eθ consistently, since the choice of subgradient matters for the identity.
Circularity Check
No significant circularity: the surrogate chain is derived from stated convex-analysis inequalities; the only self-citation is non-load-bearing, and the acknowledged over-approximation of Eq. (22) is a correctness gap rather than a circular reduction.
full rationale
The paper's derivation chain starts from the explicit assumption l(x,z) <= D_Etheta(x,z) (Eq. (8)) and standard convex-analysis tools (Bregman identity, Fenchel-Young, infimal convolution, strong-convexity/strong-smoothness duality). Definition 1 defines 'parametrized majorizer' by the upper-bound and zero-equivalence properties, so the 'without collapse' statement is a theorem about the constructed surrogates (Proposition 2), not a restatement of the data used to fit them. The descent lemma (Proposition 4) is proved in Appendix A.1.5; the authors' own reference [39] appears only as a label ('nonconvex composite majorizer') and is not used to import the monotonicity result, so the self-citation is not load-bearing. No fitted parameter is renamed as a prediction, and no uniqueness theorem from prior work is invoked to force a choice. The appendix explicitly concedes that the implemented iterative surrogate (22) inherits the descent property only approximately ('As such we expect the results of Proposition 4 to hold only approximately as stated in the main paper.'); this is a correctness limitation, as is the failure of (22) to satisfy the zero-equivalence condition of Definition 1, but it is not a circular reduction.
Assumptions & free parameters
assumptions (4)
- domain assumption Majorization condition (8): l(x,z) ≤ D_{Eθ}(x,z) for all x,z,θ
- standard math Unique minimizer x(θ) exists for all θ,y
- domain assumption E decomposes as E1+E2 with simple conjugates
- domain assumption E is m(θ,y)-strongly convex for the gradient penalty
Cite this review
Pith. "Pith review of Parametric Majorization for Data-Driven Energy Minimization Methods." pith.science (2026). https://pith.science/paper/E7BD2BEZ
@misc{pith2026190806209,
author = {Pith},
title = {Pith review of: Parametric Majorization for Data-Driven Energy Minimization Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/E7BD2BEZ}},
note = {Machine review of arXiv:1908.06209}
}
read the original abstract
Energy minimization methods are a classical tool in a multitude of computer vision applications. While they are interpretable and well-studied, their regularity assumptions are difficult to design by hand. Deep learning techniques on the other hand are purely data-driven, often provide excellent results, but are very difficult to constrain to predefined physical or safety-critical models. A possible combination between the two approaches is to design a parametric energy and train the free parameters in such a way that minimizers of the energy correspond to desired solution on a set of training examples. Unfortunately, such formulations typically lead to bi-level optimization problems, on which common optimization algorithms are difficult to scale to modern requirements in data processing and efficiency. In this work, we present a new strategy to optimize these bi-level problems. We investigate surrogate single-level problems that majorize the target problems and can be implemented with existing tools, leading to efficient algorithms without collapse of the energy function. This framework of strategies enables new avenues to the training of parameterized energy minimization models from large data.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[26]
Insights Into Analysis Operator Learning: From Patch-Based Sparse Models to Higher Order MRFs
Yunjin Chen, Ren ´e Ranftl, and Thomas Pock. Insights Into Analysis Operator Learning: From Patch-Based Sparse Models to Higher Order MRFs. IEEE Transactions on Im- age Processing, 23(3):1060–1072, Mar. 2014. 1, 2, 8, 12 13
work page 2014
-
[1]
Hidden Markov Support Vector Machines
Yasemin Altun, Ioannis Tsochantaridis, and Thomas Hof- mann. Hidden Markov Support Vector Machines. In Pro- ceedings of the Twentieth International Conference on In- ternational Conference on Machine Learning , ICML’03, pages 3–10. AAAI Press, 2003. 2
-
[2]
Zico Kolter
Brandon Amos and J. Zico Kolter. OptNet: Differentiable Optimization as a Layer in Neural Networks. In Interna- tional Conference on Machine Learning , pages 136–145, July 2017. 2
2017
-
[3]
Vegard Antun, Francesco Renna, Clarice Poon, Ben Ad- cock, and Anders C. Hansen. On instabilities of deep learn- ing in image reconstruction - Does AI come at a cost? arXiv:1902.05300 [cs], Feb. 2019. 1, 7
arXiv 1902
-
[4]
Learning real-time MRF inference for image denoising
Adrian Barbu. Learning real-time MRF inference for image denoising. 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 1574–1581, 2009. 2
2009
-
[5]
Legen- dre Functions and the Method of Random Bregman Projec- tions
Heinz H Bauschke and Jonathan (Jon Borwein. Legen- dre Functions and the Method of Random Bregman Projec- tions. Journal of Convex Analysis, 4(1):27–67, May 1997. 9
1997
-
[6]
Bauschke and Patrick L
Heinz H. Bauschke and Patrick L. Combettes. Convex Analysis and Monotone Operator Theory in Hilbert Spaces. CMS Books in Mathematics. Springer New York, New York, NY, 2011. 3, 10
2011
-
[7]
Mirror descent and nonlin- ear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle. Mirror descent and nonlin- ear projected subgradient methods for convex optimization. Operations Research Letters, 31(3):167–175, May 2003. 7, 12
2003
Show all 103 references
-
[8]
A Fast Iterative Shrinkage- Thresholding Algorithm for Linear Inverse Problems
Amir Beck and Marc Teboulle. A Fast Iterative Shrinkage- Thresholding Algorithm for Linear Inverse Problems. SIAM Journal on Imaging Sciences , 2(1):183–202, Jan
-
[9]
Bennett, Gautam Kunapuli, Jing Hu, and Jong- Shi Pang
Kristin P. Bennett, Gautam Kunapuli, Jing Hu, and Jong- Shi Pang. Bilevel Optimization and Machine Learning. In IEEE World Congress on Computational Intelligence, WCCI 2008, Lecture Notes in Computer Science, pages 25–
2008
-
[10]
Modern regularization methods for inverse problems
Martin Benning and Martin Burger. Modern regularization methods for inverse problems. Acta Numerica, 27:1–111, May 2018. 4, 9
2018
-
[11]
Bregman Distances in Inverse Problems and Partial Differential Equations
Martin Burger. Bregman Distances in Inverse Problems and Partial Differential Equations. In Advances in Mathemati- cal Modeling, Optimization and Optimal Control, Springer Optimization and Its Applications, pages 3–33. Springer In- ternational Publishing, Cham, 2016. 3, 9
2016
-
[12]
A Proximal-Projection Method for Finding Zeros of Set-Valued Operators
Dan Butnariu and Gabor Kassay. A Proximal-Projection Method for Finding Zeros of Set-Valued Operators. SIAM J. Control Optim., 47(4):2096–2136, Jan. 2008. 5, 9
2008
-
[13]
Bilevel approaches for learning of variational imaging models
Luca Calatroni, Chung Cao, Juan Carlos De Los Reyes, Carola-Bibiane Sch ¨onlieb, and Tuomo Valkonen. Bilevel approaches for learning of variational imaging models. Variational Methods: In Imaging and Geometric Control , 18:252, 2017. 2
2017
-
[14]
An introduction to total variation for image analysis
Antonin Chambolle, Vicent Caselles, Daniel Cremers, Mat- teo Novaga, and Thomas Pock. An introduction to total variation for image analysis. Theoretical foundations and numerical methods for sparse recovery , 9(263-340):227,
-
[15]
A Convex Approach to Minimal Partitions
Antonin Chambolle, Daniel Cremers, and Thomas Pock. A Convex Approach to Minimal Partitions. SIAM Journal on Imaging Sciences, 5(4):1113–1158, Oct. 2012. 7
2012
-
[16]
A First-Order Primal-Dual Algorithm for Convex Problems with Appli- cations to Imaging
Antonin Chambolle and Thomas Pock. A First-Order Primal-Dual Algorithm for Convex Problems with Appli- cations to Imaging. J Math Imaging Vis , 40(1):120–145, May 2011. 12
2011
-
[17]
On the ergodic con- vergence rates of a first-order primal–dual algorithm.Math- ematical Programming, 159(1-2):253–287, Sept
Antonin Chambolle and Thomas Pock. On the ergodic con- vergence rates of a first-order primal–dual algorithm.Math- ematical Programming, 159(1-2):253–287, Sept. 2016. 12
2016
-
[18]
Chan and Luminita A
Tony F. Chan and Luminita A. Vese. Active contours without edges. IEEE Transactions on image processing , 10(2):266–277, 2001. 1, 7
2001
-
[19]
Fast, Exact and Multi-scale Inference for Semantic Image Segmenta- tion with Deep Gaussian CRFs
Siddhartha Chandra and Iasonas Kokkinos. Fast, Exact and Multi-scale Inference for Semantic Image Segmenta- tion with Deep Gaussian CRFs. In Computer Vision – ECCV 2016 , Lecture Notes in Computer Science, pages 402–418. Springer International Publishing, 2016. 2
2016
-
[20]
Convergence Analysis of a Proximal-Like Minimization Algorithm Using Bregman Functions
Gong Chen and Marc Teboulle. Convergence Analysis of a Proximal-Like Minimization Algorithm Using Bregman Functions. SIAM J. Optim., 3(3):538–543, Aug. 1993. 6
1993
-
[21]
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Semantic Image Seg- mentation with Deep Convolutional Nets and Fully Con- nected CRFs. In International Conference on Learning Representations (ICLR), 2015. 1
2015
-
[22]
Liang-Chieh Chen, George Papandreou, Iasonas Kokki- nos, Kevin Murphy, and Alan L. Yuille. DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. arXiv:1606.00915 [cs], June 2016. 2
2016 arXiv
-
[23]
Trainable Nonlinear Re- action Diffusion: A Flexible Framework for Fast and Ef- fective Image Restoration
Yunjin Chen and Thomas Pock. Trainable Nonlinear Re- action Diffusion: A Flexible Framework for Fast and Ef- fective Image Restoration. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6):1256–1272, June
-
[24]
Learn- ing l1-based analysis and synthesis sparsity priors using bi- level optimization
Yunjin Chen, Thomas Pock, and Horst Bischof. Learn- ing l1-based analysis and synthesis sparsity priors using bi- level optimization. In Neural Information Processing Sys- tems Confercence (NIPS) 2012, 2012. 2, 8
2012
-
[25]
A bi-level view of inpainting - based image compression
Yunjin Chen, Rene Ranftl, and Thomas Pock. A bi-level view of inpainting - based image compression. InComputer Vision Winter Workshop. ., 2014. 2, 12
2014
-
[27]
On Learning Optimized Reaction Diffusion Processes for Effective Im- age Restoration
Yunjin Chen, Wei Yu, and Thomas Pock. On Learning Optimized Reaction Diffusion Processes for Effective Im- age Restoration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5261– 5269, 2015. 2
2015
-
[28]
Discriminative Training Methods for Hid- den Markov Models: Theory and Experiments with Percep- tron Algorithms
Michael Collins. Discriminative Training Methods for Hid- den Markov Models: Theory and Experiments with Percep- tron Algorithms. In Proceedings of the ACL-02 Conference on Empirical Methods in Natural Language Processing - Volume 10, EMNLP ’02, pages 1–8, Stroudsburg, PA, USA,
-
[29]
End-to-End Training of Hybrid CNN-CRF Models for Semantic Segmentation us- ing Structured Learning
Aleksander Colovic, Patrick Kn ¨obelreiter, Alexander Shekhovtsov, and Thomas Pock. End-to-End Training of Hybrid CNN-CRF Models for Semantic Segmentation us- ing Structured Learning. In Computer Vision Winter Work- shop, Feb. 2017. 2
2017
-
[30]
An overview of bilevel optimization
Beno ˆıt Colson, Patrice Marcotte, and Gilles Savard. An overview of bilevel optimization. Annals of Operations Re- search, 153(1):235–256, June 2007. 2
2007
-
[31]
The Cityscapes Dataset for Semantic Urban Scene Understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding. In Pro- ceedings of the IEEE Conference on Computer Vision and Pattern Re...
2016
-
[32]
Cham- bolle
Daniel Cremers, Thomas Pock, Kalin Kolev, and A. Cham- bolle. Convex Relaxation Techniques for Segmentation, Stereo and Multiview Reconstruction. In Markov Random Fields for Vision and Image Processing. MIT Press, Boston,
-
[33]
The structure of optimal parameters for image restoration problems
Juan Carlos De Los Reyes, Carola-Bibiane Sch ¨onlieb, and Tuomo Valkonen. The structure of optimal parameters for image restoration problems. Journal of Mathematical Anal- ysis and Applications, 434(1):464–500, Feb. 2016. 2
2016
-
[34]
Bilevel Parameter Learning for Higher- Order Total Variation Regularisation Models.J Math Imag- ing Vis, 57(1):1–25, Jan
Juan Carlos De Los Reyes, Carola-Bibiane Sch ¨onlieb, and Tuomo Valkonen. Bilevel Parameter Learning for Higher- Order Total Variation Regularisation Models.J Math Imag- ing Vis, 57(1):1–25, Jan. 2017. 2
2017
-
[35]
Foundations of Bilevel Programming
Stephan Dempe. Foundations of Bilevel Programming . Nonconvex Optimization and Its Applications. Springer US, 2002. 2
2002
-
[36]
Is bilevel program- ming a special case of a mathematical program with com- plementarity constraints? Math
Stephan Dempe and Joydeep Dutta. Is bilevel program- ming a special case of a mathematical program with com- plementarity constraints? Math. Program., 131(1):37–48, Feb. 2012. 2
2012
-
[37]
P´erez-Vald´es, and Nataliya Kalashnykova
Stephan Dempe, Vyacheslav Kalashnikov, Gerardo A. P´erez-Vald´es, and Nataliya Kalashnykova. Bilevel Pro- gramming Problems: Theory, Algorithms and Applications to Energy Networks . Energy Systems. Springer-Verlag, Berlin Heidelberg, 2015. 2
2015
-
[38]
Finlayson, Hyung Won Chung, Isaac S
Samuel G. Finlayson, Hyung Won Chung, Isaac S. Kohane, and Andrew L. Beam. Adversarial Attacks Against Medical Deep Learning Systems. arXiv:1804.05296 [cs, stat], Apr
-
[39]
Composite Optimiza- tion by Nonconvex Majorization-Minimization
Jonas Geiping and Michael Moeller. Composite Optimiza- tion by Nonconvex Majorization-Minimization. SIAM J. Imaging Sci., pages 2494–2528, Jan. 2018. 6
2018
-
[40]
Kevin Gimpel and Noah A. Smith. Softmax-Margin CRFs: Training Log-Linear Models with Cost Functions. In Hu- man Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Com- putational Linguistics, pages 733–736, Los Angeles, Cali-...
2010
-
[41]
On Differentiating Parameterized Argmin and Argmax Problems with Application to Bi-level Optimization
Stephen Gould, Basura Fernando, Anoop Cherian, Pe- ter Anderson, Rodrigo Santa Cruz, and Edison Guo. On Differentiating Parameterized Argmin and Argmax Problems with Application to Bi-level Optimization. arXiv:1607.05447 [cs, math], July 2016. 2
2016 arXiv
-
[42]
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation
Andreas Griewank. Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation. Society for In- dustrial and Applied Mathematics, Philadelphia, PA, USA,
-
[43]
Recht, Daniel K
Kerstin Hammernik, Teresa Klatzer, Erich Kobler, Michael P. Recht, Daniel K. Sodickson, Thomas Pock, and Florian Knoll. Learning a variational network for re- construction of accelerated MRI data. Magn Reson Med , 79(6):3055–3071, June 2018. 2
2018
-
[44]
A Deep Learning Architecture for Limited- Angle Computed Tomography Reconstruction
Kerstin Hammernik, Tobias W ¨urfl, Thomas Pock, and An- dreas Maier. A Deep Learning Architecture for Limited- Angle Computed Tomography Reconstruction. In Bildver- arbeitung f ¨ur die Medizin 2017 , Informatik aktuell, pages 92–97. Springer Berlin Heidelberg, 2017. 2
2017
-
[45]
Rautenberg
Michael Hinterm ¨uller and Carlos N. Rautenberg. Optimal Selection of the Regularization Function in a Weighted To- tal Variation Model. Part I: Modelling and Theory. J Math Imaging Vis, 59(3):498–514, Nov. 2017. 2
2017
-
[46]
Rautenberg, Tao Wu, and Andreas Langer
Michael Hinterm ¨uller, Carlos N. Rautenberg, Tao Wu, and Andreas Langer. Optimal Selection of the Regularization Function in a Weighted Total Variation Model. Part II: Al- gorithm, Its Analysis and Numerical Tests.J Math Imaging Vis, 59(3):515–533, Nov. 2017. 2
2017
-
[47]
Springer Berlin Heidelberg, Berlin, Heidelberg, 2008. 2
2008
-
[48]
A Tutorial on MM Algorithms
David R Hunter and Kenneth Lange. A Tutorial on MM Algorithms. The American Statistician, 58(1):30–37, Feb
-
[49]
Jinggang Huang and D. Mumford. Statistics of natural im- ages and models. In Proceedings. 1999 IEEE Computer Society Conference on Computer Vision and Pattern Recog- nition (Cat. No PR00149), volume 1, pages 541–547 V ol. 1, June 1999. 12
1999
-
[50]
McCann, Emmanuel Froustey, and Michael Unser
Kyong Hwan Jin, Michael T. McCann, Emmanuel Froustey, and Michael Unser. Deep Convolutional Neural Network for Inverse Problems in Imaging. IEEE Transactions on Image Processing, 26(9):4509–4522, Sept. 2017. 7
2017
-
[51]
FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keu- per, Alexey Dosovitskiy, and Thomas Brox. FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1647–1655, July 2017. 1
2017
-
[52]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. In International Conference on 14 Learning Representations (ICLR) , San Diego, May 2015. 12
2015
-
[53]
A deep convolutional neural network using directional wavelets for low-dose X-ray CT reconstruction
Eunhee Kang, Junhong Min, and Jong Chul Ye. A deep convolutional neural network using directional wavelets for low-dose X-ray CT reconstruction. Medical Physics , 44(10):e360–e375, Oct. 2017. 7
2017
-
[54]
End-to-End Training of Hybrid CNN-CRF Models for Stereo
Patrick Knobelreiter, Christian Reinbacher, Alexander Shekhovtsov, and Thomas Pock. End-to-End Training of Hybrid CNN-CRF Models for Stereo. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1456–1465, Honolulu, HI, July 2017. IEEE. 2
2017
-
[55]
Learning joint demosaicing and denois- ing based on sequential energy minimization
Teresa Klatzer, Kerstin Hammernik, Patrick Knobelreiter, and Thomas Pock. Learning joint demosaicing and denois- ing based on sequential energy minimization. In2016 IEEE International Conference on Computational Photography (ICCP), pages 1–11, May 2016. 2
2016
-
[56]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural net- works. In Advances in Neural Information Processing Sys- tems, pages 1097–1105, 2012. 1
2012
-
[57]
Charles. D. Kolstad and Leon S. Lasdon. Derivative evalua- tion and computational experience with large bilevel math- ematical programs. J Optim Theory Appl , 65(3):485–499, June 1990. 2
1990
-
[58]
A Projected Gradient Descent Method for CRF Inference Allowing End-to-End Training of Arbitrary Pairwise Potentials
M ˚ans Larsson, Anurag Arnab, Fredrik Kahl, Shuai Zheng, and Philip Torr. A Projected Gradient Descent Method for CRF Inference Allowing End-to-End Training of Arbitrary Pairwise Potentials. In Energy Minimization Methods in Computer Vision and Pattern Recognition, Lecture Not...
2018
-
[59]
A Bilevel Optimization Approach for Parameter Learning in Variational Models
Karl Kunisch and Thomas Pock. A Bilevel Optimization Approach for Parameter Learning in Variational Models. SIAM Journal on Imaging Sciences , 6(2):938–983, Jan
-
[60]
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. Nature, 521(7553):436–444, May 2015. 1
2015
-
[61]
Revisiting Deep Structured Models for Pixel-Level Labeling with Gradient-Based Inference.SIAM J
M ˚ans Larsson, Anurag Arnab, Shuai Zheng, Philip Torr, and Fredrik Kahl. Revisiting Deep Structured Models for Pixel-Level Labeling with Gradient-Based Inference.SIAM J. Imaging Sci., pages 2610–2628, Jan. 2018. 2
2018
-
[62]
Loss functions for discrim- inative training of energy-based models
Yann LeCun and Fu Jie Huang. Loss functions for discrim- inative training of energy-based models. In AISTATS 2005 - Proceedings of the 10th International Workshop on Arti- ficial Intelligence and Statistics , pages 206–213, 2005. 3, 4
2005
-
[63]
Ranzato, and F
Yann LeCun, Sumit Chopra, Raia Hadsell, M. Ranzato, and F. Huang. A tutorial on energy-based learning. Predicting structured data, 1(0), 2006. 3, 4
2006
-
[64]
Fully Convolutional Networks for Semantic Segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015. 1
2015
-
[65]
Efficient Piecewise Training of Deep Structured Models for Semantic Segmentation
Guosheng Lin, Chunhua Shen, Anton van den Hengel, and Ian Reid. Efficient Piecewise Training of Deep Structured Models for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition, pages 3194–3203, 2016. 2
2016
-
[66]
Optimization with First-order Surrogate Functions
Julien Mairal. Optimization with First-order Surrogate Functions. In Proceedings of the 30th International Con- ference on International Conference on Machine Learning - Volume 28, ICML’13, pages III–783–III–791, Atlanta, GA, USA, 2013. JMLR.org. 6
2013
-
[67]
Freund, and Yurii Nesterov
Haihao Lu, Robert M. Freund, and Yurii Nesterov. Rela- tively Smooth Convex Optimization by First-Order Meth- ods, and Applications. SIAM J. Optim. , pages 333–354, Jan. 2018. 4
2018
-
[68]
A Database of Human Segmented Natural Images and its Application to Evaluating Segmentation Algorithms and Measuring Ecological Statistics
David Martin, Charless Fowlkes, Doron Tal, and Jitendra Malik. A Database of Human Segmented Natural Images and its Application to Evaluating Segmentation Algorithms and Measuring Ecological Statistics. In Proceedings of 8th International Conference on Computer Vision , volume...
2001
-
[69]
Incremental Majorization-Minimization Op- timization with Application to Large-Scale Machine Learn- ing
Julien Mairal. Incremental Majorization-Minimization Op- timization with Application to Large-Scale Machine Learn- ing. SIAM Journal on Optimization , 25(2):829–855, Jan
-
[70]
A Large Dataset to Train Convolutional Networks for Dispar- ity, Optical Flow, and Scene Flow Estimation
Nikolaus Mayer, Eddy Ilg, Philip H ¨ausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A Large Dataset to Train Convolutional Networks for Dispar- ity, Optical Flow, and Scene Flow Estimation. 2016 IEEE Conference on Computer Vision and Pattern Recog...
2016
-
[71]
Andr ´e F. T. Martins, Noah A. Smith, and Eric P. Xing. Poly- hedral outer approximations with application to natural lan- guage parsing. In Proceedings of the 26th Annual Interna- tional Conference on Machine Learning - ICML ’09, pages 1–8, Montreal, Quebec, Canada, 2009. ACM...
2009
-
[72]
A Survey and Comparison of Discrete and Continuous Multi- label Optimization Approaches for the Potts Model
Claudia Nieuwenhuis, Eno T ¨oppe, and Daniel Cremers. A Survey and Comparison of Discrete and Continuous Multi- label Optimization Approaches for the Potts Model. Int J Comput Vis, 104(3):223–240, Sept. 2013. 7
2013
-
[73]
Universal Adversarial Perturbations Against Semantic Image Segmentation
Jan Hendrik Metzen, Mummadi Chaithanya Kumar, Thomas Brox, and V olker Fischer. Universal Adversarial Perturbations Against Semantic Image Segmentation. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 2774–2783, Oct. 2017. 1
2017
-
[74]
Techniques for Gradient-Based Bilevel Optimization with Non-smooth Lower Level Problems
Peter Ochs, Ren ´e Ranftl, Thomas Brox, and Thomas Pock. Techniques for Gradient-Based Bilevel Optimization with Non-smooth Lower Level Problems. J Math Imaging Vis, 56(2):175–194, Oct. 2016. 12
2016
-
[75]
Bilevel Optimization with Nonsmooth Lower Level Prob- lems
Peter Ochs, Ren ´e Ranftl, Thomas Brox, and Thomas Pock. Bilevel Optimization with Nonsmooth Lower Level Prob- lems. In Scale Space and Variational Methods in Com- puter Vision, Lecture Notes in Computer Science, pages 654–665. Springer International Publishing, 2015. 2
2015
-
[76]
Existence and Approx- imation of Fixed Points of Bregman Firmly Nonexpansive Mappings in Reflexive Banach Spaces
Simeon Reich and Shoham Sabach. Existence and Approx- imation of Fixed Points of Bregman Firmly Nonexpansive Mappings in Reflexive Banach Spaces. In Fixed-Point Al- gorithms for Inverse Problems in Science and Engineering, Springer Optimization and Its Applications, pages 301–
-
[77]
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Al- ban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in PyTorch. In NIPS 2017 Autodiff Work- shop, Long Beach, CA, 2017. 12
2017
-
[78]
Improved total variation-based CT image reconstruction applied to clinical data
Ludwig Ritschl, Frank Bergner, Christof Fleischmann, and Marc Kachelrie \s s. Improved total variation-based CT image reconstruction applied to clinical data. Phys. Med. Biol., 56(6):1545–1561, Feb. 2011. 1
2011
-
[79]
Tyrell Rockafellar
R. Tyrell Rockafellar. Convex Analysis. Princeton Univer- sity Press, Princeton, N.J, 1970. 7
1970
-
[80]
Atgv- net: Accurate depth super-resolution
Gernot Riegler, Matthias R ¨uther, and Horst Bischof. Atgv- net: Accurate depth super-resolution. In European Confer- ence on Computer Vision, pages 268–284. Springer, 2016. 2
2016
-
[81]
U- Net: Convolutional Networks for Biomedical Image Seg- mentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- Net: Convolutional Networks for Biomedical Image Seg- mentation. In Medical Image Computing and Computer- Assisted Intervention – MICCAI 2015 , Lecture Notes in Computer Science, pages 234–241. Springer International Publi...
2015
-
[82]
The perceptron: A probabilistic model for information storage and organization in the brain
Frank Rosenblatt. The perceptron: A probabilistic model for information storage and organization in the brain. Psy- chological Review, 65(6):386–408, 1958. 4
1958
-
[83]
Maurer, David A
Torsten Rohlfing, Calvin R. Maurer, David A. Bluemke, and M. A. Jacobs. V olume-preserving nonrigid registration of MR breast images using free-form deformation with an incompressibility constraint. IEEE Transactions on Medi- cal Imaging, 22(6):730–741, June 2003. 1
2003
-
[84]
Kegan G. G. Samuel and Marshall F. Tappen. Learning op- timized MAP estimates in continuously-valued MRF mod- els. In IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, June 2009. IEEE. 2
2009
-
[85]
The steepest descent direction for the nonlinear bilevel programming problem
Gilles Savard and Jacques Gauvin. The steepest descent direction for the nonlinear bilevel programming problem. Operations Research Letters, 15(5):265–272, June 1994. 2
1994
-
[86]
Rudin, Stanley Osher, and Emad Fatemi
Leonid I. Rudin, Stanley Osher, and Emad Fatemi. Nonlin- ear total variation based noise removal algorithms. Physica D: Nonlinear Phenomena, 60(1):259–268, Nov. 1992. 1, 8
1992
-
[87]
Ronny Huang, Christoph Studer, Soheil Feizi, and Tom Goldstein
Ali Shafahi, W. Ronny Huang, Christoph Studer, Soheil Feizi, and Tom Goldstein. Are adversarial examples in- evitable? In ICLR 2019, New Orleans, Sept. 2018. 1
2019
-
[88]
Ying Sun, Prabhu Babu, and Daniel P. Palomar. Majorization-Minimization Algorithms in Signal Process- ing, Communications, and Machine Learning. IEEE Trans- actions on Signal Processing, 65(3):794–816, Feb. 2017. 6
2017
-
[89]
Saxe, James L
Andrew M. Saxe, James L. McClelland, and Surya Gan- guli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv:1312.6120 [cond- mat, q-bio, stat], Dec. 2013. 12
2013 arXiv
-
[90]
Tappen, Ce Liu, Edward H
Marshall F. Tappen, Ce Liu, Edward H. Adelson, and William T. Freeman. Learning Gaussian Conditional Ran- dom Fields for Low-Level Vision. In 2007 IEEE Confer- ence on Computer Vision and Pattern Recognition , pages 1–8, June 2007. 4
2007
-
[91]
Learning Structured Prediction Models: A Large Margin Approach
Ben Taskar, Vassil Chatalbashev, Daphne Koller, and Car- los Guestrin. Learning Structured Prediction Models: A Large Margin Approach. In Proceedings of the 22nd In- ternational Conference on Machine Learning , ICML ’05, pages 896–903, New York, NY , USA, 2005. ACM. 4, 5
2005
-
[92]
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In arXiv:1312.6199 [Cs], Dec. 2013. 1
2013 arXiv
-
[93]
Jor- dan
Ben Taskar, Simon Lacoste-Julien, and Michael I. Jor- dan. Structured Prediction, Dual Extragradient and Breg- man Projections. Journal of Machine Learning Research , 7(Jul):1627–1653, 2006. 4, 5
2006
-
[94]
A simplified view of first order methods for optimization
Marc Teboulle. A simplified view of first order methods for optimization. Math. Program., pages 1–30, May 2018. 4, 6
2018
-
[95]
Max- Margin Markov Networks
Ben Taskar, Carlos Guestrin, and Daphne Koller. Max- Margin Markov Networks. In Advances in Neural Informa- tion Processing Systems 16, pages 25–32. MIT Press, 2004. 2
2004
-
[96]
Statistical Learning Theory
Vladimir Vapnik. Statistical Learning Theory. 1998 , vol- ume 3. Wiley, New York, 1998. 3, 4, 6
1998
-
[97]
Sam Wiseman and Alexander M. Rush. Sequence-to- Sequence Learning as Beam-Search Optimization. In Pro- ceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages 1296–1306, Austin, Texas, Nov. 2016. Association for Computational Linguis- tics. 2
2016
-
[98]
Large Margin Methods for Structured and Interdependent Output Variables
Ioannis Tsochantaridis, Thorsten Joachims, Thomas Hof- mann, and Yasemin Altun. Large Margin Methods for Structured and Interdependent Output Variables. Journal of Machine Learning Research , 6(Sep):1453–1484, 2005. 2
2005
-
[99]
Beyond a Gaussian Denoiser: Residual Learn- ing of Deep CNN for Image Denoising
Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. Beyond a Gaussian Denoiser: Residual Learn- ing of Deep CNN for Image Denoising. IEEE Transactions on Image Processing, 26(7):3142–3155, July 2017. 1, 12
2017
-
[100]
Shuai Zheng, Sadeep Jayasumana, Bernardino Romera- Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip H. S. Torr. Conditional Random Fields as Recurrent Neural Networks. 2015 IEEE International Conference on Computer Vision (ICCV), pages 1529–1537, Dec. 2015. 2 16
2015
-
[101]
Safe and feasible motion generation for au- tonomous driving via constrained policy net
Wei Zhan, Jiachen Li, Yeping Hu, and Masayoshi Tomizuka. Safe and feasible motion generation for au- tonomous driving via constrained policy net. In IECON 2017 - 43rd Annual Conference of the IEEE Industrial Elec- tronics Society, pages 4588–4593, Oct. 2017. 1
2017
-
[316]
Springer, New York, NY, 2011. 5, 9 15
2011
-
[2002]
Association for Computational Linguistics. 2
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.