Pith. sign in

REVIEW 2 cited by

KKT Conditions, First-Order and Second-Order Optimization, and Distributed Optimization: Tutorial and Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.01858 v1 pith:WIWSHSRG submitted 2021-10-05 math.OC cs.DCcs.LGcs.NAmath.NA

classification math.OCcs.DCcs.LGcs.NAmath.NA
keywords optimizationgradientincludingmethodsmethoddescentfirst-orderproximal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This is a tutorial and survey paper on Karush-Kuhn-Tucker (KKT) conditions, first-order and second-order numerical optimization, and distributed optimization. After a brief review of history of optimization, we start with some preliminaries on properties of sets, norms, functions, and concepts of optimization. Then, we introduce the optimization problem, standard optimization problems (including linear programming, quadratic programming, and semidefinite programming), and convex problems. We also introduce some techniques such as eliminating inequality, equality, and set constraints, adding slack variables, and epigraph form. We introduce Lagrangian function, dual variables, KKT conditions (including primal feasibility, dual feasibility, weak and strong duality, complementary slackness, and stationarity condition), and solving optimization by method of Lagrange multipliers. Then, we cover first-order optimization including gradient descent, line-search, convergence of gradient methods, momentum, steepest descent, and backpropagation. Other first-order methods are explained, such as accelerated gradient method, stochastic gradient descent, mini-batch gradient descent, stochastic average gradient, stochastic variance reduced gradient, AdaGrad, RMSProp, and Adam optimizer, proximal methods (including proximal mapping, proximal point algorithm, and proximal gradient method), and constrained gradient methods (including projected gradient method, projection onto convex sets, and Frank-Wolfe method). We also cover non-smooth and $\ell_1$ optimization methods including lasso regularization, convex conjugate, Huber function, soft-thresholding, coordinate descent, and subgradient methods. Then, we explain second-order methods including Newton's method for unconstrained, equality constrained, and inequality constrained problems....

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Fully Adaptive Frank-Wolfe Algorithm for Relatively Smooth Problems and Its Application to Centralized Distributed Optimization

    math.OC 2025-07 reject novelty 5.0 of 10

    A Frank-Wolfe method that adapts both the smoothness constant and the triangle-scaling exponent achieves sublinear convergence and a tolerance-dependent linear rate, with a centralized distributed optimization application.

  2. Constrained Non-negative Matrix Factorization for Guided Topic Modeling of Minority Topics

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A constrained NMF with a single seed word list and prevalence constraints improves detection of low-prevalence topics, at least on a small synthetic benchmark.

Pith tools