Pith. sign in

REVIEW 6 minor 153 references

Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks

T0 review · 0 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Challenge audits put real numbers on a network's optimality gap.

desk verdict A genuinely useful certificate framework with an honest separation between suite-relative passage and global-gap claims; the flagship ResNet tightness is a local checkpoint identity, but the paper says so itself. read the letter →

arxiv 2608.12655 v1 pith:FNC2L7LJ submitted 2026-08-12 cs.LG stat.ML

classification cs.LGstat.ML
keywords executablecertificatestrainingassuranceglobaloptimalitygapspectralcoveragechallengepowerReLUnetworksquantizedneuralrepresentationsufficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to turn a flat training curve into a checkable scientific claim: either there is an executable alternative model in the same class with lower loss, or the checkpoint has passed a named challenge suite, or—with an extra coverage argument—the checkpoint is within a provable gap of the global optimum. The central primitive is elementary: any complete candidate model with value $B < J(\theta_t)$ is a replayable witness that $J(\theta_t) - J^\star \ge J(\theta_t) - B$. Passing a finite suite is only suite-relative; the paper introduces challenge power $\Psi(B,s)$ and its inverse $E(B,\tau)$ to quantify the largest global gap compatible with passage under a given budget and tolerance. For squared loss, certified decrease operators $(Q_r,\epsilon_r)$ provide checkable spectral coverage, yielding bounds $J(\theta)-J^\star \le (I_{\mathrm{blk}} + \bar\epsilon)/c$ under full coverage, with a realized-residual refinement $\kappa_{\mathrm{cur}}$ that makes the certificate nonvacuous on a channel-gated ResNet-18 (within factors of 1.74–3.02 of the known gap). A ReLU construction proves coverage is necessary: infinitely many exact conditional head optima can be reached while the trajectory converges to a non-global floor.

What carries the argument

The central objects are executable challenges with certified decrease operators and the spectral coverage they induce. For a block challenge $r$, a certified decrease record $(Q_r,\epsilon_r)$ satisfies $J(\theta) - B_r(\theta) \ge \frac{1}{2n} e^\top Q_r e - \epsilon_r$; the mixture $Q_\pi = \sum_r \pi_r Q_r$ is the key identity because its minimum eigenvalue or its realized quadratic form converts executed improvements into a global-gap bound. Challenge power $\Psi(B,s)$ and its inverse $E(B,\tau)$ give the resource-indexed meaning of finite passage, and $\kappa_{\mathrm{cur}}$ measures coverage along the residual actually present. The exact non-global staircase construction is the counterexample object that shows coverage is essential, while E-optimal design and the stable-core ReLU theorem supply concrete mechanisms for establishing coverage in neural settings.

What would settle it

Run the declared eight-challenge suite on a fresh channel-gated ResNet-18 checkpoint for which the true optimum is known (teacher in the student class), recompute the certified decrease operators from verified Armijo steps and the residual $e$ from the checkpoints, and check whether $J(\theta)-J^\star \le (I_{\mathrm{blk}}+\bar\epsilon)/\kappa_{\mathrm{cur}}$ holds at every epoch; a single violation, or a checkpoint whose realized residual lies in a direction where $\kappa_{\mathrm{cur}}$ is tiny while the true gap is large, would refute the realized-residual certificate.

Watch

Extended reading notes

Core claim

The paper's central claim is that empirical global optimality of a neural network checkpoint can be audited constructively, and that the strength of the audit is fixed by a resource-qualified coverage condition. Concretely: if a declared, architecture-valid procedure materializes a complete model with value $B < J(\theta_t)$, then $J(\theta_t)-J^\star \ge J(\theta_t)-B$; if, in addition, the executed block challenges produce certified decrease operators whose mixture $Q_\pi$ satisfies $Q_\pi \succeq cI$, then $J(\theta)-J^\star \le (I_{\mathrm{blk}} + \bar\epsilon_\pi)/c$ for squared loss, and the realized-residual coefficient $\kappa_{\mathrm{cur}} = e^\top Q_\pi e / \|e\|^2$ gives a checkpoint-specific refinement. The paper also proves the converse frontier: without coverage, a first-order ReLU trainer can reach infinitely many exact conditional head optima while converging to a non-global point. The overall conclusion is a certificate ladder—witness, suite passage, coverage-qualified gap bound, representation sufficiency, task evidence—where each stronger conclusion requires a separate, explicitly justified mechanism.

Load-bearing premise

The global-gap conclusions stand or fall on the premise that each executed block challenge genuinely yields a certified decrease operator $(Q_r,\epsilon_r)$ with $J(\theta)-B_r(\theta) \ge \frac{1}{2n}e^\top Q_r e - \epsilon_r$, and that the challenge mixture $Q_\pi$ has a positive spectral gap over the residual directions relevant to the claimed gap (or, in the realized-residual certificate, that the fixed certification objective has $J^\star=0$ and $\kappa_{\mathrm{cur}}$ is computed at the current checkpoint).

Editorial extensions

If this is right

  • A flat training curve combined with a passed challenge suite is not by itself a globality claim; the paper makes the additional coverage requirement explicit and checkable.
  • With a certified coverage mechanism, passing a current suite yields a quantitative upper bound on the empirical global-optimality gap, for example $J(\theta)-J^\star \le (I_{\mathrm{blk}}+\bar\epsilon)/c$ for squared loss.
  • Challenge power gives a resource-indexed language: the inverse $E(B,\tau)$ states the largest global gap still compatible with passage under a declared budget and tolerance.
  • The exact non-global staircase proves that even exact conditional solvers, repeated stair attainment, and trajectory convergence do not imply globality unless the uncovered direction is controlled.
  • On the channel-gated ResNet-18 distillation problem with known optimum, eight internal challenges cover all 240 audited output directions and the realized-residual certificate bounds the true gap within factors of 1.74–3.02.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper is right, the realized-residual coefficient $\kappa_{\mathrm{cur}}$ could be used as a cheap online tightness gauge: when it drops, the current checkpoint is leaving the residual directions the suite can certify, signaling that new challenge directions are needed.
  • The challenge-power inverse $E(B,\tau)$ suggests a practical algorithm-selection protocol: compare trainers by the smallest budget needed to rule out a given suboptimality gap on a shared certification suite, rather than by final loss alone.
  • The paired representation certificate extends naturally to distillation audits: under squared loss, the bounded quantity is the conditional-mean deficit $\mathbb{E}\|\mu_X-\mu_Z\|^2$, so decoder under-use and representation insufficiency can be separated without access to ground-truth labels.
  • The non-global staircase construction implies that a plateau with no loss decrease should trigger a challenge that revives dead or inactive units, since the preserved representation obstruction is exactly what the construction relies on.
Share X Bluesky LinkedIn Reddit HN

Formalized claims in Lean

  1. Claim #1: The paper's central claim is that empirical global optimality of a neural network checkpoint can be audited constructively, and that the strength of the audit is fixed by a resource-qualified coverage condition. Concretely: if a declared, architecture-valid procedure materializes a complete model with value $B < J(\theta_t)$, then $J(\theta_t)-J^\star \ge J(\theta_t)-B$; if, in addition, the execu

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 6 minor

Summary. The paper introduces Training Under Challenge, an executable-certificate framework for auditing the empirical optimality of a neural-network checkpoint. A declared, architecture-valid procedure constructs complete alternative models in the same certified class and reevaluates the fixed objective; any lower-valued candidate B satisfies J(θ_t) − J* ≥ J(θ_t) − B. The paper develops a resource-indexed challenge-power modulus Ψ(B,s) and its inverse E(B,τ), a spectral coverage mechanism based on certified decrease operators (Q_r, ε_r), uniform, normal-residual, and realized-residual gap bounds under squared loss, a proof-bearing primal–dual ReLU bracket, paired representation certificates, and a finite-trial population calibration procedure. It also proves a converse result: without coverage, a first-order ReLU trainer can reach infinitely many exact conditional head optima while converging to a non-global floor. Experiments include a channel-gated ResNet-18 distillation problem with known optimum, an exact head certificate on a frozen nonlinear representation, and quantized-denoising diagnosis and repair studies. The central inequalities are proved in the text and appendices and are validated by controlled regressions against known ground truth.

Significance. If the framework is taken at face value, it provides a principled way to convert constructive counterexamples into lower bounds on empirical optimality gaps and to state precisely what finite challenge passage licenses. The main strengths are: (i) the one-sided witness inequality and the spectral coverage theorems are cleanly proved from stated premises; (ii) the paper is unusually honest about the distinction between suite-relative passage and coverage-qualified globality, including the explicit statement that the realized-residual coefficient κ_cur is checkpoint-specific; (iii) the exact non-global ReLU staircase is a valuable boundary result showing that repeated exact conditional progress does not imply global optimality; and (iv) the controlled experiments check the theorem identities against known ground truth, with reproducible code and documented artifacts. The uniform spectral bound in the ResNet-18 study is conservative (480–6449 times the true gap), and the tighter realized-residual ratios are pointwise identities rather than regional coverage; the paper discloses this, so the limitation does not undermine the central claims.

minor comments (6)
  1. [Section 1.3] The sentence "Sections 11–10 develop representation, tracking, exact-solver, proof-bearing, population, and null-relative extensions" appears to contain an ordering typo; these topics are developed in Sections 10–12.
  2. [Corollary 12 and Section 6.1, Table 6] The 1.74–3.02 realized-residual ratios are pointwise identities at the specific checkpoint, because κ_cur is defined through the same residual e that defines J(θ); the paper already states that κ_cur does not transfer without recomputation, but I recommend adding an explicit sentence in Section 6.1 and in the abstract that these factors are not regional coverage evidence.
  3. [Appendix H, Theorem 21 proof] The derivation of p_L in Eq. (67) is summarized as "inverting these nested tests"; a step-by-step argument showing monotonicity of g_{m,c}(q) and the two beta inversions would make the population-coverage guarantee easier to verify.
  4. [Figure 15, panel (b)] The y-axis label contains what appears to be a rendering artifact ("10□1"); please replace it with the intended power-of-ten notation.
  5. [Section 6.1] The statement that "the residual fraction outside its range is zero throughout" is a trivial consequence of the stated full rank 240/240 of Qπ on the audited output space; the sentence could be simplified or clarified to avoid suggesting an additional empirical check.
  6. [Theorem 9, Eq. (28)] When a valid lower floor L is not available, the positive-part term can simply be dropped with L = 0; stating this explicitly would help readers apply the partial-coverage bound.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the certificate inequalities are derived from stated certified-decrease records and validated against known-optimum benchmarks; self-citations are contextual, not load-bearing.

full rationale

The derivation chain is self-contained. Proposition 2 is the immediate feasibility consequence J* ≤ B_C,t, and the challenge-power inverse in Theorem 6 is an exact identity for the defined quantities Ψ and E, not an empirical prediction smuggled in from data. The spectral coverage bounds (Theorem 9 and Corollaries 11–12) are proven from the stated certified-decrease records of Definition 8, with Q_r and ε_r obtained from verified Armijo steps, exact affine solves, or other declared constructions; no fitted constant is later renamed as a prediction. The realized-residual coefficient κ_cur is explicitly checkpoint-specific, the paper states that it does not transfer without recomputation, and the tightness factors 1.74–3.02 are reported as pointwise numerical checks against a known optimum rather than as regional coverage statements. The only self-citations (Yeganegi et al. 2025, Eamaz et al. 2026) describe earlier layer-wise and transformer-peeling constructions as special cases or challenge generators; no load-bearing theorem in the present paper depends on the correctness of those earlier papers. All central theoretical claims are either proved in the manuscript or supported by standard external mathematical references. No circular step was found.

Assumptions & free parameters 4 free parameters · 5 assumptions · 3 invented entities

The framework's central theorems rest on certified decrease operators and a positive uniform coverage coefficient for the executed challenge bank; these are stated premises, not free parameters fitted to the target. The ResNet-18 and current-state studies verify the operator records numerically. No physical entities are introduced; the new quantities are mathematical definitions.

free parameters (4)
  • E-optimal mixture weights π_r = SDP solution per checkpoint
    Chosen to maximize the minimum eigenvalue of the coverage operator mixture (Section 6.2). This is a design choice for certificate tightness, not a physical parameter fitted to the target.
  • Green tolerance τ_G = Declared by user (e.g., 1% materiality in QNN study)
    Determines when Green status is issued; a protocol parameter, not data-fitted.
  • Stable-core width threshold p = p ≥ max(m, (16R/λ0) log(n/δ), 32m² E_max² n² R/(π λ0²))
    Sufficient width for Theorem 16; conservative and user-chosen, not fitted to data.
  • Stable margin γ = γ ≤ λ0 √(2π)/(8nR)
    Choice controlling the gate-stable region in Theorem 16.
assumptions (5)
  • domain assumption Coverage premise Q_π ⪰ cI with c > 0
    Theorem 9's global-gap bound requires a positive uniform spectral gap. In the ResNet-18 study c = λ_min(Q_π) is small (4e-5 to 1.1e-4), so the uniform bound is weak.
  • domain assumption Known optimum J* = 0 for ResNet-18
    Section 6.1: teacher belongs to the declared student class, making the ground-truth gap computable.
  • domain assumption Normal-residual orthogonality ⟨e*, d⟩ = 0 in Theorem 10
    Verified numerically to 7.64e-14 in the current-state experiment; for general nonlinear models it is conditional.
  • domain assumption Gate-stability region in Theorem 16
    Checkpoint stable neurons must stay within γ/2 of initialization and residual bounded by E_max.
  • standard math Standard mathematical tools (Sion minimax, matrix Chernoff, Eckart-Young, Fenchel-Rockafellar duality)
    Invoked in Propositions 13, 16, 20, 41 and Appendices.
invented entities (3)
  • Budgeted challenge power Ψ(B,s) and inverse E(B,τ)
    purpose: Quantifies the maximal global gap compatible with finite-suite passage
    Mathematical definitions over the declared challenge family; no external falsifiable handle.
  • Certified decrease operator (Q_r, ε_r)
    purpose: Records the proved improvement strength of each block challenge
    Defined from current gradient and Jacobian data; checkpoint-specific by construction.
  • Challenge-closed optimality
    purpose: Endpoint concept when the challenge family is exhausted
    A topological closure concept, not an empirical entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks." pith.science (2026). https://pith.science/paper/FNC2L7LJ

@misc{pith2026260812655,
  author       = {Pith},
  title        = {Pith review of: Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FNC2L7LJ}},
  note         = {Machine review of arXiv:2608.12655}
}
read the original abstract

A flat training curve does not reveal whether a neural network has reached a global optimum, is locally trapped, is representation-limited, or is mismatched to its trainer. We introduce Training Under Challenge, an executable-certificate framework in which predeclared, architecture-valid procedures construct complete alternatives in the same certified class and reevaluate the same objective. Any lower-valued candidate is a replayable witness that lower-bounds the checkpoint's empirical global-optimality gap. Passing a finite suite is only suite-relative; global-gap conclusions require a separately justified coverage mechanism. We define a resource-indexed challenge-power modulus that characterizes the largest gap compatible with passage. For squared loss, current block-decrease operators make coverage checkable and yield uniform and realized-residual bounds. We prove the converse frontier: without coverage, a first-order ReLU trainer can reach infinitely many exact conditional head optima while converging to a non-global point. On a channel-gated ResNet-18 distillation problem with known optimum, eight internal challenges cover all 240 audited output directions, and realized-residual bounds lie within factors of 1.74--3.02 of the true gap. Paired predictive certificates separate decoder under-use from representation insufficiency, while quantized-denoising studies demonstrate diagnosis, repair, and current-state recertification.

Figures

Figures reproduced from arXiv: 2608.12655 by the authors.

Figure 1
Figure 1. The proposed framework turns complete executable alternatives into a retained evidence frontier. Red and Yellow preserve explicit lower-loss models. Current Green reports passage of the completed named suite and launches the stronger examination. Each later event creates a complete executable stair strictly below the triggering loss. The endpoint becomes globally meaningful only when a coverage theorem, exact solver… view at source ↗
Figure 2
Figure 2. Intermediate targets can remove depth-induced multiplicative flatness. At equal scalar￾derivative budget, the disclosed waypoint ladder solves the product model while an end-to￾end update barely moves; the separation grows with depth, and the proved gradient-flow time required merely to double a small balanced initialization scales as α 2−K. The figure establishes the conditioning mechanism that motivates waypoint c… view at source ↗
Figure 3
Figure 3. The event-driven stronger examination. At Green-0, the stronger policy creates S1 strictly below the triggering loss; every later stair is created only after the previous stair is reached. Before Green-0, the retained Gray ceiling equals the Green frontier, so the entire region below it is displayed as an unearned floor. The final diamond records finite-policy saturation after the loss lands on the last retained sta… view at source ↗
Figures from the paper (21 more)
Figure 4
Figure 4. Figure 4: Challenge power turns a budgeted suite into a quantitative certificate. The left panel shows the guaranteed improvement against each gap severity and budget. The right panel shows the exact generalized-inverse envelope: after passage at tolerance τ , it is the largest …
Figure 5
Figure 5. Figure 5: The spectral coverage certificate chain. Finite-width geometry produces a stable coverage spectrum; executed current challenges quantify the residual directions that the suite can reduce; the certified current kernel converts that directional coverage into a global-gap…
Figure 6
Figure 6. Figure 6: Uniform and realized-residual coverage on a channel-gated ResNet-18. The current eight￾block suite covers all 240 audited output directions at every checkpoint. The minimum￾eigenvalue certificate protects a worst-case direction and is conservative; evaluating the same …
Figure 7
Figure 7. Figure 7: Challenge-basis design. E-optimal weighting preserves the weakest residual direction and can be much sparser than the available portfolio. Random subsets can remain singular even after using more blocks. Only the executed and materialized selected calls enter status. a…
Figure 8
Figure 8. Figure 8: Collective exact coverage. Four rank-two blocks jointly span a six-dimensional prediction space. The smallest fusion-frame eigenvalue gives a certified worst-case best-block fraction; observed random residuals are typically better. Approximate passage includes each sol…
Figure 9
Figure 9. Figure 9: Coverage, not stair count, determines global meaning. Left: under a valid block-coverage coefficient, exact reached stairs contract the true gap. Right: the support-preserving ReLU trainer reaches infinitely many exact current-representation head stairs but converges t…
Figure 10
Figure 10. Figure 10: One exact internal-block certificate across residual CNN and pre-LayerNorm transformer motifs. The closed-form replacement reaches the formula-predicted optimum in full-rank and rank-deficient controls. A terminal LayerNorm creates a clear affine-superposition defect,…
Figure 11
Figure 11. Figure 11: Proof-bearing pointwise coverage in a declared neural class. Complete executable adapters lower the primal ceiling while complete polar separation raises the certified floor. The resulting bracket nearly identifies the entire empirical adapter gap, whereas held-out be…
Figure 12
Figure 12. Figure 12: Calibrated population coverage for a frozen randomized challenge policy. Finite-trial correction preserves nominal confidence and strengthens with additional checkpoints and calls. Development-data selection, fresh calibration, or a declared multiplicity correction pr…
Figure 13
Figure 13. Figure 13: Two worked null-relative rarity calculations. Left: the exact noncentral-χ 2 lower-tail probability and its valid Gaussian Chernoff upper bound for the declared small-ball problem. Right: the exact beta tail for the declared rank-adjusted Gaussian alignment null. The …
Figure 14
Figure 14. Figure 14: Paired certificates isolate representation insufficiency. An outer bracket controls the best full-input predictor, while an inner representation-preserving bracket controls the best decoder on Z. Their cross-difference gives the sharp deficit interval. Under Bayes￾com…
Figure 15
Figure 15. Figure 15: Current-state excess-gap certification on a frozen nonlinear representation. (a) Exact materializable subspace challenges bound the true nonzero-optimum head gap throughout continuation. (b) The same head class can yield either a useful or a vacuous certificate depend…
Figure 16
Figure 16. Figure 16: Standard-training status and executable headroom. High-precision endpoints pass the declared Core suite, while every low-precision run contains material complete-model headroom. The bars report the registered medians over all three declared seeds; the complete per-see…
Figure 17
Figure 17. Figure 17: Intervention is actionable, task-audited, and suite-relative. Every low-precision repair improves held-out PSNR. Every endpoint passes the predeclared Core suite after current￾state recertification, whereas two binary endpoints fail the Full suite containing the prote…
Figure 18
Figure 18. Figure 18: Same-checkpoint controls separate extra compute from structural intervention. Ordinary continuation resolves W4A4. W2A4 is most reliably repaired by alternating backbone￾head training despite wall-matched continuation using more backbone-gradient batches. W1A2 stabili…
Figure 19
Figure 19. Figure 19: Green begins a trainer–challenger experiment in which tracking is measured directly. One executable target is reached by fresh same-family continuation, while another remains unattained within budget. Both are actionable complete models; only the trainer dynamics diff…
Figure 20
Figure 20. Figure 20: Nested retained portfolios are monotone even when individual routes are not. Additional waypoints can improve or worsen one executed CNN or transformer candidate, while the minimum over all no-richer configurations yields the monotone value hierarchy used by the monit…
Figure 21
Figure 21. Figure 21: Feasibility validates a realized witness; construction supplies diagnostic power. Across 4,000 noiseless linear-regression problems, exact least squares and a certified gradient step expose all untrained and early checkpoints, while arbitrary feasible proposals rarely…
Figure 22
Figure 22. Figure 22: Two challenge systems have the same complete static power and undetected-gap surfaces but different endpoint identities. The scalar modulus is exact for one-shot passage; the Bellman operator is needed for exact multistep dynamics. Theorem 39 (Scalar power surfaces do…
Figure 23
Figure 23. Figure 23: Cross-fitting separates reusable representation value from in-sample interpolation. Struc￾tured and random labels can both be fitted when the representation reaches full sample rank, while only the structured relation transfers to untouched data; permutation selectiv￾…
Figure 24
Figure 24. Figure 24: An exact conditional challenge supplies a canonical value that local restarts can miss. Across fixed-input ReLU segment problems, exact activation-pattern enumeration is initialization-independent; additional Adam restarts improve the hit rate while leaving an uncerti…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

153 extracted references · 77 canonical work pages

  1. [1]

    Introductory Lectures on Convex Optimization: A Basic Course , author=

  2. [2]

    SIAM Review , volume=

    Optimization Methods for Large-Scale Machine Learning , author=. SIAM Review , volume=. 2018 , doi=

  3. [3]

    and Pereyra, Victor , title =

    Golub, Gene H. and Pereyra, Victor , title =. SIAM Journal on Numerical Analysis , volume =. 1973 , doi =

  4. [4]

    Train Like a (

    Newman, Elizabeth and Ruthotto, Lars and Hart, Joseph and van Bloemen Waanders, Bart , journal=. Train Like a (. 2021 , doi=

  5. [5]

    Journal of Optimization Theory and Applications , volume =

    Convergence of a Block Coordinate Descent Method for Nondifferentiable Minimization , author =. Journal of Optimization Theory and Applications , volume =. 2001 , doi =

  6. [6]

    Advances in Neural Information Processing Systems , volume =

    Block Coordinate Descent for Neural Networks Provably Finds Global Minima , author =. Advances in Neural Information Processing Systems , volume =. 2025 , doi =

  7. [7]

    Closed-Form Last Layer Optimization , journal =

    Galashov, Alexandre and Da Costa, Natha. Closed-Form Last Layer Optimization , journal =. 2025 , url =

  8. [8]

    Journal of Computer and System Sciences , volume =

    How Easy Is Local Search? , author =. Journal of Computer and System Sciences , volume =. 1988 , doi =

Show all 153 references
  1. [9]

    , title =

    Rice, John R. , title =. Advances in Computers , volume =. 1976 , doi =

  2. [10]

    Variable Neighborhood Search , journal =

    Mladenovi. Variable Neighborhood Search , journal =. 1997 , doi =

  3. [11]

    and Lewis, Robert Michael and Torczon, Virginia , title =

    Kolda, Tamara G. and Lewis, Robert Michael and Torczon, Virginia , title =. SIAM Review , volume =. 2003 , doi =

  4. [12]

    Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak--

    Karimi, Hamed and Nutini, Julie and Schmidt, Mark , booktitle=. Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak--. 2016 , doi=

  5. [13]

    Mathematical Programming , volume =

    Proximal Alternating Linearized Minimization for Nonconvex and Nonsmooth Problems , author =. Mathematical Programming , volume =. 2014 , doi =

  6. [14]

    SIAM Journal on Optimization , volume =

    Nesterov, Yurii , title =. SIAM Journal on Optimization , volume =. 2012 , doi =

  7. [15]

    Journal of the American Mathematical Society , volume =

    Xu, Jinchao and Zikatanov, Ludmil , title =. Journal of the American Mathematical Society , volume =. 2002 , doi =

  8. [16]

    and Kutyniok, Gitta and Li, Shidong , title =

    Casazza, Peter G. and Kutyniok, Gitta and Li, Shidong , title =. Applied and Computational Harmonic Analysis , volume =. 2008 , doi =

  9. [17]

    2006 , doi =

    Pukelsheim, Friedrich , title =. 2006 , doi =

  10. [18]

    Advances in Neural Information Processing Systems , volume=

    Neural Tangent Kernel: Convergence and Generalization in Neural Networks , author=. Advances in Neural Information Processing Systems , volume=

  11. [19]

    Proceedings of the 36th International Conference on Machine Learning , series=

    A Convergence Theory for Deep Learning via Over-Parameterization , author=. Proceedings of the 36th International Conference on Machine Learning , series=. 2019 , publisher=

  12. [20]

    Proceedings of the 36th International Conference on Machine Learning , series=

    Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path? , author=. Proceedings of the 36th International Conference on Machine Learning , series=. 2019 , publisher=

  13. [21]

    Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep

    Nguyen, Quynh and Mondelli, Marco and Mont. Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep. Proceedings of the 38th International Conference on Machine Learning , series =

  14. [22]

    Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence , series =

    Banerjee, Arindam and Cisneros-Velarde, Pedro and Zhu, Libin and Belkin, Mikhail , title =. Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence , series =

  15. [23]

    The Positivity of the Neural Tangent Kernel , journal =

    Carvalho, Lu. The Positivity of the Neural Tangent Kernel , journal =. 2025 , doi =

  16. [25]

    Proceedings of the 37th International Conference on Machine Learning , series=

    Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks , author=. Proceedings of the 37th International Conference on Machine Learning , series=. 2020 , publisher=

  17. [26]

    Convex Geometry of Two-Layer

    Ergen, Tolga and Pilanci, Mert , booktitle=. Convex Geometry of Two-Layer. 2020 , publisher=

  18. [27]

    Fast Convex Optimization for Two-Layer

    Mishkin, Aaron and Sahiner, Arda and Pilanci, Mert , booktitle=. Fast Convex Optimization for Two-Layer. 2022 , publisher=

  19. [28]

    Convex Relaxations of

    Kim, Sungyoon and Pilanci, Mert , booktitle=. Convex Relaxations of. 2024 , publisher=

  20. [29]

    Proceedings of the 39th International Conference on Machine Learning , series =

    Unraveling Attention via Convex Duality: Analysis and Interpretations of Vision Transformers , author =. Proceedings of the 39th International Conference on Machine Learning , series =. 2022 , publisher =

  21. [30]

    Journal of Machine Learning Research , volume=

    Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification , author=. Journal of Machine Learning Research , volume=

  22. [31]

    2021 IEEE Symposium on Security and Privacy , pages=

    Proof-of-Learning: Definitions and Practice , author=. 2021 IEEE Symposium on Security and Privacy , pages=. 2021 , publisher=

  23. [32]

    Advances in Neural Information Processing Systems , volume=

    Optimistic Verifiable Training by Controlling Hardware Nondeterminism , author=. Advances in Neural Information Processing Systems , volume=. 2024 , doi=

  24. [33]

    2025 , publisher=

    Mao, Yuhao and Balauca, Stefan and Vechev, Martin , booktitle=. 2025 , publisher=

  25. [34]

    The Annals of Mathematical Statistics , volume =

    Blackwell, David , title =. The Annals of Mathematical Statistics , volume =. 1953 , doi =

  26. [35]

    , title =

    DeGroot, Morris H. , title =. The Annals of Mathematical Statistics , volume =. 1962 , doi =

  27. [36]

    , title =

    Gneiting, Tilmann and Raftery, Adrian E. , title =. Journal of the American Statistical Association , volume =. 2007 , doi =

  28. [37]

    Journal of Machine Learning Research , volume=

    Information, Divergence and Risk for Binary Experiments , author=. Journal of Machine Learning Research , volume=. 2011 , url=

  29. [38]

    and Cranko, Zac , title =

    Williamson, Robert C. and Cranko, Zac , title =. Journal of Machine Learning Research , volume =. 2024 , url =

  30. [39]

    International Conference on Learning Representations , year=

    A Theory of Usable Information Under Computational Constraints , author=. International Conference on Learning Representations , year=

  31. [40]

    and Vedantam, Ramakrishna , title =

    Dubois, Yann and Kiela, Douwe and Schwab, David J. and Vedantam, Ramakrishna , title =. Advances in Neural Information Processing Systems , volume =

  32. [41]

    The Variational Deficiency Bottleneck , booktitle =

    Banerjee, Pradeep Kumar and Mont. The Variational Deficiency Bottleneck , booktitle =. 2020 , doi =

  33. [42]

    Proceedings of the 12th International Conference on Architectural Support for Programming Languages and Operating Systems , pages=

    Combinatorial Sketching for Finite Programs , author=. Proceedings of the 12th International Conference on Architectural Support for Programming Languages and Operating Systems , pages=. 2006 , doi=

  34. [43]

    Proceedings of the 40th International Conference on Machine Learning , series =

    A Robust Optimisation Perspective on Counterexample-Guided Repair of Neural Networks , author =. Proceedings of the 40th International Conference on Machine Learning , series =. 2023 , publisher =

  35. [44]

    , title =

    Sotoudeh, Matthew and Thakur, Aditya V. , title =. Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation , pages =. 2021 , doi =

  36. [45]

    , title =

    Tao, Zhe and Nawas, Stephanie and Mitchell, Jacqueline and Thakur, Aditya V. , title =. Proceedings of the ACM on Programming Languages , volume =. 2023 , doi =

  37. [46]

    Journal of the ACM , volume=

    Distribution-Free, Risk-Controlling Prediction Sets , author=. Journal of the ACM , volume=. 2021 , doi=

  38. [47]

    The Annals of Applied Statistics , volume=

    Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control , author=. The Annals of Applied Statistics , volume=. 2025 , doi=

  39. [48]

    Journal of Computational and Graphical Statistics , volume =

    Basu, Pallavi and Brill, Barak and Yekutieli, Daniel , title =. Journal of Computational and Graphical Statistics , volume =. 2026 , doi =

  40. [49]

    Advances in Neural Information Processing Systems 28 , year=

    BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations , author=. Advances in Neural Information Processing Systems 28 , year=

  41. [50]

    Advances in Neural Information Processing Systems 29 , year=

    Binarized Neural Networks , author=. Advances in Neural Information Processing Systems 29 , year=

  42. [51]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =

    Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference , author =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2018 , doi =

  43. [53]

    International Conference on Learning Representations , year =

    Learned Step Size Quantization , author =. International Conference on Learning Representations , year =

  44. [54]

    Proceedings of the 39th International Conference on Machine Learning , series=

    Overcoming Oscillations in Quantization-Aware Training , author=. Proceedings of the 39th International Conference on Machine Learning , series=. 2022 , publisher=

  45. [55]

    2025 , publisher=

    Jin, Lisa and Ma, Jianhao and Liu, Zechun and Gromov, Andrey and Defazio, Aaron and Xiao, Lin , booktitle=. 2025 , publisher=

  46. [56]

    2025 , publisher =

    Panferov, Andrei and Chen, Jiale and Tabesh, Soroush and Nikdan, Mahdi and Alistarh, Dan , booktitle =. 2025 , publisher =

  47. [57]

    , title =

    Tropp, Joel A. , title =. Foundations of Computational Mathematics , volume =. 2012 , doi =

  48. [58]

    Journal of Machine Learning Research , volume =

    Permutation Tests for Studying Classifier Performance , author =. Journal of Machine Learning Research , volume =. 2010 , url =

  49. [59]

    The Annals of Statistics , volume =

    Time-uniform, Nonparametric, Nonasymptotic Confidence Sequences , author =. The Annals of Statistics , volume =. 2021 , doi =

  50. [60]

    Statistical Science , volume =

    Game-Theoretic Statistics and Safe Anytime-Valid Inference , author =. Statistical Science , volume =. 2023 , doi =

  51. [61]

    and Vidal, Ren

    Haeffele, Benjamin D. and Vidal, Ren. Global Optimality in Neural Network Training , booktitle =

  52. [62]

    Journal of Machine Learning Research , volume =

    Bach, Francis , title =. Journal of Machine Learning Research , volume =

  53. [63]

    Journal of Machine Learning Research , volume =

    Ergen, Tolga and Pilanci, Mert , title =. Journal of Machine Learning Research , volume =

  54. [64]

    Pacific Journal of Mathematics , volume =

    Sion, Maurice , title =. Pacific Journal of Mathematics , volume =. 1958 , doi =

  55. [65]

    Rendiconti del Circolo Matematico di Palermo , volume =

    Carath. Rendiconti del Circolo Matematico di Palermo , volume =. 1911 , doi =

  56. [66]

    , title =

    Pinsker, Mark S. , title =. 1964 , note =

  57. [67]

    Psychometrika , volume =

    Eckart, Carl and Young, Gale , title =. Psychometrika , volume =. 1936 , doi =

  58. [68]

    The Quarterly Journal of Mathematics , volume =

    Mirsky, Leon , title =. The Quarterly Journal of Mathematics , volume =. 1960 , doi =

  59. [69]

    Tyrrell , title =

    Rockafellar, R. Tyrrell , title =

  60. [70]

    Clopper, C. J. and Pearson, E. S. , title =. Biometrika , volume =. 1934 , doi =

  61. [71]

    Journal of the American Statistical Association , volume =

    Hoeffding, Wassily , title =. Journal of the American Statistical Association , volume =. 1963 , doi =

  62. [72]

    Distributed Optimization of Deeply Nested Systems , booktitle =

    Carreira-Perpi. Distributed Optimization of Deeply Nested Systems , booktitle =. 2014 , publisher =

  63. [75]

    Proceedings of the 36th International Conference on Machine Learning , series =

    Belilovsky, Eugene and Eickenberg, Michael and Oyallon, Edouard , title =. Proceedings of the 36th International Conference on Machine Learning , series =. 2019 , publisher =

  64. [76]

    Training Neural Networks with Local Error Signals , booktitle =

    N. Training Neural Networks with Local Error Signals , booktitle =. 2019 , publisher =

  65. [77]

    Neural Networks , volume =

    Baldi, Pierre and Hornik, Kurt , title =. Neural Networks , volume =. 1989 , doi =

  66. [78]

    NeurIPS 2025 Workshop on Optimization for Machine Learning (OPT 2025) , year =

    Yeganegi, Farhang and Eamaz, Arian and Soltanalian, Mojtaba , title =. NeurIPS 2025 Workshop on Optimization for Machine Learning (OPT 2025) , year =

  67. [80]

    Block coordinate descent for neural networks provably finds global minima

    Shunta Akiyama. Block coordinate descent for neural networks provably finds global minima. In Advances in Neural Information Processing Systems, volume 38, 2025. doi:10.52202/085713-5288. URL https://proceedings.neurips.cc/paper_files/paper/2025/hash/e7f81c37f330a06aed8e466a0c...

  68. [81]

    Understanding intermediate layers using linear classifier probes

    Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644, 2016. URL https://arxiv.org/abs/1610.01644

  69. [82]

    A convergence theory for deep learning via over-parameterization

    Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. A convergence theory for deep learning via over-parameterization. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 242--252. PMLR, 2019

  70. [83]

    Anderson, Ziye Ma, Jingqi Li, and Somayeh Sojoudi

    Brendon G. Anderson, Ziye Ma, Jingqi Li, and Somayeh Sojoudi. Towards optimal branching of linear and semidefinite relaxations for neural network robustness certification. Journal of Machine Learning Research, 26 0 (81): 0 1--59, 2025

  71. [84]

    Angelopoulos, Stephen Bates, Emmanuel J

    Anastasios N. Angelopoulos, Stephen Bates, Emmanuel J. Cand \`e s, Michael I. Jordan, and Lihua Lei. Learn then test: Calibrating predictive algorithms to achieve risk control. The Annals of Applied Statistics, 19 0 (2): 0 1641--1662, 2025. doi:10.1214/24-AOAS1998

  72. [85]

    Breaking the curse of dimensionality with convex neural networks

    Francis Bach. Breaking the curse of dimensionality with convex neural networks. Journal of Machine Learning Research, 18 0 (19): 0 1--53, 2017

  73. [86]

    Neural networks and principal component analysis: Learning from examples without local minima

    Pierre Baldi and Kurt Hornik. Neural networks and principal component analysis: Learning from examples without local minima. Neural Networks, 2 0 (1): 0 53--58, 1989. doi:10.1016/0893-6080(89)90014-2

  74. [87]

    Neural tangent kernel at initialization: Linear width suffices

    Arindam Banerjee, Pedro Cisneros-Velarde, Libin Zhu, and Mikhail Belkin. Neural tangent kernel at initialization: Linear width suffices. In Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, volume 216 of Proceedings of Machine Learning Resea...

  75. [88]

    The variational deficiency bottleneck

    Pradeep Kumar Banerjee and Guido Mont \'u far. The variational deficiency bottleneck. In International Joint Conference on Neural Networks, pages 1--8, 2020. doi:10.1109/IJCNN48605.2020.9206900

  76. [89]

    Exact confidence intervals for the mixing distribution from binomial mixture distribution samples

    Pallavi Basu, Barak Brill, and Daniel Yekutieli. Exact confidence intervals for the mixing distribution from binomial mixture distribution samples. Journal of Computational and Graphical Statistics, 35 0 (2): 0 880--890, 2026. doi:10.1080/10618600.2025.2573147

  77. [90]

    Angelopoulos, Lihua Lei, Jitendra Malik, and Michael I

    Stephen Bates, Anastasios N. Angelopoulos, Lihua Lei, Jitendra Malik, and Michael I. Jordan. Distribution-free, risk-controlling prediction sets. Journal of the ACM, 68 0 (6): 0 43:1--43:34, 2021. doi:10.1145/3478535

  78. [91]

    Greedy layerwise learning can scale to ImageNet

    Eugene Belilovsky, Michael Eickenberg, and Edouard Oyallon. Greedy layerwise learning can scale to ImageNet . In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 583--593. PMLR, 2019. URL https:/...

  79. [92]

    Equivalent comparisons of experiments

    David Blackwell. Equivalent comparisons of experiments. The Annals of Mathematical Statistics, 24 0 (2): 0 265--272, 1953. doi:10.1214/aoms/1177729032

  80. [93]

    A robust optimisation perspective on counterexample-guided repair of neural networks

    David Boetius, Stefan Leue, and Tobias Sutter. A robust optimisation perspective on counterexample-guided repair of neural networks. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 2712--273...

  81. [94]

    Proximal alternating linearized minimization for nonconvex and nonsmooth problems

    J \'e r \^o me Bolte, Shoham Sabach, and Marc Teboulle. Proximal alternating linearized minimization for nonconvex and nonsmooth problems. Mathematical Programming, 146 0 (1--2): 0 459--494, 2014. doi:10.1007/s10107-013-0701-9

  82. [95]

    Curtis, and Jorge Nocedal

    L \'e on Bottou, Frank E. Curtis, and Jorge Nocedal. Optimization methods for large-scale machine learning. SIAM Review, 60 0 (2): 0 223--311, 2018. doi:10.1137/16M1080173

  83. [96]

    U ber den variabilit \

    Constantin Carath \'e odory. \"U ber den variabilit \"a tsbereich der fourierschen konstanten von positiven harmonischen funktionen. Rendiconti del Circolo Matematico di Palermo, 32: 0 193--217, 1911. doi:10.1007/BF03014795

  84. [97]

    Distributed optimization of deeply nested systems

    Miguel Carreira-Perpi \ n \'a n and Weiran Wang. Distributed optimization of deeply nested systems. In Proceedings of the Seventeenth International Conference on Artificial Intelligence and Statistics, volume 33 of Proceedings of Machine Learning Research, pages 10--19. PMLR, ...

  85. [98]

    Costa, Jos \'e Mour \ a o, and Gon c alo Oliveira

    Lu \'i s Carvalho, Jo \ a o L. Costa, Jos \'e Mour \ a o, and Gon c alo Oliveira. The positivity of the neural tangent kernel. SIAM Journal on Mathematics of Data Science, 7 0 (2): 0 495--515, 2025. doi:10.1137/24M1659534

  86. [99]

    Casazza, Gitta Kutyniok, and Shidong Li

    Peter G. Casazza, Gitta Kutyniok, and Shidong Li. Fusion frames and distributed processing. Applied and Computational Harmonic Analysis, 25 0 (1): 0 114--132, 2008. doi:10.1016/j.acha.2007.10.001

  87. [100]

    Pact: Parameterized clipping activation for quantized neural networks

    Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. Pact: Parameterized clipping activation for quantized neural networks. arXiv preprint arXiv:1805.06085, 2018

  88. [101]

    C. J. Clopper and E. S. Pearson. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26 0 (4): 0 404--413, 1934. doi:10.1093/biomet/26.4.404

  89. [102]

    Binaryconnect: Training deep neural networks with binary weights during propagations

    Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in Neural Information Processing Systems 28, 2015

  90. [103]

    Morris H. DeGroot. Uncertainty, information, and sequential experiments. The Annals of Mathematical Statistics, 33 0 (2): 0 404--419, 1962. doi:10.1214/aoms/1177704567

  91. [104]

    Schwab, and Ramakrishna Vedantam

    Yann Dubois, Douwe Kiela, David J. Schwab, and Ramakrishna Vedantam. Learning optimal representations with the decodable information bottleneck. In Advances in Neural Information Processing Systems, volume 33, pages 18674--18690, 2020

  92. [105]

    Trust, but verify: Peeling low-bit transformer networks for training monitoring

    Arian Eamaz, Farhang Yeganegi, and Mojtaba Soltanalian. Trust, but verify: Peeling low-bit transformer networks for training monitoring. arXiv preprint arXiv:2605.02853, 2026. doi:10.48550/arXiv.2605.02853. URL https://arxiv.org/abs/2605.02853

  93. [106]

    The approximation of one matrix by another of lower rank

    Carl Eckart and Gale Young. The approximation of one matrix by another of lower rank. Psychometrika, 1: 0 211--218, 1936. doi:10.1007/BF02288367

  94. [107]

    Convex geometry of two-layer ReLU networks: Implicit autoencoding and interpretable models

    Tolga Ergen and Mert Pilanci. Convex geometry of two-layer ReLU networks: Implicit autoencoding and interpretable models. In Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Resear...

  95. [108]

    Convex geometry and duality of over-parameterized neural networks

    Tolga Ergen and Mert Pilanci. Convex geometry and duality of over-parameterized neural networks. Journal of Machine Learning Research, 22 0 (212): 0 1--63, 2021

  96. [109]

    Esser, Jeffrey L

    Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S. Modha. Learned step size quantization. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkgO66VKDS

  97. [110]

    Closed-form last layer optimization

    Alexandre Galashov, Natha \"e l Da Costa, Liyuan Xu, Philipp Hennig, and Arthur Gretton. Closed-form last layer optimization. arXiv preprint arXiv:2510.04606, 2025. URL https://arxiv.org/abs/2510.04606

  98. [111]

    Tilmann Gneiting and Adrian E. Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102 0 (477): 0 359--378, 2007. doi:10.1198/016214506000001437

  99. [112]

    Golub and Victor Pereyra

    Gene H. Golub and Victor Pereyra. The differentiation of pseudo-inverses and nonlinear least squares problems whose variables separate. SIAM Journal on Numerical Analysis, 10 0 (2): 0 413--432, 1973. doi:10.1137/0710036

  100. [113]

    Haeffele and Ren \'e Vidal

    Benjamin D. Haeffele and Ren \'e Vidal. Global optimality in neural network training. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7331--7339, 2017

  101. [114]

    Designing and interpreting probes with control tasks

    John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pages 2733--2743. Association...

  102. [115]

    Probability inequalities for sums of bounded random variables

    Wassily Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58 0 (301): 0 13--30, 1963. doi:10.1080/01621459.1963.10500830

  103. [116]

    Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon

    Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics, 49 0 (2): 0 1055--1080, 2021. doi:10.1214/20-AOS1991

  104. [117]

    Binarized neural networks

    Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Binarized neural networks. In Advances in Neural Information Processing Systems 29, 2016

  105. [118]

    Quantization and training of neural networks for efficient integer-arithmetic-only inference

    Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE Conference on Computer Vision an...

  106. [119]

    Neural tangent kernel: Convergence and generalization in neural networks

    Arthur Jacot, Franck Gabriel, and Cl \'e ment Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems, volume 31, 2018

  107. [120]

    Choquette-Choo, Natalie Dullerud, Anvith Thudi, Varun Chandrasekaran, and Nicolas Papernot

    Hengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud, Anvith Thudi, Varun Chandrasekaran, and Nicolas Papernot. Proof-of-learning: Definitions and practice. In 2021 IEEE Symposium on Security and Privacy, pages 1039--1056. IEEE, 2021. doi:10.1109/SP40...

  108. [121]

    PARQ : Piecewise-affine regularized quantization

    Lisa Jin, Jianhao Ma, Zechun Liu, Andrey Gromov, Aaron Defazio, and Lin Xiao. PARQ : Piecewise-affine regularized quantization. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 28044--28062. ...

  109. [122]

    Johnson, Christos H

    David S. Johnson, Christos H. Papadimitriou, and Mihalis Yannakakis. How easy is local search? Journal of Computer and System Sciences, 37 0 (1): 0 79--100, 1988. doi:10.1016/0022-0000(88)90046-3

  110. [123]

    Linear convergence of gradient and proximal-gradient methods under the polyak-- ojasiewicz condition

    Hamed Karimi, Julie Nutini, and Mark Schmidt. Linear convergence of gradient and proximal-gradient methods under the polyak-- ojasiewicz condition. In Machine Learning and Knowledge Discovery in Databases, volume 9851 of Lecture Notes in Computer Science, pages 795--811. Sprin...

  111. [124]

    Convex relaxations of ReLU neural networks approximate global optima in polynomial time

    Sungyoon Kim and Mert Pilanci. Convex relaxations of ReLU neural networks approximate global optima in polynomial time. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 24458--24485. PMLR, 2024

  112. [125]

    Kolda, Robert Michael Lewis, and Virginia Torczon

    Tamara G. Kolda, Robert Michael Lewis, and Virginia Torczon. Optimization by direct search: New perspectives on some classical and modern methods. SIAM Review, 45 0 (3): 0 385--482, 2003. doi:10.1137/S003614450242889

  113. [126]

    CTBench : A library and benchmark for certified training

    Yuhao Mao, Stefan Balauca, and Martin Vechev. CTBench : A library and benchmark for certified training. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 42984--43010. PMLR, 2025

  114. [127]

    Symmetric gauge functions and unitarily invariant norms

    Leon Mirsky. Symmetric gauge functions and unitarily invariant norms. The Quarterly Journal of Mathematics, 11 0 (1): 0 50--59, 1960. doi:10.1093/qmath/11.1.50

  115. [128]

    Fast convex optimization for two-layer ReLU networks: Equivalent model classes and cone decompositions

    Aaron Mishkin, Arda Sahiner, and Mert Pilanci. Fast convex optimization for two-layer ReLU networks: Equivalent model classes and cone decompositions. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Researc...

  116. [129]

    Variable neighborhood search

    Nenad Mladenovi \'c and Pierre Hansen. Variable neighborhood search. Computers & Operations Research, 24 0 (11): 0 1097--1100, 1997. doi:10.1016/S0305-0548(97)00031-2

  117. [130]

    Overcoming oscillations in quantization-aware training

    Markus Nagel, Marios Fournarakis, Yelysei Bondarenko, and Tijmen Blankevoort. Overcoming oscillations in quantization-aware training. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 16318--1...

  118. [131]

    Introductory Lectures on Convex Optimization: A Basic Course

    Yurii Nesterov. Introductory Lectures on Convex Optimization: A Basic Course. Kluwer Academic Publishers, 2004

  119. [132]

    Efficiency of coordinate descent methods on huge-scale optimization problems

    Yurii Nesterov. Efficiency of coordinate descent methods on huge-scale optimization problems. SIAM Journal on Optimization, 22 0 (2): 0 341--362, 2012. doi:10.1137/100802001

  120. [133]

    Train like a ( Var )pro: Efficient training of neural networks with variable projection

    Elizabeth Newman, Lars Ruthotto, Joseph Hart, and Bart van Bloemen Waanders. Train like a ( Var )pro: Efficient training of neural networks with variable projection. SIAM Journal on Mathematics of Data Science, 3 0 (4): 0 1041--1066, 2021. doi:10.1137/20M1359511

  121. [134]

    Mont \'u far

    Quynh Nguyen, Marco Mondelli, and Guido F. Mont \'u far. Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep ReLU networks. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research...

  122. [135]

    Training neural networks with local error signals

    Arild N kland and Lars Hiller Eidnes. Training neural networks with local error signals. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 4839--4850. PMLR, 2019. URL https://proceedings.mlr.pr...

  123. [136]

    Markus Ojala and Gemma C. Garriga. Permutation tests for studying classifier performance. Journal of Machine Learning Research, 11 0 (62): 0 1833--1863, 2010. URL https://jmlr.org/papers/v11/ojala10a.html

  124. [137]

    Samet Oymak and Mahdi Soltanolkotabi. Overparameterized nonlinear learning: Gradient descent takes the shortest path? In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 4951--4960. PMLR, 2019

  125. [138]

    Q u EST : Stable training of LLM s with 1-bit weights and activations

    Andrei Panferov, Jiale Chen, Soroush Tabesh, Mahdi Nikdan, and Dan Alistarh. Q u EST : Stable training of LLM s with 1-bit weights and activations. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, ...

  126. [139]

    Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks

    Mert Pilanci and Tolga Ergen. Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research...

  127. [140]

    Mark S. Pinsker. Information and Information Stability of Random Variables and Processes. Holden-Day, San Francisco, 1964. Translated and edited by Amiel Feinstein

  128. [141]

    Optimal Design of Experiments

    Friedrich Pukelsheim. Optimal Design of Experiments. Society for Industrial and Applied Mathematics, Philadelphia, 2006. doi:10.1137/1.9780898719109

  129. [142]

    Game-theoretic statistics and safe anytime-valid inference

    Aaditya Ramdas, Peter Gr \"u nwald, Vladimir Vovk, and Glenn Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38 0 (4): 0 576--601, 2023. doi:10.1214/23-STS894

  130. [143]

    Reid and Robert C

    Mark D. Reid and Robert C. Williamson. Information, divergence and risk for binary experiments. Journal of Machine Learning Research, 12 0 (22): 0 731--817, 2011. URL https://jmlr.org/papers/v12/reid11a.html

  131. [144]

    John R. Rice. The algorithm selection problem. In Advances in Computers, volume 15, pages 65--118. Elsevier, 1976. doi:10.1016/S0065-2458(08)60520-3

  132. [145]

    Tyrrell Rockafellar

    R. Tyrrell Rockafellar. Convex Analysis. Princeton University Press, Princeton, New Jersey, 1970

  133. [146]

    Unraveling attention via convex duality: Analysis and interpretations of vision transformers

    Arda Sahiner, Tolga Ergen, Batu Ozturkler, John Pauly, Morteza Mardani, and Mert Pilanci. Unraveling attention via convex duality: Analysis and interpretations of vision transformers. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Procee...

  134. [147]

    On general minimax theorems

    Maurice Sion. On general minimax theorems. Pacific Journal of Mathematics, 8 0 (1): 0 171--176, 1958. doi:10.2140/pjm.1958.8.171

  135. [148]

    Seshia, and Vijay A

    Armando Solar-Lezama, Liviu Tancau, Rastislav Bodik, Sanjit A. Seshia, and Vijay A. Saraswat. Combinatorial sketching for finite programs. In Proceedings of the 12th International Conference on Architectural Support for Programming Languages and Operating Systems, pages 404--4...

  136. [149]

    Tight worst-case bounds for the smallest eigenvalue of ReLU NTK gram matrices

    Zhao Song. Tight worst-case bounds for the smallest eigenvalue of ReLU NTK gram matrices. arXiv preprint arXiv:2608.03368, August 2026. URL https://arxiv.org/abs/2608.03368

  137. [150]

    Matthew Sotoudeh and Aditya V. Thakur. Provable repair of deep neural networks. In Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, pages 588--603. ACM, 2021. doi:10.1145/3453483.3454064

  138. [151]

    Optimistic verifiable training by controlling hardware nondeterminism

    Megha Srivastava, Simran Arora, and Dan Boneh. Optimistic verifiable training by controlling hardware nondeterminism. In Advances in Neural Information Processing Systems, volume 37, 2024. doi:10.52202/079017-3030

  139. [152]

    Zhe Tao, Stephanie Nawas, Jacqueline Mitchell, and Aditya V. Thakur. Architecture-preserving provable repair of deep neural networks. Proceedings of the ACM on Programming Languages, 7 0 (PLDI): 0 443--467, 2023. doi:10.1145/3591238

  140. [153]

    Joel A. Tropp. User-friendly tail bounds for sums of random matrices. Foundations of Computational Mathematics, 12 0 (4): 0 389--434, 2012. doi:10.1007/s10208-011-9099-z

  141. [154]

    Convergence of a block coordinate descent method for nondifferentiable minimization

    Paul Tseng. Convergence of a block coordinate descent method for nondifferentiable minimization. Journal of Optimization Theory and Applications, 109 0 (3): 0 475--494, 2001. doi:10.1023/A:1017501703105

  142. [155]

    Williamson and Zac Cranko

    Robert C. Williamson and Zac Cranko. Information processing equalities and the information--risk bridge. Journal of Machine Learning Research, 25 0 (103): 0 1--53, 2024. URL https://jmlr.org/papers/v25/22-0988.html

  143. [156]

    The method of alternating projections and the method of subspace corrections in hilbert space

    Jinchao Xu and Ludmil Zikatanov. The method of alternating projections and the method of subspace corrections in hilbert space. Journal of the American Mathematical Society, 15 0 (3): 0 573--597, 2002. doi:10.1090/S0894-0347-02-00398-3

  144. [157]

    A theory of usable information under computational constraints

    Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. A theory of usable information under computational constraints. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1eBeyHFDH

  145. [158]

    Data-aware training quality monitoring and certification for deep learning

    Farhang Yeganegi, Arian Eamaz, and Mojtaba Soltanalian. Data-aware training quality monitoring and certification for deep learning. In NeurIPS 2025 Workshop on Optimization for Machine Learning (OPT 2025), 2025. URL https://www.opt-ml.org/papers.html

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.