Pith. sign in

REVIEW 1 major objections 7 minor 191 references

When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation

T0 review · 1 major / 7 minor · reviewed 2026-07-30 · grok-4.5

Pith's one-line read Depth can replace missing numerical precision only relative to a declared low-bit library, horizon, execution arithmetic, and routing model—and a structural floor no amount of depth can cross.

desk verdict Conditionally sound resource theory that cleanly separates structural floor, pure-depth synthesis, arithmetic phase, and pre-training certificates; the tube hypothesis is the real bridge to practice, not a hidden proof gap. read the letter →

arxiv 2607.23390 v1 pith:NOXJIYPF submitted 2026-07-25 cs.LG math.OC

classification cs.LGmath.OC
keywords quantizedneuralnetworksdepth–precisiontradeoffsresidualrelaxedcontrolserrorfeedbackstructuralfloormixture-of-expertsroutingreachability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks when stacking more low-bit residual steps can stand in for the numerical precision a network is missing, for a fixed input–output map. It treats a quantized residual network as a pure schedule that picks fields from a declared low-bit operation library over a fixed horizon, and characterizes the infinite-depth limit by relaxed controls. The distance from the target map to the closed relaxed reachable set is an exact structural floor: no optimizer and no extra depth can remove it for that library. Pure schedules approach the floor at a first-order rate under ordinary time regularity, but execution arithmetic can reverse the story—full-state write-back can freeze residual updates—while increment error feedback keeps a bounded carry and an exact lattice conservation law. For coherent high-precision comparators with first-order error, matching accuracy forces student depth to scale linearly with teacher depth. Primal and dual certificates can mark a design feasible, impossible, or unresolved before training.

What carries the argument

The structural floor E_Ω,∞(F★): the distance from the target full map to the closed relaxed reachable set of the declared low-bit dictionary family. It separates what the library can express from what finite pure depth, metadata, arithmetic, and routing cost; pure schedules approach it by balanced switching / online simplex rounding, and verified primal–dual bounds turn the floor plus finite-resource radii into feasible / impossible / unresolved decisions.

What would settle it

Find a fixed target and declared low-bit residual library whose structural floor is zero, with coherent first-order high-precision error, yet whose best pure low-bit schedules either stay bounded away from the floor as depth grows or match the high-precision accuracy with depth growing much slower or much faster than linear in the comparator depth—under the paper’s execution and tube assumptions.

Watch

Extended reading notes

Core claim

For a fixed target map and a declared low-bit residual library, the exact asymptotic limit of infinite low-bit depth is the distance from the target to the closed relaxed reachable set generated by that library. That distance is a structural floor no pure schedule can cross. Finite pure depth approaches the floor at rate O(1/D) under bounded-variation time dependence (and a Hölder-adjusted rate otherwise), but only when residual increments remain numerically visible; full-state write-back can add a growing penalty and freeze updates, while increment error feedback replaces that growth by a bounded carry. When a coherent high-precision comparator also has first-order error and the floor is ze

Load-bearing premise

Everything ideal and implemented must stay inside one verified tube where the residual fields stay bounded and Lipschitz, so the theory does not cover attention or normalization that blow up, or arithmetic that overflows that tube.

Editorial extensions

If this is right

  • Before training, a dual lower bound above the tolerance certifies that no depth or optimizer can hit the target with that library.
  • Full-state activation write-back can make deeper low-bit nets worse; preserving residual increments (e.g. error-feedback carry) is required for depth to help.
  • Accuracy matching against a coherent first-order high-precision teacher forces low-bit depth on the order of teacher depth when the floor is zero.
  • Learned codebooks must be charged as metadata separate from schedule depth; logarithmic metadata bits can keep codebook error commensurate with first-order synthesis.
  • Hard routing only keeps a first-order depth law under isolated transversal events and a small-gain route–state loop, not under a frozen positive margin alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Quantization toolchains could add a pre-training screen that brackets the structural floor and rejects libraries whose dual bound already exceeds the tolerance.
  • Hardware paths that quantize the full residual state each microstep are in a different phase from increment-carry designs; kernel choice may matter as much as nominal bit width.
  • The same floor-plus-radius logic could grade looped or unrolled blocks: only refinements that stay coherent with a shared residual horizon earn a depth–precision exchange.
  • Unresolved certificates become a research queue of their own—pointing at which bound (floor, arithmetic, or route) must tighten next rather than treating failed training as non-representability.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 7 minor

Summary. The paper develops a target-specific resource theory for when low-bit residual depth can replace numerical precision. A depth-D student is modeled as a pure schedule over a declared low-bit residual-field dictionary on a fixed horizon, with the state lifted to the full input-indexed map. The distance from the target to the closed relaxed reachable set is identified as the exact structural floor; pure schedules approach it at O(D^{-1}) under bounded-variation time dependence and O(D^{-ϑ}+D^{-1}) under ϑ-Hölder dependence (Theorems 3, 5). Execution arithmetic is shown to change the phase: full-state write-back contributes a Dρ_z certificate term with an exact scalar freeze result (Proposition 7), while increment error feedback telescopes the carry (Theorem 8) and admits a bit-exact common-lattice realization with explicit register widths (Proposition 9). A fixed, D-independent binary teacher has a closed-form optimal error Θ(D^{-1}) (Theorem 10), lifted to residual-ReLU and nonuniform two-token attention realizations (Proposition 12, Theorem 13), yielding D_match = Θ(L) for coherent first-order comparators (Corollary 15). Learned codebooks add a metadata resource with upper, packing, and allocation laws (Theorems 18, 19, 57); state-dependent routing is treated by a transversal-event small-gain theorem (Theorem 23). A primal–dual stack (HJB, support, affine, occupation-measure/SOS) yields feasible/impossible/unresolved decisions (Corollary 28). Companion software (QReplace) and

Significance. If the results hold, the paper supplies a unifying conditional framework that cleanly separates library geometry, synthesis depth, metadata, execution arithmetic, and routing — resources that the literature often collapses into a nominal bit width. Several strengths deserve explicit credit: (i) an exact closed-form fixed-teacher optimum with a nonasymptotic envelope (Theorem 10), making the first-order depth price sharp for one fixed target rather than only minimax; (ii) a Lean 4 artifact with hash-locked build logs and claim-level axiom audits kernel-checking twelve discrete-core statements, plus exhaustive executable verifiers for the attention converse and common-lattice arithmetic; (iii) a certified nonlinear matrix-valued accuracy-matching depth (D_match = 8 = 2L) proved by rational piecewise-affine bounds, with a prospective falsifiable prediction (calibrated D=9 vs certified 8) that was checked; (iv) an explicit evidence hierarchy and trust-boundary discussion that is more disciplined than typical for this area. The component tools (relaxed controls, sigma–delta feedback, occupation measures, hybrid transversality) are mature, and the paper says so; the contribution is the t

major comments (1)
  1. [§3.2 Assumption 1; §9.2–9.3; Theorems 6, 8, 29] The common synthesis tube is the load-bearing bridge for the master law (1) and for every QReplace verdict, and it carries a bootstrap risk the manuscript should address more directly. Assumption 1 simultaneously asserts forward invariance of K for all measurable relaxed controls, all mixed/pure Euler states and interpolation segments, and — via Theorems 6, 8, and 29 — all implemented finite-arithmetic prefixes, together with uniform B and L_z on the enlarged tube K_ρ. For attention blocks, the QK-product Lipschitz constant is state-range dependent (§9.2.2 bounds scores through B_Q, B_K), so the constants that define the tube are valid only on a tube whose existence is part of the hypothesis. The paper resolves this constructively for the contractive soft-threshold class (§9.1, Eq. (197) gives an explicit invariant ball), but the Transformer specialization provides only componentwise err
minor comments (7)
  1. [§8.5, Theorem 27] The main text flags that primal–dual equality requires a 'closed-image qualification detailed in Appendix A; that qualification is not automatic.' Please clarify in the main text that this qualification affects only the no-duality-gap statement, not the validity of dual lower certificates: Theorem 25 is proved directly by monotonicity along trajectories, so Corollary 28's 'certified impossible' verdicts do not depend on strong duality. As written, a reader could over-discount the decision rule or, conversely, over-credit the SOS hierarchy.
  2. [§10.8] The phrase 'causal validation of the theory's central distinction' is stronger than the design supports. The coherent-target arm (depth-D Euler refinement converging to the depth-32 refinement of the same field) is essentially standard Euler convergence and is expected a priori; the informative arm is the direct-target divergence. The 4-bit study uses three QAT seeds, so the Student-t intervals have two degrees of freedom, and the layer-1 hidden-map interval crosses zero (reported, but only mid-paragraph). Please temper the causal language, state what outcome would have falsified the mechanism, and note that the fitted slopes (−1.16, −0.86) come from five depth points.
  3. [Appendix A vs. main text] Notation drift: the appendices use C_{fh,b} and C^{unif}_{fh} where the main text uses C^{end}_{syn} and C^{unif}_{syn} (e.g., Theorem 30 vs. Theorem 3; Theorem 53 vs. Theorem 18). Please harmonize or add the correspondence to Table 3. Similarly, Φ_L(T) is defined twice (Eq. (14) and after Eq. (270)), and Eq. (65) uses u = D^{-1} in the main text but x = 1/D in Appendix A.2.
  4. [§1.4 vs. §8.7] QReplace is described in §1.4 as returning five outcomes (certified feasible, certified impossible, conditionally feasible, diagnostically promising, unresolved), while Corollary 28 defines a three-way decision. Please state the mapping between the two lists and which outcomes are proof-backed versus heuristic.
  5. [§4.2, Figure 2(a)] The write-back term Dρ_z is a worst-case certificate envelope; the only exact freeze result is scalar (Proposition 7). The text acknowledges this, but the 'phase diagram' framing of Figure 2(a) may be read as asserting realized U-shaped behavior in general architectures. A sentence clarifying that no multidimensional lower bound realizing the Dρ_z growth is known would calibrate Claim 3.
  6. [Eq. (36)] The α_k√q/2 term assumes a coordinatewise uniform activation grid in a q-dimensional Euclidean state; please state this explicitly, since the surrounding development is in the lifted Banach space Z.
  7. [References] A substantial fraction of the citations are 2025–2026 arXiv preprints (e.g., Chakrabarti et al. 2026; Park et al. 2026b; Zhao et al. 2026). Please indicate which are peer-reviewed and pin versions, since several novelty-boundary comparisons depend on them.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: structural floor, synthesis rates, and converses are independently derived, not fitted or self-referential.

full rationale

The paper's load-bearing chain is definitional-plus-theorem, not circular. The structural floor E_Ω,∞(F*) is defined as dist(F*, R_Ω,rel); Theorems 3/5 then prove pure schedules approach that closed set at O(D^{-1}) or O(D^{-ϑ}+D^{-1}), so the distance is the asymptotic floor by Hausdorff convergence rather than by renaming the target. The fixed-teacher exact error (Theorem 10), residual-ReLU and attention embeddings, common-lattice conservation (Proposition 9), and packing/metadata laws are closed-form or constructive arguments under stated assumptions; they do not fit parameters from the quantities they claim to predict. Empirical sections are explicitly tiered as diagnostic and separated from theorem claims. Lean 4 checks and rational certificates are independent verification, not self-citation load-bearing. Assumption 1 (common tube) is a strong applicability hypothesis, not a circular step. No self-definitional loop, fitted-as-prediction, or uniqueness-via-author-citation pattern is present.

Assumptions & free parameters 3 free parameters · 8 assumptions · 3 invented entities

The theory is conditional on declared modeling resources: a finite low-bit field library, shared residual horizon, lifted full-map state, verified execution tubes, and stated arithmetic/routing semantics. Standard analysis and hybrid-control tools are imported; the paper does not claim unconditional replacement of precision by depth for arbitrary trained blocks.

free parameters (3)
  • Library- and tube-dependent constants (B, Lz, Lt/Vt or Ht, ϑ, Csyn, LΩ, ho z/ ho D, ηG, route u r,Δr,χ)
    These are problem-declared bounds used in rates and certificates; not universal constants. Empirical sections estimate some (e.g. stability budgets) for diagnostics only.
  • Metadata metric dimension m and entropy prefactor CΩ
    Enter the 2^{-s/m} codebook term when a compact learned dictionary family is assumed; chosen from the declared parameter geometry.
  • Per-microstep resource charge ℓσ and total budget B
    Accounting units for joint depth-metadata allocation; must be declared consistently (bits, latency, energy).
assumptions (8)
  • standard math Banach-space Carathéodory existence/uniqueness for bounded, strongly measurable, uniformly state-Lipschitz relaxed fields on a forward-invariant tube (Assumption 1 / Lemma 60).
    Standard ODE-in-Banach hypotheses; paper states them explicitly because endpoint sets require well-posedness.
  • standard math Online simplex/prefix discrepancy rounding with bound < J per coordinate (Lemma 2), used to get first-order pure-to-relaxed rates.
    Belongs to mixed-integer control / combinatorial integral approximation lineage the paper cites.
  • standard math Superposition principle for continuity equations / occupation measures (Ambrosio et al.) to equate modal measure programs with mixtures of relaxed trajectories (Theorem 27).
    Used for exactness of the infinite-dimensional primal; duality needs extra closed-image qualification the paper flags.
  • domain assumption Common verified tube containing ideal and implemented prefixes; Lipschitz and defect bounds only claimed on that tube (Assumption 1; Theorem 6).
    Load-bearing for all synthesis and hardware-transfer rates; not automatic for attention/LayerNorm globally.
  • domain assumption Coherent fixed-horizon residual refinement family shared by high-precision comparator and low-bit student when stating D=Θ(L) matching (Sections 1.1, 5, Corollary 15).
    Without coherence, subdividing unrelated blocks need not approach the original map; DistilBERT control illustrates failure.
  • domain assumption Isolated transversal top-k route events with small-gain χ<1 and isolation radii (Assumption 22 / Theorem 23); excludes simultaneous/grazing/chattering regimes.
    Hybrid necessity for first-order routed laws; frozen positive-margin arguments are explicitly insufficient for actual swaps.
  • domain assumption Increment quantizer residual radius and non-overflow of carry/state registers on declared ranges for error-feedback and common-lattice exactness (Theorem 8, Proposition 9).
    Arithmetic conclusions are semantics-relative; wrapping without invariant is excluded.
  • ad hoc to paper Deterministic global metadata + input-independent pure schedules for packing lower bound (Theorem 19); no uncharged continuous side information.
    Accounting convention matching the upper metadata theorem; other programming models need different resource ledgers.
invented entities (3)
  • Structural floor E_Ω,∞(F*) = dist(F*, closed relaxed reachable set of declared dictionary family) independent evidence
    purpose: Exact asymptotic invariant separating library expressivity from finite depth/optimization
    Defined from standard reachable-set geometry but elevated as the central neural resource invariant for one target and library.
  • Schedulewise arithmetic radius A_D and write-back vs error-feedback phase diagram independent evidence
    purpose: Quantify when execution preserves or reverses depth benefit
    Operational aggregate of transition defects; falsifiable via freeze diagnostic and register audits.
  • QReplace feasible/impossible/unresolved decision rule combining [L,U] floor bracket with finite-resource radius R_{D,s} independent evidence
    purpose: Pre-training certification interface
    Workflow/software object packaging the theorems; correctness depends on sound certificates fed in.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation." pith.science (2026). https://pith.science/paper/NOXJIYPF

@misc{pith2026260723390,
  author       = {Pith},
  title        = {Pith review of: When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NOXJIYPF}},
  note         = {Machine review of arXiv:2607.23390}
}
abstract

When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system over a fixed horizon as a pure schedule selecting fields from a declared low-bit operation library, and use relaxed controls to characterize its infinite-depth limit. The distance from the target to the closed relaxed reachable set is the exact structural floor: no increase in depth can remove it for that library. Pure schedules approach the relaxed class at rate $O(D^{-1})$ under bounded-variation time dependence and $O(D^{-\vartheta}+D^{-1})$ under Holder dependence of exponent $\vartheta$. Execution arithmetic can reverse this conclusion: full-state write-back introduces a $D\rho_z$ penalty and can freeze residual updates, whereas increment error feedback replaces this growth by a bounded carry term and obeys an exact common-lattice conservation law. A fixed-teacher converse makes this rate sharp: for coherent depth-$L$ first-order high-precision comparators, accuracy matching requires $D=\Theta(L)$. Learned codebooks add a metadata resource, while state-dependent routing introduces hybrid event conditions. Verified primal and dual bounds yield feasible, impossible, or unresolved decisions before training. Companion software implements the workflow, and Lean 4 machine-checks the exact discrete core. Depth replaces precision only relative to a declared library, horizon, execution semantics, and routing model.

Figures

Figures reproduced from arXiv: 2607.23390 by the authors.

Figure 1
Figure 1. The resource-theoretic view. The low-bit library determines the relaxed geometry and therefore the structural floor. Depth synthesizes that geometry with a pure schedule, metadata selects or describes the library, and the execution semantics determine whether the microsteps remain numerically visible. Hard routing couples state and event-time errors. Primal and dual certificates compare the resulting achievable inte… view at source ↗
Figure 2
Figure 2. Execution arithmetic creates a genuine phase diagram. (a) Ideal synthesis ap￾proaches the structural floor, whereas fixed-grid state write-back has a U-shaped envelope and a finite optimal depth. Increment error feedback preserves the first￾order law when its physical residual unit scales with the microstep. (b) In an exact scalar diagnostic, a fixed activation grid eventually erases every residual update; the error… view at source ↗
Figure 3
Figure 3. Reachability is a geometric property of the declared operation library. Pure depth￾D schedules form a discrete cloud. Balanced switching drives that cloud toward the relaxed reachable set. A compatible target has zero structural floor and only a finite-depth remainder; an incompatible target remains separated even as D → ∞. write-back, saturating integer arithmetic, two’s-complement wrapping, or an auxiliary error￾f… view at source ↗
Figures from the paper (23 more)
Figure 4
Figure 4. Figure 4: A fixed target can require a first-order depth price. (a) The exact optimum in Theorem 10 tracks its a/(2eD) asymptote and remains inside the proved nonasymptotic envelope. (b) Multiplication by D exposes the nonzero limiting constant a/(2e); the lower law is not gener…
Figure 5
Figure 5. Figure 5: Depth and global codebook metadata are distinct resources. (a) The excess-error surface Csyn/D + Cmeta2 −s/m for a representative m-dimensional family. The white-backed curve balances the two terms and is not an empirical fit. (b) Under the declared unit-cost budget in…
Figure 6
Figure 6. Figure 6: A route-changing diagnostic consistent with Theorem 23. Left: a direct score perturbation makes the implemented route switch one grid point early, but transver￾sality confines the mismatch to the shaded certified event window. Right: the resulting state error is visibl…
Figure 7
Figure 7. Figure 7: The decision interface. A primal upper bound and a verified dual lower bound define an interval for the implemented target error after adding the finite-resource radius. The interval can certify feasibility, finite-depth impossibility, or asymptotic impossibility; over…
Figure 8
Figure 8. Figure 8: An exact fixed-target attention converse. (a) The two-token state decomposes into a mean mode, which carries the binary scheduling obstruction, and a disagreement mode, which experiences a nonuniform attention contraction with eigenvalue µ = tanh(1). (b) The closed-for…
Figure 9
Figure 9. Figure 9: Exact structural split in a soft-threshold residual network. Under one finite dictio￾nary, a compatible nonlinear target converges toward zero, while an incompatible target converges to the exact infinite-depth floor αs/6. The dashed floor is analytic and the finite-de…
Figure 10
Figure 10. Figure 10: A state-dependent attention structural split. The compatible target follows a first-order decay, whereas a target with a component in a missing output subspace remains at an exact invariant floor. The right panel displays a nonuniform, input-dependent attention matrix…
Figure 11
Figure 11. Figure 11: verifies the minimax D−1 law in the signed-ray soft-threshold family. Globally solved finite programs coincide with the analytic edge-gap formula at every tested depth. At D = 128, the exact values of DED are 0.2371, 0.1185, 0.0790, and 0.0339 for binary, ternary, uni…
Figure 12
Figure 12. Figure 12: First-order synthesis around a nonlinear attention field. Compatible gain al￾phabets all produce D−1 terminal error, while DED approaches a dictionary￾dependent constant. The experiment separates the universal exponent from the nonuniversal depth price. Inside the fea…
Figure 13
Figure 13. Figure 13: Certified soft-threshold feasibility boundary and precision sweet spot. Orange cells have certified infinite-depth floors above the requested tolerance and are asymptotically infeasible. The plot does not infer finite-depth impossibility from the floor alone; isolated…
Figure 14
Figure 14. Figure 14: Attention-output codebook geometry determines the infinite-depth floor. The target is compared with the convex hull generated by each finite codebook. Binary and uniform 2-bit codebooks leave a visible geometric gap; the tested 3-bit polygon contains the target. This …
Figure 15
Figure 15. Figure 15: Attention-codebook phase boundary and matched-resource optimum. The left panel separates positive-floor asymptotic infeasibility from finite matching depths. The right panel shows that the resource-optimal precision can be interior and can change with the target toler…
Figure 16
Figure 16. Figure 16: Finite-horizon amplification controls the depth multiplier. Left: the observed large-depth constant DED grows with the empirical stability budget and with more aggressive normalization/residual scaling. Right: the first tested depth attaining error 2.5 × 10−3 . Stabil…
Figure 17
Figure 17. Figure 17: Hard routing has a margin-controlled phase transition. The dashed boundary is the quantized-router perturbation radius. Below it, route changes create a persistent error; above it, routes agree and the D−1 synthesis regime reappears. sup y∈B2 2 (1) [PITH_FULL_IMAGE:f…
Figure 18
Figure 18. Figure 18: Certified minimum accuracy-matching depth in a nonlinear matrix system. The large left panel excludes every D ≤ 7 and certifies that the selected depth-8 student is more accurate relative to the common depth-32 reference than the depth-4 high-precision comparator. The…
Figure 19
Figure 19. Figure 19: Pretrained-model causal control for depth coherence. The same DistilBERT feed-forward residual fields are evaluated under two targets. (a) Subdividing the original one-step block moves away from its map even without quantization. (b) Refining a common depth-32 residua…
Figure 20
Figure 20. Figure 20: Four-bit learned dictionaries preserve the coherent depth benefit. Means and Student-t 95% intervals are computed across three independent QAT seeds for locally hardened pure schedules. (a) Hidden-map error decreases modestly at the near-saturated first layer and more…
Figure 21
Figure 21. Figure 21: Medium-scale attention validation. In an eight-token, width-16, two-head attention–MLP system, a compatible target preserves the D−1 mechanism, while a component in a missing output direction remains at its exposed floor. Curves show held-out errors over an input ense…
Figure 22
Figure 22. Figure 22: Attention scale sweep. Left: the representative large-depth constant DED for compatible targets. Right: the exposed missing-direction floor. The horizontal labels report (n, d, H): token count, model width, and number of heads. 8 10 12 14 16 18 20 depth D 10−5 10−4 10…
Figure 23
Figure 23. Figure 23: Optimization gap for recurrent soft-threshold atoms. Globally exhaustive schedule search reveals representations that local coordinate descent misses by orders of magnitude. This experiment demonstrates that failed local schedule search is not a certificate of positiv…
Figure 24
Figure 24. Figure 24: Order and optimization matter for noncommuting attention atoms. Exhaustive global optima, balanced schedules, and local-search distributions are reported separately. The theorem supplies a constructive representability law; it does not guarantee that an arbitrary trai…
Figure 25
Figure 25. Figure 25: Relaxed-to-pure compilation gap in the pretrained experiment. Bars show (Ehard/Erelaxed − 1) × 100% for locally refined pure schedules. The small gaps demonstrate that the principal positive result is realized by pure programs rather than only by convex mixtures. The …
Figure 26
Figure 26. Figure 26: Supplementary breadth evidence. (a) The eight-teacher matrix campaign contains exact matches, certified feasible upper bounds, positive-floor impossibilities, and unresolved cases. (b) Expanding the complete dictionary from two to four atoms improves the globally opti…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

191 extracted references · 18 canonical work pages

  1. [1]

    Communications on Pure and Applied Mathematics , volume =

    Ingrid Daubechies and Michel Defrise and Christine De Mol , title =. Communications on Pure and Applied Mathematics , volume =. 2004 , doi =

  2. [2]

    SIAM Journal on Imaging Sciences , volume =

    Amir Beck and Marc Teboulle , title =. SIAM Journal on Imaging Sciences , volume =. 2009 , doi =

  3. [3]

    Proceedings of the 27th International Conference on Machine Learning , pages =

    Karol Gregor and Yann LeCun , title =. Proceedings of the 27th International Conference on Machine Learning , pages =

  4. [4]

    Advances in Neural Information Processing Systems , volume =

    Xiaohan Chen and Jialin Liu and Zhangyang Wang and Wotao Yin , title =. Advances in Neural Information Processing Systems , volume =

  5. [5]

    Eldar , title =

    Vishal Monga and Yuelong Li and Yonina C. Eldar , title =. IEEE Signal Processing Magazine , volume =. 2021 , doi =

  6. [6]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =

    Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2016 , doi =

  7. [7]

    Ricky T. Q. Chen and Yulia Rubanova and Jesse Bettencourt and David Duvenaud , title =. Advances in Neural Information Processing Systems , volume =

  8. [8]

    Inverse Problems , volume =

    Eldad Haber and Lars Ruthotto , title =. Inverse Problems , volume =. 2018 , doi =

Show all 191 references
  1. [9]

    Sander and Pierre Ablin and Gabriel Peyr

    Michael E. Sander and Pierre Ablin and Gabriel Peyr. Do Residual Neural Networks Discretize Neural Ordinary Differential Equations? , booktitle =. 2022 , doi =

  2. [10]

    Jack Warga , title =

  3. [11]

    Young , title =

    Laurence C. Young , title =

  4. [12]

    Advances in Neural Information Processing Systems , volume =

    Matthieu Courbariaux and Yoshua Bengio and Jean-Pierre David , title =. Advances in Neural Information Processing Systems , volume =

  5. [13]

    Journal of Machine Learning Research , volume =

    Itay Hubara and Matthieu Courbariaux and Daniel Soudry and Ran El-Yaniv and Yoshua Bengio , title =. Journal of Machine Learning Research , volume =

  6. [14]

    European Conference on Computer Vision , pages =

    Mohammad Rastegari and Vicente Ordonez and Joseph Redmon and Ali Farhadi , title =. European Conference on Computer Vision , pages =. 2016 , doi =

  7. [15]

    International Conference on Learning Representations , year =

    Yukun Ding and Jinglan Liu and Jinjun Xiong and Yiyu Shi , title =. International Conference on Learning Representations , year =

  8. [16]

    Advances in Neural Information Processing Systems , volume =

    Yaniv Blumenfeld and Dar Gilboa and Daniel Soudry , title =. Advances in Neural Information Processing Systems , volume =

  9. [18]

    Journal of Machine Learning Research , volume =

    Hongyu Wang and Shuming Ma and Lingxiao Ma and Lei Wang and Wenhui Wang and Li Dong and Shaohan Huang and Huaijie Wang and Jilong Xue and Ruiping Wang and Yi Wu and Furu Wei , title =. Journal of Machine Learning Research , volume =

  10. [22]

    International Conference on Learning Representations , year =

    Peter O'Connor and Max Welling , title =. International Conference on Learning Representations , year =

  11. [23]

    Annals of Mathematics , volume =

    Ingrid Daubechies and Ronald DeVore , title =. Annals of Mathematics , volume =. 2003 , doi =

  12. [24]

    C. Sinan G. One-Bit Sigma--Delta Quantization with Exponential Accuracy , journal =. 2003 , doi =

  13. [25]

    IEEE Transactions on Information Theory , volume =

    Felix Krahmer and Rayan Saab and Rachel Ward , title =. IEEE Transactions on Information Theory , volume =. 2012 , doi =

  14. [26]

    Advances in Neural Information Processing Systems , volume =

    Zechun Liu and Changsheng Zhao and Hanxian Huang and Sijia Chen and Jing Zhang and Jiawei Zhao and Scott Roy and Lisa Jin and Yunyang Xiong and Yangyang Shi and Lin Xiao and Yuandong Tian and Bilge Soran and Raghuraman Krishnamoorthi and Tijmen Blankevoort and Vikas Chandra , ...

  15. [28]

    Advances in Neural Information Processing Systems , volume =

    Adrian Bulat and Yassine Ouali and Georgios Tzimiropoulos , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =

  16. [29]

    Advances in Neural Information Processing Systems , volume =

    Yamato Arai and Yuma Ichikawa , title =. Advances in Neural Information Processing Systems , volume =

  17. [30]

    Advances in Neural Information Processing Systems , volume =

    Banseok Lee and Dongkyu Kim and Youngcheon You and Youngmin Kim , title =. Advances in Neural Information Processing Systems , volume =

  18. [31]

    Proceedings of the 43rd International Conference on Machine Learning , year =

    Shigeng Wang and Chao Li and Yangyuxuan Kang and Jiawei Fan and Anbang Yao , title =. Proceedings of the 43rd International Conference on Machine Learning , year =

  19. [32]

    Gomez and Lukasz Kaiser and Illia Polosukhin , title =

    Ashish Vaswani and Noam Shazeer and Niki Parmar and Jakob Uszkoreit and Llion Jones and Aidan N. Gomez and Lukasz Kaiser and Illia Polosukhin , title =. Advances in Neural Information Processing Systems , volume =

  20. [34]

    Advances in Neural Information Processing Systems , volume =

    Biao Zhang and Rico Sennrich , title =. Advances in Neural Information Processing Systems , volume =

  21. [35]

    Advances in Neural Information Processing Systems , volume =

    Subhabrata Dutta and Tanya Gautam and Soumen Chakrabarti and Tanmoy Chakraborty , title =. Advances in Neural Information Processing Systems , volume =

  22. [36]

    Bulletin of the American Mathematical Society , volume =

    Borjan Geshkovski and Cyril Letrouit and Yury Polyanskiy and Philippe Rigollet , title =. Bulletin of the American Mathematical Society , volume =. 2025 , doi =

  23. [37]

    Proceedings of the 38th International Conference on Machine Learning , series =

    Hyunjik Kim and George Papamakarios and Andriy Mnih , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  24. [38]

    Proceedings of the 38th International Conference on Machine Learning , series =

    George Dasoulas and Kevin Scaman and Aladin Virmaux , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  25. [39]

    How Smooth Is Attention? , booktitle =

    Val. How Smooth Is Attention? , booktitle =

  26. [40]

    Proceedings of the 38th International Conference on Machine Learning , series =

    Yihe Dong and Jean-Baptiste Cordonnier and Andreas Loukas , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  27. [41]

    Advances in Neural Information Processing Systems , volume =

    Xinyi Wu and Amir Ajorlou and Yifei Wang and Stefanie Jegelka and Ali Jadbabaie , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =

  28. [42]

    Consensus Is All You Get: The Role of Attention in Transformers , booktitle =

  29. [43]

    Mahoney and Kurt Keutzer , title =

    Sehoon Kim and Amir Gholami and Zhewei Yao and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  30. [44]

    Proceedings of the 40th International Conference on Machine Learning , series =

    Guangxuan Xiao and Ji Lin and Mickael Seznec and Hao Wu and Julien Demouth and Song Han , title =. Proceedings of the 40th International Conference on Machine Learning , series =

  31. [45]

    International Conference on Learning Representations , year =

    Elias Frantar and Saleh Ashkboos and Torsten Hoefler and Dan Alistarh , title =. International Conference on Learning Representations , year =

  32. [46]

    Proceedings of Machine Learning and Systems , volume =

    Ji Lin and Jiaming Tang and Haotian Tang and Shang Yang and Wei-Ming Chen and Wei-Chen Wang and Guangxuan Xiao and Xingyu Dang and Chuang Gan and Song Han , title =. Proceedings of Machine Learning and Systems , volume =

  33. [47]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Albert Tseng and Jerry Chee and Qingyao Sun and Volodymyr Kuleshov and Christopher De Sa , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  34. [48]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Wei Huang and Yangdong Liu and Haotong Qin and Ying Li and Shiming Zhang and Xianglong Liu and Michele Magno and Xiaojuan Qi , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  35. [49]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Shiyao Li and Xuefei Ning and Luning Wang and Tengxuan Liu and Xiangsheng Shi and Shengen Yan and Guohao Dai and Huazhong Yang and Yu Wang , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  36. [50]

    Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =

    Zifei Xu and Alexander Y. Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =. Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop , series =

  37. [51]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Harshavardhan Adepu and Zhanpeng Zeng and Li Zhang and Vikas Singh , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  38. [52]

    International Conference on Learning Representations , year =

    Noam Shazeer and Azalia Mirhoseini and Krzysztof Maziarz and Andy Davis and Quoc Le and Geoffrey Hinton and Jeff Dean , title =. International Conference on Learning Representations , year =

  39. [53]

    Journal of Machine Learning Research , volume =

    William Fedus and Barret Zoph and Noam Shazeer , title =. Journal of Machine Learning Research , volume =

  40. [54]

    Zhao and Andrew M

    Yanqi Zhou and Tao Lei and Hanxiao Liu and Nan Du and Yanping Huang and Vincent Y. Zhao and Andrew M. Dai and Zhifeng Chen and Quoc V. Le and James Laudon , title =. Advances in Neural Information Processing Systems , volume =

  41. [55]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

    Damai Dai and Li Dong and Shuming Ma and Bo Zheng and Zhifang Sui and Baobao Chang and Furu Wei , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2022 , doi =

  42. [56]

    International Conference on Learning Representations , year =

    Joan Puigcerver and Carlos Riquelme Ruiz and Basil Mustafa and Neil Houlsby , title =. International Conference on Learning Representations , year =

  43. [57]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Chaodong Xiao and Zhengqiang Zhang and Lei Zhang , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  44. [58]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Hyunha Hwang and Xuan Truong Nguyen and Hyuk-Jae Lee , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  45. [59]

    Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =

    Jiahe Qian and Peisong Wang and Zhengyang Zhuge and Qinghao Hu and Jian Cheng , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =. 2026 , doi =

  46. [60]

    Mathematical Programming , volume =

    Sebastian Sager and Gerhard Reinelt and Hans Georg Bock , title =. Mathematical Programming , volume =. 2009 , doi =

  47. [61]

    Mathematical Programming , volume =

    Sebastian Sager and Hans Georg Bock and Moritz Diehl , title =. Mathematical Programming , volume =. 2012 , doi =

  48. [62]

    Mathematical Methods of Operations Research , volume =

    Sebastian Sager and Michael Jung and Christian Kirches , title =. Mathematical Methods of Operations Research , volume =. 2011 , doi =

  49. [63]

    SIAM Journal on Control and Optimization , volume =

    Christian Kirches and Felix Lenders and Paul Manns , title =. SIAM Journal on Control and Optimization , volume =. 2020 , doi =

  50. [64]

    Shankar Sastry , title =

    Ramanarayan Vasudevan and Humberto Gonzalez and Ruzena Bajcsy and S. Shankar Sastry , title =. SIAM Journal on Control and Optimization , volume =. 2013 , doi =

  51. [66]

    Bowman , title =

    Alex Wang and Amanpreet Singh and Julian Michael and Felix Hill and Omer Levy and Samuel R. Bowman , title =. International Conference on Learning Representations , year =

  52. [67]

    Manning and Andrew Y

    Richard Socher and Alex Perelygin and Jean Wu and Jason Chuang and Christopher D. Manning and Andrew Y. Ng and Christopher Potts , title =. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , pages =. 2013 , doi =

  53. [68]

    Transformers: State-of-the-Art Natural Language Processing , booktitle =

    Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and R. Transformers: State-of-the-Art Natural Language Processing , booktitle =. 2020 , doi =

  54. [69]

    distilbert-base-uncased-finetuned-sst-2-english , year =

  55. [70]

    Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =

    Samuel Teuber and Marko Kleine B. Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =. 2021 , doi =

  56. [71]

    Mahoney and Kurt Keutzer , title =

    Zhewei Yao and Zhen Dong and Zhangcheng Zheng and Amir Gholami and Jiali Yu and Eric Tan and Leyuan Wang and Qijing Huang and Yida Wang and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  57. [73]

    Automatica , volume =

    Mathieu Claeys and Jamal Daafouz and Didier Henrion , title =. Automatica , volume =. 2016 , doi =

  58. [74]

    Burden and S

    Samuel A. Burden and S. Shankar Sastry and Daniel E. Koditschek and Shai Revzen , title =. SIAM Journal on Applied Dynamical Systems , volume =. 2016 , doi =

  59. [75]

    Kong and J

    Nathan J. Kong and J. Joe Payne and James Zhu and Aaron M. Johnson , title =. Proceedings of the IEEE , volume =. 2024 , doi =

  60. [76]

    Iris A. M. Huijben and Matthijs Douze and Matthew J. Muckley and Ruud J. G. van Sloun and Jakob Verbeek , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  61. [80]

    Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr

    Jean B. Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr. Nonlinear Optimal Control via Occupation Measures and. SIAM Journal on Control and Optimization , volume =. 2008 , doi =

  62. [81]

    Transactions on Machine Learning Research , year =

    Ian Colbert and Giuseppe Franco and Fabian Grob and Jinjie Zhang and Rayan Saab , title =. Transactions on Machine Learning Research , year =

  63. [85]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Pei Huang and Haoze Wu and Yuting Yang and Ieva Daukantas and Min Wu and Yedi Zhang and Clark Barrett , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , doi =

  64. [86]

    Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =

    Luigi Ambrosio and Nicola Gigli and Giuseppe Savar. Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =. 2008 , doi =

  65. [87]

    International Conference on Learning Representations , year =

    Shihao Zhang and Haoyu Zhang and Ian Colbert and Rayan Saab , title =. International Conference on Learning Representations , year =

  66. [89]

    Burr and Liu Liu and Meng Wang , title =

    Mohammed Nowaz Rabbani Chowdhury and Kaoutar El Maghraoui and Hsinyu Tsai and Naigang Wang and Geoffrey W. Burr and Liu Liu and Meng Wang , title =. International Conference on Learning Representations , year =

  67. [90]

    SIAM Journal on Mathematics of Data Science , volume =

    Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , volume =. 2023 , doi =

  68. [91]

    Proceedings of the 37th International Conference on Machine Learning , series =

    Markus Nagel and Rana Ali Amjad and Mart van Baalen and Christos Louizos and Tijmen Blankevoort , title =. Proceedings of the 37th International Conference on Machine Learning , series =

  69. [96]

    SIAM Journal on Mathematics of Data Science , year =

    Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , year =

  70. [97]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

    Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

  71. [98]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig and Yaman Umuroglu , title =. Proceedings of the 41st International Conference on Machine Learning , series =. 2024 , publisher =

  72. [101]

    The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =

    Moritz Br. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =. 2026 , note =

  73. [102]

    arXiv preprint arXiv:2605.25880 , year =

    Yiping Ji and Mahalakshmi Sabanayagam and Peyman Moghadam and Hemanth Saratchandran and Simon Lucey , title =. arXiv preprint arXiv:2605.25880 , year =

  74. [104]

    arXiv preprint arXiv:2606.12487 , year =

    Zimo Zhao and Maolin Wang and Bowen Yu and Bowen Liu and Xiao Han and Xiangyu Zhao , title =. arXiv preprint arXiv:2606.12487 , year =

  75. [105]

    arXiv preprint arXiv:2601.22101 , year =

    Mahdi Nikdan and Amir Zandieh and Dan Alistarh and Vahab Mirrokni , title =. arXiv preprint arXiv:2601.22101 , year =

  76. [106]

    arXiv preprint arXiv:2603.14818 , year =

    Jingyang Li and Fu Song and Guoqiang Li , title =. arXiv preprint arXiv:2603.14818 , year =

  77. [107]

    Consensus is all you get: The role of attention in transformers

    \'A lvaro Rodr \' guez Abella, Jo \ a o Pedro Silvestre, and Paulo Tabuada. Consensus is all you get: The role of attention in transformers. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 1...

  78. [108]

    Framequant: Flexible low-bit quantization for transformers

    Harshavardhan Adepu, Zhanpeng Zeng, Li Zhang, and Vikas Singh. Framequant: Flexible low-bit quantization for transformers. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 203--227, 2024

  79. [109]

    u rich. Birkh \

    Luigi Ambrosio, Nicola Gigli, and Giuseppe Savar \'e . Gradient Flows in Metric Spaces and in the Space of Probability Measures. Lectures in Mathematics ETH Z \"u rich. Birkh \"a user, Basel, 2nd edition, 2008. doi:10.1007/978-3-7643-8722-8

  80. [110]

    Quantization error propagation: Revisiting layer-wise post-training quantization

    Yamato Arai and Yuma Ichikawa. Quantization error propagation: Revisiting layer-wise post-training quantization. In Advances in Neural Information Processing Systems, volume 38, 2025

  81. [111]

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016

  82. [112]

    A fast iterative shrinkage-thresholding algorithm for linear inverse problems

    Amir Beck and Marc Teboulle. A fast iterative shrinkage-thresholding algorithm for linear inverse problems. SIAM Journal on Imaging Sciences, 2 0 (1): 0 183--202, 2009. doi:10.1137/080716542

  83. [113]

    A mean field theory of quantized deep networks: The quantization--depth trade-off

    Yaniv Blumenfeld, Dar Gilboa, and Daniel Soudry. A mean field theory of quantized deep networks: The quantization--depth trade-off. In Advances in Neural Information Processing Systems, volume 32, pages 7036--7046, 2019

  84. [114]

    The expressive power of low precision softmax transformers with (summarized) chain-of-thought

    Moritz Br \"o samle and Stephan Eckstein. The expressive power of low precision softmax transformers with (summarized) chain-of-thought. arXiv preprint arXiv:2605.18079, 2026. doi:10.48550/arXiv.2605.18079. Accepted to the 43rd International Conference on Machine Learning

  85. [115]

    QBB : Quantization with binary bases for LLMs

    Adrian Bulat, Yassine Ouali, and Georgios Tzimiropoulos. QBB : Quantization with binary bases for LLMs . In Advances in Neural Information Processing Systems, volume 37, pages 3209--3228, 2024. doi:10.52202/079017-0105

  86. [116]

    Burden, S

    Samuel A. Burden, S. Shankar Sastry, Daniel E. Koditschek, and Shai Revzen. Event-selected vector field discontinuities yield piecewise-differentiable flows. SIAM Journal on Applied Dynamical Systems, 15 0 (2): 0 1227--1267, 2016. doi:10.1137/15M1016588

  87. [117]

    How smooth is attention? In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 5817--5840, 2024

    Val \'e rie Castin, Pierre Ablin, and Gabriel Peyr \'e . How smooth is attention? In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 5817--5840, 2024

  88. [118]

    Every bit counts: A theoretical study of precision--expressivity tradeoffs in quantized transformers

    Sayak Chakrabarti, Toniann Pitassi, and Josh Alman. Every bit counts: A theoretical study of precision--expressivity tradeoffs in quantized transformers. arXiv preprint arXiv:2602.02707, 2026

  89. [119]

    Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, volume 31, pages 6571--6583, 2018 a

  90. [120]

    Theoretical linear convergence of unfolded ISTA and its practical weights and thresholds

    Xiaohan Chen, Jialin Liu, Zhangyang Wang, and Wotao Yin. Theoretical linear convergence of unfolded ISTA and its practical weights and thresholds. In Advances in Neural Information Processing Systems, volume 31, pages 9061--9071, 2018 b

  91. [121]

    Burr, Liu Liu, and Meng Wang

    Mohammed Nowaz Rabbani Chowdhury, Kaoutar El Maghraoui, Hsinyu Tsai, Naigang Wang, Geoffrey W. Burr, Liu Liu, and Meng Wang. Efficient quantization of mixture-of-experts with theoretical generalization guarantees. In International Conference on Learning Representations, 2026. ...

  92. [122]

    Modal occupation measures and LMI relaxations for nonlinear switched systems control

    Mathieu Claeys, Jamal Daafouz, and Didier Henrion. Modal occupation measures and LMI relaxations for nonlinear switched systems control. Automatica, 64: 0 143--154, 2016. doi:10.1016/j.automatica.2015.11.003

  93. [123]

    A2Q : Accumulator-aware quantization with guaranteed overflow avoidance

    Ian Colbert, Alessandro Pappalardo, and Jakoba Petri-Koenig. A2Q : Accumulator-aware quantization with guaranteed overflow avoidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16989--16998, 2023

  94. [124]

    A2Q+ : Improving accumulator-aware weight quantization

    Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig, and Yaman Umuroglu. A2Q+ : Improving accumulator-aware weight quantization. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 9275--929...

  95. [125]

    Accumulator-aware post-training quantization for large language models

    Ian Colbert, Giuseppe Franco, Fabian Grob, Jinjie Zhang, and Rayan Saab. Accumulator-aware post-training quantization for large language models. Transactions on Machine Learning Research, 2025. Originally circulated as arXiv:2409.17092

  96. [126]

    Binaryconnect: Training deep neural networks with binary weights during propagations

    Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in Neural Information Processing Systems, volume 28, pages 3123--3131, 2015

  97. [127]

    StableMoE : Stable routing strategy for mixture of experts

    Damai Dai, Li Dong, Shuming Ma, Bo Zheng, Zhifang Sui, Baobao Chang, and Furu Wei. StableMoE : Stable routing strategy for mixture of experts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7085--7095, ...

  98. [128]

    Lipschitz normalization for self-attention layers with application to graph neural networks

    George Dasoulas, Kevin Scaman, and Aladin Virmaux. Lipschitz normalization for self-attention layers with application to graph neural networks. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, page...

  99. [129]

    Approximating a bandlimited function using very coarsely quantized data: A family of stable sigma--delta modulators of arbitrary order

    Ingrid Daubechies and Ronald DeVore. Approximating a bandlimited function using very coarsely quantized data: A family of stable sigma--delta modulators of arbitrary order. Annals of Mathematics, 158 0 (2): 0 679--710, 2003. doi:10.4007/annals.2003.158.679

  100. [130]

    An iterative thresholding algorithm for linear inverse problems with a sparsity constraint

    Ingrid Daubechies, Michel Defrise, and Christine De Mol. An iterative thresholding algorithm for linear inverse problems with a sparsity constraint. Communications on Pure and Applied Mathematics, 57 0 (11): 0 1413--1457, 2004. doi:10.1002/cpa.20042

  101. [131]

    GEMQ : Global expert-level mixed-precision quantization for MoE LLM s

    Jianing Deng, Song Wang, Dongwei Wang, Zijie Liu, Tianlong Chen, Huanrui Yang, and Jingtong Hu. GEMQ : Global expert-level mixed-precision quantization for MoE LLM s. arXiv preprint arXiv:2605.23078, 2026

  102. [132]

    On the universal approximability and complexity bounds of quantized ReLU neural networks

    Yukun Ding, Jinglan Liu, Jinjun Xiong, and Yiyu Shi. On the universal approximability and complexity bounds of quantized ReLU neural networks. In International Conference on Learning Representations, 2019

  103. [133]

    Attention is not all you need: Pure attention loses rank doubly exponentially with depth

    Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: Pure attention loses rank doubly exponentially with depth. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, p...

  104. [134]

    Redesigning the transformer architecture with insights from multi-particle dynamical systems

    Subhabrata Dutta, Tanya Gautam, Soumen Chakrabarti, and Tanmoy Chakraborty. Redesigning the transformer architecture with insights from multi-particle dynamical systems. In Advances in Neural Information Processing Systems, volume 34, 2021

  105. [135]

    Unlocking efficient large inference models: One-bit unrolling tips the scales

    Arian Eamaz, Farhang Yeganegi, and Mojtaba Soltanalian. Unlocking efficient large inference models: One-bit unrolling tips the scales. arXiv preprint arXiv:2502.01908, 2025

  106. [136]

    LoopQ : Quantization for recursive transformers

    Rui Fang, Hsi-Wen Chen, and Ming-Syan Chen. LoopQ : Quantization for recursive transformers. arXiv preprint arXiv:2605.16343, 2026

  107. [137]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

    William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23 0 (120): 0 1--39, 2022

  108. [138]

    Improving quantization with post-training model expansion

    Giuseppe Franco, Pablo Monteagudo-Lago, Ian Colbert, Nicholas Fraser, and Michaela Blott. Improving quantization with post-training model expansion. arXiv preprint arXiv:2503.17513, 2025

  109. [139]

    GPTQ : Accurate post-training quantization for generative pre-trained transformers

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ : Accurate post-training quantization for generative pre-trained transformers. In International Conference on Learning Representations, 2023

  110. [140]

    A mathematical perspective on transformers

    Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. A mathematical perspective on transformers. Bulletin of the American Mathematical Society, 62 0 (3): 0 427--479, 2025. doi:10.1090/bull/1863

  111. [141]

    Learning fast approximations of sparse coding

    Karol Gregor and Yann LeCun. Learning fast approximations of sparse coding. In Proceedings of the 27th International Conference on Machine Learning, pages 399--406, 2010

  112. [142]

    Sinan G \"u nt \"u rk

    C. Sinan G \"u nt \"u rk. One-bit sigma--delta quantization with exponential accuracy. Communications on Pure and Applied Mathematics, 56 0 (11): 0 1608--1630, 2003. doi:10.1002/cpa.3044

  113. [143]

    Stable architectures for deep neural networks

    Eldad Haber and Lars Ruthotto. Stable architectures for deep neural networks. Inverse Problems, 34 0 (1): 0 014004, 2018. doi:10.1088/1361-6420/aa9a90

  114. [144]

    Lyapunov-guided training for hardware-safe neural networks under fixed-point arithmetic

    Anis Hamadouche and Amir Hussain. Lyapunov-guided training for hardware-safe neural networks under fixed-point arithmetic. arXiv preprint arXiv:2607.04531, 2026

  115. [145]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770--778, 2016. doi:10.1109/CVPR.2016.90

  116. [146]

    I-LLM : Efficient integer-only inference for fully-quantized low-bit large language models

    Xing Hu, Yuan Cheng, Dawei Yang, Zhihang Yuan, Jiangyong Yu, Chen Xu, and Sifan Zhou. I-LLM : Efficient integer-only inference for fully-quantized low-bit large language models. arXiv preprint arXiv:2405.17849, 2024

  117. [147]

    Towards efficient verification of quantized neural networks

    Pei Huang, Haoze Wu, Yuting Yang, Ieva Daukantas, Min Wu, Yedi Zhang, and Clark Barrett. Towards efficient verification of quantized neural networks. Proceedings of the AAAI Conference on Artificial Intelligence, 38 0 (19): 0 21152--21160, 2024 a . doi:10.1609/aaai.v38i19.30108

  118. [148]

    BiLLM : Pushing the limit of post-training quantization for LLMs

    Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, and Xiaojuan Qi. BiLLM : Pushing the limit of post-training quantization for LLMs . In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of...

  119. [149]

    Quantized neural networks: Training neural networks with low precision weights and activations

    Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Quantized neural networks: Training neural networks with low precision weights and activations. Journal of Machine Learning Research, 18 0 (187): 0 1--30, 2018

  120. [150]

    distilbert-base-uncased-finetuned-sst-2-english

    Hugging Face . distilbert-base-uncased-finetuned-sst-2-english. Hugging Face model repository, 2020. Model ID: distilbert/distilbert-base-uncased-finetuned-sst-2-english; commit 714eb0fa89d2f80546fda750413ed43d93601a13; DOI: 10.57967/hf/0181; accessed July 23, 2026

  121. [151]

    Iris A. M. Huijben, Matthijs Douze, Matthew J. Muckley, Ruud J. G. van Sloun, and Jakob Verbeek. Residual quantization with implicit neural codebooks. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Researc...

  122. [152]

    On expressive power of quantized neural networks under fixed-point arithmetic

    Geonho Hwang, Yeachan Park, and Sejun Park. On expressive power of quantized neural networks under fixed-point arithmetic. arXiv preprint arXiv:2409.00297, 2024. Revised 2026

  123. [153]

    LS-ViT : Least-squares hessian based block reconstruction for low-bit post-training quantization of vision transformers

    Hyunha Hwang, Xuan Truong Nguyen, and Hyuk-Jae Lee. LS-ViT : Least-squares hessian based block reconstruction for low-bit post-training quantization of vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 33588--33597, 2026

  124. [154]

    The quantization benefits of residual-free transformers

    Yiping Ji, Mahalakshmi Sabanayagam, Peyman Moghadam, Hemanth Saratchandran, and Simon Lucey. The quantization benefits of residual-free transformers. arXiv preprint arXiv:2605.25880, 2026. doi:10.48550/arXiv.2605.25880

  125. [155]

    The lipschitz constant of self-attention

    Hyunjik Kim, George Papamakarios, and Andriy Mnih. The lipschitz constant of self-attention. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5562--5571, 2021 a

  126. [156]

    Mahoney, and Kurt Keutzer

    Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. I-BERT : Integer-only BERT quantization. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5506--5518, 2021 b

  127. [157]

    Approximation properties and tight bounds for constrained mixed-integer optimal control

    Christian Kirches, Felix Lenders, and Paul Manns. Approximation properties and tight bounds for constrained mixed-integer optimal control. SIAM Journal on Control and Optimization, 58 0 (3): 0 1371--1402, 2020. doi:10.1137/18M1182917

  128. [158]

    Nathan J. Kong, J. Joe Payne, James Zhu, and Aaron M. Johnson. Saltation matrices: The essential tool for linearizing hybrid dynamical systems. Proceedings of the IEEE, 112 0 (6): 0 585--608, 2024. doi:10.1109/JPROC.2024.3440211

  129. [159]

    Root-exponential accuracy for coarse quantization of finite frame expansions

    Felix Krahmer, Rayan Saab, and Rachel Ward. Root-exponential accuracy for coarse quantization of finite frame expansions. IEEE Transactions on Information Theory, 58 0 (2): 0 1069--1079, 2012. doi:10.1109/TIT.2011.2168942

  130. [160]

    Lasserre, Didier Henrion, Christophe Prieur, and Emmanuel Tr \'e lat

    Jean B. Lasserre, Didier Henrion, Christophe Prieur, and Emmanuel Tr \'e lat. Nonlinear optimal control via occupation measures and LMI -relaxations. SIAM Journal on Control and Optimization, 47 0 (4): 0 1643--1666, 2008. doi:10.1137/070685051

  131. [161]

    Littlebit: Ultra low-bit quantization via latent factorization

    Banseok Lee, Dongkyu Kim, Youngcheon You, and Youngmin Kim. Littlebit: Ultra low-bit quantization via latent factorization. In Advances in Neural Information Processing Systems, volume 38, 2025

  132. [162]

    SimCert : Probabilistic certification for behavioral similarity in deep neural network compression

    Jingyang Li, Fu Song, and Guoqiang Li. SimCert : Probabilistic certification for behavioral similarity in deep neural network compression. arXiv preprint arXiv:2603.14818, 2026. doi:10.48550/arXiv.2603.14818

  133. [163]

    Evaluating quantized large language models

    Shiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu, Xiangsheng Shi, Shengen Yan, Guohao Dai, Huazhong Yang, and Yu Wang. Evaluating quantized large language models. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Lea...

  134. [164]

    AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration. In Proceedings of Machine Learning and Systems, volum...

  135. [165]

    Paretoq: Improving scaling laws in extremely low-bit LLM quantization

    Zechun Liu, Changsheng Zhao, Hanxian Huang, Sijia Chen, Jing Zhang, Jiawei Zhao, Scott Roy, Lisa Jin, Yunyang Xiong, Yangyang Shi, Lin Xiao, Yuandong Tian, Bilge Soran, Raghuraman Krishnamoorthi, Tijmen Blankevoort, and Vikas Chandra. Paretoq: Improving scaling laws in extreme...

  136. [166]

    The era of 1-bit LLMs : All large language models are in 1.58 bits

    Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei. The era of 1-bit LLMs : All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764, 2024

  137. [167]

    Bitnet b1.58 2b4t technical report

    Shuming Ma, Hongyu Wang, Shaohan Huang, Xingxing Zhang, Ying Hu, Ting Song, Yan Xia, and Furu Wei. Bitnet b1.58 2b4t technical report. arXiv preprint arXiv:2504.12285, 2025

  138. [168]

    Vishal Monga, Yuelong Li, and Yonina C. Eldar. Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing. IEEE Signal Processing Magazine, 38 0 (2): 0 18--44, 2021. doi:10.1109/MSP.2020.3016905

  139. [169]

    Up or down? adaptive rounding for post-training quantization

    Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or down? adaptive rounding for post-training quantization. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Researc...

  140. [170]

    Vikas Natesh, H. T. Kung, and David Kong. MGS : Markov greedy sums for accurate low-bitwidth floating-point accumulation. arXiv preprint arXiv:2504.09072, 2025

  141. [171]

    ECO : Quantized training without full-precision master weights

    Mahdi Nikdan, Amir Zandieh, Dan Alistarh, and Vahab Mirrokni. ECO : Quantized training without full-precision master weights. arXiv preprint arXiv:2601.22101, 2026. doi:10.48550/arXiv.2601.22101

  142. [172]

    Sigma delta quantized networks

    Peter O'Connor and Max Welling. Sigma delta quantized networks. In International Conference on Learning Representations, 2017

  143. [173]

    Three quantization regimes for ReLU networks

    Weigutian Ou, Philipp Schenkel, and Helmut B \"o lcskei. Three quantization regimes for ReLU networks. arXiv preprint arXiv:2405.01952, 2024

  144. [174]

    Value-and-structure alignment for routing-consistent quantization of mixture-of-experts models

    Hancheol Park, Geonho Lee, Tairen Piao, and Tae-Ho Kim. Value-and-structure alignment for routing-consistent quantization of mixture-of-experts models. arXiv preprint arXiv:2606.05688, 2026 a

  145. [175]

    Expressive power of floating-point neural networks with arbitrary reduction orders and inexact activation implementations

    Yeachan Park, Geonho Hwang, Wonyeol Lee, and Sejun Park. Expressive power of floating-point neural networks with arbitrary reduction orders and inexact activation implementations. arXiv preprint arXiv:2605.28704, 2026 b

  146. [176]

    From sparse to soft mixtures of experts

    Joan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, and Neil Houlsby. From sparse to soft mixtures of experts. In International Conference on Learning Representations, 2024

  147. [177]

    A universal self-attention enhancement for bridging low-bit quantization and vision transformers

    Jiahe Qian, Peisong Wang, Zhengyang Zhuge, Qinghao Hu, and Jian Cheng. A universal self-attention enhancement for bridging low-bit quantization and vision transformers. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 360--370, 2026. d...

  148. [178]

    XNOR-Net : Imagenet classification using binary convolutional neural networks

    Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. XNOR-Net : Imagenet classification using binary convolutional neural networks. In European Conference on Computer Vision, pages 525--542, 2016. doi:10.1007/978-3-319-46493-0_32

  149. [179]

    Sensitivity analysis of hybrid systems with state jumps with application to trajectory tracking

    Alessandro Saccon, Nathan van de Wouw, and Henk Nijmeijer. Sensitivity analysis of hybrid systems with state jumps with application to trajectory tracking. In Proceedings of the 53rd IEEE Conference on Decision and Control, pages 3065--3070. IEEE, 2014. doi:10.1109/CDC.2014.70...

  150. [180]

    Direct methods with maximal lower bound for mixed-integer optimal control problems

    Sebastian Sager, Gerhard Reinelt, and Hans Georg Bock. Direct methods with maximal lower bound for mixed-integer optimal control problems. Mathematical Programming, 118 0 (1): 0 109--149, 2009. doi:10.1007/s10107-007-0185-6

  151. [181]

    Combinatorial integral approximation

    Sebastian Sager, Michael Jung, and Christian Kirches. Combinatorial integral approximation. Mathematical Methods of Operations Research, 73 0 (3): 0 363--380, 2011. doi:10.1007/s00186-011-0355-4

  152. [182]

    The integer approximation error in mixed-integer optimal control

    Sebastian Sager, Hans Georg Bock, and Moritz Diehl. The integer approximation error in mixed-integer optimal control. Mathematical Programming, 133 0 (1--2): 0 1--23, 2012. doi:10.1007/s10107-010-0405-3

  153. [183]

    Sander, Pierre Ablin, and Gabriel Peyr \'e

    Michael E. Sander, Pierre Ablin, and Gabriel Peyr \'e . Do residual neural networks discretize neural ordinary differential equations? In Advances in Neural Information Processing Systems, volume 35, pages 36520--36532, 2022. doi:10.52202/068431-2646

  154. [184]

    DistilBERT , a distilled version of BERT : Smaller, faster, cheaper and lighter

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT , a distilled version of BERT : Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019

  155. [185]

    Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

    Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations, 2017

  156. [186]

    Manning, Andrew Y

    Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Pro...

  157. [187]

    Geometric path enumeration for equivalence verification of neural networks

    Samuel Teuber, Marko Kleine B \"u ning, Philipp Kern, and Carsten Sinz. Geometric path enumeration for equivalence verification of neural networks. In 2021 IEEE 33rd International Conference on Tools with Artificial Intelligence (ICTAI), pages 200--208. IEEE, 2021. doi:10.1109...

  158. [188]

    QuIP\# : Even better LLM quantization with hadamard incoherence and lattice codebooks

    Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa. QuIP\# : Even better LLM quantization with hadamard incoherence and lattice codebooks. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machin...

  159. [189]

    Shankar Sastry

    Ramanarayan Vasudevan, Humberto Gonzalez, Ruzena Bajcsy, and S. Shankar Sastry. Consistent approximations for the optimal control of constrained switched systems---part 1: A conceptual algorithm. SIAM Journal on Control and Optimization, 51 0 (6): 0 4463--4483, 2013 a . doi:10...

  160. [190]

    Shankar Sastry

    Ramanarayan Vasudevan, Humberto Gonzalez, Ruzena Bajcsy, and S. Shankar Sastry. Consistent approximations for the optimal control of constrained switched systems---part 2: An implementable algorithm. SIAM Journal on Control and Optimization, 51 0 (6): 0 4484--4503, 2013 b . do...

  161. [191]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017

  162. [192]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. GLUE : A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2019

  163. [193]

    Bitnet: 1-bit pre-training for large language models

    Hongyu Wang, Shuming Ma, Lingxiao Ma, Lei Wang, Wenhui Wang, Li Dong, Shaohan Huang, Huaijie Wang, Jilong Xue, Ruiping Wang, Yi Wu, and Furu Wei. Bitnet: 1-bit pre-training for large language models. Journal of Machine Learning Research, 26 0 (125): 0 1--29, 2025

  164. [194]

    CAT-Q : Cost-efficient and accurate ternary quantization for LLMs

    Shigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan, and Anbang Yao. CAT-Q : Cost-efficient and accurate ternary quantization for LLMs . In Proceedings of the 43rd International Conference on Machine Learning, 2026. Oral presentation; arXiv:2606.26650

  165. [195]

    Optimal Control of Differential and Functional Equations

    Jack Warga. Optimal Control of Differential and Functional Equations. Academic Press, New York, 1972

  166. [196]

    Featurized occupation measures for structured global search in numerical optimal control

    Qi Wei, Jianfeng Tao, Haoyang Tan, and Hongyu Nie. Featurized occupation measures for structured global search in numerical optimal control. arXiv preprint arXiv:2603.16231, 2026

  167. [197]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R \'e mi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, ...

  168. [198]

    On the role of attention masks and layernorm in transformers

    Xinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka, and Ali Jadbabaie. On the role of attention masks and layernorm in transformers. In Advances in Neural Information Processing Systems, volume 37, pages 14774--14809, 2024. doi:10.52202/079017-0472

  169. [199]

    SDQ-LLM : Sigma--delta quantization for 1-bit LLMs of any size

    Junhao Xia, Ming Zhao, Limin Xiao, and Xiujun Zhang. SDQ-LLM : Sigma--delta quantization for 1-bit LLMs of any size. arXiv preprint arXiv:2510.03275, 2025

  170. [200]

    Binaryattention: One-bit QK -attention for vision and diffusion transformers

    Chaodong Xiao, Zhengqiang Zhang, and Lei Zhang. Binaryattention: One-bit QK -attention for vision and diffusion transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12106--12117, 2026

  171. [201]

    Smoothquant: Accurate and efficient post-training quantization for large language models

    Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine...

  172. [202]

    Lan, Wanzin Yazar, Tristan Webb, Sayeh Sharify, and Xin Wang

    Zifei Xu, Alexander Y. Lan, Wanzin Yazar, Tristan Webb, Sayeh Sharify, and Xin Wang. Scaling laws for post-training quantized large language models. In Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop, volume 262 of Proceedings of Machin...

  173. [203]

    Mahoney, T

    Wanqi Yang, Yuexiao Ma, Alexander Conzelmann, Xiawu Zheng, Michael W. Mahoney, T. Konstantin Rusch, and Shiwei Liu. AlphaQ : Calibration-free bit allocation for mixture-of-experts quantization. arXiv preprint arXiv:2606.04980, 2026

  174. [204]

    Mahoney, and Kurt Keutzer

    Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W. Mahoney, and Kurt Keutzer. HAWQ-V3 : Dyadic neural network quantization. In Proceedings of the 38th International Conference on Machine Learning, volume ...

  175. [205]

    Laurence C. Young. Lectures on the Calculus of Variations and Optimal Control Theory. W. B. Saunders, Philadelphia, 1969

  176. [206]

    Root mean square layer normalization

    Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Advances in Neural Information Processing Systems, volume 32, 2019

  177. [207]

    Provable post-training quantization: Theoretical analysis of OPTQ and Qronos

    Haoyu Zhang, Shihao Zhang, Ian Colbert, and Rayan Saab. Provable post-training quantization: Theoretical analysis of OPTQ and Qronos . arXiv preprint arXiv:2508.04853, 2025

  178. [208]

    Post-training quantization for neural networks with provable guarantees

    Jinjie Zhang, Yixuan Zhou, and Rayan Saab. Post-training quantization for neural networks with provable guarantees. SIAM Journal on Mathematics of Data Science, 5 0 (2): 0 373--399, 2023. doi:10.1137/22M1511709

  179. [209]

    Corrigendum: Post-training quantization for neural networks with provable guarantees

    Jinjie Zhang, Yixuan Zhou, and Rayan Saab. Corrigendum: Post-training quantization for neural networks with provable guarantees. SIAM Journal on Mathematics of Data Science, 6 0 (3): 0 842--846, 2024. doi:10.1137/24M1635582

  180. [210]

    Qronos : Correcting the past by shaping the future

    Shihao Zhang, Haoyu Zhang, Ian Colbert, and Rayan Saab. Qronos : Correcting the past by shaping the future... in post-training quantization. In International Conference on Learning Representations, 2026. arXiv:2505.11695

  181. [211]

    DynamicPTQ : Mitigating activation quantization collapse via residual-stream dynamics

    Zimo Zhao, Maolin Wang, Bowen Yu, Bowen Liu, Xiao Han, and Xiangyu Zhao. DynamicPTQ : Mitigating activation quantization collapse via residual-stream dynamics. arXiv preprint arXiv:2606.12487, 2026. doi:10.48550/arXiv.2606.12487

  182. [212]

    RQ-MoE : Residual quantization via mixture of experts for efficient input-dependent vector compression

    Zhengjia Zhong, Shuyan Ke, Zaizhou Lin, Jiaqi Song, Hongyi Lan, and Hui Li. RQ-MoE : Residual quantization via mixture of experts for efficient input-dependent vector compression. arXiv preprint arXiv:2605.14359, 2026. To appear at ICML 2026

  183. [213]

    Zhao, Andrew M

    Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Y. Zhao, Andrew M. Dai, Zhifeng Chen, Quoc V. Le, and James Laudon. Mixture-of-experts with expert choice routing. In Advances in Neural Information Processing Systems, volume 35, pages 7103--7114, 2022

Pith tools

Reviewed July 30, 2026 · model on record in the stance chip above.