REVIEW 1 major objections 7 minor 191 references
When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation
T0 review · 1 major / 7 minor · reviewed 2026-07-30 · grok-4.5
Pith's one-line read Depth can replace missing numerical precision only relative to a declared low-bit library, horizon, execution arithmetic, and routing model—and a structural floor no amount of depth can cross.
desk verdict Conditionally sound resource theory that cleanly separates structural floor, pure-depth synthesis, arithmetic phase, and pre-training certificates; the tube hypothesis is the real bridge to practice, not a hidden proof gap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The structural floor E_Ω,∞(F★): the distance from the target full map to the closed relaxed reachable set of the declared low-bit dictionary family. It separates what the library can express from what finite pure depth, metadata, arithmetic, and routing cost; pure schedules approach it by balanced switching / online simplex rounding, and verified primal–dual bounds turn the floor plus finite-resource radii into feasible / impossible / unresolved decisions.
What would settle it
Find a fixed target and declared low-bit residual library whose structural floor is zero, with coherent first-order high-precision error, yet whose best pure low-bit schedules either stay bounded away from the floor as depth grows or match the high-precision accuracy with depth growing much slower or much faster than linear in the comparator depth—under the paper’s execution and tube assumptions.
Extended reading notes
Core claim
For a fixed target map and a declared low-bit residual library, the exact asymptotic limit of infinite low-bit depth is the distance from the target to the closed relaxed reachable set generated by that library. That distance is a structural floor no pure schedule can cross. Finite pure depth approaches the floor at rate O(1/D) under bounded-variation time dependence (and a Hölder-adjusted rate otherwise), but only when residual increments remain numerically visible; full-state write-back can add a growing penalty and freeze updates, while increment error feedback replaces that growth by a bounded carry. When a coherent high-precision comparator also has first-order error and the floor is ze
Load-bearing premise
Everything ideal and implemented must stay inside one verified tube where the residual fields stay bounded and Lipschitz, so the theory does not cover attention or normalization that blow up, or arithmetic that overflows that tube.
Editorial extensions
If this is right
- Before training, a dual lower bound above the tolerance certifies that no depth or optimizer can hit the target with that library.
- Full-state activation write-back can make deeper low-bit nets worse; preserving residual increments (e.g. error-feedback carry) is required for depth to help.
- Accuracy matching against a coherent first-order high-precision teacher forces low-bit depth on the order of teacher depth when the floor is zero.
- Learned codebooks must be charged as metadata separate from schedule depth; logarithmic metadata bits can keep codebook error commensurate with first-order synthesis.
- Hard routing only keeps a first-order depth law under isolated transversal events and a small-gain route–state loop, not under a frozen positive margin alone.
Reading between the lines
- Quantization toolchains could add a pre-training screen that brackets the structural floor and rejects libraries whose dual bound already exceeds the tolerance.
- Hardware paths that quantize the full residual state each microstep are in a different phase from increment-carry designs; kernel choice may matter as much as nominal bit width.
- The same floor-plus-radius logic could grade looped or unrolled blocks: only refinements that stay coherent with a shared residual horizon earn a depth–precision exchange.
- Unresolved certificates become a research queue of their own—pointing at which bound (floor, arithmetic, or route) must tighten next rather than treating failed training as non-representability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a target-specific resource theory for when low-bit residual depth can replace numerical precision. A depth-D student is modeled as a pure schedule over a declared low-bit residual-field dictionary on a fixed horizon, with the state lifted to the full input-indexed map. The distance from the target to the closed relaxed reachable set is identified as the exact structural floor; pure schedules approach it at O(D^{-1}) under bounded-variation time dependence and O(D^{-ϑ}+D^{-1}) under ϑ-Hölder dependence (Theorems 3, 5). Execution arithmetic is shown to change the phase: full-state write-back contributes a Dρ_z certificate term with an exact scalar freeze result (Proposition 7), while increment error feedback telescopes the carry (Theorem 8) and admits a bit-exact common-lattice realization with explicit register widths (Proposition 9). A fixed, D-independent binary teacher has a closed-form optimal error Θ(D^{-1}) (Theorem 10), lifted to residual-ReLU and nonuniform two-token attention realizations (Proposition 12, Theorem 13), yielding D_match = Θ(L) for coherent first-order comparators (Corollary 15). Learned codebooks add a metadata resource with upper, packing, and allocation laws (Theorems 18, 19, 57); state-dependent routing is treated by a transversal-event small-gain theorem (Theorem 23). A primal–dual stack (HJB, support, affine, occupation-measure/SOS) yields feasible/impossible/unresolved decisions (Corollary 28). Companion software (QReplace) and
Significance. If the results hold, the paper supplies a unifying conditional framework that cleanly separates library geometry, synthesis depth, metadata, execution arithmetic, and routing — resources that the literature often collapses into a nominal bit width. Several strengths deserve explicit credit: (i) an exact closed-form fixed-teacher optimum with a nonasymptotic envelope (Theorem 10), making the first-order depth price sharp for one fixed target rather than only minimax; (ii) a Lean 4 artifact with hash-locked build logs and claim-level axiom audits kernel-checking twelve discrete-core statements, plus exhaustive executable verifiers for the attention converse and common-lattice arithmetic; (iii) a certified nonlinear matrix-valued accuracy-matching depth (D_match = 8 = 2L) proved by rational piecewise-affine bounds, with a prospective falsifiable prediction (calibrated D=9 vs certified 8) that was checked; (iv) an explicit evidence hierarchy and trust-boundary discussion that is more disciplined than typical for this area. The component tools (relaxed controls, sigma–delta feedback, occupation measures, hybrid transversality) are mature, and the paper says so; the contribution is the t
major comments (1)
- [§3.2 Assumption 1; §9.2–9.3; Theorems 6, 8, 29] The common synthesis tube is the load-bearing bridge for the master law (1) and for every QReplace verdict, and it carries a bootstrap risk the manuscript should address more directly. Assumption 1 simultaneously asserts forward invariance of K for all measurable relaxed controls, all mixed/pure Euler states and interpolation segments, and — via Theorems 6, 8, and 29 — all implemented finite-arithmetic prefixes, together with uniform B and L_z on the enlarged tube K_ρ. For attention blocks, the QK-product Lipschitz constant is state-range dependent (§9.2.2 bounds scores through B_Q, B_K), so the constants that define the tube are valid only on a tube whose existence is part of the hypothesis. The paper resolves this constructively for the contractive soft-threshold class (§9.1, Eq. (197) gives an explicit invariant ball), but the Transformer specialization provides only componentwise err
minor comments (7)
- [§8.5, Theorem 27] The main text flags that primal–dual equality requires a 'closed-image qualification detailed in Appendix A; that qualification is not automatic.' Please clarify in the main text that this qualification affects only the no-duality-gap statement, not the validity of dual lower certificates: Theorem 25 is proved directly by monotonicity along trajectories, so Corollary 28's 'certified impossible' verdicts do not depend on strong duality. As written, a reader could over-discount the decision rule or, conversely, over-credit the SOS hierarchy.
- [§10.8] The phrase 'causal validation of the theory's central distinction' is stronger than the design supports. The coherent-target arm (depth-D Euler refinement converging to the depth-32 refinement of the same field) is essentially standard Euler convergence and is expected a priori; the informative arm is the direct-target divergence. The 4-bit study uses three QAT seeds, so the Student-t intervals have two degrees of freedom, and the layer-1 hidden-map interval crosses zero (reported, but only mid-paragraph). Please temper the causal language, state what outcome would have falsified the mechanism, and note that the fitted slopes (−1.16, −0.86) come from five depth points.
- [Appendix A vs. main text] Notation drift: the appendices use C_{fh,b} and C^{unif}_{fh} where the main text uses C^{end}_{syn} and C^{unif}_{syn} (e.g., Theorem 30 vs. Theorem 3; Theorem 53 vs. Theorem 18). Please harmonize or add the correspondence to Table 3. Similarly, Φ_L(T) is defined twice (Eq. (14) and after Eq. (270)), and Eq. (65) uses u = D^{-1} in the main text but x = 1/D in Appendix A.2.
- [§1.4 vs. §8.7] QReplace is described in §1.4 as returning five outcomes (certified feasible, certified impossible, conditionally feasible, diagnostically promising, unresolved), while Corollary 28 defines a three-way decision. Please state the mapping between the two lists and which outcomes are proof-backed versus heuristic.
- [§4.2, Figure 2(a)] The write-back term Dρ_z is a worst-case certificate envelope; the only exact freeze result is scalar (Proposition 7). The text acknowledges this, but the 'phase diagram' framing of Figure 2(a) may be read as asserting realized U-shaped behavior in general architectures. A sentence clarifying that no multidimensional lower bound realizing the Dρ_z growth is known would calibrate Claim 3.
- [Eq. (36)] The α_k√q/2 term assumes a coordinatewise uniform activation grid in a q-dimensional Euclidean state; please state this explicitly, since the surrounding development is in the lifted Banach space Z.
- [References] A substantial fraction of the citations are 2025–2026 arXiv preprints (e.g., Chakrabarti et al. 2026; Park et al. 2026b; Zhao et al. 2026). Please indicate which are peer-reviewed and pin versions, since several novelty-boundary comparisons depend on them.
Circularity Check
No significant circularity: structural floor, synthesis rates, and converses are independently derived, not fitted or self-referential.
full rationale
The paper's load-bearing chain is definitional-plus-theorem, not circular. The structural floor E_Ω,∞(F*) is defined as dist(F*, R_Ω,rel); Theorems 3/5 then prove pure schedules approach that closed set at O(D^{-1}) or O(D^{-ϑ}+D^{-1}), so the distance is the asymptotic floor by Hausdorff convergence rather than by renaming the target. The fixed-teacher exact error (Theorem 10), residual-ReLU and attention embeddings, common-lattice conservation (Proposition 9), and packing/metadata laws are closed-form or constructive arguments under stated assumptions; they do not fit parameters from the quantities they claim to predict. Empirical sections are explicitly tiered as diagnostic and separated from theorem claims. Lean 4 checks and rational certificates are independent verification, not self-citation load-bearing. Assumption 1 (common tube) is a strong applicability hypothesis, not a circular step. No self-definitional loop, fitted-as-prediction, or uniqueness-via-author-citation pattern is present.
Assumptions & free parameters
free parameters (3)
- Library- and tube-dependent constants (B, Lz, Lt/Vt or Ht, ϑ, Csyn, LΩ,
ho z/
ho D, ηG, route
u r,Δr,χ)
- Metadata metric dimension m and entropy prefactor CΩ
- Per-microstep resource charge ℓσ and total budget B
assumptions (8)
- standard math Banach-space Carathéodory existence/uniqueness for bounded, strongly measurable, uniformly state-Lipschitz relaxed fields on a forward-invariant tube (Assumption 1 / Lemma 60).
- standard math Online simplex/prefix discrepancy rounding with bound < J per coordinate (Lemma 2), used to get first-order pure-to-relaxed rates.
- standard math Superposition principle for continuity equations / occupation measures (Ambrosio et al.) to equate modal measure programs with mixtures of relaxed trajectories (Theorem 27).
- domain assumption Common verified tube containing ideal and implemented prefixes; Lipschitz and defect bounds only claimed on that tube (Assumption 1; Theorem 6).
- domain assumption Coherent fixed-horizon residual refinement family shared by high-precision comparator and low-bit student when stating D=Θ(L) matching (Sections 1.1, 5, Corollary 15).
- domain assumption Isolated transversal top-k route events with small-gain χ<1 and isolation radii (Assumption 22 / Theorem 23); excludes simultaneous/grazing/chattering regimes.
- domain assumption Increment quantizer residual radius and non-overflow of carry/state registers on declared ranges for error-feedback and common-lattice exactness (Theorem 8, Proposition 9).
- ad hoc to paper Deterministic global metadata + input-independent pure schedules for packing lower bound (Theorem 19); no uncharged continuous side information.
invented entities (3)
-
Structural floor E_Ω,∞(F*) = dist(F*, closed relaxed reachable set of declared dictionary family)
independent evidence
-
Schedulewise arithmetic radius A_D and write-back vs error-feedback phase diagram
independent evidence
-
QReplace feasible/impossible/unresolved decision rule combining [L,U] floor bracket with finite-resource radius R_{D,s}
independent evidence
Cite this review
Pith. "Pith review of When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation." pith.science (2026). https://pith.science/paper/NOXJIYPF
@misc{pith2026260723390,
author = {Pith},
title = {Pith review of: When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation},
year = {2026},
howpublished = {\url{https://pith.science/paper/NOXJIYPF}},
note = {Machine review of arXiv:2607.23390}
}
abstract
When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system over a fixed horizon as a pure schedule selecting fields from a declared low-bit operation library, and use relaxed controls to characterize its infinite-depth limit. The distance from the target to the closed relaxed reachable set is the exact structural floor: no increase in depth can remove it for that library. Pure schedules approach the relaxed class at rate $O(D^{-1})$ under bounded-variation time dependence and $O(D^{-\vartheta}+D^{-1})$ under Holder dependence of exponent $\vartheta$. Execution arithmetic can reverse this conclusion: full-state write-back introduces a $D\rho_z$ penalty and can freeze residual updates, whereas increment error feedback replaces this growth by a bounded carry term and obeys an exact common-lattice conservation law. A fixed-teacher converse makes this rate sharp: for coherent depth-$L$ first-order high-precision comparators, accuracy matching requires $D=\Theta(L)$. Learned codebooks add a metadata resource, while state-dependent routing introduces hybrid event conditions. Verified primal and dual bounds yield feasible, impossible, or unresolved decisions before training. Companion software implements the workflow, and Lean 4 machine-checks the exact discrete core. Depth replaces precision only relative to a declared library, horizon, execution semantics, and routing model.
Figures
Figures from the paper (23 more)
Reference graph
Works this paper leans on
-
[1]
Communications on Pure and Applied Mathematics , volume =
Ingrid Daubechies and Michel Defrise and Christine De Mol , title =. Communications on Pure and Applied Mathematics , volume =. 2004 , doi =
2004
-
[2]
SIAM Journal on Imaging Sciences , volume =
Amir Beck and Marc Teboulle , title =. SIAM Journal on Imaging Sciences , volume =. 2009 , doi =
2009
-
[3]
Proceedings of the 27th International Conference on Machine Learning , pages =
Karol Gregor and Yann LeCun , title =. Proceedings of the 27th International Conference on Machine Learning , pages =
-
[4]
Advances in Neural Information Processing Systems , volume =
Xiaohan Chen and Jialin Liu and Zhangyang Wang and Wotao Yin , title =. Advances in Neural Information Processing Systems , volume =
-
[5]
Eldar , title =
Vishal Monga and Yuelong Li and Yonina C. Eldar , title =. IEEE Signal Processing Magazine , volume =. 2021 , doi =
2021
-
[6]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =
Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2016 , doi =
2016
-
[7]
Ricky T. Q. Chen and Yulia Rubanova and Jesse Bettencourt and David Duvenaud , title =. Advances in Neural Information Processing Systems , volume =
-
[8]
Inverse Problems , volume =
Eldad Haber and Lars Ruthotto , title =. Inverse Problems , volume =. 2018 , doi =
2018
Show all 191 references
-
[9]
Sander and Pierre Ablin and Gabriel Peyr
Michael E. Sander and Pierre Ablin and Gabriel Peyr. Do Residual Neural Networks Discretize Neural Ordinary Differential Equations? , booktitle =. 2022 , doi =
2022
-
[10]
Jack Warga , title =
-
[11]
Young , title =
Laurence C. Young , title =
-
[12]
Advances in Neural Information Processing Systems , volume =
Matthieu Courbariaux and Yoshua Bengio and Jean-Pierre David , title =. Advances in Neural Information Processing Systems , volume =
-
[13]
Journal of Machine Learning Research , volume =
Itay Hubara and Matthieu Courbariaux and Daniel Soudry and Ran El-Yaniv and Yoshua Bengio , title =. Journal of Machine Learning Research , volume =
-
[14]
European Conference on Computer Vision , pages =
Mohammad Rastegari and Vicente Ordonez and Joseph Redmon and Ali Farhadi , title =. European Conference on Computer Vision , pages =. 2016 , doi =
2016
-
[15]
International Conference on Learning Representations , year =
Yukun Ding and Jinglan Liu and Jinjun Xiong and Yiyu Shi , title =. International Conference on Learning Representations , year =
-
[16]
Advances in Neural Information Processing Systems , volume =
Yaniv Blumenfeld and Dar Gilboa and Daniel Soudry , title =. Advances in Neural Information Processing Systems , volume =
-
[18]
Journal of Machine Learning Research , volume =
Hongyu Wang and Shuming Ma and Lingxiao Ma and Lei Wang and Wenhui Wang and Li Dong and Shaohan Huang and Huaijie Wang and Jilong Xue and Ruiping Wang and Yi Wu and Furu Wei , title =. Journal of Machine Learning Research , volume =
-
[22]
International Conference on Learning Representations , year =
Peter O'Connor and Max Welling , title =. International Conference on Learning Representations , year =
-
[23]
Annals of Mathematics , volume =
Ingrid Daubechies and Ronald DeVore , title =. Annals of Mathematics , volume =. 2003 , doi =
2003
-
[24]
C. Sinan G. One-Bit Sigma--Delta Quantization with Exponential Accuracy , journal =. 2003 , doi =
2003
-
[25]
IEEE Transactions on Information Theory , volume =
Felix Krahmer and Rayan Saab and Rachel Ward , title =. IEEE Transactions on Information Theory , volume =. 2012 , doi =
2012
-
[26]
Advances in Neural Information Processing Systems , volume =
Zechun Liu and Changsheng Zhao and Hanxian Huang and Sijia Chen and Jing Zhang and Jiawei Zhao and Scott Roy and Lisa Jin and Yunyang Xiong and Yangyang Shi and Lin Xiao and Yuandong Tian and Bilge Soran and Raghuraman Krishnamoorthi and Tijmen Blankevoort and Vikas Chandra , ...
-
[28]
Advances in Neural Information Processing Systems , volume =
Adrian Bulat and Yassine Ouali and Georgios Tzimiropoulos , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =
2024
-
[29]
Advances in Neural Information Processing Systems , volume =
Yamato Arai and Yuma Ichikawa , title =. Advances in Neural Information Processing Systems , volume =
-
[30]
Advances in Neural Information Processing Systems , volume =
Banseok Lee and Dongkyu Kim and Youngcheon You and Youngmin Kim , title =. Advances in Neural Information Processing Systems , volume =
-
[31]
Proceedings of the 43rd International Conference on Machine Learning , year =
Shigeng Wang and Chao Li and Yangyuxuan Kang and Jiawei Fan and Anbang Yao , title =. Proceedings of the 43rd International Conference on Machine Learning , year =
-
[32]
Gomez and Lukasz Kaiser and Illia Polosukhin , title =
Ashish Vaswani and Noam Shazeer and Niki Parmar and Jakob Uszkoreit and Llion Jones and Aidan N. Gomez and Lukasz Kaiser and Illia Polosukhin , title =. Advances in Neural Information Processing Systems , volume =
-
[34]
Advances in Neural Information Processing Systems , volume =
Biao Zhang and Rico Sennrich , title =. Advances in Neural Information Processing Systems , volume =
-
[35]
Advances in Neural Information Processing Systems , volume =
Subhabrata Dutta and Tanya Gautam and Soumen Chakrabarti and Tanmoy Chakraborty , title =. Advances in Neural Information Processing Systems , volume =
-
[36]
Bulletin of the American Mathematical Society , volume =
Borjan Geshkovski and Cyril Letrouit and Yury Polyanskiy and Philippe Rigollet , title =. Bulletin of the American Mathematical Society , volume =. 2025 , doi =
2025
-
[37]
Proceedings of the 38th International Conference on Machine Learning , series =
Hyunjik Kim and George Papamakarios and Andriy Mnih , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[38]
Proceedings of the 38th International Conference on Machine Learning , series =
George Dasoulas and Kevin Scaman and Aladin Virmaux , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[39]
How Smooth Is Attention? , booktitle =
Val. How Smooth Is Attention? , booktitle =
-
[40]
Proceedings of the 38th International Conference on Machine Learning , series =
Yihe Dong and Jean-Baptiste Cordonnier and Andreas Loukas , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[41]
Advances in Neural Information Processing Systems , volume =
Xinyi Wu and Amir Ajorlou and Yifei Wang and Stefanie Jegelka and Ali Jadbabaie , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =
2024
-
[42]
Consensus Is All You Get: The Role of Attention in Transformers , booktitle =
-
[43]
Mahoney and Kurt Keutzer , title =
Sehoon Kim and Amir Gholami and Zhewei Yao and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[44]
Proceedings of the 40th International Conference on Machine Learning , series =
Guangxuan Xiao and Ji Lin and Mickael Seznec and Hao Wu and Julien Demouth and Song Han , title =. Proceedings of the 40th International Conference on Machine Learning , series =
-
[45]
International Conference on Learning Representations , year =
Elias Frantar and Saleh Ashkboos and Torsten Hoefler and Dan Alistarh , title =. International Conference on Learning Representations , year =
-
[46]
Proceedings of Machine Learning and Systems , volume =
Ji Lin and Jiaming Tang and Haotian Tang and Shang Yang and Wei-Ming Chen and Wei-Chen Wang and Guangxuan Xiao and Xingyu Dang and Chuang Gan and Song Han , title =. Proceedings of Machine Learning and Systems , volume =
-
[47]
Proceedings of the 41st International Conference on Machine Learning , series =
Albert Tseng and Jerry Chee and Qingyao Sun and Volodymyr Kuleshov and Christopher De Sa , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[48]
Proceedings of the 41st International Conference on Machine Learning , series =
Wei Huang and Yangdong Liu and Haotong Qin and Ying Li and Shiming Zhang and Xianglong Liu and Michele Magno and Xiaojuan Qi , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[49]
Proceedings of the 41st International Conference on Machine Learning , series =
Shiyao Li and Xuefei Ning and Luning Wang and Tengxuan Liu and Xiangsheng Shi and Shengen Yan and Guohao Dai and Huazhong Yang and Yu Wang , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[50]
Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =
Zifei Xu and Alexander Y. Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =. Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop , series =
-
[51]
Proceedings of the 41st International Conference on Machine Learning , series =
Harshavardhan Adepu and Zhanpeng Zeng and Li Zhang and Vikas Singh , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[52]
International Conference on Learning Representations , year =
Noam Shazeer and Azalia Mirhoseini and Krzysztof Maziarz and Andy Davis and Quoc Le and Geoffrey Hinton and Jeff Dean , title =. International Conference on Learning Representations , year =
-
[53]
Journal of Machine Learning Research , volume =
William Fedus and Barret Zoph and Noam Shazeer , title =. Journal of Machine Learning Research , volume =
-
[54]
Zhao and Andrew M
Yanqi Zhou and Tao Lei and Hanxiao Liu and Nan Du and Yanping Huang and Vincent Y. Zhao and Andrew M. Dai and Zhifeng Chen and Quoc V. Le and James Laudon , title =. Advances in Neural Information Processing Systems , volume =
-
[55]
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =
Damai Dai and Li Dong and Shuming Ma and Bo Zheng and Zhifang Sui and Baobao Chang and Furu Wei , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2022 , doi =
2022
-
[56]
International Conference on Learning Representations , year =
Joan Puigcerver and Carlos Riquelme Ruiz and Basil Mustafa and Neil Houlsby , title =. International Conference on Learning Representations , year =
-
[57]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Chaodong Xiao and Zhengqiang Zhang and Lei Zhang , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[58]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Hyunha Hwang and Xuan Truong Nguyen and Hyuk-Jae Lee , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[59]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =
Jiahe Qian and Peisong Wang and Zhengyang Zhuge and Qinghao Hu and Jian Cheng , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =. 2026 , doi =
2026
-
[60]
Mathematical Programming , volume =
Sebastian Sager and Gerhard Reinelt and Hans Georg Bock , title =. Mathematical Programming , volume =. 2009 , doi =
2009
-
[61]
Mathematical Programming , volume =
Sebastian Sager and Hans Georg Bock and Moritz Diehl , title =. Mathematical Programming , volume =. 2012 , doi =
2012
-
[62]
Mathematical Methods of Operations Research , volume =
Sebastian Sager and Michael Jung and Christian Kirches , title =. Mathematical Methods of Operations Research , volume =. 2011 , doi =
2011
-
[63]
SIAM Journal on Control and Optimization , volume =
Christian Kirches and Felix Lenders and Paul Manns , title =. SIAM Journal on Control and Optimization , volume =. 2020 , doi =
2020
-
[64]
Shankar Sastry , title =
Ramanarayan Vasudevan and Humberto Gonzalez and Ruzena Bajcsy and S. Shankar Sastry , title =. SIAM Journal on Control and Optimization , volume =. 2013 , doi =
2013
-
[66]
Bowman , title =
Alex Wang and Amanpreet Singh and Julian Michael and Felix Hill and Omer Levy and Samuel R. Bowman , title =. International Conference on Learning Representations , year =
-
[67]
Manning and Andrew Y
Richard Socher and Alex Perelygin and Jean Wu and Jason Chuang and Christopher D. Manning and Andrew Y. Ng and Christopher Potts , title =. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , pages =. 2013 , doi =
2013
-
[68]
Transformers: State-of-the-Art Natural Language Processing , booktitle =
Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and R. Transformers: State-of-the-Art Natural Language Processing , booktitle =. 2020 , doi =
2020
-
[69]
distilbert-base-uncased-finetuned-sst-2-english , year =
-
[70]
Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =
Samuel Teuber and Marko Kleine B. Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =. 2021 , doi =
2021
-
[71]
Mahoney and Kurt Keutzer , title =
Zhewei Yao and Zhen Dong and Zhangcheng Zheng and Amir Gholami and Jiali Yu and Eric Tan and Leyuan Wang and Qijing Huang and Yida Wang and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[73]
Automatica , volume =
Mathieu Claeys and Jamal Daafouz and Didier Henrion , title =. Automatica , volume =. 2016 , doi =
2016
-
[74]
Burden and S
Samuel A. Burden and S. Shankar Sastry and Daniel E. Koditschek and Shai Revzen , title =. SIAM Journal on Applied Dynamical Systems , volume =. 2016 , doi =
2016
-
[75]
Kong and J
Nathan J. Kong and J. Joe Payne and James Zhu and Aaron M. Johnson , title =. Proceedings of the IEEE , volume =. 2024 , doi =
2024
-
[76]
Iris A. M. Huijben and Matthijs Douze and Matthew J. Muckley and Ruud J. G. van Sloun and Jakob Verbeek , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[80]
Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr
Jean B. Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr. Nonlinear Optimal Control via Occupation Measures and. SIAM Journal on Control and Optimization , volume =. 2008 , doi =
2008
-
[81]
Transactions on Machine Learning Research , year =
Ian Colbert and Giuseppe Franco and Fabian Grob and Jinjie Zhang and Rayan Saab , title =. Transactions on Machine Learning Research , year =
-
[85]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Pei Huang and Haoze Wu and Yuting Yang and Ieva Daukantas and Min Wu and Yedi Zhang and Clark Barrett , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , doi =
2024
-
[86]
Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =
Luigi Ambrosio and Nicola Gigli and Giuseppe Savar. Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =. 2008 , doi =
2008
-
[87]
International Conference on Learning Representations , year =
Shihao Zhang and Haoyu Zhang and Ian Colbert and Rayan Saab , title =. International Conference on Learning Representations , year =
-
[89]
Burr and Liu Liu and Meng Wang , title =
Mohammed Nowaz Rabbani Chowdhury and Kaoutar El Maghraoui and Hsinyu Tsai and Naigang Wang and Geoffrey W. Burr and Liu Liu and Meng Wang , title =. International Conference on Learning Representations , year =
-
[90]
SIAM Journal on Mathematics of Data Science , volume =
Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , volume =. 2023 , doi =
2023
-
[91]
Proceedings of the 37th International Conference on Machine Learning , series =
Markus Nagel and Rana Ali Amjad and Mart van Baalen and Christos Louizos and Tijmen Blankevoort , title =. Proceedings of the 37th International Conference on Machine Learning , series =
-
[96]
SIAM Journal on Mathematics of Data Science , year =
Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , year =
-
[97]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[98]
Proceedings of the 41st International Conference on Machine Learning , series =
Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig and Yaman Umuroglu , title =. Proceedings of the 41st International Conference on Machine Learning , series =. 2024 , publisher =
2024
-
[101]
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =
Moritz Br. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =. 2026 , note =
2026
-
[102]
arXiv preprint arXiv:2605.25880 , year =
Yiping Ji and Mahalakshmi Sabanayagam and Peyman Moghadam and Hemanth Saratchandran and Simon Lucey , title =. arXiv preprint arXiv:2605.25880 , year =
-
[104]
arXiv preprint arXiv:2606.12487 , year =
Zimo Zhao and Maolin Wang and Bowen Yu and Bowen Liu and Xiao Han and Xiangyu Zhao , title =. arXiv preprint arXiv:2606.12487 , year =
-
[105]
arXiv preprint arXiv:2601.22101 , year =
Mahdi Nikdan and Amir Zandieh and Dan Alistarh and Vahab Mirrokni , title =. arXiv preprint arXiv:2601.22101 , year =
-
[106]
arXiv preprint arXiv:2603.14818 , year =
Jingyang Li and Fu Song and Guoqiang Li , title =. arXiv preprint arXiv:2603.14818 , year =
-
[107]
Consensus is all you get: The role of attention in transformers
\'A lvaro Rodr \' guez Abella, Jo \ a o Pedro Silvestre, and Paulo Tabuada. Consensus is all you get: The role of attention in transformers. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 1...
2025
-
[108]
Framequant: Flexible low-bit quantization for transformers
Harshavardhan Adepu, Zhanpeng Zeng, Li Zhang, and Vikas Singh. Framequant: Flexible low-bit quantization for transformers. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 203--227, 2024
2024
-
[109]
u rich. Birkh \
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savar \'e . Gradient Flows in Metric Spaces and in the Space of Probability Measures. Lectures in Mathematics ETH Z \"u rich. Birkh \"a user, Basel, 2nd edition, 2008. doi:10.1007/978-3-7643-8722-8
2008 doi
-
[110]
Quantization error propagation: Revisiting layer-wise post-training quantization
Yamato Arai and Yuma Ichikawa. Quantization error propagation: Revisiting layer-wise post-training quantization. In Advances in Neural Information Processing Systems, volume 38, 2025
2025
-
[111]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016
2016 arXiv
-
[112]
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle. A fast iterative shrinkage-thresholding algorithm for linear inverse problems. SIAM Journal on Imaging Sciences, 2 0 (1): 0 183--202, 2009. doi:10.1137/080716542
2009 doi
-
[113]
A mean field theory of quantized deep networks: The quantization--depth trade-off
Yaniv Blumenfeld, Dar Gilboa, and Daniel Soudry. A mean field theory of quantized deep networks: The quantization--depth trade-off. In Advances in Neural Information Processing Systems, volume 32, pages 7036--7046, 2019
2019
-
[114]
The expressive power of low precision softmax transformers with (summarized) chain-of-thought
Moritz Br \"o samle and Stephan Eckstein. The expressive power of low precision softmax transformers with (summarized) chain-of-thought. arXiv preprint arXiv:2605.18079, 2026. doi:10.48550/arXiv.2605.18079. Accepted to the 43rd International Conference on Machine Learning
-
[115]
QBB : Quantization with binary bases for LLMs
Adrian Bulat, Yassine Ouali, and Georgios Tzimiropoulos. QBB : Quantization with binary bases for LLMs . In Advances in Neural Information Processing Systems, volume 37, pages 3209--3228, 2024. doi:10.52202/079017-0105
2024 doi
-
[116]
Burden, S
Samuel A. Burden, S. Shankar Sastry, Daniel E. Koditschek, and Shai Revzen. Event-selected vector field discontinuities yield piecewise-differentiable flows. SIAM Journal on Applied Dynamical Systems, 15 0 (2): 0 1227--1267, 2016. doi:10.1137/15M1016588
2016 doi
-
[117]
How smooth is attention? In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 5817--5840, 2024
Val \'e rie Castin, Pierre Ablin, and Gabriel Peyr \'e . How smooth is attention? In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 5817--5840, 2024
2024
-
[118]
Every bit counts: A theoretical study of precision--expressivity tradeoffs in quantized transformers
Sayak Chakrabarti, Toniann Pitassi, and Josh Alman. Every bit counts: A theoretical study of precision--expressivity tradeoffs in quantized transformers. arXiv preprint arXiv:2602.02707, 2026
2026
-
[119]
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, volume 31, pages 6571--6583, 2018 a
2018
-
[120]
Theoretical linear convergence of unfolded ISTA and its practical weights and thresholds
Xiaohan Chen, Jialin Liu, Zhangyang Wang, and Wotao Yin. Theoretical linear convergence of unfolded ISTA and its practical weights and thresholds. In Advances in Neural Information Processing Systems, volume 31, pages 9061--9071, 2018 b
2018
-
[121]
Burr, Liu Liu, and Meng Wang
Mohammed Nowaz Rabbani Chowdhury, Kaoutar El Maghraoui, Hsinyu Tsai, Naigang Wang, Geoffrey W. Burr, Liu Liu, and Meng Wang. Efficient quantization of mixture-of-experts with theoretical generalization guarantees. In International Conference on Learning Representations, 2026. ...
2026 arXiv
-
[122]
Modal occupation measures and LMI relaxations for nonlinear switched systems control
Mathieu Claeys, Jamal Daafouz, and Didier Henrion. Modal occupation measures and LMI relaxations for nonlinear switched systems control. Automatica, 64: 0 143--154, 2016. doi:10.1016/j.automatica.2015.11.003
2016 doi
-
[123]
A2Q : Accumulator-aware quantization with guaranteed overflow avoidance
Ian Colbert, Alessandro Pappalardo, and Jakoba Petri-Koenig. A2Q : Accumulator-aware quantization with guaranteed overflow avoidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16989--16998, 2023
2023
-
[124]
A2Q+ : Improving accumulator-aware weight quantization
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig, and Yaman Umuroglu. A2Q+ : Improving accumulator-aware weight quantization. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 9275--929...
2024
-
[125]
Accumulator-aware post-training quantization for large language models
Ian Colbert, Giuseppe Franco, Fabian Grob, Jinjie Zhang, and Rayan Saab. Accumulator-aware post-training quantization for large language models. Transactions on Machine Learning Research, 2025. Originally circulated as arXiv:2409.17092
2025 arXiv
-
[126]
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in Neural Information Processing Systems, volume 28, pages 3123--3131, 2015
2015
-
[127]
StableMoE : Stable routing strategy for mixture of experts
Damai Dai, Li Dong, Shuming Ma, Bo Zheng, Zhifang Sui, Baobao Chang, and Furu Wei. StableMoE : Stable routing strategy for mixture of experts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7085--7095, ...
2022 doi
-
[128]
Lipschitz normalization for self-attention layers with application to graph neural networks
George Dasoulas, Kevin Scaman, and Aladin Virmaux. Lipschitz normalization for self-attention layers with application to graph neural networks. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, page...
2021
-
[129]
Approximating a bandlimited function using very coarsely quantized data: A family of stable sigma--delta modulators of arbitrary order
Ingrid Daubechies and Ronald DeVore. Approximating a bandlimited function using very coarsely quantized data: A family of stable sigma--delta modulators of arbitrary order. Annals of Mathematics, 158 0 (2): 0 679--710, 2003. doi:10.4007/annals.2003.158.679
2003 doi
-
[130]
An iterative thresholding algorithm for linear inverse problems with a sparsity constraint
Ingrid Daubechies, Michel Defrise, and Christine De Mol. An iterative thresholding algorithm for linear inverse problems with a sparsity constraint. Communications on Pure and Applied Mathematics, 57 0 (11): 0 1413--1457, 2004. doi:10.1002/cpa.20042
2004 doi
-
[131]
GEMQ : Global expert-level mixed-precision quantization for MoE LLM s
Jianing Deng, Song Wang, Dongwei Wang, Zijie Liu, Tianlong Chen, Huanrui Yang, and Jingtong Hu. GEMQ : Global expert-level mixed-precision quantization for MoE LLM s. arXiv preprint arXiv:2605.23078, 2026
2026 arXiv
-
[132]
On the universal approximability and complexity bounds of quantized ReLU neural networks
Yukun Ding, Jinglan Liu, Jinjun Xiong, and Yiyu Shi. On the universal approximability and complexity bounds of quantized ReLU neural networks. In International Conference on Learning Representations, 2019
2019
-
[133]
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: Pure attention loses rank doubly exponentially with depth. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, p...
2021
-
[134]
Redesigning the transformer architecture with insights from multi-particle dynamical systems
Subhabrata Dutta, Tanya Gautam, Soumen Chakrabarti, and Tanmoy Chakraborty. Redesigning the transformer architecture with insights from multi-particle dynamical systems. In Advances in Neural Information Processing Systems, volume 34, 2021
2021
-
[135]
Unlocking efficient large inference models: One-bit unrolling tips the scales
Arian Eamaz, Farhang Yeganegi, and Mojtaba Soltanalian. Unlocking efficient large inference models: One-bit unrolling tips the scales. arXiv preprint arXiv:2502.01908, 2025
2025
-
[136]
LoopQ : Quantization for recursive transformers
Rui Fang, Hsi-Wen Chen, and Ming-Syan Chen. LoopQ : Quantization for recursive transformers. arXiv preprint arXiv:2605.16343, 2026
2026 arXiv
-
[137]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23 0 (120): 0 1--39, 2022
2022
-
[138]
Improving quantization with post-training model expansion
Giuseppe Franco, Pablo Monteagudo-Lago, Ian Colbert, Nicholas Fraser, and Michaela Blott. Improving quantization with post-training model expansion. arXiv preprint arXiv:2503.17513, 2025
2025 arXiv
-
[139]
GPTQ : Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ : Accurate post-training quantization for generative pre-trained transformers. In International Conference on Learning Representations, 2023
2023
-
[140]
A mathematical perspective on transformers
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. A mathematical perspective on transformers. Bulletin of the American Mathematical Society, 62 0 (3): 0 427--479, 2025. doi:10.1090/bull/1863
2025 doi
-
[141]
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun. Learning fast approximations of sparse coding. In Proceedings of the 27th International Conference on Machine Learning, pages 399--406, 2010
2010
-
[142]
Sinan G \"u nt \"u rk
C. Sinan G \"u nt \"u rk. One-bit sigma--delta quantization with exponential accuracy. Communications on Pure and Applied Mathematics, 56 0 (11): 0 1608--1630, 2003. doi:10.1002/cpa.3044
2003 doi
-
[143]
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto. Stable architectures for deep neural networks. Inverse Problems, 34 0 (1): 0 014004, 2018. doi:10.1088/1361-6420/aa9a90
2018 doi
-
[144]
Lyapunov-guided training for hardware-safe neural networks under fixed-point arithmetic
Anis Hamadouche and Amir Hussain. Lyapunov-guided training for hardware-safe neural networks under fixed-point arithmetic. arXiv preprint arXiv:2607.04531, 2026
2026 arXiv
-
[145]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770--778, 2016. doi:10.1109/CVPR.2016.90
2016 doi
-
[146]
I-LLM : Efficient integer-only inference for fully-quantized low-bit large language models
Xing Hu, Yuan Cheng, Dawei Yang, Zhihang Yuan, Jiangyong Yu, Chen Xu, and Sifan Zhou. I-LLM : Efficient integer-only inference for fully-quantized low-bit large language models. arXiv preprint arXiv:2405.17849, 2024
2024 arXiv
-
[147]
Towards efficient verification of quantized neural networks
Pei Huang, Haoze Wu, Yuting Yang, Ieva Daukantas, Min Wu, Yedi Zhang, and Clark Barrett. Towards efficient verification of quantized neural networks. Proceedings of the AAAI Conference on Artificial Intelligence, 38 0 (19): 0 21152--21160, 2024 a . doi:10.1609/aaai.v38i19.30108
2024 doi
-
[148]
BiLLM : Pushing the limit of post-training quantization for LLMs
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, and Xiaojuan Qi. BiLLM : Pushing the limit of post-training quantization for LLMs . In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of...
2024
-
[149]
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Quantized neural networks: Training neural networks with low precision weights and activations. Journal of Machine Learning Research, 18 0 (187): 0 1--30, 2018
2018
-
[150]
distilbert-base-uncased-finetuned-sst-2-english
Hugging Face . distilbert-base-uncased-finetuned-sst-2-english. Hugging Face model repository, 2020. Model ID: distilbert/distilbert-base-uncased-finetuned-sst-2-english; commit 714eb0fa89d2f80546fda750413ed43d93601a13; DOI: 10.57967/hf/0181; accessed July 23, 2026
2020 doi
-
[151]
Iris A. M. Huijben, Matthijs Douze, Matthew J. Muckley, Ruud J. G. van Sloun, and Jakob Verbeek. Residual quantization with implicit neural codebooks. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Researc...
2024
-
[152]
On expressive power of quantized neural networks under fixed-point arithmetic
Geonho Hwang, Yeachan Park, and Sejun Park. On expressive power of quantized neural networks under fixed-point arithmetic. arXiv preprint arXiv:2409.00297, 2024. Revised 2026
2024
-
[153]
LS-ViT : Least-squares hessian based block reconstruction for low-bit post-training quantization of vision transformers
Hyunha Hwang, Xuan Truong Nguyen, and Hyuk-Jae Lee. LS-ViT : Least-squares hessian based block reconstruction for low-bit post-training quantization of vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 33588--33597, 2026
2026
- [154]
-
[155]
The lipschitz constant of self-attention
Hyunjik Kim, George Papamakarios, and Andriy Mnih. The lipschitz constant of self-attention. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5562--5571, 2021 a
2021
-
[156]
Mahoney, and Kurt Keutzer
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. I-BERT : Integer-only BERT quantization. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5506--5518, 2021 b
2021
-
[157]
Approximation properties and tight bounds for constrained mixed-integer optimal control
Christian Kirches, Felix Lenders, and Paul Manns. Approximation properties and tight bounds for constrained mixed-integer optimal control. SIAM Journal on Control and Optimization, 58 0 (3): 0 1371--1402, 2020. doi:10.1137/18M1182917
2020 doi
-
[158]
Nathan J. Kong, J. Joe Payne, James Zhu, and Aaron M. Johnson. Saltation matrices: The essential tool for linearizing hybrid dynamical systems. Proceedings of the IEEE, 112 0 (6): 0 585--608, 2024. doi:10.1109/JPROC.2024.3440211
2024
-
[159]
Root-exponential accuracy for coarse quantization of finite frame expansions
Felix Krahmer, Rayan Saab, and Rachel Ward. Root-exponential accuracy for coarse quantization of finite frame expansions. IEEE Transactions on Information Theory, 58 0 (2): 0 1069--1079, 2012. doi:10.1109/TIT.2011.2168942
2012
-
[160]
Lasserre, Didier Henrion, Christophe Prieur, and Emmanuel Tr \'e lat
Jean B. Lasserre, Didier Henrion, Christophe Prieur, and Emmanuel Tr \'e lat. Nonlinear optimal control via occupation measures and LMI -relaxations. SIAM Journal on Control and Optimization, 47 0 (4): 0 1643--1666, 2008. doi:10.1137/070685051
2008 doi
-
[161]
Littlebit: Ultra low-bit quantization via latent factorization
Banseok Lee, Dongkyu Kim, Youngcheon You, and Youngmin Kim. Littlebit: Ultra low-bit quantization via latent factorization. In Advances in Neural Information Processing Systems, volume 38, 2025
2025
-
[162]
SimCert : Probabilistic certification for behavioral similarity in deep neural network compression
Jingyang Li, Fu Song, and Guoqiang Li. SimCert : Probabilistic certification for behavioral similarity in deep neural network compression. arXiv preprint arXiv:2603.14818, 2026. doi:10.48550/arXiv.2603.14818
2026 doi
-
[163]
Evaluating quantized large language models
Shiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu, Xiangsheng Shi, Shengen Yan, Guohao Dai, Huazhong Yang, and Yu Wang. Evaluating quantized large language models. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Lea...
2024
-
[164]
AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration. In Proceedings of Machine Learning and Systems, volum...
2024
-
[165]
Paretoq: Improving scaling laws in extremely low-bit LLM quantization
Zechun Liu, Changsheng Zhao, Hanxian Huang, Sijia Chen, Jing Zhang, Jiawei Zhao, Scott Roy, Lisa Jin, Yunyang Xiong, Yangyang Shi, Lin Xiao, Yuandong Tian, Bilge Soran, Raghuraman Krishnamoorthi, Tijmen Blankevoort, and Vikas Chandra. Paretoq: Improving scaling laws in extreme...
2025
-
[166]
The era of 1-bit LLMs : All large language models are in 1.58 bits
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei. The era of 1-bit LLMs : All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764, 2024
2024 arXiv
-
[167]
Bitnet b1.58 2b4t technical report
Shuming Ma, Hongyu Wang, Shaohan Huang, Xingxing Zhang, Ying Hu, Ting Song, Yan Xia, and Furu Wei. Bitnet b1.58 2b4t technical report. arXiv preprint arXiv:2504.12285, 2025
2025 arXiv
-
[168]
Vishal Monga, Yuelong Li, and Yonina C. Eldar. Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing. IEEE Signal Processing Magazine, 38 0 (2): 0 18--44, 2021. doi:10.1109/MSP.2020.3016905
2021
-
[169]
Up or down? adaptive rounding for post-training quantization
Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or down? adaptive rounding for post-training quantization. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Researc...
2020
-
[170]
Vikas Natesh, H. T. Kung, and David Kong. MGS : Markov greedy sums for accurate low-bitwidth floating-point accumulation. arXiv preprint arXiv:2504.09072, 2025
2025 arXiv
-
[171]
ECO : Quantized training without full-precision master weights
Mahdi Nikdan, Amir Zandieh, Dan Alistarh, and Vahab Mirrokni. ECO : Quantized training without full-precision master weights. arXiv preprint arXiv:2601.22101, 2026. doi:10.48550/arXiv.2601.22101
2026 doi
-
[172]
Sigma delta quantized networks
Peter O'Connor and Max Welling. Sigma delta quantized networks. In International Conference on Learning Representations, 2017
2017
-
[173]
Three quantization regimes for ReLU networks
Weigutian Ou, Philipp Schenkel, and Helmut B \"o lcskei. Three quantization regimes for ReLU networks. arXiv preprint arXiv:2405.01952, 2024
2024 arXiv
-
[174]
Value-and-structure alignment for routing-consistent quantization of mixture-of-experts models
Hancheol Park, Geonho Lee, Tairen Piao, and Tae-Ho Kim. Value-and-structure alignment for routing-consistent quantization of mixture-of-experts models. arXiv preprint arXiv:2606.05688, 2026 a
2026 arXiv
-
[175]
Expressive power of floating-point neural networks with arbitrary reduction orders and inexact activation implementations
Yeachan Park, Geonho Hwang, Wonyeol Lee, and Sejun Park. Expressive power of floating-point neural networks with arbitrary reduction orders and inexact activation implementations. arXiv preprint arXiv:2605.28704, 2026 b
2026 arXiv
-
[176]
From sparse to soft mixtures of experts
Joan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, and Neil Houlsby. From sparse to soft mixtures of experts. In International Conference on Learning Representations, 2024
2024
-
[177]
A universal self-attention enhancement for bridging low-bit quantization and vision transformers
Jiahe Qian, Peisong Wang, Zhengyang Zhuge, Qinghao Hu, and Jian Cheng. A universal self-attention enhancement for bridging low-bit quantization and vision transformers. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 360--370, 2026. d...
2026
-
[178]
XNOR-Net : Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. XNOR-Net : Imagenet classification using binary convolutional neural networks. In European Conference on Computer Vision, pages 525--542, 2016. doi:10.1007/978-3-319-46493-0_32
2016 doi
-
[179]
Sensitivity analysis of hybrid systems with state jumps with application to trajectory tracking
Alessandro Saccon, Nathan van de Wouw, and Henk Nijmeijer. Sensitivity analysis of hybrid systems with state jumps with application to trajectory tracking. In Proceedings of the 53rd IEEE Conference on Decision and Control, pages 3065--3070. IEEE, 2014. doi:10.1109/CDC.2014.70...
2014
-
[180]
Direct methods with maximal lower bound for mixed-integer optimal control problems
Sebastian Sager, Gerhard Reinelt, and Hans Georg Bock. Direct methods with maximal lower bound for mixed-integer optimal control problems. Mathematical Programming, 118 0 (1): 0 109--149, 2009. doi:10.1007/s10107-007-0185-6
2009 doi
-
[181]
Combinatorial integral approximation
Sebastian Sager, Michael Jung, and Christian Kirches. Combinatorial integral approximation. Mathematical Methods of Operations Research, 73 0 (3): 0 363--380, 2011. doi:10.1007/s00186-011-0355-4
2011 doi
-
[182]
The integer approximation error in mixed-integer optimal control
Sebastian Sager, Hans Georg Bock, and Moritz Diehl. The integer approximation error in mixed-integer optimal control. Mathematical Programming, 133 0 (1--2): 0 1--23, 2012. doi:10.1007/s10107-010-0405-3
2012 doi
-
[183]
Sander, Pierre Ablin, and Gabriel Peyr \'e
Michael E. Sander, Pierre Ablin, and Gabriel Peyr \'e . Do residual neural networks discretize neural ordinary differential equations? In Advances in Neural Information Processing Systems, volume 35, pages 36520--36532, 2022. doi:10.52202/068431-2646
2022 doi
-
[184]
DistilBERT , a distilled version of BERT : Smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT , a distilled version of BERT : Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019
1910 arXiv
-
[185]
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations, 2017
2017
-
[186]
Manning, Andrew Y
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Pro...
2013 doi
-
[187]
Geometric path enumeration for equivalence verification of neural networks
Samuel Teuber, Marko Kleine B \"u ning, Philipp Kern, and Carsten Sinz. Geometric path enumeration for equivalence verification of neural networks. In 2021 IEEE 33rd International Conference on Tools with Artificial Intelligence (ICTAI), pages 200--208. IEEE, 2021. doi:10.1109...
2021
-
[188]
QuIP\# : Even better LLM quantization with hadamard incoherence and lattice codebooks
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa. QuIP\# : Even better LLM quantization with hadamard incoherence and lattice codebooks. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machin...
2024
-
[189]
Shankar Sastry
Ramanarayan Vasudevan, Humberto Gonzalez, Ruzena Bajcsy, and S. Shankar Sastry. Consistent approximations for the optimal control of constrained switched systems---part 1: A conceptual algorithm. SIAM Journal on Control and Optimization, 51 0 (6): 0 4463--4483, 2013 a . doi:10...
2013 doi
-
[190]
Shankar Sastry
Ramanarayan Vasudevan, Humberto Gonzalez, Ruzena Bajcsy, and S. Shankar Sastry. Consistent approximations for the optimal control of constrained switched systems---part 2: An implementable algorithm. SIAM Journal on Control and Optimization, 51 0 (6): 0 4484--4503, 2013 b . do...
2013 doi
-
[191]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017
2017
-
[192]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. GLUE : A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2019
2019
-
[193]
Bitnet: 1-bit pre-training for large language models
Hongyu Wang, Shuming Ma, Lingxiao Ma, Lei Wang, Wenhui Wang, Li Dong, Shaohan Huang, Huaijie Wang, Jilong Xue, Ruiping Wang, Yi Wu, and Furu Wei. Bitnet: 1-bit pre-training for large language models. Journal of Machine Learning Research, 26 0 (125): 0 1--29, 2025
2025
-
[194]
CAT-Q : Cost-efficient and accurate ternary quantization for LLMs
Shigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan, and Anbang Yao. CAT-Q : Cost-efficient and accurate ternary quantization for LLMs . In Proceedings of the 43rd International Conference on Machine Learning, 2026. Oral presentation; arXiv:2606.26650
2026 arXiv
-
[195]
Optimal Control of Differential and Functional Equations
Jack Warga. Optimal Control of Differential and Functional Equations. Academic Press, New York, 1972
1972
-
[196]
Featurized occupation measures for structured global search in numerical optimal control
Qi Wei, Jianfeng Tao, Haoyang Tan, and Hongyu Nie. Featurized occupation measures for structured global search in numerical optimal control. arXiv preprint arXiv:2603.16231, 2026
2026 arXiv
-
[197]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R \'e mi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, ...
2020
-
[198]
On the role of attention masks and layernorm in transformers
Xinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka, and Ali Jadbabaie. On the role of attention masks and layernorm in transformers. In Advances in Neural Information Processing Systems, volume 37, pages 14774--14809, 2024. doi:10.52202/079017-0472
2024 doi
-
[199]
SDQ-LLM : Sigma--delta quantization for 1-bit LLMs of any size
Junhao Xia, Ming Zhao, Limin Xiao, and Xiujun Zhang. SDQ-LLM : Sigma--delta quantization for 1-bit LLMs of any size. arXiv preprint arXiv:2510.03275, 2025
2025
-
[200]
Binaryattention: One-bit QK -attention for vision and diffusion transformers
Chaodong Xiao, Zhengqiang Zhang, and Lei Zhang. Binaryattention: One-bit QK -attention for vision and diffusion transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12106--12117, 2026
2026
-
[201]
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine...
2023
-
[202]
Lan, Wanzin Yazar, Tristan Webb, Sayeh Sharify, and Xin Wang
Zifei Xu, Alexander Y. Lan, Wanzin Yazar, Tristan Webb, Sayeh Sharify, and Xin Wang. Scaling laws for post-training quantized large language models. In Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop, volume 262 of Proceedings of Machin...
2024
-
[203]
Mahoney, T
Wanqi Yang, Yuexiao Ma, Alexander Conzelmann, Xiawu Zheng, Michael W. Mahoney, T. Konstantin Rusch, and Shiwei Liu. AlphaQ : Calibration-free bit allocation for mixture-of-experts quantization. arXiv preprint arXiv:2606.04980, 2026
2026 arXiv
-
[204]
Mahoney, and Kurt Keutzer
Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W. Mahoney, and Kurt Keutzer. HAWQ-V3 : Dyadic neural network quantization. In Proceedings of the 38th International Conference on Machine Learning, volume ...
2021
-
[205]
Laurence C. Young. Lectures on the Calculus of Variations and Optimal Control Theory. W. B. Saunders, Philadelphia, 1969
1969
-
[206]
Root mean square layer normalization
Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Advances in Neural Information Processing Systems, volume 32, 2019
2019
-
[207]
Provable post-training quantization: Theoretical analysis of OPTQ and Qronos
Haoyu Zhang, Shihao Zhang, Ian Colbert, and Rayan Saab. Provable post-training quantization: Theoretical analysis of OPTQ and Qronos . arXiv preprint arXiv:2508.04853, 2025
2025 arXiv
-
[208]
Post-training quantization for neural networks with provable guarantees
Jinjie Zhang, Yixuan Zhou, and Rayan Saab. Post-training quantization for neural networks with provable guarantees. SIAM Journal on Mathematics of Data Science, 5 0 (2): 0 373--399, 2023. doi:10.1137/22M1511709
2023 doi
-
[209]
Corrigendum: Post-training quantization for neural networks with provable guarantees
Jinjie Zhang, Yixuan Zhou, and Rayan Saab. Corrigendum: Post-training quantization for neural networks with provable guarantees. SIAM Journal on Mathematics of Data Science, 6 0 (3): 0 842--846, 2024. doi:10.1137/24M1635582
2024 doi
-
[210]
Qronos : Correcting the past by shaping the future
Shihao Zhang, Haoyu Zhang, Ian Colbert, and Rayan Saab. Qronos : Correcting the past by shaping the future... in post-training quantization. In International Conference on Learning Representations, 2026. arXiv:2505.11695
2026
-
[211]
DynamicPTQ : Mitigating activation quantization collapse via residual-stream dynamics
Zimo Zhao, Maolin Wang, Bowen Yu, Bowen Liu, Xiao Han, and Xiangyu Zhao. DynamicPTQ : Mitigating activation quantization collapse via residual-stream dynamics. arXiv preprint arXiv:2606.12487, 2026. doi:10.48550/arXiv.2606.12487
-
[212]
RQ-MoE : Residual quantization via mixture of experts for efficient input-dependent vector compression
Zhengjia Zhong, Shuyan Ke, Zaizhou Lin, Jiaqi Song, Hongyi Lan, and Hui Li. RQ-MoE : Residual quantization via mixture of experts for efficient input-dependent vector compression. arXiv preprint arXiv:2605.14359, 2026. To appear at ICML 2026
2026 arXiv
-
[213]
Zhao, Andrew M
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Y. Zhao, Andrew M. Dai, Zhifeng Chen, Quoc V. Le, and James Laudon. Mixture-of-experts with expert choice routing. In Advances in Neural Information Processing Systems, volume 35, pages 7103--7114, 2022
2022
Reviewed July 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.