REVIEW 1 major objections 7 minor 191 references
Depth can replace missing numerical precision only relative to a declared low-bit library, horizon, execution arithmetic, and routing model—and a structural floor no amount of depth can cross.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 23:29 UTC pith:NOXJIYPF
load-bearing objection Conditionally sound resource theory that cleanly separates structural floor, pure-depth synthesis, arithmetic phase, and pre-training certificates; the tube hypothesis is the real bridge to practice, not a hidden proof gap. the 1 major comments →
When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For a fixed target map and a declared low-bit residual library, the exact asymptotic limit of infinite low-bit depth is the distance from the target to the closed relaxed reachable set generated by that library. That distance is a structural floor no pure schedule can cross. Finite pure depth approaches the floor at rate O(1/D) under bounded-variation time dependence (and a Hölder-adjusted rate otherwise), but only when residual increments remain numerically visible; full-state write-back can add a growing penalty and freeze updates, while increment error feedback replaces that growth by a bounded carry. When a coherent high-precision comparator also has first-order error and the floor is ze
What carries the argument
The structural floor E_Ω,∞(F★): the distance from the target full map to the closed relaxed reachable set of the declared low-bit dictionary family. It separates what the library can express from what finite pure depth, metadata, arithmetic, and routing cost; pure schedules approach it by balanced switching / online simplex rounding, and verified primal–dual bounds turn the floor plus finite-resource radii into feasible / impossible / unresolved decisions.
Load-bearing premise
Everything ideal and implemented must stay inside one verified tube where the residual fields stay bounded and Lipschitz, so the theory does not cover attention or normalization that blow up, or arithmetic that overflows that tube.
What would settle it
Find a fixed target and declared low-bit residual library whose structural floor is zero, with coherent first-order high-precision error, yet whose best pure low-bit schedules either stay bounded away from the floor as depth grows or match the high-precision accuracy with depth growing much slower or much faster than linear in the comparator depth—under the paper’s execution and tube assumptions.
If this is right
- Before training, a dual lower bound above the tolerance certifies that no depth or optimizer can hit the target with that library.
- Full-state activation write-back can make deeper low-bit nets worse; preserving residual increments (e.g. error-feedback carry) is required for depth to help.
- Accuracy matching against a coherent first-order high-precision teacher forces low-bit depth on the order of teacher depth when the floor is zero.
- Learned codebooks must be charged as metadata separate from schedule depth; logarithmic metadata bits can keep codebook error commensurate with first-order synthesis.
- Hard routing only keeps a first-order depth law under isolated transversal events and a small-gain route–state loop, not under a frozen positive margin alone.
Where Pith is reading between the lines
- Quantization toolchains could add a pre-training screen that brackets the structural floor and rejects libraries whose dual bound already exceeds the tolerance.
- Hardware paths that quantize the full residual state each microstep are in a different phase from increment-carry designs; kernel choice may matter as much as nominal bit width.
- The same floor-plus-radius logic could grade looped or unrolled blocks: only refinements that stay coherent with a shared residual horizon earn a depth–precision exchange.
- Unresolved certificates become a research queue of their own—pointing at which bound (floor, arithmetic, or route) must tighten next rather than treating failed training as non-representability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a target-specific resource theory for when low-bit residual depth can replace numerical precision. A depth-D student is modeled as a pure schedule over a declared low-bit residual-field dictionary on a fixed horizon, with the state lifted to the full input-indexed map. The distance from the target to the closed relaxed reachable set is identified as the exact structural floor; pure schedules approach it at O(D^{-1}) under bounded-variation time dependence and O(D^{-ϑ}+D^{-1}) under ϑ-Hölder dependence (Theorems 3, 5). Execution arithmetic is shown to change the phase: full-state write-back contributes a Dρ_z certificate term with an exact scalar freeze result (Proposition 7), while increment error feedback telescopes the carry (Theorem 8) and admits a bit-exact common-lattice realization with explicit register widths (Proposition 9). A fixed, D-independent binary teacher has a closed-form optimal error Θ(D^{-1}) (Theorem 10), lifted to residual-ReLU and nonuniform two-token attention realizations (Proposition 12, Theorem 13), yielding D_match = Θ(L) for coherent first-order comparators (Corollary 15). Learned codebooks add a metadata resource with upper, packing, and allocation laws (Theorems 18, 19, 57); state-dependent routing is treated by a transversal-event small-gain theorem (Theorem 23). A primal–dual stack (HJB, support, affine, occupation-measure/SOS) yields feasible/impossible/unresolved decisions (Corollary 28). Companion software (QReplace) and
Significance. If the results hold, the paper supplies a unifying conditional framework that cleanly separates library geometry, synthesis depth, metadata, execution arithmetic, and routing — resources that the literature often collapses into a nominal bit width. Several strengths deserve explicit credit: (i) an exact closed-form fixed-teacher optimum with a nonasymptotic envelope (Theorem 10), making the first-order depth price sharp for one fixed target rather than only minimax; (ii) a Lean 4 artifact with hash-locked build logs and claim-level axiom audits kernel-checking twelve discrete-core statements, plus exhaustive executable verifiers for the attention converse and common-lattice arithmetic; (iii) a certified nonlinear matrix-valued accuracy-matching depth (D_match = 8 = 2L) proved by rational piecewise-affine bounds, with a prospective falsifiable prediction (calibrated D=9 vs certified 8) that was checked; (iv) an explicit evidence hierarchy and trust-boundary discussion that is more disciplined than typical for this area. The component tools (relaxed controls, sigma–delta feedback, occupation measures, hybrid transversality) are mature, and the paper says so; the contribution is the t
major comments (1)
- [§3.2 Assumption 1; §9.2–9.3; Theorems 6, 8, 29] The common synthesis tube is the load-bearing bridge for the master law (1) and for every QReplace verdict, and it carries a bootstrap risk the manuscript should address more directly. Assumption 1 simultaneously asserts forward invariance of K for all measurable relaxed controls, all mixed/pure Euler states and interpolation segments, and — via Theorems 6, 8, and 29 — all implemented finite-arithmetic prefixes, together with uniform B and L_z on the enlarged tube K_ρ. For attention blocks, the QK-product Lipschitz constant is state-range dependent (§9.2.2 bounds scores through B_Q, B_K), so the constants that define the tube are valid only on a tube whose existence is part of the hypothesis. The paper resolves this constructively for the contractive soft-threshold class (§9.1, Eq. (197) gives an explicit invariant ball), but the Transformer specialization provides only componentwise err
minor comments (7)
- [§8.5, Theorem 27] The main text flags that primal–dual equality requires a 'closed-image qualification detailed in Appendix A; that qualification is not automatic.' Please clarify in the main text that this qualification affects only the no-duality-gap statement, not the validity of dual lower certificates: Theorem 25 is proved directly by monotonicity along trajectories, so Corollary 28's 'certified impossible' verdicts do not depend on strong duality. As written, a reader could over-discount the decision rule or, conversely, over-credit the SOS hierarchy.
- [§10.8] The phrase 'causal validation of the theory's central distinction' is stronger than the design supports. The coherent-target arm (depth-D Euler refinement converging to the depth-32 refinement of the same field) is essentially standard Euler convergence and is expected a priori; the informative arm is the direct-target divergence. The 4-bit study uses three QAT seeds, so the Student-t intervals have two degrees of freedom, and the layer-1 hidden-map interval crosses zero (reported, but only mid-paragraph). Please temper the causal language, state what outcome would have falsified the mechanism, and note that the fitted slopes (−1.16, −0.86) come from five depth points.
- [Appendix A vs. main text] Notation drift: the appendices use C_{fh,b} and C^{unif}_{fh} where the main text uses C^{end}_{syn} and C^{unif}_{syn} (e.g., Theorem 30 vs. Theorem 3; Theorem 53 vs. Theorem 18). Please harmonize or add the correspondence to Table 3. Similarly, Φ_L(T) is defined twice (Eq. (14) and after Eq. (270)), and Eq. (65) uses u = D^{-1} in the main text but x = 1/D in Appendix A.2.
- [§1.4 vs. §8.7] QReplace is described in §1.4 as returning five outcomes (certified feasible, certified impossible, conditionally feasible, diagnostically promising, unresolved), while Corollary 28 defines a three-way decision. Please state the mapping between the two lists and which outcomes are proof-backed versus heuristic.
- [§4.2, Figure 2(a)] The write-back term Dρ_z is a worst-case certificate envelope; the only exact freeze result is scalar (Proposition 7). The text acknowledges this, but the 'phase diagram' framing of Figure 2(a) may be read as asserting realized U-shaped behavior in general architectures. A sentence clarifying that no multidimensional lower bound realizing the Dρ_z growth is known would calibrate Claim 3.
- [Eq. (36)] The α_k√q/2 term assumes a coordinatewise uniform activation grid in a q-dimensional Euclidean state; please state this explicitly, since the surrounding development is in the lifted Banach space Z.
- [References] A substantial fraction of the citations are 2025–2026 arXiv preprints (e.g., Chakrabarti et al. 2026; Park et al. 2026b; Zhao et al. 2026). Please indicate which are peer-reviewed and pin versions, since several novelty-boundary comparisons depend on them.
Circularity Check
No significant circularity: structural floor, synthesis rates, and converses are independently derived, not fitted or self-referential.
full rationale
The paper's load-bearing chain is definitional-plus-theorem, not circular. The structural floor E_Ω,∞(F*) is defined as dist(F*, R_Ω,rel); Theorems 3/5 then prove pure schedules approach that closed set at O(D^{-1}) or O(D^{-ϑ}+D^{-1}), so the distance is the asymptotic floor by Hausdorff convergence rather than by renaming the target. The fixed-teacher exact error (Theorem 10), residual-ReLU and attention embeddings, common-lattice conservation (Proposition 9), and packing/metadata laws are closed-form or constructive arguments under stated assumptions; they do not fit parameters from the quantities they claim to predict. Empirical sections are explicitly tiered as diagnostic and separated from theorem claims. Lean 4 checks and rational certificates are independent verification, not self-citation load-bearing. Assumption 1 (common tube) is a strong applicability hypothesis, not a circular step. No self-definitional loop, fitted-as-prediction, or uniqueness-via-author-citation pattern is present.
Axiom & Free-Parameter Ledger
free parameters (3)
- Library- and tube-dependent constants (B, Lz, Lt/Vt or Ht, ϑ, Csyn, LΩ,
ho z/
ho D, ηG, route
u r,Δr,χ)
- Metadata metric dimension m and entropy prefactor CΩ
- Per-microstep resource charge ℓσ and total budget B
axioms (8)
- standard math Banach-space Carathéodory existence/uniqueness for bounded, strongly measurable, uniformly state-Lipschitz relaxed fields on a forward-invariant tube (Assumption 1 / Lemma 60).
- standard math Online simplex/prefix discrepancy rounding with bound < J per coordinate (Lemma 2), used to get first-order pure-to-relaxed rates.
- standard math Superposition principle for continuity equations / occupation measures (Ambrosio et al.) to equate modal measure programs with mixtures of relaxed trajectories (Theorem 27).
- domain assumption Common verified tube containing ideal and implemented prefixes; Lipschitz and defect bounds only claimed on that tube (Assumption 1; Theorem 6).
- domain assumption Coherent fixed-horizon residual refinement family shared by high-precision comparator and low-bit student when stating D=Θ(L) matching (Sections 1.1, 5, Corollary 15).
- domain assumption Isolated transversal top-k route events with small-gain χ<1 and isolation radii (Assumption 22 / Theorem 23); excludes simultaneous/grazing/chattering regimes.
- domain assumption Increment quantizer residual radius and non-overflow of carry/state registers on declared ranges for error-feedback and common-lattice exactness (Theorem 8, Proposition 9).
- ad hoc to paper Deterministic global metadata + input-independent pure schedules for packing lower bound (Theorem 19); no uncharged continuous side information.
invented entities (3)
-
Structural floor E_Ω,∞(F*) = dist(F*, closed relaxed reachable set of declared dictionary family)
independent evidence
-
Schedulewise arithmetic radius A_D and write-back vs error-feedback phase diagram
independent evidence
-
QReplace feasible/impossible/unresolved decision rule combining [L,U] floor bracket with finite-resource radius R_{D,s}
independent evidence
read the original abstract
When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system over a fixed horizon as a pure schedule selecting fields from a declared low-bit operation library, and use relaxed controls to characterize its infinite-depth limit. The distance from the target to the closed relaxed reachable set is the exact structural floor: no increase in depth can remove it for that library. Pure schedules approach the relaxed class at rate $O(D^{-1})$ under bounded-variation time dependence and $O(D^{-\vartheta}+D^{-1})$ under Holder dependence of exponent $\vartheta$. Execution arithmetic can reverse this conclusion: full-state write-back introduces a $D\rho_z$ penalty and can freeze residual updates, whereas increment error feedback replaces this growth by a bounded carry term and obeys an exact common-lattice conservation law. A fixed-teacher converse makes this rate sharp: for coherent depth-$L$ first-order high-precision comparators, accuracy matching requires $D=\Theta(L)$. Learned codebooks add a metadata resource, while state-dependent routing introduces hybrid event conditions. Verified primal and dual bounds yield feasible, impossible, or unresolved decisions before training. Companion software implements the workflow, and Lean 4 machine-checks the exact discrete core. Depth replaces precision only relative to a declared library, horizon, execution semantics, and routing model.
Figures
Reference graph
Works this paper leans on
-
[1]
Communications on Pure and Applied Mathematics , volume =
Ingrid Daubechies and Michel Defrise and Christine De Mol , title =. Communications on Pure and Applied Mathematics , volume =. 2004 , doi =
2004
-
[2]
SIAM Journal on Imaging Sciences , volume =
Amir Beck and Marc Teboulle , title =. SIAM Journal on Imaging Sciences , volume =. 2009 , doi =
2009
-
[3]
Proceedings of the 27th International Conference on Machine Learning , pages =
Karol Gregor and Yann LeCun , title =. Proceedings of the 27th International Conference on Machine Learning , pages =
-
[4]
Advances in Neural Information Processing Systems , volume =
Xiaohan Chen and Jialin Liu and Zhangyang Wang and Wotao Yin , title =. Advances in Neural Information Processing Systems , volume =
-
[5]
Eldar , title =
Vishal Monga and Yuelong Li and Yonina C. Eldar , title =. IEEE Signal Processing Magazine , volume =. 2021 , doi =
2021
-
[6]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =
Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2016 , doi =
2016
-
[7]
Ricky T. Q. Chen and Yulia Rubanova and Jesse Bettencourt and David Duvenaud , title =. Advances in Neural Information Processing Systems , volume =
-
[8]
Inverse Problems , volume =
Eldad Haber and Lars Ruthotto , title =. Inverse Problems , volume =. 2018 , doi =
2018
-
[9]
Sander and Pierre Ablin and Gabriel Peyr
Michael E. Sander and Pierre Ablin and Gabriel Peyr. Do Residual Neural Networks Discretize Neural Ordinary Differential Equations? , booktitle =. 2022 , doi =
2022
-
[10]
Jack Warga , title =
-
[11]
Young , title =
Laurence C. Young , title =
-
[12]
Advances in Neural Information Processing Systems , volume =
Matthieu Courbariaux and Yoshua Bengio and Jean-Pierre David , title =. Advances in Neural Information Processing Systems , volume =
-
[13]
Journal of Machine Learning Research , volume =
Itay Hubara and Matthieu Courbariaux and Daniel Soudry and Ran El-Yaniv and Yoshua Bengio , title =. Journal of Machine Learning Research , volume =
-
[14]
European Conference on Computer Vision , pages =
Mohammad Rastegari and Vicente Ordonez and Joseph Redmon and Ali Farhadi , title =. European Conference on Computer Vision , pages =. 2016 , doi =
2016
-
[15]
International Conference on Learning Representations , year =
Yukun Ding and Jinglan Liu and Jinjun Xiong and Yiyu Shi , title =. International Conference on Learning Representations , year =
-
[16]
Advances in Neural Information Processing Systems , volume =
Yaniv Blumenfeld and Dar Gilboa and Daniel Soudry , title =. Advances in Neural Information Processing Systems , volume =
-
[18]
Journal of Machine Learning Research , volume =
Hongyu Wang and Shuming Ma and Lingxiao Ma and Lei Wang and Wenhui Wang and Li Dong and Shaohan Huang and Huaijie Wang and Jilong Xue and Ruiping Wang and Yi Wu and Furu Wei , title =. Journal of Machine Learning Research , volume =
-
[22]
International Conference on Learning Representations , year =
Peter O'Connor and Max Welling , title =. International Conference on Learning Representations , year =
-
[23]
Annals of Mathematics , volume =
Ingrid Daubechies and Ronald DeVore , title =. Annals of Mathematics , volume =. 2003 , doi =
2003
-
[24]
C. Sinan G. One-Bit Sigma--Delta Quantization with Exponential Accuracy , journal =. 2003 , doi =
2003
-
[25]
IEEE Transactions on Information Theory , volume =
Felix Krahmer and Rayan Saab and Rachel Ward , title =. IEEE Transactions on Information Theory , volume =. 2012 , doi =
2012
-
[26]
Advances in Neural Information Processing Systems , volume =
Zechun Liu and Changsheng Zhao and Hanxian Huang and Sijia Chen and Jing Zhang and Jiawei Zhao and Scott Roy and Lisa Jin and Yunyang Xiong and Yangyang Shi and Lin Xiao and Yuandong Tian and Bilge Soran and Raghuraman Krishnamoorthi and Tijmen Blankevoort and Vikas Chandra , title =. Advances in Neural Information Processing Systems , volume =
-
[28]
Advances in Neural Information Processing Systems , volume =
Adrian Bulat and Yassine Ouali and Georgios Tzimiropoulos , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =
2024
-
[29]
Advances in Neural Information Processing Systems , volume =
Yamato Arai and Yuma Ichikawa , title =. Advances in Neural Information Processing Systems , volume =
-
[30]
Advances in Neural Information Processing Systems , volume =
Banseok Lee and Dongkyu Kim and Youngcheon You and Youngmin Kim , title =. Advances in Neural Information Processing Systems , volume =
-
[31]
Proceedings of the 43rd International Conference on Machine Learning , year =
Shigeng Wang and Chao Li and Yangyuxuan Kang and Jiawei Fan and Anbang Yao , title =. Proceedings of the 43rd International Conference on Machine Learning , year =
-
[32]
Gomez and Lukasz Kaiser and Illia Polosukhin , title =
Ashish Vaswani and Noam Shazeer and Niki Parmar and Jakob Uszkoreit and Llion Jones and Aidan N. Gomez and Lukasz Kaiser and Illia Polosukhin , title =. Advances in Neural Information Processing Systems , volume =
-
[34]
Advances in Neural Information Processing Systems , volume =
Biao Zhang and Rico Sennrich , title =. Advances in Neural Information Processing Systems , volume =
-
[35]
Advances in Neural Information Processing Systems , volume =
Subhabrata Dutta and Tanya Gautam and Soumen Chakrabarti and Tanmoy Chakraborty , title =. Advances in Neural Information Processing Systems , volume =
-
[36]
Bulletin of the American Mathematical Society , volume =
Borjan Geshkovski and Cyril Letrouit and Yury Polyanskiy and Philippe Rigollet , title =. Bulletin of the American Mathematical Society , volume =. 2025 , doi =
2025
-
[37]
Proceedings of the 38th International Conference on Machine Learning , series =
Hyunjik Kim and George Papamakarios and Andriy Mnih , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[38]
Proceedings of the 38th International Conference on Machine Learning , series =
George Dasoulas and Kevin Scaman and Aladin Virmaux , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[39]
How Smooth Is Attention? , booktitle =
Val. How Smooth Is Attention? , booktitle =
-
[40]
Proceedings of the 38th International Conference on Machine Learning , series =
Yihe Dong and Jean-Baptiste Cordonnier and Andreas Loukas , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[41]
Advances in Neural Information Processing Systems , volume =
Xinyi Wu and Amir Ajorlou and Yifei Wang and Stefanie Jegelka and Ali Jadbabaie , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =
2024
-
[42]
Consensus Is All You Get: The Role of Attention in Transformers , booktitle =
-
[43]
Mahoney and Kurt Keutzer , title =
Sehoon Kim and Amir Gholami and Zhewei Yao and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[44]
Proceedings of the 40th International Conference on Machine Learning , series =
Guangxuan Xiao and Ji Lin and Mickael Seznec and Hao Wu and Julien Demouth and Song Han , title =. Proceedings of the 40th International Conference on Machine Learning , series =
-
[45]
International Conference on Learning Representations , year =
Elias Frantar and Saleh Ashkboos and Torsten Hoefler and Dan Alistarh , title =. International Conference on Learning Representations , year =
-
[46]
Proceedings of Machine Learning and Systems , volume =
Ji Lin and Jiaming Tang and Haotian Tang and Shang Yang and Wei-Ming Chen and Wei-Chen Wang and Guangxuan Xiao and Xingyu Dang and Chuang Gan and Song Han , title =. Proceedings of Machine Learning and Systems , volume =
-
[47]
Proceedings of the 41st International Conference on Machine Learning , series =
Albert Tseng and Jerry Chee and Qingyao Sun and Volodymyr Kuleshov and Christopher De Sa , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[48]
Proceedings of the 41st International Conference on Machine Learning , series =
Wei Huang and Yangdong Liu and Haotong Qin and Ying Li and Shiming Zhang and Xianglong Liu and Michele Magno and Xiaojuan Qi , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[49]
Proceedings of the 41st International Conference on Machine Learning , series =
Shiyao Li and Xuefei Ning and Luning Wang and Tengxuan Liu and Xiangsheng Shi and Shengen Yan and Guohao Dai and Huazhong Yang and Yu Wang , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[50]
Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =
Zifei Xu and Alexander Y. Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =. Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop , series =
-
[51]
Proceedings of the 41st International Conference on Machine Learning , series =
Harshavardhan Adepu and Zhanpeng Zeng and Li Zhang and Vikas Singh , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[52]
International Conference on Learning Representations , year =
Noam Shazeer and Azalia Mirhoseini and Krzysztof Maziarz and Andy Davis and Quoc Le and Geoffrey Hinton and Jeff Dean , title =. International Conference on Learning Representations , year =
-
[53]
Journal of Machine Learning Research , volume =
William Fedus and Barret Zoph and Noam Shazeer , title =. Journal of Machine Learning Research , volume =
-
[54]
Zhao and Andrew M
Yanqi Zhou and Tao Lei and Hanxiao Liu and Nan Du and Yanping Huang and Vincent Y. Zhao and Andrew M. Dai and Zhifeng Chen and Quoc V. Le and James Laudon , title =. Advances in Neural Information Processing Systems , volume =
-
[55]
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =
Damai Dai and Li Dong and Shuming Ma and Bo Zheng and Zhifang Sui and Baobao Chang and Furu Wei , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2022 , doi =
2022
-
[56]
International Conference on Learning Representations , year =
Joan Puigcerver and Carlos Riquelme Ruiz and Basil Mustafa and Neil Houlsby , title =. International Conference on Learning Representations , year =
-
[57]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Chaodong Xiao and Zhengqiang Zhang and Lei Zhang , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[58]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Hyunha Hwang and Xuan Truong Nguyen and Hyuk-Jae Lee , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[59]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =
Jiahe Qian and Peisong Wang and Zhengyang Zhuge and Qinghao Hu and Jian Cheng , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =. 2026 , doi =
2026
-
[60]
Mathematical Programming , volume =
Sebastian Sager and Gerhard Reinelt and Hans Georg Bock , title =. Mathematical Programming , volume =. 2009 , doi =
2009
-
[61]
Mathematical Programming , volume =
Sebastian Sager and Hans Georg Bock and Moritz Diehl , title =. Mathematical Programming , volume =. 2012 , doi =
2012
-
[62]
Mathematical Methods of Operations Research , volume =
Sebastian Sager and Michael Jung and Christian Kirches , title =. Mathematical Methods of Operations Research , volume =. 2011 , doi =
2011
-
[63]
SIAM Journal on Control and Optimization , volume =
Christian Kirches and Felix Lenders and Paul Manns , title =. SIAM Journal on Control and Optimization , volume =. 2020 , doi =
2020
-
[64]
Shankar Sastry , title =
Ramanarayan Vasudevan and Humberto Gonzalez and Ruzena Bajcsy and S. Shankar Sastry , title =. SIAM Journal on Control and Optimization , volume =. 2013 , doi =
2013
-
[66]
Bowman , title =
Alex Wang and Amanpreet Singh and Julian Michael and Felix Hill and Omer Levy and Samuel R. Bowman , title =. International Conference on Learning Representations , year =
-
[67]
Manning and Andrew Y
Richard Socher and Alex Perelygin and Jean Wu and Jason Chuang and Christopher D. Manning and Andrew Y. Ng and Christopher Potts , title =. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , pages =. 2013 , doi =
2013
-
[68]
Transformers: State-of-the-Art Natural Language Processing , booktitle =
Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and R. Transformers: State-of-the-Art Natural Language Processing , booktitle =. 2020 , doi =
2020
-
[69]
distilbert-base-uncased-finetuned-sst-2-english , year =
-
[70]
Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =
Samuel Teuber and Marko Kleine B. Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =. 2021 , doi =
2021
-
[71]
Mahoney and Kurt Keutzer , title =
Zhewei Yao and Zhen Dong and Zhangcheng Zheng and Amir Gholami and Jiali Yu and Eric Tan and Leyuan Wang and Qijing Huang and Yida Wang and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =
-
[73]
Automatica , volume =
Mathieu Claeys and Jamal Daafouz and Didier Henrion , title =. Automatica , volume =. 2016 , doi =
2016
-
[74]
Burden and S
Samuel A. Burden and S. Shankar Sastry and Daniel E. Koditschek and Shai Revzen , title =. SIAM Journal on Applied Dynamical Systems , volume =. 2016 , doi =
2016
-
[75]
Kong and J
Nathan J. Kong and J. Joe Payne and James Zhu and Aaron M. Johnson , title =. Proceedings of the IEEE , volume =. 2024 , doi =
2024
-
[76]
Iris A. M. Huijben and Matthijs Douze and Matthew J. Muckley and Ruud J. G. van Sloun and Jakob Verbeek , title =. Proceedings of the 41st International Conference on Machine Learning , series =
-
[80]
Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr
Jean B. Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr. Nonlinear Optimal Control via Occupation Measures and. SIAM Journal on Control and Optimization , volume =. 2008 , doi =
2008
-
[81]
Transactions on Machine Learning Research , year =
Ian Colbert and Giuseppe Franco and Fabian Grob and Jinjie Zhang and Rayan Saab , title =. Transactions on Machine Learning Research , year =
-
[85]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Pei Huang and Haoze Wu and Yuting Yang and Ieva Daukantas and Min Wu and Yedi Zhang and Clark Barrett , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , doi =
2024
-
[86]
Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =
Luigi Ambrosio and Nicola Gigli and Giuseppe Savar. Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =. 2008 , doi =
2008
-
[87]
International Conference on Learning Representations , year =
Shihao Zhang and Haoyu Zhang and Ian Colbert and Rayan Saab , title =. International Conference on Learning Representations , year =
-
[89]
Burr and Liu Liu and Meng Wang , title =
Mohammed Nowaz Rabbani Chowdhury and Kaoutar El Maghraoui and Hsinyu Tsai and Naigang Wang and Geoffrey W. Burr and Liu Liu and Meng Wang , title =. International Conference on Learning Representations , year =
-
[90]
SIAM Journal on Mathematics of Data Science , volume =
Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , volume =. 2023 , doi =
2023
-
[91]
Proceedings of the 37th International Conference on Machine Learning , series =
Markus Nagel and Rana Ali Amjad and Mart van Baalen and Christos Louizos and Tijmen Blankevoort , title =. Proceedings of the 37th International Conference on Machine Learning , series =
-
[96]
SIAM Journal on Mathematics of Data Science , year =
Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , year =
-
[97]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[98]
Proceedings of the 41st International Conference on Machine Learning , series =
Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig and Yaman Umuroglu , title =. Proceedings of the 41st International Conference on Machine Learning , series =. 2024 , publisher =
2024
-
[101]
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =
Moritz Br. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =. 2026 , note =
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.