REVIEW 3 major objections 6 minor 60 references
This paper claims that a CESTAC-based software tool, noisefloat, can detect numerically unstable deep-learning operators by estimating significant digits from synchronized stochastic samples during training and inference.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 02:15 UTC pith:E7H57YZT
load-bearing objection A genuinely useful CESTAC-in-DL integration whose detection claim is only supported for Python-visible instabilities, not the fused GEMM kernels where many real problems live. the 3 major comments →
Automated Numerical Stability Analysis of Deep Learning Operators
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The author's central claim is that noisefloat is the first software tool to integrate CESTAC into the numerical validation of deep-learning deployment, and that it can effectively detect numerically unstable operators. The method represents every deterministic value as several synchronized stochastic samples, quantizes each operation with randomly directed rounding, and computes an estimated number of significant decimal digits from the sample mean and dispersion. Unstable computations—catastrophic cancellation, overflow in softmax, near-zero variance in normalization, near-tie attention logits—report very low or zero digits, while algebraically stabilized variants retain high digits. A sync
What carries the argument
The central mechanism is CESTAC-style stochastic arithmetic lifted to deep-learning operators: each tensor is carried as a batch of s stochastic samples (default 3), backend-native randomized directed rounding is applied at instrumented boundaries, and the significant-digit estimate C_Y = log10( sqrt(s) |mean| / (tau_beta * sigma) ) is computed elementwise. Operator-boundary wrappers quantize only operator outputs; arithmetic-level propagation rounds after Python-visible primitives. GEMM-like operators (matmul, linear, convolution lowered to matmul) are treated as trusted: when stochastic operands are insufficiently separated, inputs are perturbed at the unit-roundoff scale eta = 2^{-p} befo
Load-bearing premise
The central claim depends on treating GEMM-like kernels as trusted operators, so instabilities inside their fused accumulation are only probed by perturbing insufficiently separated inputs; if the perturbation surrogate misses an internal hazard, the tool will not flag it.
What would settle it
Construct a neural network whose only instability is inside a fused matrix-multiply kernel's internal accumulation order—for instance, an inner product that sums alternating large and near-equal terms so that the true result is small—and ask noisefloat's operator-level report to flag that operator. If the report stays high because the kernel is treated as trusted, the detection claim is limited to Python-visible operations.
If this is right
- Developers can localize numerically unstable operators during training or inference without changing the model's optimization trajectory.
- The significant-digit reports separate stable and unstable formulations of the same operation—for example, shifted versus naive softmax, or rationalized versus cancellation-prone expressions.
- Operator-level digit rankings provide a concrete signal for where to increase precision, reformulate an expression, or replace a numerically fragile operator.
- Local loss of digits does not necessarily change the final decision; the paper shows that downstream margins can absorb operator-level instability, so the diagnostics are per-operator reliability measures rather than end-to-end failure predictions.
- GEMM-like kernels are deliberately treated as trusted; their reports are data-perturbation estimates whose reliability depends on the cited probabilistic validation hypotheses.
Where Pith is reading between the lines
- The same per-operator digit estimates could be used as a precision-tuning signal: operators with chronically low C_Y values are natural candidates for higher precision or stable reformulation, and this could be tested by correlating digit drops with actual reduced-precision training failures.
- Because the tool does not instrument arithmetic inside fused GEMM-like kernels, instabilities that live entirely in internal accumulation order—such as a matmul whose inner products cancel—may escape detection; extending CESTAC to kernel internals or using randomized rounding inside the kernel would close that gap.
- The reported overhead (roughly three sample evaluations plus quantization, potentially far more in practice) suggests that scalable monitoring will likely use calibration subsets or sparse iteration sampling rather than full-data tracing; a sampling strategy that preserves the localization power could be a natural follow-up.
- The observed decoupling of low operator digits from high decision-level agreement points toward a stability-aware margin criterion: a network could be considered numerically safe when its minimum digit level stays above a threshold that depends on downstream logit margins.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes 'noisefloat', a Python framework that wraps NumPy/PyTorch/JAX/TensorFlow tensor programs with CESTAC-style stochastic arithmetic. It represents each value by three synchronous samples with randomized directed rounding, estimates significant decimal digits, and reports diagnostics per operator or per intercepted primitive. For GEMM-like operators it does not instrument inside the kernel; instead it perturbs inputs by eta=2^{-p} when samples are insufficiently separated (Eqs. 15-17). The authors evaluate the method on classical numerical examples, controlled stable/unstable operator pairs, injected pathology studies on Fashion-MNIST/CIFAR-10/AG News, and activation/normalization reliability on Fashion-MNIST, reporting perfect separation of the controlled pairs and localization of injected hazards.
Significance. If the central claim holds, the paper contributes a practical, open-source diagnostic layer for mixed-precision deep learning: backend-native stochastic quantization, straight-through estimator compatibility, operator-level reports, and explicit error bounds (Theorems 2-3) that quantify the gap between the GEMM surrogate and a finer stochastic execution. The controlled S/U pairs separate cleanly and the theoretical results are internally consistent. However, the evidence does not yet establish the stated claim for fused GEMM-like kernels, and the evaluation relies on hand-set detection thresholds and single-run training studies. With narrowed claims or targeted additional experiments, the tool would be a useful complement to CADNA/Verificarlo in deep-learning deployment.
major comments (3)
- [§3.4, Eqs. (15)-(17), and Theorem 3] The 'trusted GEMM-like operator' path does not apply randomized directed rounding inside matrix multiplication, linear layers, or convolutions lowered to matmul; it only perturbs inputs by eta=2^{-p} when samples are insufficiently separated. For a reduction of length n, internal accumulation rounding can grow as gamma_n * sum|a_i| ~ n*eps*sum|a_i|, whereas the input perturbation changes the exact sum by O(eta*sum|a_i|). With n large and eta=eps, the former can exceed the latter by orders of magnitude, so three samples can remain nearly identical and Eq. (8) reports high digits despite kernel-internal instability. Theorem 3 bounds the effect only of the data perturbations already injected (Delta_p), not of rounding inside the kernel. Since §4.2-4.4 place all injected hazards at Python-visible boundaries, the central claim of §6 that noisefloat detects unstable operators is not demonstrat
- [Eq. (53) and Table 3] The reported perfect classification (accuracy=precision=recall=1.000) uses hand-set, precision-dependent thresholds gamma_23=3 and gamma_52=10. The paper does not justify these values or report how detection accuracy varies with them; the thresholds are free parameters and no selection procedure is described. Since the unstable constructions are deliberately extreme (about 0 digits vs. about 3.5+ in FP32), the benchmark demonstrates ranking/separation, not a robust automatic detection criterion. A principled threshold procedure or a sensitivity/ROC-style analysis is needed to support the 'effectively detect' claim without user calibration.
- [§4.3-4.6, Figs. 9-13] The end-to-end studies use a single seed, 2,048 training examples, at most 32 optimizer steps per epoch, and 3 epochs. The conclusions about localization and about activation/normalization ordering are based on one trajectory each; no confidence intervals, repeated seeds, or ablations are provided. Moreover, the injected pathologies are all of the cancellation/overflow/near-zero-variance types that the CADNA-style source counters are designed to flag, so success on these controlled cases is weaker evidence for detecting unknown instability modes. The experiments support the tool's utility as an operator-level diagnostic, but they are not yet strong enough for the general detection claim in §6.
minor comments (6)
- [Throughout] The paper alternates between 'noisefloat' and 'noisyfloat' (title/abstract vs. §3); standardize the software name.
- [§2] Typo: 'deep learnin models' in the Fuzzy PyTorch paragraph.
- [§4.1] The session enumeration is garbled: 'the second session ... the second and third sessions'; clarify which experiments correspond to which session.
- [§1] The 'first work' claim should be positioned more carefully with respect to Fuzzy PyTorch [49] and Verificarlo-based stochastic arithmetic; the distinction from CESTAC is clear but should be stated explicitly.
- [Eq. (53)] The notation gamma_23 and gamma_52 is not defined; state explicitly that these correspond to simulated FP32 and FP64 significand precisions.
- [Figs. 12-13] Figure 12 panel (b) is labeled 'deterministic batch accuracy' while the text discusses representative accuracy; clarify which quantity is plotted. Figure 13 has duplicate panel labels for pre-normalization digits.
Circularity Check
No derivation-level circularity; the core estimator is the independent CESTAC formula, while the GEMM surrogate and empirical thresholds are acknowledged approximations rather than fitted predictions.
full rationale
The paper's central significant-digit estimator is the standard CESTAC formula (Eq. 8) with s=3 stochastic samples and a Student-t confidence factor; it is a fixed external stochastic-arithmetic definition, not a parameter fitted to the stable/unstable labels. The GEMM-like operator treatment in Section 3.4 explicitly states it is not instruction-level CESTAC inside the kernel and relies on the cited Jézéquel–Mary probabilistic inner-product analysis; the associated Lemma 1 and Theorems 2–3 are bounding statements about the surrogate's effect on the same estimator, not an identity between the surrogate and its inputs. Section 3.6 also disclaims instrumentation of fused vendor kernels, which is a coverage limitation, not a circular derivation. The empirical evaluation in Section 4.2–4.4 uses controlled constructions whose instability is grounded in independent mechanisms (cancellation ratio, condition number, overflow, near-zero normalization), and Eq. (53) applies a fixed digit threshold; although the threshold is manually chosen and the benchmark is partly self-referent in that the detector and the instability definition both involve loss of significant digits, no fitted parameter is renamed as a prediction and no derivation is forced by definition. Minor self-citations in the reference list ([9], [11]) appear only in related-work and precision-tuning contexts and are not load-bearing. Overall, the derivation chain is self-contained modulo standard CESTAC theory; the identified weaknesses are validation-coverage and threshold-choice concerns rather than circularity.
Axiom & Free-Parameter Ledger
free parameters (3)
- detection digit thresholds gamma_p (gamma_23=3, gamma_52=10) =
3 (FP32), 10 (FP64)
- GEMM data-perturbation amplitude eta =
eta = 2^{-p} for simulated significand precision p
- number of stochastic samples s =
3 (default)
axioms (4)
- domain assumption CESTAC significance assumption: digits common to three synchronous directed-rounding samples are the reliable digits, and the Student-t critical value at confidence 0.95 converts sample spread into a digit count.
- domain assumption GEMM-like backend kernels are backward stable, so input perturbation at amplitude eta=2^{-p} (Eq. 16) can stand in for internal rounding-error analysis.
- ad hoc to paper Controlled injected hazards (large-offset cancellation, near-zero variance, near-tie attention, unshifted exponentials) are representative of real DL operator instabilities.
- domain assumption Local Lipschitz constants L_p exist along the segment between ideal and surrogate trajectories for compositions that include activations, normalizations, softmax, attention, or recurrences.
read the original abstract
Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or an improper formulation, which leads to numerical instability. In this paper, we introduce the first unified software tool that integrates CESTAC for detecting the numerical stability of deep learning operators. Our developed software not only enables numerical validation with a single computation pass but also detects the sources of numerical instability and provides numerical stability monitoring during deep learning training and inference. We verified its effectiveness on the detection of polluted operators with injected numerical instabilities across various tasks. We believe that our developed method and tool provide valuable insights into developing numerically stable computing kernels, which are particularly critical for numerically stable and efficient deep learning training and inference.
Figures
Reference graph
Works this paper leans on
-
[1]
Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A System for...
2016
-
[2]
El-Mehdi El Arar, Silviu-Ioan Filip, Theo Mary, and Elisa Riccietti. 2025. Mixed precision accumulation for neural network inference guided by componentwise forward error analysis. arXiv:2503.15568 [cs.LG]
arXiv 2025
-
[3]
Dorra Ben Khalifa and Matthieu Martel. 2024. Efficient Implementation of Neural Networks Usual Layers on Fixed-Point Architectures. InProceedings of the 25th International Workshop on Software and Compilers for Embedded Systems. doi:10.1145/3652032.3657578
arXiv 2024
-
[4]
Dorra Ben Khalifa, Matthieu Martel, and Assalé Adjé. 2020. POP: A Tuning Assistant for Mixed-Precision Floating-Point Computations. InFormal Techniques for Safety-Critical Systems, Osman Hasan and Frédéric Mallet (Eds.). Springer, Cham, 77–94
2020
-
[5]
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013. Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation. InAdvances in Neural Information Processing Systems, Vol. 26
2013
-
[6]
Théo Beuzeville, Alfredo Buttari, Serge Gratton, and Theo Mary. 2026. Deterministic and probabilistic rounding error analysis of neural networks in floating-point arithmetic.IMA J. Numer. Anal.(2026), draf130. doi:10.1093/imanum/draf130
-
[7]
2018.JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Yash Katariya, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018.JAX: composable transformations of Python+NumPy programs. http: //github.com/jax-ml/jax
2018
-
[8]
Stanislav Budzinskiy, Wenyi Fang, Longbin Zeng, and Philipp Petersen. 2026. Numerical stability analysis of large language models. arXiv:2503.10251 [math.NA] https://arxiv.org/abs/2503.10251
Pith/arXiv arXiv 2026
-
[9]
Erin Carson, Xinye Chen, and Xiaobo Liu. 2026. Computing k-means in mixed precision. arXiv:2407.12208 [math.NA]
Pith/arXiv arXiv 2026
-
[10]
Erin Carson and Nicholas J. Higham. 2017. A New Analysis of Iterative Refinement and Its Application to Accurate Solution of Ill-Conditioned Sparse Linear Systems.SIAM Journal on Scientific Computing39, 6 (2017), A2834–A2856. doi:10.1137/17M1122918
-
[11]
Xinye Chen, Thibault Hilaire, and Fabienne Jézéquel. 2026. Floating-point autotuning with customized precisions. arXiv:2606.08339 [cs.MS] https://arxiv.org/abs/2606.08339
Pith/arXiv arXiv 2026
-
[12]
Stefano Cherubin and Giovanni Agosta. 2020. Tools for Reduced Precision Computation: A Survey.ACM Comput. Surv.53, 2 (April 2020), 35 pages. doi:10.1145/3381039
doi:10.1145/3381039 2020
-
[13]
Christophe Denis, Pablo de Oliveira Castro, and Eric Petit. 2015. Verificarlo: checking floating point accuracy through Monte Carlo Arithmetic. CoRRabs/1509.01347 (2015). http://arxiv.org/abs/1509.01347
Pith/arXiv arXiv 2015
-
[14]
Christophe Denis, Pablo de Oliveira Castro, and Eric Petit. 2016. Verificarlo: Checking Floating Point Accuracy through Monte Carlo Arithmetic. In IEEE 23rd Symposium on Computer Arithmetic (ARITH). IEEE, 55–62. doi:10.1109/ARITH.2016.31
-
[15]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1. Association for Computational Linguistics, Minnea...
-
[16]
Pacôme Eberhart, Julien Brajard, Pierre Fortin, and Fabienne Jézéquel. 2015. High Performance Numerical Validation using Stochastic Arithmetic. Reliable Computing21 (2015), 35–52. https://interval.louisiana.edu/reliable-computing-journal/volume-21/reliable-computing-21-pp-035-052.pdf
2015
-
[17]
Julian Faraone and Philip Leong. 2019. Monte Carlo Deep Neural Network Arithmetic. InInternational Conference on Learning Representations. https://openreview.net/forum?id=HyePberFvH
2019
-
[18]
Quentin Ferro, Stef Graillat, Thibault Hilaire, Fabienne Jézéquel, and Basile Lewandowski. 2022. Neural Network Precision Tuning Using Stochastic Arithmetic. InComputer Arithmetic (Lecture Notes in Computer Science, Vol. 13253). Springer, 36–53. https://hal.science/hal-03682645 31
2022
-
[19]
Quentin Ferro, Stef Graillat, Thibault Hilaire, and Fabienne Jézéquel. 2023. Performance of precision auto-tuned neural networks. In2023 IEEE 16th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC). 592–599. doi:10.1109/MCSoC60832.2023.00092
arXiv 2023
-
[20]
Michael Frechtling and Philip H. W. Leong. 2015. MCALIB: Measuring Sensitivity to Rounding Error with Monte Carlo Programming.ACM Transactions on Programming Languages and Systems37, 2 (2015), 5:1–5:25. doi:10.1145/2665073
-
[21]
Cédric Gernigon, Clément Coggiola, Silviu-Ioan Filip, Olivier Sentieys, and Mickaël Bruno. 2024. AdaQAT: Adaptive Bit-Width Quantization-Aware Training. arXiv:2404.16876 [cs.LG]
Pith/arXiv arXiv 2024
-
[22]
David Goldberg. 1991. What Every Computer Scientist Should Know About Floating-Point Arithmetic.Comput. Surveys23, 1 (1991), 5–48. doi:10.1145/103162.103163
arXiv 1991
-
[23]
Stef Graillat, Fabienne Jézéquel, Romain Picot, François Févotte, and Bruno Lathuilière. 2019. Auto-tuning for floating-point precision with Discrete Stochastic Arithmetic.Journal of Computational Science36 (2019), 101017. doi:10.1016/j.jocs.2019.07.004
-
[24]
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Y. Ng. 2014. Deep Speech: Scaling up End-to-End Speech Recognition.arXiv preprint arXiv:1412.5567(2014). arXiv:1412.5567 [cs.CL]
Pith/arXiv arXiv 2014
-
[25]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 770–778. doi:10.1109/CVPR.2016.90
-
[26]
Nicholas J. Higham. 2002.Accuracy and Stability of Numerical Algorithms(2 ed.). SIAM, Philadelphia, PA. doi:10.1137/1.9780898718027
-
[27]
Nicholas J. Higham and Theo Mary. 2019. A New Approach to Probabilistic Rounding Error Analysis.SIAM Journal on Scientific Computing41, 5 (2019), A2815–A2835. doi:10.1137/18M1226312
-
[28]
Dahl, Abdel rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N
Geoffrey Hinton, Li Deng, Dong Yu, George E. Dahl, Abdel rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N. Sainath, and Brian Kingsbury. 2012. Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups.IEEE Signal Processing Magazine29, 6 (2012), 82–97. doi:10.1109/MSP...
arXiv 2012
-
[29]
IEEE. 2019. IEEE Standard for Floating-Point Arithmetic. doi:10.1109/IEEESTD.2019.8766229
arXiv 2019
-
[30]
2024.Interim Report on Binary Floating-point Formats for Machine Learning
IEEE P3109 Working Group. 2024.Interim Report on Binary Floating-point Formats for Machine Learning. Technical Report. IEEE. Version 0.9.1
2024
-
[31]
Ioualalen and M
A. Ioualalen and M. Martel. 2019. Neural Network Precision Tuning. InQuantitative Evaluation of Systems (Lecture Notes in Computer Science, Vol. 11785). Springer, 129–143
2019
-
[32]
F. Jézéquel and J.-M. Chesneaux. 2008. CADNA: A Library for Estimating Round-Off Error Propagation.Computer Physics Communications178, 12 (2008), 933–955. doi:10.1016/j.cpc.2008.02.013
-
[33]
Fabienne Jézéquel, Stef Graillat, Daichi Mukunoki, Toshiyuki Imamura, and Roman Iakymchuk. 2020. Can We Avoid Rounding-Error Estimation in HPC Codes and Still Get Trustworthy Results?. InSoftware Verification (Lecture Notes in Computer Science, Vol. 12549). Springer, 163–177. doi:10.1007/978-3-030-63618-0_10
-
[34]
Fabienne Jézéquel and Theo Mary. 2024. Probabilistic Estimation of the Accuracy of Inner Products and Application to Stochastic Validation. https://hal.science/hal-04554459
2024
-
[35]
F. Jézéquel, J.-L. Lamotte, and O. Chubach. 2013. Parallelization of discrete stochastic arithmetic on multicore architectures. In2013 10th International Conference on Information Technology: New Generations. 160–166. doi:10.1109/ITNG.2013.28
-
[36]
2009.Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky. 2009.Learning Multiple Layers of Features from Tiny Images. Technical Report. University of Toronto
2009
-
[37]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. InAdvances in Neural Information Processing Systems, Vol. 25. 1097–1105
2012
-
[38]
Christoph Lauter and Anastasia Volkova. 2020. A Framework for Semi-Automatic Precision and Accuracy Analysis for Fast and Rigorous Deep Learning. arXiv:2002.03869 [cs.LG]
Pith/arXiv arXiv 2020
-
[39]
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep Learning.Nature521, 7553 (2015), 436–444. doi:10.1038/nature14539
-
[40]
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. Gradient-Based Learning Applied to Document Recognition.Proc. IEEE86, 11 (1998), 2278–2324. doi:10.1109/5.726791
doi:10.1109/5.726791 1998
-
[41]
Joonhyung Lee, Jeongin Bae, Byeongwook Kim, Se Jung Kwon, and Dongsoo Lee. 2024. To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability.arXiv preprint arXiv:2405.18710(2024). doi:10.48550/arXiv.2405.18710
-
[42]
Hanqing Liu, Jianjun Cao, Yuanze Li, and Zijian Zhou. 2026. Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes. InICML 2026 Workshop on High-Dimensional Learning Dynamics. arXiv:2605.06152 [cs.LG] doi:10.48550/arXiv.2605.06152 Spotlight
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2605.06152 2026
-
[43]
Debasmita Lohar, Clothilde Jeangoudoux, Anastasia Volkova, and Eva Darulova. 2023. Sound Mixed Fixed-Point Quantization of Neural Networks. ACM Transactions on Embedded Computing Systems22, 5s (2023). doi:10.1145/3609118
-
[44]
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. Mixed Precision Training. InInternational Conference on Learning Representations. https://openreview.net/ forum?id=r1gs9JgRZ
2018
- [45]
-
[46]
1997.Monte Carlo Arithmetic: Exploiting Randomness in Floating-Point Arithmetic
Douglass Stott Parker. 1997.Monte Carlo Arithmetic: Exploiting Randomness in Floating-Point Arithmetic. Technical Report CSD-970002. UCLA Computer Science Department
1997
-
[47]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, 32 Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, H...
2019
-
[48]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research12 (2011), 2825–2830
2011
-
[49]
Inés Gonzalez Pepe, Hiba Akhaddar, Tristan Glatard, and Yohan Chatelain. 2026. Fuzzy PyTorch: Rapid Numerical Variability Evaluation for Deep Learning Models.Transactions on Machine Learning Research(2026). https://openreview.net/forum?id=0ogq232VGP
2026
-
[50]
Haiquan Qiu and Quanming Yao. 2026. Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention. InThe 114th International Conference on Learning Representations. arXiv:2510.04212 [cs.LG] https://openreview.net/forum?id=0jHyEKHDyx
Pith/arXiv arXiv 2026
-
[51]
Reyna Cruz, Kristalys Ruiz-Rohena, Yahriel I
Maria L. Reyna Cruz, Kristalys Ruiz-Rohena, Yahriel I. Guel, Henry Salgado, Lisa Taldir, Elian Pena Ramos, Natalia Cervantes, Tzetzaith Rivero, Martine Ceberio, Christoph Lauter, and Anastasia Volkova. 2025. Contribution to Error Analysis of Deep Neural Networks: Case of the Activation Functions. (2025). https://inria.hal.science/hal-05367563 working pape...
2025
-
[52]
Jackson Vanover, Alper Altuntas, and Cindy Rubio-González. 2024. Toward Automated Precision Tuning of Weather and Climate Models: A Case Study. InSC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. 148–159. doi:10.1109/SCW63240.2024.00026
arXiv 2024
-
[53]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems, Vol. 30
2017
-
[54]
Jean Vignes. 1993. A stochastic arithmetic for reliable scientific computation.Mathematics and Computers in Simulation35, 3 (1993), 233–261. doi:10.1016/0378-4754(93)90003-D
-
[55]
Jean Vignes. 2004. Discrete Stochastic Arithmetic for Validating Results of Numerical Software.Numerical Algorithms37, 1 (2004), 377–390. doi:10.1023/B:NUMA.0000049483.75679.ce
arXiv 2004
-
[56]
La Porte
Jean Vignes and M. La Porte. 1974. Error Analysis in Computing. InInformation Processing: Proceedings of the 6th IFIP Congress. North-Holland, 610–614
1974
-
[57]
Zhongzhen Wen, Hongyu Liu, Tingwei Zhu, Minxue Pan, Shaohua Wang, Yuanyi Lin, Kairui Liu, Tian Zhang, and Xuandong Li. 2025. A Study of Floating-Point Precision Tuning in Deep Learning Operators Implementations.ACM Trans. Softw. Eng. Methodol.(2025). doi:10.1145/3773992
-
[58]
Morcos, Ali Farhadi, and Ludwig Schmidt
Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari S. Morcos, Ali Farhadi, and Ludwig Schmidt. 2023. Stable and Low-Precision Training for Large-Scale Vision-Language Models. InAdvances in Neural Information Processing Systems, Vol. 36. Curran Associates, Inc., 10271–10298. https://proceedings.neurips.cc/paper_files/paper/2023/hash/20bd42d82998bc61732...
2023
-
[59]
Han Xiao, Kashif Rasul, and Roland Vollgraf. 2017. Fashion-MNIST: A Novel Image Dataset for Benchmarking Machine Learning Algorithms. arXiv:1708.07747 [cs.LG]
Pith/arXiv arXiv 2017
-
[60]
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level Convolutional Networks for Text Classification. InAdvances in Neural Information Processing Systems, Vol. 28. 33
2015
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.