REVIEW 3 major objections 5 minor 78 references
Across 16 DNN workloads and three IEEE 754 formats, the lowest fraction bits can be flipped without crossing a 1% task-quality bound, and this threshold can be turned into a cheaper unequal-error-protection memory design.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:13 UTC pith:YNS3ZPBP
load-bearing objection Solid empirical bit-sensitivity study, but the BF16 UEP design uses X=5 while the paper's own floor is X=4, undercutting the headline hardware savings. the 3 major comments →
From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Flipping any of the least-significant fraction bits up to a data-type-specific threshold, Xsafe, degrades task metrics by under 1% in deterministic single-bit stress tests; sensitivity then ramps through the upper fraction bits and spikes at the exponent-mantissa boundary, where a single-bit flip is catastrophic. The paper argues this transition is a property of the IEEE 754 data type, not of any particular model architecture: attention-free CNNs show the same pattern as transformers, and statistical tests place significance only at exponent and sign positions. On that basis it claims conservative floors FP16 Xsafe=6, BF16 Xsafe=4, FP32 Xsafe=15 are stable across the 16 evaluated workloads,
What carries the argument
The central object is Xsafe, the number of least-significant fraction bits below which flips are benign; it is read off the sharp 'cliff' in per-bit-position sensitivity curves where quality holds flat then collapses near the exponent boundary. The carrying mechanism is a three-tier unequal-error-protection codec: sign/exponent bits get full SECDED (odd-parity codes), upper fraction bits get single-error correction (even-parity codes), and the Xsafe low bits are stored with no code, physically separated into a low-voltage SRAM partition selected by a per-cacheline 3-bit tag.
Load-bearing premise
The argument rests on assuming that measuring sensitivity at one injection site per model—pre-softmax QK^T scores for attention models and FFN/MLP outputs otherwise—reveals the most fault-sensitive tensor; if some other tensor (embeddings, normalization statistics, residual stream, or an untested layer) is more sensitive at low fraction bits, the Xsafe floor would not be a true worst-case bound.
What would settle it
Run the deterministic single-bit-position sweep (pword = 1, every word's bit toggled) on every tensor of a text-conditioned diffusion model—embeddings, normalization statistics, residual stream, and attention outputs—in BF16, and check whether any fraction bit at or below bit 3 degrades PSNR by more than the allowed bound (first bit before a single-step PSNR drop > 5 dB); if such a bit exists, the BF16 floor of 4 is not worst-case.
If this is right
- Bit-protection thresholds can be fixed in hardware at design time because they vary by at most one bit across datasets, injection sites, and sample counts.
- A single 136-bit physical word and one codec serve FP16, BF16, and FP32; only the per-cacheline tag changes routing.
- Multi-cell upsets that straddle the Xsafe boundary contract the safe region by at most one bit per additional flipped bit, so the floor remains usable under MCU-2 and MCU-3 faults.
- Leaving the low partition unprotected at reduced voltage yields ~17% gross BF16 read-energy savings with a ~4% macro-area overhead recovered by the energy savings, and the most aggressive dual-voltage operating point keeps MTTF around 900 years.
- Workload-aware tiers widen the unprotected region (up to X=10 for FP16) for vision, NLU, and resilient LLM workloads, boosting ECC savings to 37.5–62.5% without retraining.
Where Pith is reading between the lines
- The paper leaves implicit that the same bit-sensitivity logic applies to weights once protection cost is amortized over a load-once lifetime; since it notes weight LSBs are tolerant, a weight-side UEP is a natural extension.
- The per-layer sensitivity spreads, which exceed an order of magnitude across LLM layers, suggest that per-tensor protection tiers could reclaim more energy than the per-model tiers reported here.
- Because Xsafe is defined against task-specific quality bounds, stricter safety contracts would tighten the floors monotonically; applications with hard quality requirements should treat the quoted floors as ceilings, not defaults.
- The codec's reduced double-error detection means a safety-critical deployment should pay the extra check bit to restore full SECDED; the paper enables this as an option, making the throughput-versus-safety tradeoff a policy decision.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper characterizes per-bit-position fault sensitivity in FP16/BF16/FP32 values across 16 DNN workloads, reports a sharp transition between benign low-order fraction bits and catastrophic exponent-bit faults, and uses the observed threshold Xsafe (FP16=6, BF16=4, FP32=15) to motivate a three-tier Unequal Error Protection (UEP) codec and a dual-partition SRAM. The claimed contributions are: a data-type-level Xsafe floor stable across workloads; workload-aware protection tiers; an RTL-synthesized UEP codec with per-cacheline format tags; validation under random, persistent, and multi-bit correlated upsets; and area/energy estimates showing a 27.8% codec-area reduction and roughly 17% gross BF16 read-energy savings.
Significance. If the empirical floors and their generality hold, this is a practically relevant result for ML inference memories: a fixed hardware partition, a small per-cacheline tag, and no retraining could make selective ECC deployable across FP16/BF16/FP32 accelerators. The paper has genuine strengths: 16 models across 6 modalities, deterministic sweeps complemented by stochastic, persistent, and MCU-2/3 campaigns, statistical tests (McNemar, Wilcoxon), RTL verification with 146k+ vectors, and an honest discussion of limitations in §7. However, the load-bearing generalization from a few injection sites to all tensors, and the internal mismatch between the certified Xsafe floor and the implemented RTL bypass width, must be resolved before the hardware claims are credible.
major comments (3)
- [§3.3, §4.1.1, Table 5, Table 6] Xsafe is defined as the number of fraction LSBs that can safely remain unprotected. Under that definition, for FP16 (10 fraction bits), BF16 (7 fraction bits), and FP32 (23 fraction bits), Xsafe cannot exceed the fraction width. Table 5 lists FP16 values 11, 12, 14 (ViT-L/16, BERT-base, Swin-B), BF16 values 8 (BERT, ViT, Swin), and FP32 values 24, 25, 27, all of which exceed the corresponding fraction widths. Table 6 deepens the contradiction: for ViT-L/16 it reports Xsafe=11 but labels b10 as the first bit above Xsafe, and b11 (which should be below the threshold) already shows a 3.94pp top-1 drop, crossing the 1pp bound. Since Table 5 feeds the workload tiers in Table 8, the tier boundaries inherit this inconsistency. The authors should correct the per-model Xsafe values, clarify the bit indexing used, or revise the definition so that Table 5 and Table 6 are mutually consistent.
- [§5.1.5, Table 11, Table 14, §6.4] The hardware implementation and validation do not use the certified BF16 floor Xsafe=4. Table 11 and Table 14 size BF16 Partition B as 40 data bits (5×8 words), i.e., X=5 unprotected fraction LSBs, and §6.4 confirms the RTL operating point is k=11 for BF16 with '5 unprotected LSBs, bits 0–4.' The text calls this a conservative boundary distinct from Xsafe, but the paper's headline energy and area claims (17% gross BF16 read-energy reduction, 27.8% codec-area reduction) and the end-to-end task validation in Tables 20–21 are all computed at this wider X=5 point. At the actual floor X=4, Partition B would be 32 bits/word, f_nc would drop from 40/136 to 32/136, and the energy saving would shrink accordingly. The paper needs to either validate and justify X=5 as its own design point, or re-derive the hardware claims and validation at Xsafe=4.
- [§3.3, Table 4, §4.1.1] The claim that Xsafe is a data-type-level property, not a model artifact, is load-bearing for the hardware design, because a single floor is applied to all inference-produced values. However, per §3.3, the primary injection site is the pre-Softmax QK^T score tensor for attention models and FFN/MLP outputs otherwise, with additional loci tested only 'for selected models when a within-model sensitivity ranking is needed.' Table 4 reports at most a one-bit shift across injection sites, but this axis is exercised on a subset of models only. The floor-setting models (e.g., SD3.5-Medium, PixArt-α) are not shown with per-tensor Xsafe sweeps covering residual streams, normalization statistics, embeddings, or KV-cache tensors. The authors should report per-site sweeps for at least the floor-setting models, or temper the universality conclusion.
minor comments (5)
- [Table 8] The tier table is hard to read and appears internally inconsistent. The 'Models in tier' entries (4, 7, 16) seem to mix non-cumulative and cumulative counts, and Tier 1 at X=9 cannot include 'all below' if models with Xsafe=6 are in the same list. Please recast the table so coverage counts and membership are unambiguous.
- [§4.1.1, Table 6] The prose says Llama-3 8B perplexity 'remains stable through bit 11' while Table 6 reports 7.01 at b11 against a baseline of 6.64, which is already above the 1% bound used to define Xsafe. Please align the narrative with the table values.
- [§3.2, §4.1.3] The C4/HellaSwag external-validity check for Llama-3 8B is mentioned in §3.2 but no result is reported. Either add the data or remove the claim.
- [§5.3.2, Eq. (3)–(4)] The read-energy saving is reported as 'gross' and 'net positive after level shifters,' but the dynamic/leakage split and the read vs. write energy mix are not stated. Clarifying the energy model would make the 17% figure easier to reproduce.
- [§6.4] The RTL operating points k=11/12/24 are described as 'conservative boundaries distinct from the Xsafe floors,' but the relationship between k and Xsafe is never formalized. A short derivation or table mapping k to the number of unprotected LSBs per format would remove ambiguity.
Circularity Check
No significant circularity: Xsafe thresholds are empirically calibrated and cross-validated under independent fault patterns; the BF16 X=5 vs Xsafe=4 mismatch is an internal consistency issue, not circularity.
full rationale
The paper's central claims are empirical rather than deductive. Xsafe is defined in §3.3 as the number of fraction LSBs that can remain unprotected while staying within a modality-specific quality bound, and then measured in §4.1.1 via deterministic single-bit-position sweeps. The later validations (§4.1.2–§4.1.5, §6.4) use different fault families—persistent stuck-at faults, stochastic BER sweeps, MCU-2/3, and tiered-BER end-to-end runs—so the safety claim is cross-validated, not obtained by restating the calibration. The hardware equations (§5.1.5, §5.3) are arithmetic consequences of the chosen code parameters, and the ECC area/energy numbers follow from the codeword construction rather than from the sensitivity data alone. No author-overlap citations are invoked as load-bearing uniqueness arguments, and no ansatz is imported via self-citation. The one notable issue is an internal inconsistency: the BF16 RTL operating point k=11 exposes 5 unprotected LSBs (§6.4) while the certified floor is Xsafe=4 (Table 5), and the headline BF16 energy saving is computed at f_nc=40/136 (§5.3.2). This is a correctness/consistency concern about whether the evaluated configuration matches the certified floor, not a circularity: the k=11 experiment is an independent (if mismatched) measurement, not a reduction of the conclusion to its premise.
Axiom & Free-Parameter Ledger
free parameters (5)
- Xsafe floors (FP16/BF16/FP32) =
6 / 4 / 15
- FP16 workload tiers X =
6 / 9 / 10
- Modality quality bounds =
PPL≤1.01x; acc≥-1pp; WER≤+1pp; first bit before PSNR drop >5dB
- Low-voltage partition operating point Vlow =
1.2 V (sky130, Vnom=1.8 V)
- RTL validation boundaries k =
BF16 k=11 (5 unprotected LSBs); FP16 k=12 (4); FP32 k=24 (8)
axioms (5)
- domain assumption The measured injection loci (QK^T scores for attention; FFN/MLP outputs otherwise) are worst-case representatives for all tensors in each model.
- domain assumption Deterministic single-bit stress at p_word=1 plus random/MCU sweeps adequately captures SRAM soft-error behavior relevant to the design.
- standard math The odd/even parity-class capacity split of Hsiao SECDED supports Tier 1/Tier 2 code assignment as encoded in Eqs. (1)-(2).
- domain assumption SRAM bit-error rate at Vlow follows the Stutz mapping, and CACTI 7.0 plus published 7nm scaling equations give valid area/energy projections.
- domain assumption The terrestrial-neutron SER model (1000 FIT/Mbit) and lambda_SEU=1e-12/bit/hr are valid for the FIT/MTTF calculations.
invented entities (2)
-
Three-tier UEP codec (Tier 1 SECDED / Tier 2 SEC / Tier 3 bypass)
no independent evidence
-
Dual-partition SRAM with per-cacheline format tag
no independent evidence
Cite this review
Pith. "Pith review of From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory." pith.science (2026). https://pith.science/paper/YNS3ZPBP
@misc{pith2026260719623,
author = {Pith},
title = {Pith review of: From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory},
year = {2026},
howpublished = {\url{https://pith.science/paper/YNS3ZPBP}},
note = {Machine review of arXiv:2607.19623}
}
read the original abstract
We characterize per-bit-position fault sensitivity in ML inference across 16 workloads -- spanning transformer-based models and attention-free CNNs -- and across three floating-point formats. Our central empirical finding is a sharp bit-sensitivity transition: flipping any of the least-significant fraction bits up to a data-type-specific threshold, Xsafe, degrades task metrics by less than 1% under deterministic single-bit stress tests. Sensitivity rises through the upper fraction bits and spikes at the exponent-mantissa boundary, where a single-bit flip causes catastrophic collapse. Because low-order bits are largely inconsequential while high-order and exponent bits are critical, uniform SECDED protection -- which guards every bit equally at 12.5% storage overhead -- is unnecessarily conservative. We derive per-data-type Xsafe floors (FP16: 6, BF16: 4, FP32: 15) and workload-aware tiers that widen the unprotected region for resilient model classes, raising ECC savings to 37.5-62.5% without retraining. Text-conditioned diffusion models dictate the conservative floor; vision encoders, NLU models, and resilient LLMs tolerate wider bypass regions. These floors and tiers drive an Unequal Error Protection (UEP) codec with per-cacheline data-type tags and a dual-partition SRAM architecture for ML accelerators. Validation across 870+ fault-injection runs confirms selective protection holds under contiguous 2- and 3-bit upsets. The codec reduces ECC area by 27.8% relative to uniform SECDED; dual-voltage operation of the non-critical partition lowers gross BF16 read energy by about 17%, with a roughly 4% dual-partition macro-area overhead.
Figures
Reference graph
Works this paper leans on
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, et al. 2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. arXiv:2404.14219 https://arxiv.org/abs/2404.14219
Pith/arXiv arXiv 2024
-
[2]
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 4895–4901. doi:10.18653/v1/2023.emnlp-main.298 12 From Bit-Position Sen...
-
[3]
Irina Alam and Puneet Gupta. 2022. COMET: On-Die and In-Controller Collabora- tive Memory ECC Technique for Safer and Stronger Correction of DRAM Errors. InDSN. IEEE, Baltimore, MD, USA, 124–136. doi:10.1109/DSN53405.2022.00024
arXiv 2022
-
[4]
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, et al . 2023. The Falcon Series of Open Language Models. arXiv:2311.16867 https://arxiv.org/abs/ 2311.16867
Pith/arXiv arXiv 2023
-
[5]
1999.Design and Analysis of Fast Low-Power SRAMs
Bharadwaj S Amrutur. 1999.Design and Analysis of Fast Low-Power SRAMs. Stanford University
1999
-
[6]
2019.Armv8.5-A Memory Tagging Extension White Paper
Arm. 2019.Armv8.5-A Memory Tagging Extension White Paper. Arm. https://developer.arm.com/-/media/Arm%20Developer%20Community/PDF/ Arm_Memory_Tagging_Extension_Whitepaper.pdf
2019
-
[7]
Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas
Rajeev Balasubramonian, Andrew B. Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas. 2017. CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories.ACM Trans. Archit. Code Optim.14, 2 (2017), 14:1–14:25. doi:10.1145/3085572
doi:10.1145/3085572 2017
-
[8]
Robert C. Baumann. 2005. Radiation-Induced Soft Errors in Advanced Semicon- ductor Technologies.IEEE Transactions on Device and Materials Reliability5, 3 (2005), 305–316. doi:10.1109/TDMR.2005.853449
arXiv 2005
-
[9]
Izan Catalán, José Flich, and Carles Hernández. 2025. Exploiting Neural Networks Bit-Level Redundancy to Mitigate the Impact of Faults at Inference.The Journal of Supercomputing81 (2025), 183. doi:10.1007/s11227-024-06693-7
-
[10]
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhong- dao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. 2023. PixArt-𝛼: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthe- sis. arXiv:2310.00426 https://arxiv.org/abs/2310.00426
Pith/arXiv arXiv 2023
-
[11]
C. J. Clopper and E. S. Pearson. 1934. The Use of Confidence or Fiducial Limits Illustrated in the Case of the Binomial.Biometrika26, 4 (1934), 404–413. https: //barestatistics.nl/uploads/1/1/7/9/11797954/clopper__pearson_1934.pdf
1934
-
[12]
Tri Dao and Albert Gu. 2025. FlashAttention-4: Hardware-Friendly Fused At- tention with Warp-Specialization and Pingpong Scheduling.arXiv preprint arXiv:2603.05451(2025)
arXiv 2025
-
[13]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InNAACL-HLT. https://arxiv.org/abs/1810.04805
Pith/arXiv arXiv 2019
-
[14]
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 1286–...
2021
-
[15]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. InICLR. https://arxiv.org/abs/2010.11929
Pith/arXiv arXiv 2021
-
[16]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. 2024. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. InProceedings of the 41st International Conference on Machine Learning...
2024
-
[17]
David Goldberg. 1991. What Every Computer Scientist Should Know About Floating-Point Arithmetic.Comput. Surveys23, 1 (1991), 5–48. doi:10.1145/ 103162.103163
arXiv 1991
-
[18]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 https://arxiv.org/abs/2407.21783
Pith/arXiv arXiv 2024
-
[19]
Taeuk Ha, Gyuri Kim, Chanki Kim, and Sang-Hyo Kim. 2025. Chipkill-Level ECC Using 4-Bit Symbol Reed–Solomon Codes for DDR5 DRAM. InICTC. IEEE, Jeju, South Korea, 371–375. doi:10.1109/ICTC66702.2025.11387968
arXiv 2025
-
[20]
Richard W. Hamming. 1950. Error Detecting and Error Correcting Codes.Bell System Technical Journal29, 2 (1950), 147–160. doi:10.1002/j.1538-7305.1950. tb00463.x
arXiv 1950
-
[21]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InCVPR. http://openaccess.thecvf.com/content_ cvpr_2016/html/He_Deep_Residual_Learning_CVPR_2016_paper.html
2016
-
[22]
Sanghyun Hong, Pietro Frigo, Yigitcan Kaya, Cristiano Giuffrida, and Tudor Dumitras. 2019. Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault Attacks. InUSENIX Security Symposium. https://www.usenix.org/conference/usenixsecurity19/presentation/ hong
2019
-
[23]
Le, and Hartwig Adam
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasude- van, Quoc V. Le, and Hartwig Adam. 2019. Searching for MobileNetV3. InICCV. http://openaccess.thecvf.com/content_ICCV_2019/html/Howard_ Searching_for_MobileNetV3_ICCV_2019_paper.html
2019
-
[24]
M. Y. Hsiao. 1970. A Class of Optimal Minimum Odd-Weight-Column SEC- DED Codes.IBM Journal of Research and Development14, 4 (1970), 395–401. doi:10.1147/rd.144.0395
-
[25]
Ibe, Hitoshi Taniguchi, Yasuo Yahagi, Ken-ichi Shimbo, and Tadanobu Toba
Eishi H. Ibe, Hitoshi Taniguchi, Yasuo Yahagi, Ken-ichi Shimbo, and Tadanobu Toba. 2010. Impact of Scaling on Neutron-Induced Soft Error in SRAMs From a 250 nm to a 22 nm Design Rule.IEEE Transactions on Electron Devices57, 7 (2010), 1527–1538. doi:10.1109/TED.2010.2047907
arXiv 2010
-
[26]
2019.IEEE Standard for Floating-Point Arithmetic
IEEE Computer Society. 2019.IEEE Standard for Floating-Point Arithmetic. Insti- tute of Electrical and Electronics Engineers, New York, NY, USA. doi:10.1109/ IEEESTD.2019.8766229 IEEE Std 754-2019 (Revision of IEEE 754-2008)
arXiv 2019
-
[27]
Jiang, Alexandre Sablayrolles, Arthur Mensch, et al
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, et al. 2023. Mistral 7B. arXiv:2310.06825 https://arxiv.org/abs/2310.06825
Pith/arXiv arXiv 2023
-
[28]
Vinay Joshi, Manuel Le Gallo, Simon Haefeli, Irem Boybat, S. R. Nandakumar, Christophe Piveteau, Martino Dazzi, Bipin Rajendran, Abu Sebastian, and Evan- gelos Eleftheriou. 2020. Accurate Deep Neural Network Inference Using Com- putational Phase-Change Memory.Nature Communications11 (2020), 2473. doi:10.1038/s41467-020-16108-9
-
[29]
Sullivan, Seong-Lyong Gong, and Mattan Erez
Jungrae Kim, Michael B. Sullivan, Seong-Lyong Gong, and Mattan Erez. 2015. Frugal ECC: Efficient and Versatile Memory Error Protection Through Fine- Grained Compression. InSC. ACM, Austin, TX, USA, 12:1–12:12. doi:10.1145/ 2807591.2807659
arXiv 2015
-
[30]
Tae Hyun Kim, Hanwool Jeong, Juhyun Park, Hoonki Kim, Taejoong Song, and Seong-Ook Jung. 2020. An Embedded Level-Shifting Dual-Rail SRAM for High- Speed and Low-Power Cache.IEEE Access8 (2020), 187126–187139. doi:10.1109/ ACCESS.2020.3030099
arXiv 2020
-
[31]
Giray Yağlıkçı, Roknoddin Azizi, Taha Shahroodi, Konstantinos Kanellopoulos, and Onur Mutlu
Skanda Koppula, Lois Orosa, A. Giray Yağlıkçı, Roknoddin Azizi, Taha Shahroodi, Konstantinos Kanellopoulos, and Onur Mutlu. 2019. EDEN: Enabling Energy- Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM. InMICRO. ACM, Columbus, OH, USA, 218–231. doi:10.1145/3352460. 3358280
doi:10.1145/3352460 2019
-
[32]
2023.Unequal Error Protection Source–Channel Coding
Andres Kwasinski and Vinay Chande. 2023.Unequal Error Protection Source–Channel Coding. 135–172. doi:10.1002/9781118693803.ch5
-
[33]
Young Seo Lee, Gunjae Koo, Young-Ho Gong, and Sung Woo Chung. 2022. Stealth ECC: A Data-Width Aware Adaptive ECC Scheme for DRAM Error Resilience. In DATE. IEEE, Antwerp, Belgium, 382–387. doi:10.23919/DATE54114.2022.9774775
arXiv 2022
-
[34]
Daiqing Li, Aleks Kamko, Ehsan Akhgari, Ali Sabet, Linmiao Xu, and Suhail Doshi
-
[35]
Guanpeng Li, Siva Kumar Sastry Hari, Michael Sullivan, Timothy Tsai, Karthik Pattabiraman, Joel Emer, and Stephen W. Keckler. 2017. Understanding Error Propagation in Deep Learning Neural Network (DNN) Accelerators and Applica- tions. InSC. Denver, CO, USA. doi:10.1145/3126908.3126964
arXiv 2017
-
[36]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. InICCV. https://arxiv.org/abs/2103.14030
Pith/arXiv arXiv 2021
-
[37]
2014.Tagged Memory and Minion Cores in the lowRISC SoC
lowRISC. 2014.Tagged Memory and Minion Cores in the lowRISC SoC. Technical Report memo-2014-001. lowRISC. https://lowrisc.org/docs/memo-2014-001- tagged-memory-and-minion-cores/
2014
-
[38]
Mathew, Martin Schultheis, Carl C
Deepak M. Mathew, Martin Schultheis, Carl C. Rheinländer, Chirag Sudarshan, Christian Weis, Norbert Wehn, and Matthias Jung. 2018. An Analysis on Reten- tion Error Behavior and Power Consumption of Recent DDR4 DRAMs. InDATE. IEEE, Dresden, Germany, 293–296. doi:10.23919/DATE.2018.8342023
arXiv 2018
-
[39]
Quinn McNemar. 1947. Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages.Psychometrika12, 2 (1947), 153–157. doi:10.1007/BF02295996
-
[40]
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer Sentinel Mixture Models. InICLR. https://openreview.net/forum?id= Byj72udxe¬eId=Byj72udxe
2017
-
[41]
Justin Meza, Qiang Wu, Sanjeev Kumar, and Onur Mutlu. 2015. Revisiting Mem- ory Errors in Large-Scale Production Data Centers: Analysis and Modeling of New Trends from the Field. InDSN. https://research.facebook.com/publications/ revisiting-memory-errors-in-large-scale-production-data-centers-analysis- and-modeling-of-new-trends-from-the-field/
2015
-
[42]
Vishesh Mishra, Marcello Traiola, Angeliki Kritikakou, Olivier Sentieys, and Urbi Chatterjee. 2025. SERA-Float: A Soft Error Resilient Approximate Floating-Point Computing Format. InICCAD. IEEE, Munich, Germany, 1–9. https://inria.hal. science/hal-05333255v1/file/SERA_FLOAT_v2.pdf
2025
-
[43]
Modal Team. 2025. Reverse-Engineering FlashAttention-4. https://modal.com/ blog/reverse-engineer-flash-attention-4
2025
-
[44]
Sanches, Kohji Hosokawa, Scott C
Pritish Narayanan, Alessandro Fumarola, Lucas L. Sanches, Kohji Hosokawa, Scott C. Lewis, Robert M. Shelby, and Geoffrey W. Burr. 2017. To- ward On-Chip Acceleration of the Backpropagation Algorithm Using Non- volatile Memory.IBM Journal of Research and Development61, 4/5 (2017). https://research.ibm.com/publications/toward-on-chip-acceleration- of-the-ba...
2017
-
[45]
Duy Thanh Nguyen, Hyun Kim, Hyuk-Jae Lee, and Ik Joon Chang. 2018. An Approximate Memory Architecture for a Reduction of Refresh Power Con- sumption in Deep Learning Applications. InISCAS. IEEE, Florence, Italy, 1–5. doi:10.1109/ISCAS.2018.8351021
arXiv 2018
-
[46]
Byoungchan Oh, Nilmini Abeyratne, Jeongseob Ahn, Ronald G. Dreslinski, and Trevor Mudge. 2016. Enhancing DRAM Self-Refresh for Idle Power Reduction. 13 Muhammad Husnain Mubarik, Karthik Mohan Kumar, Pedro Antonio Pena, Keshavan Varadarajan, Kunal Tyagi InISLPED. ACM, San Francisco, CA, USA, 254–259. doi:10.1145/2934583.2934632
arXiv 2016
-
[47]
Palframan, Nam Sung Kim, and Mikko H
David J. Palframan, Nam Sung Kim, and Mikko H. Lipasti. 2014. Precision- Aware Soft Error Protection for GPUs. InHPCA. IEEE, Orlando, FL, USA, 49–59. doi:10.1109/HPCA.2014.6835966
arXiv 2014
-
[48]
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. LibriSpeech: An ASR Corpus Based on Public Domain Audio Books. InICASSP. https://us.openslr.org/resources/12/about.html
2015
-
[49]
William Peebles and Saining Xie. 2023. Scalable Diffusion Models with Trans- formers. InICCV. https://www.wpeebles.com/DiT.html
2023
-
[50]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv:2307.01952 https: //arxiv.org/abs/2307.01952
Pith/arXiv arXiv 2023
-
[51]
Qureshi, Dae-Hyun Kim, Samira Manabi Khan, Prashant J
Moinuddin K. Qureshi, Dae-Hyun Kim, Samira Manabi Khan, Prashant J. Nair, and Onur Mutlu. 2015. AVATAR: A Variable-Retention-Time (VRT) Aware Refresh for DRAM Systems. InDSN. IEEE, Rio de Janeiro, Brazil, 427–437. doi:10.1109/ DSN.2015.58
2015
-
[52]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356 https://arxiv.org/abs/2212.04356
Pith/arXiv arXiv 2022
-
[53]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. OpenAI technical report. https://cdn.openai.com/better-language-models/language_ models_are_unsupervised_multitask_learners.pdf
2019
-
[54]
Brandon Reagen, Udit Gupta, Lillian Pentecost, Paul Whatmough, Sae Kyu Lee, Niamh Mulholland, David Brooks, and Gu-Yeon Wei. 2018. Ares: A Framework for Quantifying the Resilience of Deep Neural Networks. InDAC. ACM, San Francisco, CA, USA, 17:1–17:6. doi:10.1145/3195970.3195997
arXiv 2018
-
[55]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. InCVPR. https://arxiv.org/abs/2112.10752
Pith/arXiv arXiv 2022
-
[56]
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexan- der C. Berg, and Li Fei-Fei. 2015. ImageNet Large Scale Visual Recognition Challenge.International Journal of Computer Vision115, 3 (2015), 211–252. doi:10.1007/s11263-015-0816-y
-
[57]
Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. 2024. Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation. arXiv:2403.12015 https://arxiv.org/abs/2403. 12015
Pith/arXiv arXiv 2024
-
[58]
Bianca Schroeder, Eduardo Pinheiro, and Wolf-Dietrich Weber. 2009. DRAM Errors in the Wild: A Large-Scale Field Study. InSIGMETRICS. https://research. google/pubs/dram-errors-in-the-wild-a-large-scale-field-study/
2009
-
[59]
Noam Shazeer. 2019. Fast Transformer Decoding: One Write-Head is All You Need. arXiv:1911.02150 https://arxiv.org/abs/1911.02150
Pith/arXiv arXiv 2019
-
[60]
Manning, Andrew Ng, and Christopher Potts
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive Deep Models for Seman- tic Compositionality Over a Sentiment Treebank. InEMNLP. Association for Computational Linguistics, 1631–1642. https://aclanthology.org/D13-1170/
2013
-
[61]
Vilas Sridharan, Nathan Debardeleben, Sean Blanchard, Kurt Ferreira, Sudhanva Gurumurthi, and John Shalf. 2015. Memory Errors in Modern Systems: The Good, The Bad, and The Ugly. InASPLOS. doi:10.1145/2694344.2694348
arXiv 2015
-
[62]
Aaron Stillmaker and Bevan M. Baas. 2017. Scaling Equations for the Accu- rate Prediction of CMOS Device Performance from 180 nm to 7 nm.Integra- tion, the VLSI Journal58 (2017), 74–81. http://vcl.ece.ucdavis.edu/pubs/2017.02. VLSIintegration.TechScale/
2017
-
[63]
David Stutz, Nandhini Chandramoorthy, Matthias Hein, and Bernt Schiele. 2021. Bit Error Robustness for Energy-Efficient DNN Accelerators. InProceedings of MLSys. MLSys, Virtual, 1–21. https://proceedings.mlsys.org/paper_files/paper/ 2021/file/2f9b1b6b29361118f630783c19891ea0-Paper.pdf
2021
-
[64]
Sullivan, Mohamed Tarek Ibn Ziad, Aamer Jaleel, and Stephen W
Michael B. Sullivan, Mohamed Tarek Ibn Ziad, Aamer Jaleel, and Stephen W. Keckler. 2023. Implicit Memory Tagging: No-Overhead Memory Safety Using Alias-Free Tagged ECC. InISCA. ACM, Orlando, FL, USA, 1–14. doi:10.1145/ 3579371.3589079
arXiv 2023
-
[65]
Michael B. Sullivan, Nirmal Saxena, Mike O’Connor, Donghyuk Lee, Paul Racu- nas, Saurabh Hukerikar, Timothy Tsai, Siva Hari, and Stephen W. Keckler. 2021. Characterizing and Mitigating Soft Errors in GPU DRAM. InMICRO. ACM. doi:10.1145/3466752.3480111
arXiv 2021
-
[66]
Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. InICML. https://research.google/pubs/ efficientnet-rethinking-model-scaling-for-convolutional-neural-networks/
2019
-
[67]
Hoyoung Tang and Jongsun Park. 2016. Unequal-Error-Protection Error Cor- rection Codes for the Embedded Memories in Digital Signal Processors.IEEE Transactions on Very Large Scale Integration (VLSI) Systems24, 6 (2016), 2397–2401. doi:10.1109/TVLSI.2015.2497368
arXiv 2016
-
[68]
Kim, Juan Gómez Luna, Mohammad Sadrosadati, Nika Mansouri Ghiasi, and Onur Mutlu
Yaohua Wang, Lois Orosa, Xiangjun Peng, Yang Guo, Saugata Ghose, Minesh Patel, Jeremie S. Kim, Juan Gómez Luna, Mohammad Sadrosadati, Nika Mansouri Ghiasi, and Onur Mutlu. 2020. FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and Caching. InMICRO. IEEE, Virtual (originally Athens, Greece), 235–249. https://ieeexplore.ieee.org...
2020
-
[69]
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004. Image Quality Assessment: From Error Visibility to Structural Similarity.IEEE Transactions on Image Processing13, 4 (2004), 600–612. doi:10.1109/TIP.2003. 819861
doi:10.1109/tip.2003 2004
-
[70]
Frank Wilcoxon. 1945. Individual Comparisons by Ranking Methods.Biomet- rics Bulletin1, 6 (1945), 80–83. https://sci2s.ugr.es/keel/pdf/algorithm/articulo/ wilcoxon1945.pdf
1945
-
[71]
Rui Xie, Yunhua Fang, Asad Ul Haq, Linsen Ma, Sanchari Sen, Swagath Venkatara- mani, Liu Liu, and Tong Zhang. 2025. Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference. arXiv:2512.18152 doi:10.48550/arXiv. 2512.18152
-
[72]
Rui Xie, Asad Ul Haq, Yunhua Fang, Linsen Ma, Sanchari Sen, Swagath Venkatara- mani, Liu Liu, and Tong Zhang. 2025. Breaking the HBM Bit Cost Bar- rier: Domain-Specific ECC for AI Inference Infrastructure. arXiv:2507.02654 doi:10.48550/arXiv.2507.02654
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2507.02654 2025
-
[73]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence?. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 4791–4800. doi:10.18653/v1/P19-1472
-
[74]
Jeff Zhang, Kartheek Rangineni, Zahra Ghodsi, and Siddharth Garg. 2018. ThUn- derVolt: Enabling Aggressive Voltage Underscaling and Timing Error Resilience for Energy Efficient Deep Learning Accelerators. InDAC. ACM, San Francisco, CA, USA, 19:1–19:6. doi:10.1145/3195970.3196129
arXiv 2018
-
[75]
Efros, Eli Shechtman, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang
-
[76]
Ruohuang Zheng and Michael C. Huang. 2017. Redundant Memory Array Architecture for Efficient Selective Protection. InISCA. ACM, Toronto, ON, Canada, 214–227. doi:10.1145/3079856.3080213 14
arXiv 2017
-
[2018]
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InCVPR. http://richzhang.github.io/PerceptualSimilarity/
-
[2024]
arXiv:2402.17245 https://arxiv.org/abs/2402.17245
Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation. arXiv:2402.17245 https://arxiv.org/abs/2402.17245
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.