REVIEW 4 major objections 5 minor 89 references
The paper claims a lightweight probe, sketch, and topology-ranking pipeline can detect fail-slow cores and links in many-core DNN accelerators, reducing trace storage by 115.9x while keeping detection accuracy near 87% at a 12% false-positi
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A simulation-based framework using compiler-inserted probes, a two-stage sketch, and a PageRank-style ranking detects on-chip fail-slow cores/links at ~86.8% accuracy with ~116x trace compression.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection A plausible integrated design for on-chip fail-slow detection that is undercut by a stylized synthetic evaluation; worth a serious referee, but the headline accuracy numbers shouldn't be taken at face value. the 4 major comments →
SLOTH: Lightweight Detection and Localization of On-Chip Fail-Slow Failures for DNN Accelerators
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that a hardware-aware monitoring pipeline composed of (1) compiler-inserted probe fragments that record compute and communication events, (2) a two-stage Fail-Slow Sketch that retains frequently recurring trace patterns while filtering noise, and (3) a multi-level communication graph that maps software dependencies onto the physical mesh and ranks candidates with the FailRank algorithm, can localize fail-slow cores and links accurately enough for practice under strict SRAM limits. The authors demonstrate this in a simulator with five workloads (four DNNs plus a binary-tree benchmark), reporting per-workload accuracy from 80.4% to 94.98% and an average FPR of 12.1
What carries the argument
The load-bearing piece is the Fail-Slow Sketch, a two-stage streaming data structure: Stage-1 uses d hash tables with per-bucket frequency counters (increment on match, decrement on collision) to identify recurring trace patterns; Stage-2 keeps a bounded FIFO list of candidate fail-slow patterns with aggregated statistics such as data volume, timestamps, and duration. Around it, the SL-Compiler defines probes by a five-tuple (fragment, type, location, level, structure) to insert monitoring pseudo-instructions, and the SL-Tracer builds a multi-level communication graph in which nodes are cores at time windows, edges carry normalized propagation weights derived from traffic volume, and virtual
Load-bearing premise
The results rest on the assumption that the simulator's injected fail-slow behavior—a fixed 10x slowdown lasting 0-10 seconds on top of normal performance variance—represents how real on-chip cores and links actually fail; if real failures are more intermittent, milder, or noisier, the measured accuracy and overhead may not transfer to silicon.
What would settle it
Run the same probe, compression, and ranking pipeline on a real many-core accelerator (or a cycle-accurate simulator calibrated to silicon) with a controlled 10x slowdown injected into a specific core and link; if root-cause accuracy falls well below the 86.77% average, or if trace storage exceeds the KB budget, the central claim is contradicted. A cheaper check: simulate slowdown factors of 2x, 5x, and intermittent patterns and observe whether accuracy collapses, which would show the method is brittle to realistic failure signatures.
If this is right
- On-chip fail-slow detection becomes feasible in KB-scale SRAM rather than MB-scale tracing, so continuous monitoring of every core and link is possible without dedicated off-chip logging.
- Root-cause localization can distinguish the true slow component from neighbors that only appear slow due to propagated stalls, by fusing the software dependency graph with the physical topology.
- The framework generalizes across mesh sizes and workloads, and the probe overhead stays below 10%, suggesting it can ride along in production DNN accelerators.
- With a ranked list of fail-slow candidates, the chip can trigger mitigation actions—e.g., voltage/frequency scaling, task migration, or re-mapping—before a slowdown becomes a full stall.
Where Pith is reading between the lines
- The reported accuracy is tied to the injected failure model (fixed 10x slowdown, durations 0-10 s, normally distributed core capacity, Gamma link latency). A natural testable extension is to sweep slower degradation factors (2x-5x) and intermittent recovery to see where accuracy degrades.
- Because FailRank operates on a hardware-agnostic multi-level graph, the same pipeline likely extends to torus or dragonfly topologies, as long as deterministic routing can be assumed for the link-inference equations.
- The paper's own admission that DarkNet-19's uniform mapping produces correlated traces and lowers accuracy suggests a tuning lever: the EM-based link inference could be stabilized with prior knowledge of the mapping or regularization.
- The sketch's parameters (hash count, bucket count, threshold) show clear trade-offs; an online parameter-adaptation scheme could maintain accuracy when workload characteristics drift over time.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SLOTH/SlowPoke, a framework for detecting and localizing fail-slow failures in many-core DNN accelerators. It comprises compiler-inserted probes (SL-Compiler), a sketch-based streaming compressor (SL-Recorder) that retains high-frequency trace patterns in a two-stage structure, and a graph-based diagnosis pipeline (SL-Tracer) that builds a multi-level communication graph and runs a PageRank-like FailRank algorithm. The framework is evaluated in a SimPy event-driven simulation on five workloads (binary tree, GoogLeNet, DarkNet-19, VGG, ResNet-50) with injected core and link fail-slow faults. The central claims are an average 115.9x reduction in trace storage, 86.77% average fail-slow detection accuracy, and 12.11% FPR.
Significance. If the reported results hold, the work would provide a genuinely lightweight, topology-aware mechanism for on-chip fail-slow detection, an area where prior distributed-systems methods are not directly applicable. The paper is clearly structured, releases source code and failure datasets, and includes a mathematical retention bound for its compressor. However, the headline numbers come exclusively from a custom simulation with a narrow injected fault model and with parameters selected on the same workloads used for evaluation; as presented, they do not yet substantiate the practical claims.
major comments (4)
- [§4.1, Table 2] The core evaluation protocol makes the headline numbers hard to trust. The fault model fixes the slowdown rate at 10x and duration uniform in [0,10]s, with negative samples generated by 'modifying DNN structures' rather than by observing healthy runs; no sensitivity to milder/intermittent fail-slow behavior is reported. The normal/Gamma variance parameters of §2.3 are never given. Table 2's denominators are also inconsistent with the stated 152-failure dataset (e.g., DarkNet-19 FPR is 22/174), and no error bars or repeated-run statistics are shown despite §4.3 claiming repetitions. Please clarify the metric definitions, reconcile sample counts, and add sensitivity analysis over slowdown rate/duration/noise.
- [§3.4.3, §4.5] The reported detection accuracy is not an out-of-sample result. SL-Recorder's parameters (H, B, S, T) are selected by DSE on the same five workloads that appear in Table 2, and the FailRank coefficients α=0.1, β=0.3, γ=0.6 are fixed by hand tuning. There is no training/test split, cross-validation, or hold-out workload, so the average 86.77% could reflect overfitting to these graphs. Please report performance on held-out workloads or nested cross-validation, and show sensitivity of accuracy to the FailRank coefficients and SL-Recorder configuration.
- [§3.3, Lemma 3.1] The retention bound is not proven as stated. The argument equates the event F_{i,j,k} ≤ f_i − H with the bucket recording t_i, but in Algorithm 1 a collision with another key decrements the current occupant's counter rather than incrementing t_i's counter; whether t_i reaches H depends on arrival order, not just aggregate counts. The proof also ignores Stage-2 eviction (MAX_LENGTH/FIFO), and the sentence 'there must be a j' has the wrong quantifier. The lemma may be recoverable with additional assumptions (e.g., random ordering, no eviction), but as written it does not support the retention guarantee claimed in §3.3.
- [§3.5, §4.3] All results come from a custom SimPy simulator with no hardware, RTL, cycle-accurate NoC, or FPGA validation. Probe costs are modeled as a fixed 10-cycle clock-read latency, and the EM link-inference and FailRank are exercised only under the simulator's own injected variance distributions. For a paper claiming 'practical on-chip' operation and KB-scale memory, at least a cycle-accurate NoC model or a mapping to a realistic RTL budget is needed to make the overhead and accuracy claims credible.
minor comments (5)
- [Title/Abstract/Conclusion] The title and the body disagree: the manuscript is titled 'SLOTH' but the full text and conclusion consistently say 'SlowPoke' (the GitHub URL uses sloth). The abstract also reports accuracy 'from 69.68% to 86.69%' while the conclusion gives 86.77%; these numbers must be harmonized.
- [§2.3] The distributions for healthy core capacity and link latency are stated qualitatively, but the actual parameter values (μ_c, σ_c, α, β) are never specified, making the simulation non-reproducible.
- [§3.4.2] The EM-based link bandwidth inference is described only in prose; no update equations, initialization, convergence criterion, or handling of the underdetermined system is given. Please provide the formal derivation or pseudocode.
- [§3.4.3, §4.5] The symbols α, β, γ are reused for the FailRank edge-update coefficients and for the DSE objective COST=ACC^α×R^β×M^γ with conflicting meanings. Rename one set to avoid confusion.
- [§4.4, Figure 12] The heatmap labels such as 'S=8192,T=10' are not fully defined in the caption; please state the default values of H, B, S, T and the meaning of axes.
Circularity Check
No significant circularity: the evaluation is an empirical simulator study, not a result forced by self-citation or by construction.
full rationale
The paper's central claims are empirical measurements from a SimPy-based simulator: probe overhead, trace compression ratio, and fail-slow detection accuracy against a generated ground-truth dataset. These claims are not derived from the inputs by construction. Lemma 3.1 is an independent mathematical bound on the sketch's retention probability, derived from the insertion algorithm and Markov's inequality, not from the target detection results. The detector pipeline (outlier detection, EM link inference, FailRank) does not read the injected failure labels during inference; it processes simulated traces and is then compared with the injected ground truth, which is a legitimate evaluation design. The DSE of SL-Recorder parameters on the same workloads and the hand-set FailRank coefficients (alpha=0.1, beta=0.3, gamma=0.6) are generalization/overfitting limitations rather than circular reductions: the reported accuracy is a measured quantity, not an algebraic consequence of those tuned values. The self-citations in the related work (ScalAna, Vapro) are comparison references, not load-bearing support for SLOTH's claimed results. The fixed 10x slowdown, 0-10s duration injected failure model is an external-validity threat for real milder fail-slow signatures, but it is an input modeling assumption, not a circularity in the paper's derivation chain.
Axiom & Free-Parameter Ledger
free parameters (6)
- SL-Recorder parameters (H hash functions, B buckets, S stage-2 size, T threshold) =
Not reported; selected via DSE on the same workloads
- FailRank coefficients (α, β, γ) =
0.1, 0.3, 0.6
- FailRank damping constant λ =
Not specified
- Injected fail-slow slowdown rate =
10×
- Injected failure duration distribution =
Uniform [0, 10] s
- Core/link performance-variance distribution parameters =
Normal(μ_c, σ_c²) for cores; Gamma(α, β) for links; values not reported
axioms (7)
- standard math Markov inequality and independence of the d hash functions in the Lemma 3.1 retention bound
- domain assumption Deterministic routing on a 2D mesh maps each communication event to a fixed set of links
- domain assumption Core computing capacity follows a normal distribution and link latency follows a Gamma distribution
- domain assumption The SimPy simulator faithfully captures execution timing, resource contention, and fail-slow propagation
- ad hoc to paper Fail-slow failures are representable as a fixed 10× slowdown over a contiguous window
- ad hoc to paper Negative samples can be synthesized by modifying DNN structures instead of observing healthy runs
- ad hoc to paper Stage-2 FIFO eviction and MAX_LENGTH heuristics preserve detection accuracy
invented entities (1)
-
Virtual DRAM nodes in the Multi-Level Communication Graph
no independent evidence
Cite this review
Pith. "Pith review of SLOTH: Lightweight Detection and Localization of On-Chip Fail-Slow Failures for DNN Accelerators." pith.science (2026). https://pith.science/paper/O2XS2KDX
@misc{pith2026251024112,
author = {Pith},
title = {Pith review of: SLOTH: Lightweight Detection and Localization of On-Chip Fail-Slow Failures for DNN Accelerators},
year = {2026},
howpublished = {\url{https://pith.science/paper/O2XS2KDX}},
note = {Machine review of arXiv:2510.24112}
}
abstract
Spatial DNN accelerators are essential for high-performance inference, but their performance is undermined by widespread fail-slow failures. Detecting such failures on-chip is challenging, as prior methods from distributed systems are unsuitable due to strict memory limits and their inability to track failures across the hardware topology. We present SLOTH, a lightweight, hardware-aware framework for practical on-chip fail-slow detection in DNN accelerators. SLOTH combines workload-aware instrumentation for operator-level monitoring with minimal overhead, on-the-fly trace compression to operate within kilobytes of memory, and a novel topology-aware ranking algorithm to pinpoint a failure's root cause. We evaluate SLOTH on a wide range of representative DNN workloads. The results demonstrate that SLOTH reduces the storage overhead by an average of 115.9$\times$, while achieving an average fail-slow detection accuracy from 69.68\% to 86.69\%.
Figures
Reference graph
Works this paper leans on
-
[1]
Mridul Agarwal, Bipul C Paul, Ming Zhang, and Subhasish Mitra. 2007. Circuit failure prediction and its application to transistor aging. In 25th IEEE VLSI Test Symposium (VTS’07). IEEE, IEEE, Berkeley, CA, USA, 277–286
2007
-
[2]
Armin Ahmadzadeh and Hamid Sarbazi-Azad. 2023. Fast and scal- able quantum computing simulation on multi-core and many-core platforms.Quantum Information Processing22 (05 2023). doi:10.1007/ s11128-023-03955-w
2023
-
[3]
Ramnatthan Alagappan, Aishwarya Ganesan, Yuvraj Patel, Thanu- malayan Sankaranarayana Pillai, Andrea C Arpaci-Dusseau, and Remzi H Arpaci-Dusseau. 2016. Correlated crash vulnerabilities. In 12th USENIX Symposium on Operating Systems Design and Implemen- tation (OSDI 16). USENIX Association, Savannah, GA, USA, 151–167
2016
-
[4]
Rajeshwari Banakar, Stefan Steinke, Bo-Sik Lee, Mahesh Balakrishnan, and Peter Marwedel. 2002. Scratchpad memory: design alternative for cache on-chip memory in embedded systems. InProceedings of the tenth international symposium on Hardware/software codesign. ACM, Estes Park, Colorado, USA, 73–78
2002
-
[5]
Kshitij Bhardwaj, Koushik Chakraborty, and Sanghamitra Roy. 2012. Towards graceful aging degradation in NoCs through an adaptive routing algorithm. InProceedings of the 49th Annual Design Automation Conference. ACM, San Francisco, California USA, 382–391
2012
-
[6]
Tobias Bjerregaard and Shankar Mahadevan. 2006. A survey of re- search and practices of network-on-chip.ACM Computing Surveys (CSUR)38, 1 (2006), 1–es
2006
-
[7]
Jingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei, Zhanhong Tan, Guiming Shi, Mingyu Gao, and Kaisheng Ma. 2024. Gemini: Mapping and architecture co-exploration for large-scale dnn chiplet accelerators. In2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, IEEE, Edinburgh, United Kingdom, 156– 171
2024
-
[8]
Yung-Chang Chang, Ching-Te Chiu, Shih-Yin Lin, and Chung-Kai Liu
-
[9]
Yu-Hsin Chen, Joel Emer, and Vivienne Sze. 2016. Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks.ACM SIGARCH computer architecture news44, 3 (2016), 367–379
2016
-
[10]
Allen Clement, Edmund Wong, Lorenzo Alvisi, Mike Dahlin, Mirco Marchetti, et al. 2009. Making Byzantine fault tolerant systems tolerate Byzantine faults. InProceedings of the 6th USENIX symposium on Net- worked systems design and implementation. The USENIX Association, USENIX Association, Boston, MA, 153–168
2009
-
[11]
Guojing Cong and Konstantin Makarychev. 2012. Optimizing Large- scale Graph Analysis on Multithreaded, Multicore Platforms. In2012 IEEE 26th International Parallel and Distributed Processing Symposium. IEEE, Shanghai, China, 414–425. doi:10.1109/IPDPS.2012.46
-
[12]
Jack B Dennis and David P Misunas. 1974. A preliminary architec- ture for a basic data-flow processor. InProceedings of the 2nd annual symposium on Computer architecture. ACM, Barcelona, Spain, 126–132
1974
-
[13]
Pratyush Dhingra, Jana Doppa, and Partha Pratim Pande. 2024. HeT- raX: Energy Efficient 3D Heterogeneous Manycore Architecture for Transformer Acceleration. InProceedings of the 29th ACM/IEEE Inter- national Symposium on Low Power Electronics and Design(Newport Beach, CA, USA)(ISLPED ’24). Association for Computing Machinery, New York, NY, USA, 1–6. doi:1...
arXiv 2024
-
[14]
Pratyush Dhingra, Janardhan Rao Doppa, and Partha Pratim Pande
-
[15]
Bernhard Egger, Jaejin Lee, and Heonshik Shin. 2006. Scratchpad memory management for portable systems with a memory man- agement unit. InProceedings of the 6th ACM & IEEE International Conference on Embedded Software(Seoul, Korea)(EMSOFT ’06). As- sociation for Computing Machinery, New York, NY, USA, 321–330. doi:10.1145/1176887.1176933
arXiv 2006
-
[16]
Stijn Eyerman, Wim Heirman, Kristof Du Bois, Joshua B. Fryman, and Ibrahim Hur. 2018. Many-Core Graph Workload Analysis. InSC18: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, Texas, Dallas, 282–292. doi:10.1109/SC. 2018.00025
arXiv 2018
-
[17]
Yinxiao Feng, Wei Li, and Kaisheng Ma. 2024. Ring Road: A Scalable Polar-Coordinate-based 2D Network-on-Chip Architecture. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, IEEE, Austin, TX, USA, 871–884
2024
-
[18]
Yinxiao Feng, Dong Xiang, and Kaisheng Ma. 2023. A scalable method- ology for designing efficient interconnection network of chiplets. In 2023 IEEE International Symposium on High-Performance Computer 12 SlowPoke: Understanding and Detecting On-Chip Fail-Slow Failures in Many-Core Systems Conference’17, July 2017, Washington, DC, USA Architecture (HPCA). ...
2023
-
[19]
Yu Gan, Mingyu Liang, Sundar Dev, David Lo, and Christina Delim- itrou. 2021. Sage: practical and scalable ML-driven performance debug- ging in microservices. InProceedings of the 26th ACM International Con- ference on Architectural Support for Programming Languages and Oper- ating Systems(Virtual, USA)(ASPLOS ’21). Association for Computing Machinery, Ne...
arXiv 2021
-
[20]
Yu Gan, Guiyang Liu, Xin Zhang, Qi Zhou, Jiesheng Wu, and Jiangwei Jiang. 2024. Sleuth: A Trace-Based Root Cause Analysis System for Large-Scale Microservices with Graph Neural Networks. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4(Vancouver, BC, Canada)(ASPLOS ’2...
arXiv 2024
-
[21]
Paul Gratz, Changkyu Kim, Robert McDonald, Stephen W Keckler, and Doug Burger. 2006. Implementation and evaluation of on-chip network architectures. In2006 International Conference on Computer Design. IEEE, IEEE, San Jose, CA, USA, 477–484
2006
-
[22]
Haryadi S Gunawi, Riza O Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Xing Lin, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, et al. 2018. Fail-slow at scale: Evidence of hardware performance faults in large production systems. ACM Transactions on Storage (TOS)14, 3 (2018), 1–26
2018
-
[23]
Mohammad-Hashem Haghbayan, Antonio Miele, Zhuo Zou, Hannu Tenhunen, and Juha Plosila. 2020. Thermal-cycling-aware dynamic reliability management in many-core system-on-chip. In2020 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, IEEE, Grenoble, France, 1229–1234
2020
-
[24]
Edward Hanson, Shiyu Li, Guanglei Zhou, Feng Cheng, Yitu Wang, Rohan Bose, Hai Li, and Yiran Chen. 2023. Si-kintsugi: Towards recov- ering golden-like performance of defective many-core spatial architec- tures for ai. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture. ACM, ON, Toronto, Canada, 972–985
2023
-
[25]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Las Vegas, NV, USA, 770–778. doi:10.1109/CVPR.2016.90
-
[26]
Yi He, Mike Hutton, Steven Chan, Robert De Gruijl, Rama Govindaraju, Nishant Patil, and Yanjing Li. 2023. Understanding and mitigating hardware failures in deep learning training systems. InProceedings of the 50th Annual International Symposium on Computer Architecture. ACM, FL, Orlando, USA, 1–16
2023
-
[27]
Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, and Christopher W Fletcher. 2021. Mind mappings: enabling efficient algorithm-accelerator mapping space search. InProceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. ACM, Virtual, USA, 943–958
2021
-
[28]
Carles Hernández, Federico Silla, and José Duato. 2010. A methodology for the characterization of process variation in NoC links. In2010 Design, Automation & Test in Europe Conference & Exhibition (DATE 2010). IEEE, IEEE, Dresden, Germany, 685–690
2010
-
[29]
Haiyu Huang, Cheng Chen, Kunyi Chen, Pengfei Chen, Guangba Yu, Zilong He, Yilun Wang, Huxing Zhang, and Qi Zhou. 2025. Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(...
arXiv 2025
-
[30]
Zicheng Huang, Pengfei Chen, Guangba Yu, Hongyang Chen, and Zibin Zheng. 2021. Sieve: Attention-based Sampling of End-to-End Trace Data in Distributed Microservice Systems. In2021 IEEE Inter- national Conference on Web Services (ICWS). IEEE, Chicago, IL, USA, 436–446. doi:10.1109/ICWS53863.2021.00063
arXiv 2021
-
[31]
Yuichi Inadomi, Tapasya Patki, Koji Inoue, Mutsumi Aoyagi, Barry Rountree, Martin Schulz, David Lowenthal, Yasutaka Wada, Keiichiro Fukazawa, Masatsugu Ueda, et al . 2015. Analyzing and mitigating the impact of manufacturing variability in power-constrained super- computing. InProceedings of the international conference for high per- formance computing, n...
2015
-
[32]
Yuyang Jin, Haojie Wang, Teng Yu, Xiongchao Tang, Torsten Hoefler, Xu Liu, and Jidong Zhai. 2020. ScalAna: Automating scaling loss detection with graph analysis. InSC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, IEEE, Atlanta, GA, USA, 1–14
2020
-
[33]
2018.Performance Analysis of OpenSHMEM Applications with TAU Commander
Samuel Khuvis, Sameer Shende, Allen Malony, Neena Imam, and Man- junath Gorentla Venkata. 2018.Performance Analysis of OpenSHMEM Applications with TAU Commander. Springer, Cham, Cham, Switzer- land, 161–179. doi:10.1007/978-3-319-73814-7_11
-
[34]
Yudhishthira Kundu, Manroop Kaur, Tripty Wig, Kriti Kumar, Push- panjali Kumari, Vivek Puri, and Manish Arora. 2025. A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU- based Systems for Artificial Intelligence. arXiv:2503.11698 [cs.AR] https://arxiv.org/abs/2503.11698
Pith/arXiv arXiv 2025
-
[35]
Hyoukjun Kwon, Prasanth Chatarasi, Vivek Sarkar, Tushar Krishna, Michael Pellauer, and Angshuman Parashar. 2020. MAESTRO: A Data- Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings.IEEE Micro40, 3 (2020), 20–29. doi:10.1109/ MM.2020.2985963
arXiv 2020
-
[36]
Weihe Li and Paul Patras. 2024. Stable-Sketch: A Versatile Sketch for Accurate, Fast, Web-Scale Data Stream Processing. InProceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ’24). Association for Computing Machinery, New York, NY, USA, 4227–4238. doi:10.1145/3589334.3645581
arXiv 2024
-
[37]
Zeyan Li, Junjie Chen, Rui Jiao, Nengwen Zhao, Zhijun Wang, Shuwei Zhang, Yanjun Wu, Long Jiang, Leiqin Yan, Zikai Wang, Zhekang Chen, Wenchi Zhang, Xiaohui Nie, Kaixin Sui, and Dan Pei. 2021. Practical Root Cause Localization for Microservice Systems via Trace Analysis. In2021 IEEE/ACM 29th International Symposium on Quality of Service (IWQOS). IEEE, Tok...
arXiv 2021
-
[38]
Jinjin Lin, Pengfei Chen, and Zibin Zheng. 2018. Microscope: Pin- point Performance Issues with Causal Graphs in Micro-service Envi- ronments. InService-Oriented Computing: 16th International Confer- ence, ICSOC 2018, Hangzhou, China, November 12-15, 2018, Proceed- ings(Hangzhou, China). Springer-Verlag, Berlin, Heidelberg, 3–20. doi:10.1007/978-3-030-03596-9_1
-
[39]
Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Chenyuan Wang, Zuocheng Shi, Xiang Shi, Wei Jia, et al. 2025. Understanding Stragglers in Large Model Training Us- ing What-if Analysis.arXiv preprint arXiv:2505.05713(2025), 19 pages
Pith/arXiv arXiv 2025
-
[40]
Shu-Yen Lin and Jin-Yi Lin. 2017. Thermal-and performance-aware address mapping for the multi-channel three-dimensional DRAM sys- tems.IEEE Access5 (2017), 5566–5577
2017
-
[41]
Hatem Ltaief, Yuxi Hong, Leighton Wilson, Mathias Jacquelin, Matteo Ravasi, and David Elliot Keyes. 2023. Scaling the “Memory Wall” for Multi-Dimensional Seismic Processing with Algebraic Compression on Cerebras CS-2 Systems. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, CO, USA)(SC...
arXiv 2023
-
[42]
Ruiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Jiwu Shu, Minglu Li, and Jiesheng Wu. 2023. PERSEUS: a fail-slow detection framework for cloud storage systems. InProceedings of the 21st USENIX Conference on File and Storage Technologies(Santa Clara, CA, USA)(FAST’23). 13 Conference’17, July 2017, Wa...
2023
-
[43]
Ruiming Lu, Erci Xu, Yiming Zhang, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Minglu Li, and Jiesheng Wu. 2022. {NVMe}{ SSD} failures in the field: the {Fail-Stop} and the{Fail- Slow}. In2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA, USA, 1005–1020
2022
-
[44]
Jonathan Mace, Ryan Roelke, and Rodrigo Fonseca. 2015. Pivot tracing: dynamic causal monitoring for distributed systems. InProceedings of the 25th Symposium on Operating Systems Principles(Monterey, California)(SOSP ’15). Association for Computing Machinery, New York, NY, USA, 378–393. doi:10.1145/2815400.2815415
arXiv 2015
-
[45]
Aniruddha Marathe, Yijia Zhang, Grayson Blanks, Nirmal Kumbhare, Ghaleb Abdulla, and Barry Rountree. 2017. An empirical survey of performance and energy efficiency variation on intel processors. In Proceedings of the 5th International Workshop on Energy Efficient Su- percomputing. ACM, CO, Denver, USA, 1–8
2017
-
[46]
2025.Architectures for Scientific Computing
Farhad Merchant. 2025.Architectures for Scientific Computing. Springer Nature Singapore, Singapore, 401–414. doi:10.1007/978-981-97-9314- 3_16
-
[47]
K. G. Müller, T.Vignaux, O. Lünsdorf, and S. Scherfke. 2002.SimPy: Discrete Event Simulation for Python. Team SimPy.https://simpy. readthedocs.io/Accessed: 2023-11-13
2002
-
[48]
Rishiyur S Nikhil et al. 2002. Executing a program on the MIT tagged- token dataflow architecture.IEEE Transactions on computers39, 3 (2002), 300–318
2002
-
[49]
2019.{IASO}: A{Fail-Slow} Detection and Mitigation Framework for Distributed Storage Services
Biswaranjan Panda, Deepthi Srinivasan, Huan Ke, Karan Gupta, Vinayak Khot, and Haryadi S Gunawi. 2019.{IASO}: A{Fail-Slow} Detection and Mitigation Framework for Distributed Storage Services. In2019 USENIX Annual Technical Conference (USENIX ATC 19). USENIX Association, Renton, WA, 47–62
2019
-
[50]
Shailja Pandey, Sayam Sethi, and Preeti Ranjan Panda. 2024. 3D- TemPo: Optimizing 3D DRAM performance under temperature and power constraints.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems43, 8 (2024), 2263 – 2276
2024
-
[51]
Thanumalayan Sankaranarayana Pillai, Ramnatthan Alagappan, Lanyue Lu, Vijay Chidambaram, Andrea C Arpaci-Dusseau, and Remzi H Arpaci-Dusseau. 2017. Application crash consistency and performance with CCFS.ACM Transactions on Storage (TOS)13, 3 (2017), 1–29
2017
-
[52]
Joseph Redmon and Ali Farhadi. 2017. YOLO9000: Better, Faster, Stronger. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Honolulu, HI, USA, 6517–6525. doi:10.1109/ CVPR.2017.690
2017
-
[53]
Bogdan F Romanescu, Sule Ozev, and Daniel J Sorin. 2006. Quanti- fying the impact of process variability on microprocessor behavior. InWorkshop on Architectural Reliability. Citeseer, Orlando, Florida, 10 pages
2006
-
[54]
Siva Kumar Sastry Hari, Man-Lap Li, Pradeep Ramachandran, Byn Choi, and Sarita V Adve. 2009. mSWAT: Low-cost hardware fault detection and diagnosis for multicore systems. InProceedings of the 42nd Annual IEEE/ACM International Symposium on Microarchitecture. IEEE, New York, NY, USA, 122–132
2009
-
[55]
Prachi Shukla, Ayse K Coskun, Vasilis F Pavlidis, and Emre Salman
-
[56]
Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv:1409.1556 [cs.CV]https://arxiv.org/abs/1409.1556
Pith/arXiv arXiv 2015
-
[57]
Slota, Sivasankaran Rajamanickam, and Kamesh Madduri
George M. Slota, Sivasankaran Rajamanickam, and Kamesh Madduri
-
[58]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015. Going deeper with convolutions. InProceedings of the IEEE conference on computer vision and pattern recognition. IEEE, Boston, MA, USA, 1–9
2015
-
[59]
Ebadollah Taheri, Sudeep Pasricha, and Mahdi Nikdast. 2022. DeFT: A deadlock-free and fault-tolerant routing algorithm for 2.5 D chiplet networks. In2022 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, IEEE, Antwerp, Belgium, 1047–1052
2022
-
[60]
Xiongchao Tang, Jidong Zhai, Xuehai Qian, Bingsheng He, Wei Xue, and Wenguang Chen. 2018. vSensor: leveraging fixed-workload snip- pets of programs for performance variance detection. InProceedings of the 23rd ACM SIGPLAN symposium on principles and practice of parallel programming. ACM, Vienna Austria, 124–136
2018
-
[61]
M Bedford Taylor, Walter Lee, Saman Amarasinghe, and Anant Agar- wal. 2003. Scalar operand networks: On-chip interconnect for ILP in partitioned architectures. InThe Ninth International Symposium on High-Performance Computer Architecture, 2003. HPCA-9 2003. Proceed- ings.IEEE, IEEE, Anaheim, CA, USA, 341–353
2003
-
[62]
Min Tian, Junjie Wang, Zanjun Zhang, Wei Du, Jingshan Pan, and Tao Liu. 2022. swSuperLU: A highly scalable sparse direct solver on Sunway manycore architecture.J. Supercomput.78, 9 (June 2022), 11441–11463. doi:10.1007/s11227-021-04270-w
-
[63]
Vasileios Tsoutsouras, Dimosthenis Masouros, Sotirios Xydis, and Dim- itrios Soudris. 2017. SoftRM: Self-organized fault-tolerant resource management for failure detection and recovery in NoC based many- cores.ACM Transactions on Embedded Computing Systems (TECS)16, 5s (2017), 1–19
2017
-
[64]
Hanzhang Wang, Zhengkai Wu, Huai Jiang, Yichao Huang, Jiamu Wang, Selcuk Kopru, and Tao Xie. 2022. Groot: an event-graph-based approach for root cause analysis in industrial settings. InProceedings of the 36th IEEE/ACM International Conference on Automated Software Engineering(Melbourne, Australia)(ASE ’21). IEEE Press, Melbourne, Australia, 419–429. doi:...
arXiv 2022
-
[65]
Lihuan Wang, Shuyan Jiang, Shuyu Chen, Junshi Wang, and Letian Huang. 2019. Optimized mapping algorithm to extend lifetime of both NoC and cores in many-core system.Integration67 (2019), 82–94
2019
-
[66]
Pengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang, Peng Zhang, Tao Tang, and Zheng Wang. 2023. Optimizing Direct Convolutions on ARM Multi-Cores. InSC23: International Conference for High Performance Computing, Networking, Storage and Analysis. ACM, CO, Denver, USA, 1–14
2023
-
[67]
Xiaodong Wang and José F. Martínez. 2015. XChange: A market-based approach to scalable dynamic multi-resource allocation in multicore architectures. In2015 IEEE 21st International Symposium on High Per- formance Computer Architecture (HPCA). IEEE, Burlingame, CA, USA, 113–125. doi:10.1109/HPCA.2015.7056026
arXiv 2015
-
[68]
Yipeng Wang, Tong Yang, Ren Wang, and Charlie Tai. 2019. Dynamic Sketch: Efficient and Adjustable Heavy Hitter Detection for Software Packet Processing. In2019 IEEE 8th International Conference on Cloud Networking (CloudNet). IEEE, Coimbra, Portugal, 1–7. doi:10.1109/ CloudNet47604.2019.9064148
arXiv 2019
-
[69]
Felix Wolf, Brian JN Wylie, Erika Abraham, Daniel Becker, Wolfgang Frings, Karl Fürlinger, Markus Geimer, Marc-André Hermanns, Bernd Mohr, Shirley Moore, et al. 2008. Usage of the SCALASCA toolset for scalable performance analysis of large-scale parallel applications. In Tools for High Performance Computing: Proceedings of the 2nd Inter- national Workshop...
2008
-
[70]
Tianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang, Wenchao Wu, Qinkai Duan, Guodong Yang, Jiamang Wang, Lin Qu, and Liping Zhang
-
[71]
Siyuan Xiao, Xiaohang Wang, Maurizio Palesi, Amit Kumar Singh, and Terrence Mak. 2019. ACDC: An accuracy-and congestion-aware dynamic traffic control method for networks-on-chip. In2019 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, IEEE, Florence, Italy, 630–633
2019
-
[72]
Zheng Xu, Dehao Kong, Jiaxin Liu, Jinxi Li, Jingxiang Hou, Xu Dai, Chao Li, Shaojun Wei, Yang Hu, and Shouyi Yin. 2025. WSC-LLM: Efficient LLM Service and Architecture Co-exploration for Wafer-scale Chips. InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association for Computing Machin- ery, New York, NY, USA,...
arXiv 2025
-
[73]
Zikang Xu, Yiming Zhang, and Zhirong Shen. 2025. A Fail-Slow Detection Framework for HBM Devices. InProceedings of the 30th Asia and South Pacific Design Automation Conference. USENIX Association, Santa Clara, CA, 491–497
2025
-
[74]
Lingxiang Yin, Amir Ghazizadeh, Ahmed Louri, and Hao Zheng. 2023. ARIES: Accelerating Distributed Training in Chiplet-Based Systems via Flexible Interconnects. In2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD). IEEE, San Francisco, CA, USA, 1–9. doi:10.1109/ICCAD57390.2023.10323955
arXiv 2023
-
[75]
Xin You, Zhibo Xuan, Hailong Yang, Zhongzhi Luan, Yi Liu, and Depei Qian. 2024. GVARP: Detecting Performance Variance on Large-Scale Heterogeneous Systems. InSC24: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, IEEE, Atlanta, GA, USA, 1–16
2024
-
[76]
Guangba Yu, Pengfei Chen, Hongyang Chen, Zijie Guan, Zicheng Huang, Linxiao Jing, Tianjun Weng, Xinmeng Sun, and Xiaoyun Li
-
[77]
Jidong Zhai, Wenguang Chen, and Weimin Zheng. 2010. PHANTOM: predicting performance of parallel applications on large-scale parallel machines using a single node. InProceedings of the 15th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming(Banga- lore, India)(PPoPP ’10). Association for Computing Machinery, New York, NY, USA, 305–314...
arXiv 2010
-
[78]
Lei Zhang, Zhiqiang Xie, Vaastav Anand, Ymir Vigfusson, and Jonathan Mace. 2023. The Benefit of Hindsight: Tracing Edge-Cases in Dis- tributed Systems. In20th USENIX Symposium on Networked Systems De- sign and Implementation (NSDI 23). USENIX Association, Boston, MA, 321–339.https://www.usenix.org/conference/nsdi23/presentation/ zhang-lei
2023
-
[79]
Nathan Zhang, Rubens Lacouture, Gina Sohn, Paul Mure, Qizheng Zhang, Fredrik Kjolstad, and Kunle Olukotun. 2024. The Dataflow Abstract Machine Simulator Framework. InProceedings of the 51st Annual International Symposium on Computer Architecture (ISCA ’24). IEEE Press, Buenos Aires, Argentina, 532–547. doi:10.1109/ISCA59077. 2024.00046
arXiv 2024
-
[80]
Hao Zheng, Ke Wang, and Ahmed Louri. 2021. Adapt-noc: A flexible network-on-chip design for heterogeneous manycore architectures. In2021 IEEE international symposium on high-performance computer architecture (HPCA). IEEE, IEEE, Seoul, Korea (South), 723–735
2021
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.